You measure AI visibility the way anyone measures something unstable: ask a fixed set of real buyer questions, on a schedule, across the engines your buyers actually use, and report the share of answers your business is named in, with the date and the denominator attached. That share is the number. Everything else on the slide is decoration. At Probably Genius we run that panel monthly across five engines and keep every answer word for word, because a number you cannot re-derive next month is not a measurement.
This question usually arrives with an invoice sitting next to it. You are three months into a program, the update says the AI work is on track, and there is nothing on the page you could check even if you wanted to. So you scroll for a screenshot, find one where a competitor is named, and feel worse without knowing anything new.
That is not a failure of attention. You were handed a category that arrived without an instrument, in the same years that gave everyone an instrument for everything else. Rankings had a number. Traffic had a number. This has a number too. It is just a different kind of number, and almost nobody hands it to you correctly.
Here is what "working" means in leading and lagging terms, why one answer proves nothing, the single figure worth putting on a report, how to handle more than one engine, what to ask a vendor before you pay them, and when a flat result is genuinely a reason to change course.
What "Working" Actually Means
Start by separating four things that get collapsed into one word. Being mentioned. Being cited as a source. Being recommended by name when the buyer asked who to hire. And a visit or a call you can actually trace back. They are different events with different value, and a report that blurs them can call almost anything progress.
Recommendation is the one you are paying for. Citation is pleasant and not the same thing, which is worth sitting with for a moment: an engine can quote your article and then name someone else as the person to call. We wrote about that gap in why AI can cite you and still recommend someone else.
Then split your indicators into leading and lagging. Leading indicators move first and tell you whether the machinery is working: whether the engines can reach and read your pages, whether you appear at all on the buyer questions that matter, how accurately you are described when you do appear, and which competitors are named instead. Lagging indicators are the money: inbound calls, qualified leads, closed work, and whatever traffic your analytics can honestly attribute.
Judge the first quarter on leading indicators only. Attribution for assistant-driven work is a lower bound by nature, because some of those buyers read the answer, look you up separately, and arrive with no referrer at all. Your analytics will file them under direct. They were not direct. They were briefed.
- Count appearances, not impressions of effort. Articles published is an activity number. Answers you appear in is an outcome number.
- Record how you are described. Being named as "a general marketing firm" when you are a specialist is a visibility result and a positioning problem at the same time.
- Keep the competitor list. The names that appear instead of yours are the most useful free clue you get about what the engines can currently verify.
Why One Screenshot Proves Nothing
The single most useful fact in this whole subject is that these systems do not repeat themselves. In research published in January 2026, Rand Fishkin of SparkToro and Patrick O'Donnell of Gumshoe had 600 volunteers run 12 identical prompts through ChatGPT, Claude and Google's AI 2,961 times. The odds of any two responses returning the same list of brands came in under 1 in 100. The odds of the same list in the same order were closer to 1 in 1,000.
Fishkin's conclusion was blunt: any tool selling you a ranking position in AI is, in his words, full of baloney. What he endorsed instead was a visibility percentage across dozens to hundreds of prompts, run multiple times.
Two 2026 papers put statistical shape under that. In "Quantifying Uncertainty in AI Visibility," posted March 9, 2026, Ronald Sielinski sampled three generative search platforms across three topics, daily for nine days and again at ten minute intervals, and found that many apparent differences between domains fall within the noise floor of the measurement process. Single run metrics, he concluded, give a misleadingly precise picture. A second paper that April, "Don't Measure Once," argued the same case from the other end: visibility should be reported as a distribution, not a single point.
The most practical finding is about how much sampling actually buys you. In a July 2026 variance study of 12,933 responses across 20 brands, eight languages and three models, scoring how each answer talked about each brand, Dmitrij Zatuchin found that a single answer carries almost no brand-discriminating signal, and that a repeat past the fifth run reduces error variance by roughly three ten-thousandths. Reliability, he concluded, is bought by spreading across models rather than by hammering one prompt. On that evidence, roughly five runs per question across several engines is where the sampling budget stops paying. Treat it as a starting design rather than a law.
So the screenshot in your inbox is not evidence and neither is the flattering one. Anyone quoting your exact rank is selling noise. If you want to run a small check yourself before anyone gets paid, the method is in how to check whether ChatGPT recommends your business.
The One Number Worth Reporting
Here is the sentence a report should be able to write. Named in 6 of 50 answers on July 1, 2026, across five engines, on ten buyer questions, signed out, live web on. Numerator, denominator, date, engines, questions, conditions. Six facts. You can re-run it, and you can argue with it, which is the whole point.
Fix the question set and never quietly change it. The panel is your ruler, and a ruler that gets re-cut every month measures nothing. Ours is ten questions a real buyer would type, kept the same month to month so the comparison survives. Add experimental questions if you like, and report them separately.
Then publish your own number, including when it is bad. On August 11, 2026 we put our ten buyer questions to five models with live web access through OpenRouter, fifty answers in total, and asked each model to account for the names it gave. Probably Genius was named in zero of the fifty. At least forty other companies were named across those same answers. That is our baseline on the day this piece went up, and we would rather hand you a dated zero than a number with no method behind it.
The follow-ups were instructive in a different way. Two of the five models, asked to account for their own answer, said that most of what they had named came from memory rather than from live search. Treat that as the models' stated criteria and not as a retrieval log, because a model cannot replay its own lookups. Read as criteria, it explains a lot: an engine that answers partly from training memory will move slowly and unevenly, which is exactly why you sample rather than screenshot.
Two more things belong in the record every month: the exact wording used to describe you, and the full list of names that appeared instead. Reading the second list properly is its own skill, covered in how to find out which competitors ChatGPT recommends and why. Everything we ship carries a score and a reason on a ledger, and the books reconcile, because a measurement practice that exempts its own work is a marketing practice.
One Engine Is a Slice, Not the Market
Owners usually ask this question as "can I see how often ChatGPT mentions my business." You can, roughly, by sampling. The harder problem is that ChatGPT alone is a shrinking share of the thing you are trying to measure. Similarweb's July 29, 2026 analysis put ChatGPT's estimated share of worldwide visits to AI platforms at roughly 76 percent in June 2025 and around 53 percent by May 2026, with Gemini climbing to about 27 percent and Claude going from barely 2 percent to close to 9 percent. That is web visits rather than app usage or your own buyers, and the direction is still the point: a ChatGPT-only tracker was watching three quarters of that traffic last year and about half of it now.
The differences are not only between companies. Google's own documentation states that AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary. Two surfaces, one company, different answers. Any single number that claims to summarize "AI" is quietly averaging across systems that do not agree, which is the argument we made at length in one question, five engines.
Treat the engines as a portfolio rather than an average. A gain in Perplexity and a loss in Google's AI Mode are not offsetting events if your buyers use one and ignore the other. Report each engine on its own line, then look at the spread.
Google has also started reporting some of this itself. Search Console now carries a generative AI performance report covering AI Overviews and AI Mode, and its help documentation describes it as showing how many times links to your site were shown in a generative AI feature, rolling out to a subset of website owners. Useful, and narrow. It counts impressions on Google's surfaces. It does not tell you whether you were the name recommended, and it says nothing at all about the other four engines. Take it as one instrument in the drawer.
What to Ask Before You Buy a Dashboard
There is a real category of AI visibility analytics tools, and some of them are good. The distinction worth holding is a sequence, not a rivalry: a dashboard shows you the gap, and operations close it. Buying the first and expecting the second is how a year goes by.
Whether the number comes from a vendor, your agency or a spreadsheet you keep yourself, the same six questions separate a measurement from a mood.
- What exactly was asked? The full prompt list, in the buyer's words, not a category keyword with a question mark bolted on.
- Which engines and which model versions? Named, with dates. Model versions change under the same product name.
- How many runs per question, and when? Repetition is the entire defense against noise. One run per question is a screenshot with better formatting.
- Signed in or signed out, and from where? Personalization and location change answers. Both must be stated and held constant.
- Are the raw answers kept? If you cannot read the actual text months later, you cannot audit the score, and neither can they.
- What counts as an appearance? A brand mention, a linked citation and a direct recommendation are three different things. A score that merges them can rise while your recommendations fall.
A vendor who answers all six is worth listening to at any price. A score with no denominator, no prompt list and no raw answers is a mood ring with a login. That test is not adversarial, by the way. Ask us the same six.
How Long to Wait, and When to Change Course
Patience here has a documented floor. Google's documentation puts recrawling at a few days to a few weeks, and separately says changes can take several months to fully take effect across its systems. A page an engine has not fetched cannot be quoted by it, and that first leg runs on somebody else's schedule.
So set the review windows before the work starts, and set them as a rhythm rather than a promise. Take a baseline in week one, when nothing has shipped and the number is whatever it is. Re-run the identical panel at ninety days and read the leading indicators only: coverage of the question set, appearance rate, how you are described, who appears instead. Judge lagging indicators at two comparable quarters, never sooner.
Change course when three things are true together. The technical layer checks out, meaning the engines can reach and read your pages. The panel is being covered, meaning the work has actually addressed the questions buyers ask rather than the topics that were easy to write. And two comparable quarters show no directional movement in appearance rate across engines. That is a genuine signal. One disappointing answer, on one engine, on one afternoon, is not, and firing a useful partner over it is the expensive version of this mistake.
One boundary, stated plainly. Nobody outside the model companies controls what an assistant says, and any promise of a fixed place in an answer is a promise nobody can keep. What you control is whether the evidence exists, agrees with itself, and holds up when it is checked. What you can prove is movement, measured against the same questions over time.
If you want a baseline before you judge anyone's work, our free Recommendation Check asks the questions your buyers ask across five engines and scores your record against 109 checkpoints in eight zones. Scores run 0 to 109. Ninety and up means AI treats you as a primary source. Under 35 means you are effectively invisible. You get the number and the method, and there is no discovery call attached to it.
Want to Learn More?
Probably Genius builds and operates the recommendation layer around expert-led businesses, then measures it in the open. That is the job inside The Answer: twelve done-for-you thought-leadership articles a month, your expertise structured so machines can read it, corroboration beyond your own pages, and the monthly five-engine probe with every answer saved word for word. The published methodologies show how it runs, including the 109-Point Diagnostic behind the free check.
You're probably a genius at what you do. We make sure AI gets the memo.