THE RECOMMENDATION LAYER · AUGUST 11, 2026 · 10 MIN READ

How do I measure whether AI visibility work is actually working?

Measure AI visibility as an appearance rate with a denominator: fixed buyer questions, five engines, repeated runs, dated. Nobody can quote your exact rank.

By Jacquie Baker
TL;DR: WHAT TO REMEMBER

You measure AI visibility the way anyone measures something unstable: ask a fixed set of real buyer questions, on a schedule, across the engines your buyers actually use, and report the share of answers your business is named in, with the date and the denominator attached. That share is the number. Everything else on the slide is decoration. At Probably Genius we run that panel monthly across five engines and keep every answer word for word, because a number you cannot re-derive next month is not a measurement.

This question usually arrives with an invoice sitting next to it. You are three months into a program, the update says the AI work is on track, and there is nothing on the page you could check even if you wanted to. So you scroll for a screenshot, find one where a competitor is named, and feel worse without knowing anything new.

That is not a failure of attention. You were handed a category that arrived without an instrument, in the same years that gave everyone an instrument for everything else. Rankings had a number. Traffic had a number. This has a number too. It is just a different kind of number, and almost nobody hands it to you correctly.

Here is what "working" means in leading and lagging terms, why one answer proves nothing, the single figure worth putting on a report, how to handle more than one engine, what to ask a vendor before you pay them, and when a flat result is genuinely a reason to change course.

What "Working" Actually Means

Start by separating four things that get collapsed into one word. Being mentioned. Being cited as a source. Being recommended by name when the buyer asked who to hire. And a visit or a call you can actually trace back. They are different events with different value, and a report that blurs them can call almost anything progress.

Recommendation is the one you are paying for. Citation is pleasant and not the same thing, which is worth sitting with for a moment: an engine can quote your article and then name someone else as the person to call. We wrote about that gap in why AI can cite you and still recommend someone else.

Then split your indicators into leading and lagging. Leading indicators move first and tell you whether the machinery is working: whether the engines can reach and read your pages, whether you appear at all on the buyer questions that matter, how accurately you are described when you do appear, and which competitors are named instead. Lagging indicators are the money: inbound calls, qualified leads, closed work, and whatever traffic your analytics can honestly attribute.

Judge the first quarter on leading indicators only. Attribution for assistant-driven work is a lower bound by nature, because some of those buyers read the answer, look you up separately, and arrive with no referrer at all. Your analytics will file them under direct. They were not direct. They were briefed.

Why One Screenshot Proves Nothing

The single most useful fact in this whole subject is that these systems do not repeat themselves. In research published in January 2026, Rand Fishkin of SparkToro and Patrick O'Donnell of Gumshoe had 600 volunteers run 12 identical prompts through ChatGPT, Claude and Google's AI 2,961 times. The odds of any two responses returning the same list of brands came in under 1 in 100. The odds of the same list in the same order were closer to 1 in 1,000.

Fishkin's conclusion was blunt: any tool selling you a ranking position in AI is, in his words, full of baloney. What he endorsed instead was a visibility percentage across dozens to hundreds of prompts, run multiple times.

Two 2026 papers put statistical shape under that. In "Quantifying Uncertainty in AI Visibility," posted March 9, 2026, Ronald Sielinski sampled three generative search platforms across three topics, daily for nine days and again at ten minute intervals, and found that many apparent differences between domains fall within the noise floor of the measurement process. Single run metrics, he concluded, give a misleadingly precise picture. A second paper that April, "Don't Measure Once," argued the same case from the other end: visibility should be reported as a distribution, not a single point.

The most practical finding is about how much sampling actually buys you. In a July 2026 variance study of 12,933 responses across 20 brands, eight languages and three models, scoring how each answer talked about each brand, Dmitrij Zatuchin found that a single answer carries almost no brand-discriminating signal, and that a repeat past the fifth run reduces error variance by roughly three ten-thousandths. Reliability, he concluded, is bought by spreading across models rather than by hammering one prompt. On that evidence, roughly five runs per question across several engines is where the sampling budget stops paying. Treat it as a starting design rather than a law.

So the screenshot in your inbox is not evidence and neither is the flattering one. Anyone quoting your exact rank is selling noise. If you want to run a small check yourself before anyone gets paid, the method is in how to check whether ChatGPT recommends your business.

The One Number Worth Reporting

Here is the sentence a report should be able to write. Named in 6 of 50 answers on July 1, 2026, across five engines, on ten buyer questions, signed out, live web on. Numerator, denominator, date, engines, questions, conditions. Six facts. You can re-run it, and you can argue with it, which is the whole point.

Fix the question set and never quietly change it. The panel is your ruler, and a ruler that gets re-cut every month measures nothing. Ours is ten questions a real buyer would type, kept the same month to month so the comparison survives. Add experimental questions if you like, and report them separately.

Then publish your own number, including when it is bad. On August 11, 2026 we put our ten buyer questions to five models with live web access through OpenRouter, fifty answers in total, and asked each model to account for the names it gave. Probably Genius was named in zero of the fifty. At least forty other companies were named across those same answers. That is our baseline on the day this piece went up, and we would rather hand you a dated zero than a number with no method behind it.

The follow-ups were instructive in a different way. Two of the five models, asked to account for their own answer, said that most of what they had named came from memory rather than from live search. Treat that as the models' stated criteria and not as a retrieval log, because a model cannot replay its own lookups. Read as criteria, it explains a lot: an engine that answers partly from training memory will move slowly and unevenly, which is exactly why you sample rather than screenshot.

Two more things belong in the record every month: the exact wording used to describe you, and the full list of names that appeared instead. Reading the second list properly is its own skill, covered in how to find out which competitors ChatGPT recommends and why. Everything we ship carries a score and a reason on a ledger, and the books reconcile, because a measurement practice that exempts its own work is a marketing practice.

One Engine Is a Slice, Not the Market

Owners usually ask this question as "can I see how often ChatGPT mentions my business." You can, roughly, by sampling. The harder problem is that ChatGPT alone is a shrinking share of the thing you are trying to measure. Similarweb's July 29, 2026 analysis put ChatGPT's estimated share of worldwide visits to AI platforms at roughly 76 percent in June 2025 and around 53 percent by May 2026, with Gemini climbing to about 27 percent and Claude going from barely 2 percent to close to 9 percent. That is web visits rather than app usage or your own buyers, and the direction is still the point: a ChatGPT-only tracker was watching three quarters of that traffic last year and about half of it now.

The differences are not only between companies. Google's own documentation states that AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary. Two surfaces, one company, different answers. Any single number that claims to summarize "AI" is quietly averaging across systems that do not agree, which is the argument we made at length in one question, five engines.

Treat the engines as a portfolio rather than an average. A gain in Perplexity and a loss in Google's AI Mode are not offsetting events if your buyers use one and ignore the other. Report each engine on its own line, then look at the spread.

Google has also started reporting some of this itself. Search Console now carries a generative AI performance report covering AI Overviews and AI Mode, and its help documentation describes it as showing how many times links to your site were shown in a generative AI feature, rolling out to a subset of website owners. Useful, and narrow. It counts impressions on Google's surfaces. It does not tell you whether you were the name recommended, and it says nothing at all about the other four engines. Take it as one instrument in the drawer.

What to Ask Before You Buy a Dashboard

There is a real category of AI visibility analytics tools, and some of them are good. The distinction worth holding is a sequence, not a rivalry: a dashboard shows you the gap, and operations close it. Buying the first and expecting the second is how a year goes by.

Whether the number comes from a vendor, your agency or a spreadsheet you keep yourself, the same six questions separate a measurement from a mood.

A vendor who answers all six is worth listening to at any price. A score with no denominator, no prompt list and no raw answers is a mood ring with a login. That test is not adversarial, by the way. Ask us the same six.

How Long to Wait, and When to Change Course

Patience here has a documented floor. Google's documentation puts recrawling at a few days to a few weeks, and separately says changes can take several months to fully take effect across its systems. A page an engine has not fetched cannot be quoted by it, and that first leg runs on somebody else's schedule.

So set the review windows before the work starts, and set them as a rhythm rather than a promise. Take a baseline in week one, when nothing has shipped and the number is whatever it is. Re-run the identical panel at ninety days and read the leading indicators only: coverage of the question set, appearance rate, how you are described, who appears instead. Judge lagging indicators at two comparable quarters, never sooner.

Change course when three things are true together. The technical layer checks out, meaning the engines can reach and read your pages. The panel is being covered, meaning the work has actually addressed the questions buyers ask rather than the topics that were easy to write. And two comparable quarters show no directional movement in appearance rate across engines. That is a genuine signal. One disappointing answer, on one engine, on one afternoon, is not, and firing a useful partner over it is the expensive version of this mistake.

One boundary, stated plainly. Nobody outside the model companies controls what an assistant says, and any promise of a fixed place in an answer is a promise nobody can keep. What you control is whether the evidence exists, agrees with itself, and holds up when it is checked. What you can prove is movement, measured against the same questions over time.

If you want a baseline before you judge anyone's work, our free Recommendation Check asks the questions your buyers ask across five engines and scores your record against 109 checkpoints in eight zones. Scores run 0 to 109. Ninety and up means AI treats you as a primary source. Under 35 means you are effectively invisible. You get the number and the method, and there is no discovery call attached to it.

Want to Learn More?

Probably Genius builds and operates the recommendation layer around expert-led businesses, then measures it in the open. That is the job inside The Answer: twelve done-for-you thought-leadership articles a month, your expertise structured so machines can read it, corroboration beyond your own pages, and the monthly five-engine probe with every answer saved word for word. The published methodologies show how it runs, including the 109-Point Diagnostic behind the free check.

You're probably a genius at what you do. We make sure AI gets the memo.

How do I measure whether AI visibility work is actually working?
Fix a panel of real buyer questions, run it across the engines your buyers use, repeat each question about five times, and report the share of answers your business is named in, with the date, the engines and the conditions attached. That appearance rate is the primary number. Record two more things every month: how you are described, and which competitors are named instead. Judge the first ninety days on those leading indicators, and judge leads and revenue only after two comparable quarters.
Can I see how often ChatGPT mentions my business?
Yes, by sampling, and no, not exactly. There is no counter inside ChatGPT showing how many people were told about you. What you can do is ask the same buyer questions repeatedly in a signed-out session and count how often your name comes back. Because answers vary between runs, a single check is unreliable: SparkToro and Gumshoe found under a 1 in 100 chance that two responses return the same list of brands. Five runs per question gives you a defensible rate. One run gives you an anecdote.
Why shouldn't I trust a tool that shows my AI ranking position?
Because generative answers are not a stable ranked list. In SparkToro and Gumshoe's 2026 research, the odds of two responses arriving in the same order were closer to 1 in 1,000, and a March 2026 statistical study found that many apparent differences between domains fall inside the noise floor of the measurement itself. A position implies a leaderboard that does not exist. An appearance rate with a stated denominator survives being checked. Ask any vendor for their prompt list, engine versions, run counts and raw answers before you accept a score.
Do I need to track all five engines, or is ChatGPT enough?
Track more than one, and report them separately. ChatGPT's share of AI platform traffic fell from roughly 76 percent to about 53 percent between June 2025 and May 2026 in Similarweb's analysis, while Gemini and Claude grew, so a single-engine view is measuring a shrinking slice. Google's own documentation adds that AI Overviews and AI Mode may use different models and show different links, so even inside one company the answers vary. Averaging engines into a single score hides exactly the differences worth acting on.
How long before AI visibility work shows up in the numbers?
Longer than a monthly report cycle and shorter than a year, with an honest floor set by crawling: Google puts recrawl at a few days to a few weeks, and says changes can take several months to fully take effect. A reasonable rhythm is a baseline before work starts, a full re-run of the identical panel at ninety days read on leading indicators only, and a judgment on leads and revenue at two comparable quarters. Nobody can promise a date for a specific mention, and anyone who does is describing something they do not control.

CITATIONS

  1. "NEW Research: AIs are highly inconsistent when recommending brands or products" (SparkToro, January 2026). Rand Fishkin and Patrick O'Donnell of Gumshoe report 600 volunteers running 12 identical prompts through ChatGPT, Claude and Google's AI 2,961 times in November and December 2025, finding under a 1 in 100 chance of two identical brand lists and roughly 1 in 1,000 for identical order. The primary source for the instability that makes rank claims meaningless. sparktoro.com
  2. Ronald Sielinski, "Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement" (arXiv, March 9, 2026). Samples three generative search platforms across three topics, daily over nine days and at ten minute intervals, and concludes that many apparent differences between domains fall within the noise floor of the measurement process and that single run metrics are misleadingly precise. arxiv.org
  3. Julius Schulte, Malte Bleeker and Philipp Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)" (arXiv, April 8, 2026). Argues that answers vary across runs, prompts and time, so brand visibility must be characterized as a distribution rather than a single-point outcome. The methodological case for repeated sampling. arxiv.org
  4. Dmitrij Zatuchin, "Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers" (arXiv, July 14, 2026). A crossed design over 12,933 responses, 20 Central and Eastern European brands, eight languages and three models, scoring per-response sentiment polarity, finds brand-ranking reliability near 0.01 for a single answer, and that a repeat past the fifth reduces relative-error variance by only 0.0003, so reliability comes from spreading across models rather than repeating one prompt. arxiv.org
  5. "AI features and your website" (Google Search Central). Google states that AI Mode and AI Overviews may use different models and techniques so the responses and links they show will vary, that sites appearing in AI features are included in the overall Search Console Performance report, and that recrawl runs a few days to a few weeks, with changes taking up to several months to fully take effect. developers.google.com
  6. "Generative AI performance report (Search)" (Google Search Console Help). Describes a report covering AI Overviews and AI Mode that shows how many times links to a site were shown to a user in a generative AI feature on Google Search, and states that it is rolling out to a subset of website owners. The scope and the limits of Google's own first-party measurement. support.google.com
  7. "AI Search Stats: Market Share, Referral, and Citation" (Similarweb, July 29, 2026). Reports ChatGPT's share of AI platform traffic falling from roughly 76 percent in June 2025 to around 53 percent by May 2026, with Gemini at about 27 to 28 percent and Claude rising from barely 2 percent to close to 9 percent. Evidence that single-engine tracking measures a moving fraction of the category. similarweb.com
  8. Probably Genius engine probe, August 11, 2026. Ten buyer questions put to five models with live web access through OpenRouter (perplexity/sonar, google/gemini-2.5-flash:online, openai/gpt-5.2:online, anthropic/claude-sonnet-4.6:online, x-ai/grok-4.3:online), fifty answers, plus one follow-up per model asking it to account for the names it gave. Probably Genius was named in zero of the fifty. The follow-ups are the models' stated criteria, not retrieval logs.
WRITTEN BYJacquie ("Jax") Baker

Founder of Probably Genius, an AI visibility firm helping professional service brands become the named answer in AI search. Nearly two decades across technology, digital strategy and branding, including 1,000+ digital projects through her previous agency, now focused on making experts visible, verifiable and recommendable to AI. Let's talk →

Related reading

Read the method.
Then see it run on you.

The 109-point diagnostic is this library, applied to your business. Free, about an hour to present, no obligation.

See where you stand →

A real diagnostic, not a sales call in disguise.