THE RECOMMENDATION LAYER · MAY 20, 2026 · 9 MIN READ

Why Your Website Is Invisible to AI (and the 10-Minute Test That Proves It)

Your site is invisible to AI for one of three reasons. A free ten-minute test: three buyer questions, three engines, nine answers, and the fix each needs.

By Jax Baker
TL;DR — WHAT TO REMEMBER

Your website is invisible to AI for one of three reasons: the engines are not sure who you are, they cannot lift a clean answer from your pages, or nobody else backs up what you say about yourself. You can find out which one is yours in about ten minutes, with three free chat windows and the three questions your buyers already ask. This article gives you the test, step by step, and the fix each result points to.

If you have a decent website, this probably sounds like someone else’s problem. Your site ranks for your name. Your traffic is steady. Your analytics dashboard has never once flashed red.

And yet there is a version of your market you have never looked at: the answers an AI assistant reads out loud when one of your buyers asks who to hire. Most owners have never asked. Not once.

The cut you are trying to survive is steep. When Ahrefs analyzed 1.4 million ChatGPT prompts, the model went on to cite only 49.98 percent of the URLs it had already retrieved. Half the candidates that make it into the room never get quoted. The business websites we test often fail to make it into the room at all.

Rankings and traffic measure a different scoreboard. AI visibility is whether the assistants mention you, describe you accurately, and recommend you when a buyer asks in a full sentence, and none of your current stats can see it. At Probably Genius we measure that scoreboard for a living, with a scored panel of buyer questions run across five engines and checked by humans. What follows is the honest ten-minute version of that instrument, free, because a test you can run yourself beats a claim you have to take on faith.

The Ten-Minute Test, Step by Step

This is the self-serve version of what we call the recommend-test: the same instrument we run at panel scale for clients, cut down to what one owner can do before their coffee goes cold. No signup, no tool, no cost. Three steps.

  1. Minutes 0 to 2: write down the three questions your buyers actually ask. Not keywords. Full sentences, the way a stressed human types them. Three rules: each question must carry buying intent (who should I hire, who is the best, what does this cost and who is worth it), each must name your real market (your city or region if you serve one, your specialty if you serve a niche), and your business name must appear in none of them. If you get stuck, use these three shapes: “Who should I hire for [the problem you solve] in [your area]?” · “What is the best [your kind of firm] for [your ideal client’s situation]?” · the exact question your last good client asked before they found you.
  2. Minutes 2 to 8: ask all three questions in three AI assistants. Use ChatGPT, Gemini, and Perplexity. Open a fresh chat for each question so earlier answers do not color later ones, and paste each question word for word. Do not coach the assistant, do not mention your business, do not rephrase when you dislike the answer. You are eavesdropping on the conversation your buyer has without you. Nine answers total.
  3. Minutes 8 to 10: read the nine answers and score three things. First: were you named, anywhere, at all? Second: if you were named, was the description accurate, or is it an old version of you, or a confusion with someone else? Third: were you recommended, or merely listed? Then do the one thing almost nobody does: look at what got cited. Most of these assistants show their sources. Note whose pages the answers were built from, because that list is the room where the decision got made.

Write the results down, including the date. If you are like many established firms we test, you already know the feeling coming: nine answers about your own market, full of names, and yours is not one of them.

What the Nine Answers Will Tell You

The results we score keep landing in one of three failure patterns. Reading yours against them turns a bad feeling into a diagnosis.

Pattern one: not mentioned. The most common result. The answers are confident, structured, and full of your competitors, directories, and roundup articles, and you simply do not exist in them. No error, no warning. Silence.

If this is you, note what did fill the space. In our experience it is usually third-party list pages and platforms rather than any single rival, which means the seat is not taken so much as unclaimed.

Pattern two: described wrongly. You appear, and the description makes you wince. The old office. The service you stopped offering years ago. Your name attached to a business two towns over, or your practice blended with a similarly named one.

This pattern feels smaller than absence. It is often worse, because a wrong answer travels with the machine’s full confidence, and your buyer has no way to know the record is stale.

Pattern three: mentioned, never recommended. The subtlest one. The engine knows you exist. Ask it about you by name and it recites your story warmly. But in the nine buyer-question answers, where a recommendation was the entire assignment, you show up as a listing, a footnote, an “other options include” entry. Understood, never endorsed.

There is a whole layer between being online and being the answer, and this pattern means you are living one rung below it. We wrote about that layer in The Recommendation Layer.

One caution before you conclude anything: a probe is a sample, not a scoreboard. Three engines on one afternoon is nine data points. If one assistant names you and two do not, you have learned something real about the spread, not a final grade. The engines disagree with each other constantly, which is why any serious version of this test runs across more engines, more questions, and more than one moment in time.

Why Each Failure Happens

The three patterns have three different mechanics, and none of them is fixed by publishing harder. You are not invisible because you’re bad. You’re invisible because the proof is scattered.

Not mentioned is usually an identity and corroboration problem. Before an engine risks recommending you, it has to be sure it knows who you are: one coherent business, described the same way everywhere, confirmed by sources other than your own website. Most small firms fail that check invisibly, with their identity scattered across profiles that never confirm each other. The machine does not distrust you. It just cannot assemble enough of you to say your name.

We publish our method for fixing this as Entity Identity, and the plain-English version as a companion guide for owners. One blunt mechanical check while you are at it: the AI companies run separate crawlers for model training and for live answer retrieval, and plenty of sites block the retrieval crawlers by accident, with a robots.txt rule someone copied in 2023 to opt out of training.

Described wrongly is a stale-record problem. Engines assemble you from evidence with timestamps you do not control: old directory entries, an unedited profile, a news mention from three rebrands ago. Your current website is one witness among many, and it gets outvoted by whatever the rest of the record keeps repeating.

The fix is not louder marketing. It is correcting the record the machines actually consult, surface by surface, until the current you outweighs the archived one.

Mentioned but never recommended is a proof problem. The engine can read you fine. What it cannot find is a reason to put its name behind you: claimed positions, a named method, specific outcomes, someone else vouching for the story.

The evidence on what gets cited is unusually clear. Semrush’s study of 11,882 prompts found clear summarization (+32.83 percent) and strong E-E-A-T signals (+30.64 percent) were the qualities most correlated with citation, and the peer-reviewed GEO study presented at KDD 2024 found adding citations, quotations and statistics can lift a source’s visibility in AI answers by up to 40 percent.

Machines quote specifics. A wall of adjectives gives a cautious engine nothing to hold, and cautious is the engines’ entire personality.

What the Test Cannot Tell You

Here is the honest part, and we would rather say it ourselves than have you discover it. The ten-minute test answers one question: do I have a problem? It is genuinely good at that, and you should run it quarterly, free, forever, whether or not you ever talk to us.

What it cannot do is tell you why, how bad, or in what order to fix things. Nine answers is a snapshot from three engines on one day, scored by the least objective person available: you.

It cannot separate an identity failure from a proof failure when both are present, which they usually are. It cannot see whether your best expertise is locked in formats the machines cannot read. It cannot run the hundred-plus other checks a full examination covers, and it cannot tell you whether last month’s result was a trend or noise.

That is the difference between a smoke alarm and a building inspection. The full version of this instrument is our 109-Point Diagnostic: the same recommend-test run as a scored panel of buyer questions across five engines, inside a human-verified examination of eight zones, from entity identity to proof architecture to how liftable your answers are. We publish the methodology, including our multi-model approach, because an audit you cannot inspect is just a quote with a rubric attached.

And if the ten-minute version stings, hold on to this: every pattern above is buildable. Identity can be reconciled. Records can be corrected. Proof can be counted, claimed, and published.

The experts losing these answers are almost never the weakest ones; they are the ones too busy doing the work to write it down. The people AI will never replace are the ones AI should be recommending. The whole point of the test is to find out, while the seat in your market is still unclaimed, whether the machines can make that case for you yet.

Want to Learn More?

Probably Genius was built by Jacquie Baker and Christopher Shaw, who spent more than 40 combined years translating what makes an expert the best in the room for the audiences that decide who gets chosen. We built the ten-minute test to give away because it is the same first question we ask about every business we examine: when your buyer asks the machines, are you the answer? Everything we sell starts from measuring that honestly.

Frequently asked questions

How do I test whether AI can see my business?
Write the three questions your buyers actually ask, in full sentences with buying intent and without your business name in them. Ask all three, word for word, in fresh chats in ChatGPT, Gemini, and Perplexity. Then score the nine answers for three things: were you named, was the description accurate, and were you recommended rather than merely listed. Check what sources each answer cited. The whole exercise takes about ten minutes and costs nothing.
Why doesn’t ChatGPT mention my business?
Usually one of three reasons: the engine cannot be sure who you are because your identity is inconsistent or thinly corroborated across the web, its record of you is stale or confused with someone else, or it can read you but finds no claimed positions, named methods, or specific proof strong enough to justify a recommendation. Ahrefs’ analysis of 1.4 million prompts found ChatGPT cites only about half of the pages it retrieves, so even being found is not enough to be quoted.
Do I need to run the test on more than one AI assistant?
Yes. One engine’s answer is a data point, not a verdict. The assistants differ in how they browse, what they cite, and how cautious they are, and they disagree about businesses constantly. Three engines is the minimum for a useful spread, which is why the ten-minute test uses ChatGPT, Gemini, and Perplexity, and why our full diagnostic panel runs the same buyer questions across five engines and reports results as percentages of sampled answers rather than a single verdict.
I failed the test. What should I do first?
First, identify your pattern: not mentioned, described wrongly, or mentioned but never recommended. Not mentioned points to identity and corroboration work. Described wrongly points to correcting the stale record the engines keep consulting. Never recommended points to claiming and publishing your proof. If more than one pattern applies, or you want the why scored properly before you spend anything on fixes, a full diagnostic maps all three across eight zones and hands you the order of operations.

CITATIONS

  1. “Why ChatGPT Cites One Page Over Another (Study of 1.4M Prompts)” (Ahrefs, April 2026). Large-scale study of 1.4 million ChatGPT prompts: only 49.98 percent of retrieved URLs were cited, with title-to-question relevance the strongest survival factor. ahrefs.com
  2. “How We Built a Content Optimization Tool for AI Search” (Semrush, January 2026). Correlation study of 11,882 prompts across ChatGPT Search, Google AI Mode and Perplexity: clear summarization (+32.83 percent) and E-E-A-T signals (+30.64 percent) most associated with getting cited. semrush.com
  3. “GEO: Generative Engine Optimization” (KDD 2024, arXiv:2311.09735). The peer-reviewed study showing optimization built on citations, quotations and statistics can boost a source’s visibility in generative engine responses by up to 40 percent. arxiv.org
WRITTEN BYJacquie ("Jax") Baker

Founder of Probably Genius, an AI visibility firm helping professional service brands become the named answer in AI search. Nearly two decades across technology, digital strategy and branding — including 1,000+ digital projects through her previous agency — now focused on making experts visible, verifiable and recommendable to AI. Let's talk →

Related reading

Read the method.
Then see it run on you.

The 109-point diagnostic is this library, applied to your business. Free, about an hour to present, no obligation.

See where you stand →

A real diagnostic, not a sales call in disguise.