GEO / THE RECOMMENDATION LAYER · JULY 4, 2026 · 15 MIN READ

How Does ChatGPT Decide Who to Recommend?

ChatGPT recommends like a cautious editor: verifiable, corroborated, consistently described. How the audition works, and how to watch it run on yours.

By Jax Baker
TL;DR — WHAT TO REMEMBER

ChatGPT decides who to recommend by running a fast, ruthless audition. When it judges that a question needs current facts, it searches the live web, pulls in dozens of candidate sources, discards roughly half of them on nothing more than their titles, URLs and snippets, and then writes its answer from the survivors it can verify and cross-reference. That is not a guess about how it works. Ahrefs watched the process across 1.4 million ChatGPT prompts and found the model ultimately cites only 49.98 percent of the URLs it retrieves. Half the field is cut before the machine reads a full page.

If you have asked ChatGPT who does what you do and read a list of competitors, some of them objectively worse at the work, you know how personal this feels. It is worth being precise about what actually happened. You weren’t rejected. You were omitted. The machine did not weigh your twenty years against their twelve and pick them. It never assembled enough verifiable material about you to say your name at all.

At Probably Genius we watch this decision from both sides: we run live probe panels across five AI engines to see who gets named, and we build the identity, proof and content layers that change the answer. The pattern is consistent. ChatGPT recommends the way a cautious editor does. It checks who is verifiable, who is corroborated, and who is consistently described everywhere it looks. Most businesses fail all three checks without ever knowing the checks exist.

The reality is that the decision has almost nothing to do with who is best and almost everything to do with who is legible. That is uncomfortable, and it is also the good news, because legibility can be built. This guide walks through the decision step by step: how the answer gets assembled, the three checks that decide who is quotable, why most firms fail them, how to watch the decision happen on your own category, and where to start.

The Decision, Step by Step

Start with the fork in the road. Every question you type gets triaged: answer from memory, or go look. When ChatGPT answers from memory, it is drawing on its training data, and only businesses already established in that corpus exist to it at all. When it decides the question needs current information, which buyer questions like “who should I hire” usually do, it runs a live search, and the audition begins. Two different games, one consequence: a business that is invisible in both the training data and the live web is invisible twice.

The search itself does not work the way you search. The model rewrites your question into a set of narrower sub-queries, what Ahrefs calls fan-out queries, and retrieves candidates for each. Ask “who is the best estate planning attorney for blended families” and the machine may quietly also ask who publishes on second-marriage inheritance disputes, what credentials matter, and who serves your city. You are not competing on one query. You are competing on the hidden five.

Then comes the cut that should change how you think about your website. Across Ahrefs’ 1.4 million analyzed prompts, ChatGPT went on to cite only 49.98 percent of what it retrieved, and the study’s authors found the title, snippet and URL were doing the heavy lifting in that first decision. Half of the candidates never get properly read. A page titled “Home” with a URL full of query strings loses the audition before it starts, no matter what wisdom sits below the fold.

What survives gets read for corroboration, and here the machine behaves less like a fan and more like a fact-checker. The same study found Reddit accounts for 67.8 percent of everything ChatGPT retrieves and then declines to cite, and Reddit URLs get cited at a rate of just 1.93 percent. Read that again: the model uses the crowd to gauge consensus, then quotes somebody more accountable. It reads the room, and cites the record.

Only then does synthesis happen, and this is where the way we map it at Probably Genius departs from how most of the industry talks. The final answer is not a ranked list of pages with the top one winning. The model is resolving entities: deciding whether the firm on the website, the firm in the directory, the firm in the reviews and the firm in the news story are the same firm, and whether the assembled picture is solid enough to stand behind. AI doesn’t rank pages. It understands entities. If that sentence rearranges how you think about your marketing, it should. It is the difference between decorating a page and building an identity, and it is the premise behind everything in what GEO actually is.

The Three Checks: Verifiable, Corroborated, Consistently Described

Strip the pipeline down and the recommendation decision reduces to three questions a cautious editor would ask before printing your name. Can I verify who this is? Does anyone else confirm it? Is the story the same everywhere I look? We treat those as the three checks, and each one has data behind it.

Check one: verifiable. The machine needs claims it can trace. Named people with credentials, numbers with sources, services with plain descriptions, structured data that states outright who you are and what you do. Semrush’s January 2026 study of 11,882 prompts across ChatGPT Search, Google AI Mode and Perplexity found strong E-E-A-T signals carried a +30.64 percent correlation with getting cited, and structured data elements +21.60 percent. This is the machine-readable identity layer, what we build as Entity Identity: the schema, the entity definitions, the digital fingerprint that lets a model confirm you are a specific, real, credentialed thing rather than a plausible string of words.

Check two: corroborated. One surface saying you are excellent is a claim. Several unconnected surfaces agreeing is evidence, and the synthesis step is built to notice the difference. Your website, your professional profiles, your directory listings, your reviews and the places that mention you either assemble into one coherent picture or they do not. We run this as the Coherence Test: does every surface AI consults describe the same business, the same people, the same specialty, in compatible terms? The connective tissue is what we call the Golden Thread, a single consistent narrative running through every surface, so the machine assembles one authoritative entity instead of four fragments that might be the same company.

Check three: consistently described, in words a machine can lift. Passing the identity checks gets you considered. Being quotable gets you named. In the Semrush data, the single strongest positive correlation with citation was clear summarization at +32.83 percent, with question-and-answer formatting at +25.45 percent. The peer-reviewed GEO study presented at KDD 2024 found the same direction from the other side: adding citations, quotations and statistics lifted a source’s visibility in generative engine responses by up to 40 percent, while keyword stuffing, the reflex of old search, actively reduced it. Write direct answers with real numbers and attributable sources, and you are handing the model sentences it can reuse. The buyer on the other end of that reused sentence behaves differently too, which is the story of being AI-chosen.

Why Most Firms Fail All Three Without Knowing

Here is the uncomfortable arithmetic: the firms with the deepest expertise are often the least legible, because the expertise never made it out of their heads. The judgment calls, the war stories, the specific way a founder sizes up a problem in the first ten minutes of a call, none of it is on the record. It lives in conversations, in PDFs nobody structured, in testimonials trapped in screenshots. We call this the Black Box Effect, and it is check-one failure in its purest form. The machine cannot verify what was never published.

The second failure is the opposite reflex: publishing plenty, none of it yours. A firm hears it needs content, points an AI tool at its blog, and produces articles that read exactly like the material the model was already trained on. We call the output Training Data Echo. The machine ingests it, recognizes its own reflection, and moves on. Your original expertise is a high-resolution photograph. Sharp. Detailed. Unique. Generic AI content is a copy of a copy, and you can’t photocopy a photocopy and expect it to be worth citing. Volume does not fix this. A hundred echo articles fail check three a hundred times.

The third failure is the quietest one: modesty. Professionals trained to let the work speak for itself write websites that claim almost nothing. No numbers, no named methodology, no stated specialty, because specifics feel like bragging. To a human who meets you, the understatement reads as confidence. To a machine running the verifiable check, it reads as absence. AI can’t cite what you won’t claim. The gaps in your EEAT aren’t information problems. They’re claiming problems. They get fixed in your own voice, on your own record.

Notice what is missing from this list of failures: your platform. Owners regularly arrive convinced their website builder is hiding them from AI, and the evidence says otherwise, as we unpacked in the Wix question. The blockers are almost never the software. They are the unclaimed expertise, the unstructured identity and the echo content, and every one of them is fixable. When a brand is present but the engines keep reaching past it, there is a specific playbook for earning the citation, which we laid out in force the mention.

The Recommend-Test: Watch the Decision Yourself

You do not have to take any of this on faith, because the decision is observable. The recommend-test is simple: ask the engines the questions your buyers actually ask, and record who gets named. Not “tell me about [your firm]”, which almost any model will answer politely. Buyer questions. “Who should I hire to do X in my city?” “Best firm for Y?” The branded question tests whether you are known. The buyer question tests whether you are chosen, and plenty of firms discover they are known by name and never chosen.

We run this on ourselves, and we publish the result even when it stings. In July 2026 we ran our own baseline panel: 10 buyer questions about AI visibility services, asked across five engines (GPT, Gemini, Perplexity, Claude and Grok), 50 probes in total. Probably Genius appeared in zero of the 50 answers. Worth saying plainly: we sell this, and we started at zero, because in AI search everyone starts at zero. The finding that matters more is the second one. In 45 of the 50 answers, none of the competitors we track appeared either. Two software platforms surfaced across four of the ten questions; the other six questions came back with no tracked name on any engine. On the evidence of our sample, the recommendation slot in this category is not owned by a competitor. It is mostly sitting empty.

This is also why your keyword tools have been telling you a comforting lie. Buyer questions are conversational, and conversational queries barely register in keyword volume data. We have watched high-value niches show near-zero monthly search volume in the standard tools while the same questions get asked of AI engines in a steady stream. The old instruments measure the old behavior. The recommend-test measures the new one.

Run your own five-minute version today. Write down five questions a real buyer would ask in your category, ask them in ChatGPT, Gemini and Perplexity, and log three things: who was named, what sources were cited, and whether you appeared at all. Then hold the results with appropriate humility, because a probe is a sample, not a scoreboard. Answers vary across engines, across phrasings and across weeks, which is exactly why we run structured multi-engine panels on a schedule rather than one-off spot checks, using the same multi-model methodology we build content with. One answer tells you an anecdote. Fifty answers tell you where you stand.

The First Step: Measure, Then Build

Everything above compresses into one instruction: measure before you build, and build against the checks. The recommend-test tells you whether you are being named. What it cannot tell you is why not, and the why lives in the layers: identity, corroboration, extractable proof. That is what a structured diagnostic is for. Our 109-Point AI Visibility Diagnostic scores a business across eight zones, from the Identity Layer (17 points) through Share of Answer (20 points) to answer extractability. Scores run 0 to 109, where ninety and up means AI treats you as a primary source and under 35 means you are effectively invisible. Most firms are not losing the audition narrowly. They have not entered it.

The build side has to clear the same bar the machine sets, which is why we gate our own output. Every piece we publish passes the AI Integrity Standard, a 100-point assessment where anything below 85 does not publish. A cautious machine deciding whose name to say out loud is, in the end, an argument for doing the work properly: real claims, real sources, real structure, verified before it ships. How long the flip takes depends on the proof you already have and the category you compete in, so we do not promise windows. We publish the method, run the panel again, and let the before-and-after speak.

One more thing worth holding onto. The caution that makes these engines hard to impress is the same caution that makes their recommendation worth having. When a system built to be sceptical says your name, the buyer arrives already convinced, and that is a fundamentally better first conversation than any ad ever bought you. The people AI will never replace are the ones AI should be recommending. The machine is ready to do it. The record it needs is yours to build.

Want to Learn More?

Probably Genius was built by Jacquie Baker and Christopher Shaw, who spent more than 40 combined years and two agencies translating what makes an expert the best in the room for the system that decides who gets chosen. The audience changed from investors to AI. The skill did not. We map genius, build the infrastructure that makes it verifiable, and publish the proof that gets it recommended.

Frequently asked questions

How does ChatGPT decide who to recommend?
ChatGPT decides who to recommend by retrieving candidate sources from the live web, filtering them hard, and synthesising an answer from what it can verify. Ahrefs’ analysis of 1.4 million prompts found it cites only 49.98 percent of the URLs it retrieves, with titles, snippets and URLs driving the first cut. The businesses that survive to be named are the ones that are verifiable, corroborated across independent surfaces, and consistently described in machine-liftable language.
Why does ChatGPT recommend my competitor and not me?
Almost always because your competitor is more legible, not more capable. If their site states specifics, their identity is structured, and third-party surfaces corroborate their story, they pass checks you are silently failing. It is also common that nobody is being recommended: in our July 2026 panel of 50 buyer-question probes, 45 answers surfaced none of the competitors we track. Run the buyer questions yourself before assuming the slot is taken.
Do ChatGPT, Gemini and Perplexity recommend the same businesses?
No, and the differences are the point. Each engine has different training data, retrieval behavior and trust weighting, so the same buyer question routinely produces different names on different engines, and different names on the same engine weeks apart. That is why a single spot check tells you very little, and why we measure with repeated multi-engine panels and report results as rates across sampled questions rather than fixed rankings.
How do I find out if AI recommends my business right now?
Start with the five-minute recommend-test in this guide: five real buyer questions, three engines, log who gets named and what gets cited. For the full picture, our 109-Point AI Visibility Diagnostic scores your business across eight zones covering identity, corroboration, share of answer and extractability. Scores run 0 to 109. Ninety and up means AI treats you as a primary source. Under 35 means you are effectively invisible. The honest starting point for most firms is lower than they expect, and more fixable than they fear.

CITATIONS

  1. “Why ChatGPT Cites One Page Over Another (Study of 1.4M Prompts)” (Ahrefs, April 2026). The large-scale empirical study behind this article’s retrieval figures: ChatGPT cited 49.98 percent of retrieved URLs, titles and snippets drove the first cut, and Reddit made up 67.8 percent of retrieved-but-uncited URLs with a 1.93 percent citation rate. ahrefs.com
  2. “How We Built a Content Optimization Tool for AI Search” (Semrush, January 2026). Correlation study of 11,882 prompts across ChatGPT Search, Google AI Mode and Perplexity, finding clear summarization (+32.83 percent), E-E-A-T signals (+30.64 percent), Q&A formatting (+25.45 percent) and structured data (+21.60 percent) most associated with getting cited. semrush.com
  3. “GEO: Generative Engine Optimization” (KDD 2024, arXiv:2311.09735). The peer-reviewed study showing optimization methods built on citations, quotations and statistics can boost a source’s visibility in generative engine responses by up to 40 percent, while keyword stuffing reduces it. arxiv.org
  4. These findings line up with the identity standards the engines lean on in practice: Schema.org structured data for machine-readable identity and Google’s E-E-A-T quality framework for judging who deserves to be cited, even though no engine publishes the exact rules of its decision.
WRITTEN BYJacquie ("Jax") Baker

Founder of Probably Genius, an AI visibility firm helping professional service brands become the named answer in AI search. Nearly two decades across technology, digital strategy and branding — including 1,000+ digital projects through her previous agency — now focused on making experts visible, verifiable and recommendable to AI. Let's talk →

Related reading

Read the method.
Then see it run on you.

The 109-point diagnostic is this library, applied to your business. Free, about an hour to present, no obligation.

See where you stand →

A real diagnostic, not a sales call in disguise.