You check whether ChatGPT recommends your business by asking it the questions your buyers ask, in a chat with no memory of you, more than once, and writing down what comes back. It takes twenty or thirty minutes and costs nothing. The writing down isn't admin. It is the part that turns a bad feeling into something you can compare next month.
Most owners have already run the fast version. You typed your category into ChatGPT one evening, read three names that were not yours, and closed the tab. Then someone else in your office asked something similar on Thursday and got a different list, and now there are two screenshots on your desk that disagree.
That disagreement is normal, and it isn't the machine being unfair to you. These systems answer probabilistically. The wording of the question, the settings on the account, and whether the assistant went looking on the live web all move the result. A screenshot argues. A record settles.
So here is the honest version of the check: how to set up a clean session, what to ask, the four things worth scoring, how many answers you need before you believe any of it, and where a free self-run check stops being enough.
Set Up a Chat With No Memory of You
The most common way owners get a wrong reading is the friendliest one. You run the test inside your own account, where the assistant has months of context about you, your company and your work, and it obligingly says your name. You feel better for an hour. Your buyer, who has never mentioned you to anything, gets a different answer entirely.
OpenAI is clear that ChatGPT can personalize using saved memories and details from past chats, and that a Temporary Chat neither uses existing memories nor creates new ones. So the cleanest condition available to you is a Temporary Chat with custom instructions turned off, because those instructions keep applying to new conversations until you change or remove them. Signing out or opening a private window reduces some account context too, without being the same controlled condition.
Three more setup rules, all of which sound fussy until the first time they save you from a false conclusion:
- Open a fresh chat for every question. Answers earlier in a thread color the ones after them, so a single long conversation gives you one reading, not five.
- Note whether the assistant searched the web. Search on and search off are different tests. In a study of 1,000 product prompts run 20,000 times, reported by Search Engine Land in June 2026, only 19.8 percent of the products recommended without search also appeared once search was enabled. That was retail, not professional services, so treat it as a warning about conditions rather than a number about your category.
- Write your location into the question the way a buyer would. "In Tampa" and "near me" are not the same prompt, and a session without your account behind it may not carry the location you assume. Name the city rather than relying on the assistant to guess it.
Then leave it alone. Do not coach, do not correct, do not rephrase because you dislike the answer. You are eavesdropping on a conversation that happens without you every day.
Ask the Questions Your Buyers Ask
Now the questions, and this is where most self-run checks quietly break. The instinct is to type your own company name. A branded prompt tests what the assistant says once you have handed it the answer, which is a different thing from whether your name arrives on its own. Ask hiring questions instead, the way a stranger would.
Write five questions with three properties: buying intent, your real market, and no mention of your business. Full sentences, the way a worried person types at eleven at night. Useful shapes include "Who should I hire for [the problem you solve] in [your city]?", "What is the best [your kind of firm] for [your ideal client's situation]?", "Who handles [the specific hard case you are known for]?", and the exact question your last good client asked before they found you.
Keep those five fixed. Small rewording is not a harmless variation: in roughly 6,000 paraphrase runs across OpenAI and Anthropic models, reported in a May 2026 preprint, two cosmetic rewordings of the same buying intent produced recommendation sets that overlapped by 0.288, and adding a constraint dropped it to 0.135. The same prompt run twice overlapped by 0.50 to 0.61, so even repetition is not a photocopy. Read that as permission to stop redrafting and start repeating. Your question set is the instrument, and an instrument you change every month compares nothing. It will not tell you what every buyer sees, since real people phrase things a thousand ways. It will tell you honestly whether the same five questions return your name more often than they did last quarter.
One more discipline worth borrowing from how we run this at panel scale. Ask about the category, not the brand, and put the same five questions to the same engines every time. That is the difference between watching the weather and owning a thermometer.
Score Four Things, Not One
Read each answer once, then score it. Four columns, no interpretation yet.
- Named. Your business appears anywhere in the answer, in any form, or it does not.
- Accurate. If it appears, the description has to be current. Old address, retired service and confusion with a similarly named firm all count as a miss, and a confident wrong answer travels further than silence.
- Recommended or merely listed. There is a real gap between "other options include" and "I would start with." Being mentioned is not the same as being the name the assistant puts forward, and owners routinely score a listing as a win.
- Who got named instead. Write down every business, directory and roundup page that filled the space, plus any sources the assistant showed. That list is the shortlist your buyer actually saw, and it's the most useful thing on the page.
Then total it as a fraction: named in 3 of 20 answers, accurate in 1, recommended in 0. A fraction is comparable next month. "It didn't mention us" is not.
The fourth column is the one people skip, and it is the one that changes the meeting. Often the space is not held by a rival at all. Semrush tracked 1,094 US categories in ChatGPT from January to June 2026, using five prompts per category, and found only 15.2 percent had a clear owner, meaning a brand that appeared in at least four of the five prompts with a five point lead over the runner up. In 53.7 percent of categories, no brand appeared in even three of five. A seat nobody has claimed is a very different problem from a seat your competitor is sitting in, and you cannot tell which one you have until you write down what came back.
How Many Answers Before You Believe It
Five questions asked once gives you five answers, and five answers is a mood. The same research that showed how much paraphrasing moves recommendations also showed that the identical prompt, rerun the same day, returns sets that overlap by only about a half to three fifths. Variation is the normal behavior of the system, not evidence of a change in your standing.
A practical starter panel: five questions, run twice each, in one sitting. Twenty answers if you also run a second engine. That is a small sample and you should say so, always with the denominator attached: "we appeared in 2 of 20 answers on July 25, signed out, search on." What it does not license is rank language. An assistant can list you fourth in one answer, and that position is not a standing you can quote next week.
Log six things with every run, or the whole exercise expires within the week: the engine and model shown, the date, the exact prompt text, whether the assistant searched, whether you were signed in, and the location wording you used. Those six turn a folder of screenshots into a baseline. It's the same discipline we hold ourselves to when we run the same buyer question across five engines for a client, and the reason a month-over-month comparison means anything.
Then be patient with the trend. Notice how much evidence the researchers demanded before calling a category owned: four appearances out of five prompts, with a clear margin. Hold your own results to something like that. One good answer is not arrival. One bad answer is not a verdict. Two comparable quarters is the first thing worth calling a signal, and even then you keep the raw answers.
Do Gemini, Claude and Perplexity Need a Different Check?
The method travels. The conditions do not. You can run the identical five questions in Gemini, Claude, Perplexity or Grok, and you should. In every panel we run, the engines disagree with each other about who to name, and your buyers are not all using the same one. What changes engine to engine is the setup: what personalization is on by default, whether the assistant browses live, and how it shows its sources.
What you may not do is average them into a sentence about "AI." Each engine is its own reading, recorded under its own name and date. When we measure this for clients we ask five engines the same buyer questions on a schedule and report each one separately, because a result blended across engines hides exactly the disagreement you needed to see.
What This Check Cannot Tell You
Run this quarterly, free, forever, whether or not you ever speak to us. It answers one question well: does your business appear in a recorded sample of buyer questions, under conditions you wrote down. That's worth twenty minutes of anyone's year.
What it cannot do is tell you why. It cannot separate an identity problem, where the engines are not certain you are one specific real business, from an evidence problem, where they know you exist and find nothing they can stand behind. It cannot see whether your best expertise is sitting in formats a machine cannot read, which is often the quieter half of the story on your own website. It cannot tell you what to fix first, and order matters here more than effort. Once you do know the answer, the work of getting recommended is a different project with a different sequence.
That is the gap the free Recommendation Check closes. It scores your business across 109 checkpoints in eight zones, hand-scored, and asks the buyer questions of the five engines we track: ChatGPT, Gemini, Perplexity, Claude and Grok. You get a plain map of what is missing and what order to repair it in. There is no discovery call to sit through first, the report does the talking, and the method is published so you can check how it was scored.
Whatever it shows, keep the boundary in view. No outside firm controls how a model answers a stranger's next question, and anyone promising you a fixed position in an AI answer is selling a certainty these systems do not produce. What you control is whether the record about you is verifiable, corroborated and worth quoting. Then you measure, honestly, and watch the fraction move.
Want to Learn More?
Probably Genius is a done-for-you operation for expert-led firms: we get what you know onto the public record in a form machines can verify, then measure what the engines say about you every month. If the work is excellent and the introductions have gone quiet, the problem is probably not the work. You're probably a genius at what you do. We make sure AI gets the memo.