THE RECOMMENDATION LAYER · JULY 17, 2026 · 10 MIN READ

What questions should I ask ChatGPT when testing my business's visibility?

The five questions to ask ChatGPT when testing your visibility: what buying intent looks like, why your own brand name ruins the test, and worked examples.

By Jacquie Baker
TL;DR: WHAT TO REMEMBER

Ask ChatGPT the questions your buyers ask before they have ever heard of you: five real hiring situations, written in your market's own words, with your business name nowhere in the text. Five questions is enough to start. What makes them work is not how clever they are. It is that a stranger could have typed them, and that you did not.

Most owners begin somewhere else. You open ChatGPT, type the name of your firm, and read a paragraph that gets you roughly right. The relief lasts about four seconds, because that is not the question your next client is going to ask. They are going to describe a mess, a town, a budget and a fear, and they are never going to say your name. You just tested the one scenario that cannot happen.

This is the part of a self-run check that nobody teaches, and it is the part that decides whether the whole exercise is worth anything. The setup matters, the scoring matters, the number of runs matters. All of it sits downstream of the questions. A clean session asking the wrong five questions produces a tidy, well-documented answer to something you did not want to know.

So here is the craft: why a branded question quietly ruins the test, the three properties every question needs, and the five shapes worth writing. Then the same five filled in for four very different practices, which shapes give steadier readings, and why the wording has to stop moving once you are happy with it. For how to run them cleanly and what to score, our companion piece on how to check whether ChatGPT recommends your business has the method. This one is about what to put in the box.

Why Typing Your Own Name Tells You Nothing

When you name your business in the prompt, you have already done the assistant's job for it. The hard part of a recommendation is selection: out of everyone in the category, which names are worth putting in front of this person. Hand over the name and there is nothing left to select. All that remains is description, and these systems are generous describers.

That produces the most common false comfort in this whole subject. The answer comes back accurate and warm, you conclude you are visible, and you stop looking. Meanwhile the question your buyer asked, which contained a problem and no company names at all, returned four other firms. That is the answer that matters.

Both questions are worth asking eventually. They just answer different things. A branded question tells you what the record says about you once someone already knows you exist, which is a real thing to check, especially if you suspect the machine has you confused with a similarly named business. An unbranded question tells you whether you get considered at all. Keep the two on separate sheets of paper and never average them into one score.

The other tempting question is the grudge: naming a competitor and asking how you compare. It feels like the sharpest possible test and it is one of the least useful, because it drops you into a shortlist that already exists rather than measuring whether you would have made it. If a rival's name is what is bothering you, there is a better way to get at that, which we walk through in how to find out which competitors ChatGPT recommends and why.

Three Things Every Test Question Needs

Here is the checklist we hold our own panels to. A usable question has three properties, and if one is missing the answer still arrives looking exactly as authoritative as a good one, which is what makes this worth being careful about.

Buying intent. The question has to be one where the answer could plausibly change who gets hired. "What is a revocable trust?" is a fine question and it tests nothing commercial, because the honest answer is an explanation with no names in it. "Who should I talk to about setting up a trust for a blended family?" is the same territory with a decision attached. If an answer could reasonably be given with no business named at all, the question is mostly measuring education, and it belongs on a different list.

Your real market. Write the words your buyers use, not the words on your services page. Practice-group labels, internal service names and the vocabulary of your professional body are all invisible to the person typing at eleven at night. Their language is sitting in your intake notes, your consultation calls, the subject lines of your inbound emails and the reviews people wrote about you. That is where the phrasing comes from. A question built out of your own website is a mirror. Mirrors always say you are covered.

No mention of your business. Covered above, and it is worth repeating because it is the property people break by accident. Location is not a brand cue and should be there. Specialty is not a brand cue either, as long as it is a specialty a buyer would name rather than a phrase only you use.

One more thing about how the sentence should read. People do not type keywords at an assistant, at least not when they are thinking something through. Semrush's study of more than 80 million lines of clickstream data from the second half of 2024 found that prompts submitted without search enabled averaged 23 words, while queries with search enabled looked like ordinary search behavior at 4.2 words. Both behaviors are real, and the long one is the one that carries a buying situation. Write full sentences with the mess left in.

There is a mechanical reason the detail helps rather than hurts. Google's documentation on its AI features says that both AI Overviews and AI Mode may use a "query fan-out" technique, issuing multiple related searches across subtopics and data sources to build a response. That is Google describing its own surfaces, not a universal account of assistants; ChatGPT and Perplexity retrieve on their own terms and document them separately. Read as one worked example, it is still a useful picture of what a rich question can set off, and a reason to think a vague question gets handled vaguely.

Five Shapes Worth Writing

Five questions, five different jobs. The point of using shapes rather than five rewrites of "best accountant near me" is that a buyer does not ask the same thing five times. They circle a decision from different angles, and you want to know which angles you are present for. Five angles, not five rewrites.

Five shapes, five questions, one each. Resist the urge to write three versions of your favorite. A panel with five near-duplicates measures one thing five times and tells you nothing about the other four moments in the decision.

The Same Five Shapes, Four Different Practices

The shapes are portable. What they get filled with is not, and abstract advice about buyer intent tends to evaporate the moment somebody opens a blank document. So here they are, filled in, for four practices that have almost nothing else in common.

An estate attorney in a mid-size city.

A pediatric occupational therapist.

An architect who works on heritage buildings.

A bookkeeper who works with medical practices.

Read those four sets and the pattern shows up on its own. The hire questions look almost interchangeable. The situation questions could not belong to anybody else. That gap is the whole argument for writing all five, because the crowded question and the question you can win are rarely the same sentence.

Some Shapes Wobble More Than Others

Different shapes give readings of different steadiness, and knowing which is which will save you from over-reading a bad afternoon. Conductor, an SEO software vendor, published a study dated July 17, 2026 covering 14,000 API calls across ten industries, seven intent types and four assistants, at fifty runs per combination. It measured brand overlap, meaning the share of brands that any two runs had in common. The type of prompt predicted consistency better than the industry did: comparison prompts came in at 63 percent overlap and purchase prompts at 40 percent. It is vendor-produced research on a set of industries the vendor chose, so hold the exact figures loosely. Keep the ordering.

The practical reading for your panel is not that you should only ask the steady kind. It is that a wobbly shape needs more patience before you conclude anything from it. Our five shapes are not the same taxonomy Conductor used, so treat the mapping as rough: an either-or question sits near their comparison intent, a hire question nearer their purchase intent. Read that way, a situation question that loses you one month and finds you the next may be the shape moving rather than your record.

Which is also the answer to the engine question. The five shapes travel: a buying situation is a buying situation whether you type it into ChatGPT, Gemini, Claude or Perplexity, and you should put the identical five to each of them. What does not travel is the conditions around them, or the answers, which on the evidence of the panels we run disagree more than owners expect. We took that apart separately in why five engines answer the same question differently about you. Record each engine under its own name and date, and never blend the five into one sentence about "AI."

Freeze the Wording, Then Leave It Alone

Now the discipline that turns five good questions into an instrument. Once you are happy with them, they stop changing. Not the phrasing, not the city, not the order of the clauses. Write them into a document, date it, and treat any edit as an event you have to note, the way you would note replacing a scale in a kitchen.

This is stricter than it sounds necessary, and there is evidence for the strictness. A May 2026 preprint tested roughly 6,000 paraphrase runs against 6,000 same-prompt reruns on OpenAI and Anthropic models, and found that two cosmetic rewordings of the same buying intent produced recommendation sets overlapping by 0.288, while adding a constraint pushed it down to 0.135. The same prompt run twice overlapped by 0.50 to 0.61. That study is a preprint rather than peer-reviewed work, and it is about paraphrases rather than calendars, so read it as a warning rather than a law: a question you tidy up between quarters is close enough to a different question that the comparison you were building gets much harder to trust.

It also means you can stop agonizing over the perfect wording. There is no perfect wording. There is a defensible set of five that a real buyer might plausibly have typed, and the value comes almost entirely from asking those same five again in three months. Good enough and repeated beats excellent and rewritten, every time.

The wording is not the only thing worth holding still. The conditions you run under move results too, which is why the companion piece asks you to record the engine, the date, the session and whether the assistant searched. Give the panel one scheduled review a year, so a genuine change in your market gets into it deliberately rather than by drift. If you want to add questions in between, add them as a second list and report them separately. That is how we run our own measurement, and it is the difference between improving the instrument and breaking it. The other half of that discipline, the denominators and the honest way to state a result, sits in how to measure whether AI visibility work is actually working.

One last thing, for the owner who already has a marketing partner and is quietly wondering what this exercise is going to start. Five written questions are the most useful thing you can bring to that conversation, because they turn a vague worry into a specific list your agency can be asked about directly. We wrote about how to have that discussion without it turning into an audit in what "we handle the AI stuff" actually covers. Nobody has to be wrong for the panel to be worth having.

Want to Learn More?

Writing the five is free, and you should do it whether or not you ever speak to us. What the panel cannot tell you is why the answers came back the way they did, or which repair to make first, and the order matters more than the effort. Our free Recommendation Check is the same idea run properly: your buyer questions put to the five engines we track, alongside a score of your record against 109 checkpoints across eight zones, hand-scored. Scores run 0 to 109. Ninety and up means AI treats you as a primary source. Under 35 means you are effectively invisible. There is no discovery call attached to it, because the report does the talking.

Beyond that sits The Answer, which is the operated version: twelve done-for-you thought-leadership articles a month, planned to create approximately 60 clear Knowledge Entries, your expertise structured so machines can read it, corroboration built beyond your own pages, and the five-engine probe run monthly with every answer saved word for word. No one can promise you a fixed place in an AI answer, so what we build is the record the assistant reads, and then we measure it in the open.

You're probably a genius at what you do. We make sure AI gets the memo.

What questions should I ask ChatGPT when testing my business's visibility?
Ask five questions a real buyer would type before they knew your business existed. Each one needs buying intent, your market's actual words, and no mention of your business name. Cover five different shapes rather than five versions of one: the plain hire question, the buyer's situation described with no category named, an either-or between two kinds of provider, a constrained request naming the industry or regulation that decides fit, and a diligence question about what to look for before hiring. Write them once, date them, and ask the identical five again next quarter.
Why shouldn't I just type my business name into ChatGPT?
Because naming your business removes the only hard part of the task. A recommendation is a selection: out of everyone in the category, which names get put in front of this person. Once you supply the name, nothing is being selected and the assistant simply describes what it can find, which usually reads well and proves nothing. A branded question is still worth running separately, particularly to check whether your details are accurate or whether you are being confused with a similarly named business. Keep it on its own list, and never merge it into the unbranded results.
How specific should my test questions be?
Specific enough that a stranger in your market could have typed it, and no more. Include the city, the situation and the one or two constraints that genuinely decide who is right for the job. Leave out anything that only describes you: your firm's phrasing, your practice-group labels, and the full list of credentials you would like acknowledged. A useful test is to read the question aloud and ask whether a worried person with your problem could have written it without knowing you exist. If the answer is no, it is too specific and you are testing whether the assistant can follow a description.
Do I need different questions for Gemini, Claude and Perplexity?
No. Question design travels across assistants, because a buying situation does not change with the software, so put the identical five to each engine. What changes is everything around the questions: default personalization, whether the assistant browses the live web, how sources are shown, and above all the answers themselves, which disagree between engines more than most owners expect. Record each engine separately with its own date, and resist averaging them into a single claim about AI in general, since blending hides the disagreement that was the most useful part of the reading.
How many questions should I write?
Five is where we start for a single-service business, one question for each shape, and it is small enough that you will actually repeat it. Treat it as a starting design rather than a validated sample size. Add a question when you add a genuinely different service, a second location or a distinct buyer type, rather than to cover another phrasing of something already on the list. Growth in a panel should track real coverage gaps, not enthusiasm. If you do add questions, keep the original five untouched and report the new ones on a separate line so the comparison you have been building survives.

CITATIONS

  1. "AI Brand Recommendation Study: Why Intent Type Predicts AI Output Consistency" (Conductor, July 17, 2026). A study of 14,000 API calls across 10 industries, 7 intent types and 4 assistants, at 50 stateless runs per combination, finding that prompt intent predicted brand-set consistency better than industry, with comparison prompts at 63 percent brand overlap between runs and purchase prompts at 40 percent. Vendor-produced research on its own customer categories, and the clearest public evidence that the shape of a question changes what the reading can tell you. conductor.com
  2. "Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation" (Jack, Lehman, Maloney and Xu, arXiv preprint, not peer reviewed, submitted May 22, 2026). Roughly 6,000 paraphrase runs and 6,000 same-prompt rerun controls across OpenAI and Anthropic models, finding recommendation-set overlap of 0.288 for cosmetic rewordings and 0.135 for constraint-adding rewordings, against a 0.50 to 0.61 same-prompt baseline. The evidence behind freezing your wording once the five questions are written. arxiv.org
  3. "AI features and your website" (Google Search Central documentation). Google's own statement that AI Overviews and AI Mode may use a "query fan-out" technique, issuing multiple related searches across subtopics and data sources to develop a response, and that AI Mode is particularly helpful where further exploration, reasoning or complex comparison is needed. A description of Google's surfaces specifically, and a useful picture of what a detailed buyer question sets in motion. developers.google.com
  4. "New Semrush Study Reveals ChatGPT Search Trends: Insights from 80 Million Clickstream Records" (Semrush, February 3, 2025). An analysis of more than 80 million lines of clickstream data from the second half of 2024, reporting that prompts submitted without search enabled averaged 23 words while queries with search enabled averaged 4.2 words. A clickstream sample rather than a controlled study, and the clearest public evidence that people bring whole situations to an assistant rather than keywords. semrush.com
WRITTEN BYJacquie ("Jax") Baker

Founder of Probably Genius, an AI visibility firm helping professional service brands become the named answer in AI search. Nearly two decades across technology, digital strategy and branding, including 1,000+ digital projects through her previous agency, now focused on making experts visible, verifiable and recommendable to AI. Let's talk →

Related reading

Read the method.
Then see it run on you.

The 109-point diagnostic is this library, applied to your business. Free, about an hour to present, no obligation.

See where you stand →

A real diagnostic, not a sales call in disguise.