You still need AI visibility work even with a good SEO agency, because the two disciplines cover different ground: SEO earns your place in ranked results, while AI recommendation runs on verification and evidence that sit mostly outside an SEO retainer's scope. In a July 2025 Ahrefs study of 15,000 prompts, only 12 percent of the URLs cited by AI assistants ranked in Google's top 10 for the same query. Ranking well and being the answer have quietly become different achievements.
This is an awkward article to need. The relationship with your agency is good, the work has been solid for years, and when they said "we handle the AI stuff" in the last quarterly review, you nodded, because you had no way to check. Nobody on the call could show you the checklist that proves the AI layer is covered, and you did not want to turn a good meeting into an audit.
So let's be precise instead of suspicious. Nothing here argues for firing anyone. Your agency is probably doing exactly what it was hired to do. The question is whether what it was hired to do includes the layer where buyers now ask an assistant who to trust, and that question has a checkable answer.
Here is what a strong SEO program already earns, where the evidence says rankings stop transferring, what "we handle the AI stuff" can actually mean, and the questions that map the gap while keeping the relationship intact.
What Your SEO Agency Already Covers
Start with credit where it belongs, because the foundation matters more in the AI era, not less. A good SEO agency keeps your site technically healthy: crawlable, fast, indexed, structurally sound. It builds pages that rank for the searches your buyers still run, maintains the local presence that feeds your maps results, and earns links that tell Google your site matters. It captures demand you would otherwise lose, and it usually earns its retainer in traffic you can see and measure.
None of that is wasted work for AI visibility. Assistants retrieve from the open web your agency has been optimizing, and a site that is findable and readable is the entry ticket. If your agency has done its job, you are standing on the first rungs of the ladder we walk in The Recommendation Layer: findable and readable are real achievements.
The stall comes higher up. Trusted and recommended run on different evidence than ranked, and that evidence was never in the SEO scope of work. Not because your agency is negligent. Because the deliverable did not exist when the contract was written.
SEO is not obsolete. It is incomplete. The distinction decides everything about what you do next.
Where Rankings Stop: The 12 Percent Problem
The cleanest measurement of the gap comes from Ahrefs, which tested 15,000 long-tail prompts across ChatGPT, Gemini, Copilot and Perplexity in early July 2025 and checked whether the URLs the assistants cited also ranked in Google's top 10 for the original query. Overall, 12 percent did. For ChatGPT the overlap was about 8 percent. Perplexity, the most search-grounded of the four, still came in under a third, and roughly 80 percent of cited URLs did not rank anywhere in Google for that query at all.
Read that carefully, because it is not saying rankings are worthless. It is saying rankings do not transfer. The assistants break a buyer's question into their own research queries, fan out across the web, and assemble an answer from sources your rank tracker has never watched. Your position-three ranking is real. It just is not the input the assistant was using.
There is a second, harsher finding about what happens even when your content does get used. In a June 2026 Search Engine Land analysis by Lily Ray of 100 B2B software queries, Google's AI Overviews cited a brand's own listicle and still did not recommend that brand in 69 percent of the cases studied. Being read is not being chosen. Citation and recommendation are different jobs, and the second one runs on evidence about your business, not the quality of your pages.
There is a content trap hiding in that gap too. The classic SEO response to weak visibility is a bigger version of the page that already ranks. But an assistant assembling a single answer has no use for the fifth restatement of the same advice; in practice, duplicated advice tends to get one summary line while the sources with something distinct to say get the citations. What earns a place in the answer is knowledge the other sources do not carry, which is an editorial problem, not a length problem.
This is why the shortlist can go to a competitor whose site ranks below yours. The machine is not reading your rankings. It is checking what it can verify, what others say, and what your pages know that nobody else's do. Your rankings compete for a click. A recommendation replaces the click, and the click itself is thinning: Pew Research Center's study of 68,879 real searches found that with an AI summary on the page, clicks on traditional results fell from 15 percent of visits to 8.
What "We Handle the AI Stuff" Can Mean
When an agency says the AI stuff is handled, they are usually telling the truth about one of three very different jobs. Knowing which one you are getting is the whole conversation.
Version one: tracking. They have added an AI visibility dashboard, one of the monitoring platforms that reports whether your brand appears in AI answers. Genuinely useful, and worth having. But a dashboard is measurement, and measurement changes nothing by itself. Observed citation sources can also change substantially month to month, so a one-time screenshot is not even reliable measurement.
Version two: formatting. They have started structuring content for answer engines: schema, FAQ blocks, direct-answer openings, the tactics that travel under GEO or AEO. Real work, and it overlaps with what we describe in our plain-English GEO guide. But formatting improves pages you already have. It does not create the off-site corroboration, the coherent entity, or the demonstrated expertise the engines check before they name you.
Version three: recommendation readiness. The full job: a verifiable entity, independent third-party evidence, published expertise carrying knowledge the machine cannot get elsewhere, and scheduled measurement across the engines your buyers actually ask. This is operational work, month over month, and it is rarely inside an SEO retainer, because it is not SEO.
Most "we handle it" answers describe version one or two, and both are worth keeping, but neither one is the third job. Measurement shows the gap. Operations close it.
What to Ask Your Agency This Quarter
You can establish which version you have in one collaborative conversation. Bring these to the next review, framed as scope questions rather than accusations.
- Ask what they measure: which AI engines they check, with what buyer questions, on what schedule, and whether you can see last month's results beside this month's.
- Ask what the engines answer today: not whether you are mentioned somewhere, but who actually gets recommended when a buyer describes your kind of problem without naming you.
- Ask what exists off your own site: independent citations, directory consistency and third-party mentions, the corroboration an assistant checks that no amount of on-site work creates.
- Ask who maintains your entity: whether your name, people, credentials and locations agree everywhere the machine looks, and who fixes the record when the details drift.
- Ask what new expertise reached the record this quarter: pages that answer buyer questions with knowledge only your firm has, as opposed to refreshed category content.
A good agency will answer some of these well, and the honest ones will tell you plainly which items are outside their scope. That plain answer is the win. Now the gap has edges, and nobody got fired to find them.
If the Answers Come Back Thin
If the conversation shows the third job unowned, the move is addition, not replacement. Keep the agency on the foundation they are good at, and add the recommendation layer beside it, built by an operator whose whole job is making your expertise verifiable.
In practice that looks like the Genius Interview getting your judgment out of your head and onto the record, then 12 done-for-you thought-leadership articles a month planned to produce approximately 60 clear Knowledge Entries, plus external validation beyond your own pages, with every piece scored 0 to 100 through the Integrity Gate, where nothing publishes under 80. None of it duplicates what your agency bills for, which is exactly the point.
The two programs do not compete. Your SEO agency makes you findable and keeps the site healthy. The recommendation work gives the engines something to verify once they have found you. Buyers meanwhile keep moving upstream of the click, into shortlists you never see, which is exactly why the layer cannot stay unowned.
The zero-risk first step is a baseline reading beside your current program. The free Recommendation Check tests the five engines your buyers actually ask, ChatGPT, Gemini, Perplexity, Claude and Grok, with the questions they actually ask, scores your visibility across 109 checkpoints, and shows you in plain language which gaps sit outside normal SEO scope. Take the results to your agency if you like. Good partners tend to appreciate a map.
No one can promise what a model will answer, and this work does not pretend to. It makes you easier to verify, harder to overlook, and measured on a schedule. The rest is earned.
Want to Learn More?
Probably Genius was built by Jacquie Baker and Christopher Shaw, who spent more than 40 combined years translating what makes an expert the best in the room for the systems that decide who gets chosen. They have watched good businesses almost fire good agencies over a gap neither side had words for. The gap has words now, and a measurement. You're probably a genius at what you do. We make sure AI gets the memo.