RESEARCH · AUGUST 27, 2026 · 11 MIN READ

We Scored 169 Professional-Service Articles With the Gate Our Client Work Must Pass. The Average Was 46.

We scored 169 professional-service articles from 45 domains with the Integrity Gate, the same instrument our client work must pass. The average was 45.7 out of 100. Zero reached our publishing floor of 80.

By Jacquie Baker
TL;DR — WHAT TO REMEMBER

The average professional-service article we scored measures 45.7 out of 100 on the Integrity Gate, the editorial quality scorecard our own client work at Probably Genius, the AI visibility company we run from St. Petersburg, Florida, must pass before it publishes. We measured 169 of them, from 45 real firm websites across nine industries, and not one reached the score we require before an article goes live.

That sentence is the whole study. The rest of this piece is the method, the numbers, and what they mean for anyone paying for content right now.

Here is the part nobody tells you. Nobody failed on purpose. The articles we scored were written in good faith, published proudly, and most of them read fine. Reading fine is no longer the test. AI assistants now decide which articles get quoted when a buyer asks who to trust, and they check things a casual reader never checks: whether the claims carry sources, whether the article says anything its neighbors do not, whether each section would survive being lifted out alone. Fluent is not the same as verifiable. The market just took that exam, and this is the grade.

How we built the corpus

A benchmark is only as honest as its sample, so here is exactly what we did.

We harvested articles from the sitemaps of 45 domains across nine professional-service niches: car buying, dementia and senior finance, financial advisors, CPA firms, MSP security, dental, chiropractic, heritage renovation, and family law. We capped the take at eight articles per domain so no single publisher could tilt the results, and we validated every URL live before scoring.

The judge scored 172 pages. Three of them turned out to be date-archive index pages our URL filter missed, not articles, so we removed them before computing any statistic. We document that exclusion because a benchmark that quietly drops pages invites the accusation of juicing the average. These three were not articles, the judge correctly zeroed them, and leaving them in would have dragged the market average down unfairly.

Every article was scored by the same calibrated AI judge, on the same rubric, on the same day, and every score is logged. One standing law governs the whole system: the judge is never the writer. The model that scores an article is never the model that wrote it, because a writer grading its own homework will always find a way to pass.

One more discipline matters here. Single judge runs wobble by three to four points on mid-range articles, so any score sitting within two points of a claim boundary gets re-measured as the median of three independent runs before we say it out loud. Every boundary number in this piece survived that check.

What 169 articles actually scored

The distribution, in plain numbers:

The strongest article on a single judge run scored 77: NerdWallet's guide to choosing a financial advisor. We want to say that respectfully. NerdWallet is a national editorial operation with real standards, and it still produced one of the strongest pieces in the whole corpus. Re-measured under the boundary protocol, that article settled at 73, which is exactly what the protocol is for. Single runs wobble; medians hold. The market's best article on a median of three runs was a different piece, and it landed at 78.

Our floor is 80. Every article we publish for a client is scored 0-100 through the Integrity Gate: nothing publishes under 80. So the honest, measured sentence reads like this: the best professional-service article we could find on the open internet would not have cleared the minimum bar our client work must clear. Not the average article. The best one.

You were probably never told your content had a score. It does. Everything published on the open web is now being read, weighed, and quoted or skipped by machines every day. The score existed before we measured it. We just wrote it down.

The dimension almost everyone fails: citations

The Integrity Gate scores five public dimensions: Intent Match out of 20, Information Gain out of 25, Entity Density out of 20, Citation Quality out of 15, and Extractability out of 20. Across 169 articles, the market means were:

Market means across 169 articles, out of each dimension's maximum
Citation Quality3.5 / 15
Information Gain9.2 / 25
Extractability10.0 / 20
Entity Density10.3 / 20
Intent Match12.4 / 20

Citation Quality is the market's floor: 3.5 of a possible 15.

Means are rounded to one decimal, so the list sums within rounding of the 45.7 average.

Look at the floor of that chart. Citation Quality at 3.5 out of 15 means the typical professional-service article supports its claims with almost nothing a machine can check: no named sources, no dates, no data with provenance. The article asserts. It does not evidence.

That is also the most fixable number on the list, and the most consequential. An AI assistant deciding whether to quote you is doing exactly one thing: looking for reasons to trust the sentence. AI can't recommend what it can't verify. A specific claim with a named, checkable source is quotable. The same claim naked is filler.

Information Gain, the second-weakest score, is the quieter problem. At 9.2 out of 25, the average article contributes little that the ten articles beside it do not already say. It is not wrong. It is interchangeable, and interchangeable content gives a recommendation engine no reason to pick you over anyone.

The gate you have never heard of: sections

The number that surprised us most was not the average. It was the 84 percent.

Beyond the composite score, the Integrity Gate scores every section of an article separately, and any single section scoring below 6 out of 10 fails the whole piece regardless of the total. We call it the per-section gate. The logic: machines quote sections, not articles. When an assistant lifts your weakest section as the answer to someone's question, that section is your whole reputation in that moment. An article is only as strong as the worst passage a machine might quote.

142 of the 169 articles failed it. The pattern is nearly always the same: a strong opening written with care, then sections that thin out as the writer runs out of things to say. Padding to hit a word count. A conclusion that restates the introduction. The composite score forgives that. A buyer's AI assistant does not.

The test we apply to our own drafts is blunt: could this section, lifted out alone at a couple of hundred words, still answer a real question with something specific in it? A section that only makes sense as connective tissue between two better sections is a section a machine will one day quote as your answer, and it will be empty. The per-section discipline is the strictest rule in the instrument, and the one we defend hardest. It converts "the article is pretty good overall" into "every quotable passage holds," and quotable passages are the unit of the recommendation economy.

We scored our own work the same day

A benchmark from a company that sells content should make you suspicious, so we pointed the instrument at ourselves. Same judge, same rubric, same day.

Every article published under our current 80-point standard, 18 of them at the time of measurement, was scored on the identical protocol. The cohort averaged 83.1, with a median of 82.5, and all 18 measured at or above 80.

Which produces the sentence this whole study earns: the lowest-scoring article published under our current standard still outscores the best article we found in the market, 80 against 78.

We are not claiming our writing is magic. We are claiming something more boring and more repeatable: a floor, enforced by an independent judge, changes what ships. Most content fails not because writers lack talent but because nothing stands between a fluent draft and the publish button. Our articles clear 80 because the ones that do not are sent back. That is the entire trick, and we have now measured what the absence of that trick looks like, 169 times.

Why the gate is now public

I built this instrument for our own pipeline, and for a while it lived where most agency machinery lives: inside, invisible, working. Making it public was not my idea. It was my co-founder Christopher Shaw's, and his framing was better than any positioning document we could have written. Give it to the people. Let anyone put an article in and watch the same judge our client work faces come back with the score, because that is your chance to run a piece of our engine yourself, not read a brochure about it.

He was right, and not only about the generosity of it. A published method invites checking, and we would rather be checked than believed on charm. The judge model auditioned for its seat against two rivals on four fixed anchor articles before it was trusted with a single score. The entire measurement program behind this article, the calibrations, the audition, all 169 corpus articles, our own portfolio, cost $22.44 in judge fees, every cent logged. Rigor at that price has no excuse to stay private.

So it is public. The Integrity Gate, an AI content quality checker, now runs free at probablygenius.com/integrity-gate/. Paste a URL or the text, wait about two minutes while the judges come back, and your score appears on the page. The full report, every dimension, every section and every fix, lands in your inbox.

What to do with your score

If you run a professional-service firm, the study compresses to three moves.

First, measure before you spend. You now know the market average is 46 and the passing bar is 80. Score what your current writer, agency, or AI tool produced last month before commissioning the next batch, because the study says the honest prior is somewhere in the mid-40s, and you deserve to know whether your money is buying the average.

Second, fix citations first. It is the market's weakest dimension, at 3.5 of 15, and the cheapest to repair. The repair is mechanical: for every claim in the article, name the source, date it, and link something a machine can resolve and read. A vague line like "clients often overpay at this stage" becomes a dated number from your own case files with the method attached. The first is filler. The second is a receipt with your name on it, and receipts are the difference between being quoted and being skipped.

Third, adopt a floor. Any floor honestly enforced beats no floor, because the floor is what stands between a fluent draft and the publish button. Ours is 80, in public, with the instrument attached, and the free score is where every engagement of ours starts anyway. You are probably sitting on real expertise. The score tells you whether the record shows it.

Frequently asked questions

Who checks whether an article is AI-ready?
Almost nobody, which is what the 169-article study measured. Most firms rely on the writer's own judgment, and most agencies have no independent scoring step at all. The working answer is an external judge: an instrument that scores the article against what AI assistants actually check, run by someone who did not write the piece. That is the role the Integrity Gate plays in our pipeline, and the public version at probablygenius.com/integrity-gate/ lets you run the same check on any article, free.
What is a good Integrity Gate score?
The market average across 169 professional-service articles was 45.7, and the best article we found measured 78. Our publishing floor is 80: nothing we produce for a client publishes under it. So read your score against those two lines. Above the mid-40s, you are beating the average article in your field. At 80 or above, you have cleared a bar that, in our study, zero of 169 market articles reached.
Why did you score your own articles with your own instrument?
Because a benchmark from a content company invites exactly that question. We scored all 18 articles published under our current standard with the same judge, same rubric, same day as the market corpus: mean 83.1, median 82.5, all 18 at or above 80. And the judge is never the writer: the model that scores our work never wrote it. The method is published so you can re-run any of it.
How do we get started with Probably Genius?
Start free, and start with a measurement. Run an article through the Integrity Gate at probablygenius.com/integrity-gate/, or take the Recommendation Check, our free 109-Point AI Visibility Diagnostic, for the wider picture of what AI can currently verify about your business. Every engagement of ours starts with one of those numbers, because there is no point discussing a program until you know what your record actually says. If you want the programs and what they cost, they are all public at probablygenius.com/pricing/.

CITATIONS

  1. "The Integrity Gate" (Probably Genius). The public scoring tool and method record: the 0-100 scale, the five dimensions, the 80 floor, and the three-run median protocol used for every number in this article. probablygenius.com/integrity-gate/
  2. "Our Methodologies: The Public Ledger of Probably Genius Thinking" (Probably Genius). The published methodology hub, including the AI Integrity Standard behind the Integrity Gate's rubric. probablygenius.com/methodologies/
  3. "Pricing" (Probably Genius). The public price list for every program tier. probablygenius.com/pricing/
  4. Probably Genius measurement program, 2026-08-12 to 2026-08-13. First-party logged data: 169-article market corpus (45 domains, nine niches, max eight articles per domain, three non-article index pages excluded and documented), the 18-article PG cohort, judge calibration and audition records, and the $22.44 program cost. Summarized on the tool page above.
WRITTEN BYJacquie ("Jax") Baker

Technical co-founder and CEO of Probably Genius, an AI visibility firm helping expert-led businesses become the named answer in AI search. She built the Integrity Gate, the scoring instrument behind this study. Let's talk →

Related reading

Read the method.
Then see it run on you.

The 109-point diagnostic is this library, applied to your business. Free, about an hour to present, no obligation.

See where you stand →

A real diagnostic, not a sales call in disguise.