Measurement
The schema study everyone is quoting excluded every page that needed schema
By Arnav Mukherjee, founder of TofuBofu · August 20, 2026
Three numbers are circulating as the end of an argument. Adding schema markup moved AI citations by plus 2.2 percent on ChatGPT, plus 2.4 percent on Google AI Mode and minus 4.6 percent on AI Overviews. I went and read the study at source, because our own product hands customers a Fixes list made largely of schema, directory listings and technical plumbing, and if that work is worthless I would rather find out from Ahrefs than from a customer.
The three numbers are correct. The statistics are careful, the design is better than most things published in this category, and I am not going to spend a post picking at it. The problem is one line, and it is not buried and it is not hidden. It sits under a heading that reads Caveat, in Ahrefs' own words: every page in the dataset had 100 or more AI Overview citations in February 2025, before any schema was added.
Read that as an entry requirement, because that is what it is. To be in this study your page had to already be one of the pages Google's AI was quoting hundreds of times. If you are a B2B firm that no engine has ever named, your pages were not eligible for the sample, the control group or the conclusion. The study is a measurement of whether the winners can win harder. It was never a measurement of you.
What the study did, and why it deserves the respect
Start with the thing it killed, because that part is genuinely useful and it kills a claim our own category has been selling.
Ahrefs first looked at 6 million URLs and found that AI-cited pages were almost three times more likely to carry JSON-LD than pages that were not cited. That is the stat that has been travelling around conference slides for two years as proof that schema is a visibility lever. Ahrefs then did the honest thing and refused to trust their own number, on the grounds that schema tends to live on better-maintained sites which also publish stronger content, earn more links and do everything else that gets a page cited.
So they built a causal test. Using their crawler's HTML history they found 1,885 pages where JSON-LD flipped from absent to present between August 2025 and March 2026, and dated the flip. For each one they picked three control pages on different domains with similar prior citation levels that never added JSON-LD, and the matched control set came to 4,000 pages. They counted citations 30 days before and 30 days after, then ran a matched difference-in-differences test to strip out the platform-wide trends that were shaking everything at the time, with AI Overviews contracting and AI Mode exploding. That last part matters more than it sounds: raw before-and-after growth for AI Mode came in at plus 43 percent, and almost all of it turned out to be the platform lifting everybody. Strip the trend and you get plus 2.4 percent.
Four analyses, all agreeing. The AI Overviews decline of 4.6 percent is statistically significant and the authors put the odds of seeing it by chance at roughly 1 in 2,500, then immediately tell you not to read it as schema hurting you: the absolute size is about 12 daily citations on pages that were mostly getting hundreds, and both groups were already falling before anyone touched anything. They label the decline real and unexplained rather than dressing it as a finding. That is what good work looks like, and the conclusion drawn from it is defensible: on a page already inside the consideration set, adding markup is not the unlock.
The sentence that decides whether any of it is about you
Here is the author, Louise Linehan, drawing the boundary in her own post. I am quoting both sentences because the second one is what decides whether the study is about you, and it is the one that disappears the moment the finding gets compressed into a headline.
“If a page is already getting picked up, our data suggests that adding schema isn’t going to push it higher. But for pages that aren’t being seen by AI systems at all, schema markup might still play a role in helping them get crawled, parsed, or indexed in the first place.”
Louise Linehan, Ahrefs, 11 May 2026. Emphasis in the original.
She goes further in the surrounding paragraphs, writing that the pages studied were already inside the consideration set, being crawled and surfaced by LLMs, and that the study cannot speak directly to pages that are not. This is not a grudging footnote extracted by a critic. It is the author setting the edge of her own claim, which is the behaviour you want from anyone publishing data.
The tell is in their own instructions for replicating the study on your own site. Pick 5 to 10 test pages, they say, ideally pages already getting some AI citations, so you have a baseline, because pages with zero citations make it harder to tell whether schema did nothing or whether the page just was not going to get cited either way. That is a completely reasonable protocol for a controlled experiment. It also means the DIY version reproduces the same selection: everyone running this test is testing it on pages that were already visible. The population that most needs the answer is methodologically inconvenient, so it keeps getting excluded, and the absence of evidence keeps getting reported as evidence of absence.
Conceptual. The two blocks are not drawn to any measured proportion. The only figures shown are the study's own sample sizes.
Nobody hid the caveat. Half the people repeating it threw it away anyway
I expected a game of telephone, where the scope limit fell out somewhere between the study and the founder quoting it at me. I read four write-ups of it. That is not what happened, and what did happen is more useful.
Two of the four handled it properly, and the category deserves the credit when it gets something right. Search Engine Journal's news write-up on 16 May states the threshold plainly, that every page in the dataset already had more than 100 AI Overview citations before any schema was added, repeats it in its own limitations section, and then scopes its closing line to pages already visible in AI Overviews rather than to everyone. It also states plainly that reading the study as proof that schema does not work overstates what the data showed. A second analysis quotes Linehan's caveat verbatim and carries it into its own conclusion, that the study does not dismantle schema, it dismantles schema as a booster for pages that are already visible. Neither of those is the version anyone is quoting at me.
The third is where it goes wrong, and it is worth reading closely because the failure is not dishonesty. A July post titled Schema Will Not Save AI Visibility contains this sentence: every page in the sample already had a meaningful AI Overview citation baseline before treatment. It sits in the second paragraph, directly beneath an opening line that says schema markup does not increase AI citations and Ahrefs proved it, under a headline that generalises to every site on the internet.
That is the failure mode. The scope limit gets transcribed accurately and then plays no part in the conclusion. It is treated as a disclaimer, the thing you print so nobody can complain, rather than as what it actually is: the definition of who the finding is about.
The fourth drops it entirely and goes furthest. A curated brief from 15 May reports zero effect across the three platforms, never mentions the citation threshold anywhere on the page, and concludes that the technical-markup-as-visibility-lever era is over. Same numbers, no sample, universal verdict. That is the version that reaches a founder, and it is the version that will be quoted back at anyone selling technical AEO work this quarter.
Find out which population you are actually in
A free scan puts real buying questions to six engines and shows you the verbatim answers, so you can see whether you are inside the consideration set or outside it before you decide what to fix.
Run your free scanThe population that was excluded is the one we spend all day measuring
A page with 100 AI Overview citations is not a normal page. It is an outlier by a wide margin, and it is nowhere near the situation of the companies that come to us.
Across 46 completed reports covering 34 brands and 419 distinct buying questions, on 2,421 answered engine cells, our best-performing engine named the scanned brand on 7.4 percent of buying questions, and that is measured on a corpus which is not a random sample, because nearly every firm in it ran a scan already suspecting it was missing from AI answers. The rest of the field sits below that. Gemini reads 0.4 percent on that same corpus of firms who already suspected they were missing, and we publish it as a floor rather than a clean reading, because our own token budget was truncating its answers and our brand-matching guard could not rescue a mention until we fixed it in August. ChatGPT reads 1.7 percent across the same firms who suspected they were invisible, and I will not lean on that either, because every one of those calls was made before we forced web search on that probe, so the figure describes what a model remembered rather than what the engine finds.
Caveat every number, including your own. That is the entire point of this post, so it would be absurd to publish a range without saying which parts of it I distrust and why.
What survives all of those caveats is the shape. Single digits at the top, fractions of a percent at the bottom, on the questions where a buyer is choosing a vendor. Not one of those brands could have entered the Ahrefs sample. Their pages have no citation history to form a baseline from, which is precisely why the study could not include them and precisely why the conclusion cannot be extended to them. Applying a finding about heavily cited pages to a company that has never been cited is not a small stretch. It is a different question with a different answer that nobody has measured yet.
This is not me arguing that schema works
It would be convenient to end here with a defence of the plumbing, and I am not going to, because our own data does not support it and I would be doing exactly what I just criticised.
Earlier this month we crawled 14 managed IT firms our scans had already measured as absent from AI vendor answers, against the specific rivals those same engines named instead in the same regions. Organization schema sat on 86 percent of the invisible firms and 88 percent of the named ones. Three axes ran backwards, with more of the invisible firms serving an llms.txt file and publishing pricing pages than the named ones. One axis out of eight moved in the right direction, published FAQ markup, and the sample is small enough that I called it thin at the time. The full crawl is here. If schema were the differentiator, that table would look completely different.
There is a second piece of evidence worth handling carefully, because it is being over-read the same way. Ahrefs point to a searchVIU experiment that put prices on a test page in JSON-LD, Microdata and RDFa and asked five systems to read them back. During direct fetch, ChatGPT, Claude, Perplexity, Gemini and Google AI Mode all extracted visible content only, and ignored the structured data. That gets repeated as proof that engines do not read schema. The author's own conclusion is narrower: he writes that his tests primarily show the direct-fetch phase, that schema markup could very well be used in the phases before it, and that during indexing schema markup is very likely extracted. He tested one stage of a pipeline and said so. His readers announced a verdict on the pipeline.
Then there is Google retiring the FAQ rich result. Their changelog records the deprecation notice on 8 May 2026, with the feature no longer appearing in Google Search from 7 May, and the documentation removed on 15 June. That is a real thing that happened and it is bundled into the same obituary. It is also a fact about a blue-link search appearance, decided by the Search team for Search reasons. It is not a statement about how a generative engine parses a page, and treating those two systems as one is the same category error, committed at a different layer.
So here is my position, labelled as a position rather than a measurement. Schema is necessary plumbing and nowhere near sufficient. It is a fixed week of work that you finish once and stop thinking about, not a retainer and not a strategy. Anyone selling it as the thing that gets you into AI answers is selling you a checklist because a checklist is easy to invoice. And anyone telling you to rip it out on the strength of this study is over-reading a paper that studied the opposite situation. Both readings are wrong, and they are wrong for the same reason: neither one checked who was in the sample.
Four questions to ask the next study, and there will be one next week
This category now publishes a headline finding roughly every fortnight, and most of them will be quoted at you by someone trying to sell you something or trying to avoid doing something. These four questions cost about ten minutes and settle most of them.
1. How did a page or brand qualify to be in the sample? Not how many were in it. How they got in. The Ahrefs study says it in one sentence and that sentence is the whole scope of the result. If a study cannot tell you its entry criteria, it does not have a defined population and the finding has no address.
2. Is the outcome the thing you care about, or a proxy for it? This study counts AI Overview and ChatGPT citations of a URL. That is a real outcome and it is not the same as an engine recommending your company by name in an answer to a buying question. A page can be cited as a source in an answer that recommends someone else, which is a distinction worth holding onto whichever tool you buy.
3. Which stage of the pipeline was measured? Crawling, indexing, entity resolution, retrieval and answer generation are five different systems, and a result from one of them is not a result about the others. The searchVIU work is a clean example: correct at the retrieval stage, routinely quoted as if it covered all five.
4. What did the authors refuse to claim? Read the caveats before the headline. Good researchers mark their own boundaries, and Ahrefs did: they flag that pages adding JSON-LD often change other things at the same time, that all schema types were pooled together so FAQ and Product and Organization cannot be separated, that only 30 days were measured, that JavaScript-injected schema was not tested, and that the AI Overviews decline is unexplained. Any one of those, taken seriously, softens the headline that is being sold.
We got this wrong ourselves and in public, which is why I am confident it is a general failure rather than someone else's carelessness. This site published engine mention rates computed over a question set that included questions containing the customer's own company name, which every engine answers with that company most of the time because the name is sitting in the question. Every number was inflated by roughly 2.6 times, and the wrong ones reached all eight of our engine guides. The statistics were fine. The denominator was the claim, and we had not looked at it. The correction now sits on the pages that carry the numbers rather than in an apology nobody would read.
Frequently asked questions
Does schema markup help AI citations?
On pages that are already cited heavily by AI, the best available evidence says no. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 matched controls, and found citation changes of plus 2.2 percent on ChatGPT, plus 2.4 percent on Google AI Mode and minus 4.6 percent on AI Overviews. The first two are statistically indistinguishable from zero. But every page in that dataset had passed a screen of 100 or more AI Overview citations before any schema was added, so the finding describes pages already inside the consideration set. For a page no engine has ever named, the study does not have an answer, and its own author says so.
What did the Ahrefs schema study actually measure?
Whether adding JSON-LD to a page that is already being cited a lot makes it cited even more. The design is a matched difference-in-differences test: for each of 1,885 pages that flipped from no JSON-LD to JSON-LD, three control pages on different domains with similar prior citation levels were picked, and citations were compared for 30 days either side of the change. Four separate analyses agreed. That is careful work and the answer it produced is credible for the question it asked. The question it asked is a narrow one, and the narrowness comes from the sample, not from the statistics.
Why does the Ahrefs schema study not apply to my site?
Because of one line in the caveat section: every page in the dataset had 100 or more AI Overview citations before any schema was added. If your company is missing from AI answers, your pages could not have entered that sample. Author Louise Linehan draws the boundary herself, writing that for pages that are not being seen by AI systems at all, schema markup might still play a role in helping them get crawled, parsed, or indexed in the first place. A study on pages that are already winning cannot tell you what moves a page that is losing, in the same way that a study of marathon finishers cannot tell you why someone did not start.
Should I remove my schema markup after this study?
No, and nothing in the study suggests it. The largest effect it found was a 4.6 percent relative decline in AI Overview citations, which the authors describe as real, small and unexplained, and which they explicitly warn against reading as evidence that schema hurts. Removing markup that already works is a cost you pay today against a benefit nobody has measured. The correct action after reading this study is to stop treating markup as your growth plan and to stop paying anyone a retainer to maintain it, not to spend a sprint tearing out working code.
Do AI engines read JSON-LD at all?
At the moment they fetch a page live, apparently not. A searchVIU experiment put prices on a test page in JSON-LD, Microdata and RDFa and asked ChatGPT, Claude, Perplexity, Gemini and Google AI Mode to read them back, and during direct retrieval every system extracted only visible content. That result is narrower than the way it gets repeated. The same author writes that his tests primarily show the direct-fetch phase, that schema markup could very well be used in the phases before it, and that in the indexing phase schema markup is very likely extracted. Retrieval is one stage of a pipeline. Measuring one stage and announcing a verdict on the pipeline is the same error as measuring already-cited pages and announcing a verdict on every page.
Google deprecated FAQ rich results. Does that mean FAQ schema is dead for AI?
Those are two different systems and the second does not follow from the first. Google's own changelog records the deprecation on 8 May 2026, with the feature no longer appearing in Google Search from 7 May 2026, and the documentation removed the following month. That retires a blue-link search appearance. It says nothing about how a generative engine parses a page, and the two decisions are made by different teams for different reasons. In our own matched-pair crawl of firms AI names against firms it ignores, published FAQ markup was the one axis out of eight that moved in the right direction. Publishing the questions your buyers actually ask, in plain visible text, is the part that was doing the work anyway.
If schema is not the lever, what is?
Having said something worth quoting, on a page an engine can reach, before your competitor did. Markup makes an existing answer machine-readable. It does not create an answer, and no amount of it rescues a site that has never addressed the buying question in visible prose. We crawled 14 firms that engines never named against the specific rivals named instead, and the technical checklist barely separated the two groups: Organization schema sat on 86 percent of the invisible firms and 88 percent of the named ones. Treat the plumbing as a fixed week of work you finish and never think about again, then spend the rest of the quarter on the answers themselves.
Sources and further reading
- Ahrefs, We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved. By Louise Linehan, contributor Xibeijia Guan, reviewed by Ryan Law, 11 May 2026. Source of every study figure, the sample construction and the caveat quoted above. Read it before you quote anyone quoting it.
- searchVIU, Schema Markup and AI: what ChatGPT, Claude, Perplexity and Gemini really see. December 2025. The direct-fetch experiment, and the author's own statement that the test covers the retrieval phase rather than the indexing phases.
- Google Search Central, documentation changelog. The 8 May 2026 entry deprecating the FAQ rich result, and the 15 June entry removing its documentation. Primary source for the deprecation dates.
- Search Engine Journal, Matt G. Southern, 16 May 2026. Included as a positive example: the news write-up carries the citation threshold and repeats the scope limit more than once.
- Our matched-pair crawl of 14 invisible firms against the rivals AI names. The first-party source for the 86 versus 88 percent Organization schema figure and the three axes that ran backwards.