Foundations
Is AEO and GEO just snake oil?
By Arnav Mukherjee, founder of TofuBofu · July 31, 2026
TL;DR
The mechanism is real. The certainty is often sold. I pulled 4,463 query-engine observations from 75 of our own scans and ran an identical prompt through three engines ten times each. Two findings do most of the work: companies are named 74.9% of the time when an engine is asked about them by name and 0% on general discovery questions, which is how visibility scores get inflated; and across ten identical ChatGPT calls, 61 companies were named and not one appeared in all ten. If a vendor gives you a single number without a denominator, they are selling you the variance as a result.
The first thing I ever did with our own scanner was point it at my own company. It came back zero. Not low: zero. Across every buying question we generated, no engine named us once. I had built the thing, so I knew the queries were sane and the matching worked. The number was just true.
That is the moment a founder either becomes a believer or gets suspicious. I did both, and the suspicion is the useful half. Because the accusation that this whole category is snake oil is not stupid, and the people making it are usually the people who got burned by content marketing's last three certainty cycles. They deserve a real answer instead of a brochure. So I went and looked at our own data, including the parts that are unflattering to us.
Stating the accusation fairly
"Snake oil" is doing a lot of work in one phrase, so it is worth splitting into the three separate claims people actually mean. They have very different answers.
The metric is unfalsifiable. Every vendor has an index, every index is proprietary, and no two agree. You cannot check the number, reproduce it, or find out what happens to it if you change nothing. A metric you cannot audit is a metric you have to take on faith, and taking numbers on faith is how marketing budgets get spent on nothing.
The vendor grades its own homework. Several tools in this category publish leaderboards of "most visible brands in AI search," computed with their own methodology, on which they rank themselves at or near the top. That is not evidence, it is a mirror. It is also very effective, which is why it keeps happening.
The causal claim is unproven. Score went up, revenue went up, therefore the score caused the revenue. There is no control group anywhere in this category. Nobody runs a holdout. And in a period when almost every B2B company is also changing its content, its pricing and its outbound, attributing a revenue move to a citation count is a strong claim on thin evidence.
I think claim one is partly fair, claim two is entirely fair, and claim three is fair as stated but proves less than the skeptic thinks. Here is the data.
Where the numbers get inflated
I pulled every completed scan we have run: 75 scans across 44 distinct domains between June 24 and July 27, 2026, producing 1,348 distinct question rows and 4,463 individual question-by-engine observations. Every observation is a record of whether a specific engine named a specific company in response to a specific question. Sample caveats are at the bottom of this section, and they matter.
The single most important thing in that dataset is not the average. It is how violently the average moves depending on what you ask.
| Question type | What it asks | Named | n |
|---|---|---|---|
| Brand-name | "What do you know about Acme Corp?" | 74.9% | 358 |
| Decision | "Is Acme or a rival better for X?" | 15.2% | 434 |
| Consideration | "What should I look for in a vendor?" | 14.3% | 147 |
| Category buying | "Best managed IT provider in Canada" | 4.6% | 2,115 |
| Long-tail tool | "Tool for X specific job" | 2.7% | 404 |
| Category discovery | "How does X work?" | 0.0% | 276 |
| Long-tail informational | "Explain Y problem" | 0.0% | 366 |
| Awareness | "Why does Z matter?" | 0.0% | 102 |
TofuBofu first-party data. 75 completed scans, 44 domains, June 24 to July 27, 2026. "Named" is the share of question-by-engine observations in which the company appeared in the answer.
Read the top row and the bottom row together. When you ask an engine about a company by name, it names the company three times out of four. When you ask the question a buyer actually types when they do not yet know who to call, it names them zero times out of 744.
That gap is the whole trick. A brand-name question does not measure whether an engine recommends you. It measures whether the engine can retrieve a company that exists. It is close to a lookup test, and almost everyone passes it. Put those questions in the denominator of a visibility score and the score inflates enormously, in a way that feels like good news and predicts nothing. We hit this ourselves and it is why our scanner now excludes brand-name questions from the visibility number entirely: we had a scan grading an engine's answer about a 2021 film that shared a client's name. If a vendor will not tell you which question types are inside their score, you cannot tell the difference between a company AI recommends and a company AI can merely find.
The honest caveats on this dataset, because a post accusing others of hiding denominators had better show its own. It is small: 44 companies, not 44,000. It is self-selected, since people who run a visibility scan usually suspect they have a problem, which biases the whole distribution downward. It skews to B2B services and SaaS. And the zero rows are zero out of a few hundred observations, not zero out of infinity. None of that changes the shape of the gradient, which is very large and consistent, but it does mean you should treat these as our numbers rather than the industry's.
The noise problem nobody prices in
Here is the experiment I would want to see from any vendor before believing their score, so we ran it on ourselves. Take one question. Send it to one model through the API. Change absolutely nothing. Send it ten times. Then look at how much the answer set moves.
| Engine | Distinct firms named | In all 10 runs | In 1 run only | Mean overlap |
|---|---|---|---|---|
| ChatGPT | 61 | 0 | 45 | 0.20 |
| Claude | 122 | 0 | 100 | 0.09 |
| Gemini | 47 | 1 | 30 | 0.23 |
TofuBofu first-party test, July 2026. One fixed prompt ("best managed IT services providers for a mid-sized company in Canada"), 10 identical API calls per engine, default parameters. Mean overlap is the average Jaccard similarity between the set of companies named in any two runs, where 1.0 means identical and 0 means no companies in common.
On ChatGPT, ten identical calls produced 61 different company names, and not a single company survived all ten runs. On Claude it was worse: 122 companies, none in all ten, and 100 of them appeared exactly once. Two randomly chosen Claude runs share about nine percent of their answer.
Sit with what that does to a scan. If a tool asks each question once and reports a score, a meaningful part of what you are looking at is which coin flips landed that afternoon. Your score can move several points between Tuesday and Thursday with nothing changed on your website and nothing changed at the engine. That is not the tool lying. It is the tool reporting a single draw from a wide distribution as though it were a measurement. This is the strongest version of the skeptic's case, and it is correct.
It is also fixable, which is the part skeptics usually miss. You beat variance with repeated sampling and frozen question sets, the same way every noisy measurement problem has ever been handled. Our own scans sample more than once and only compare like-for-like question sets between runs, and even then I would tell you to trust the direction across three scans over the absolute number in any one of them. The presence of noise is not evidence of fraud. Presenting a noisy number as a precise one is.
See your own numbers, with the denominators shown
A free scan runs your real buying questions across six AI engines and reports per engine and per question, not one blended grade. Brand-name questions are excluded from the visibility number by design.
Run your free scanWhat survives the scrutiny
Strip out the inflated scores and the vendor leaderboards and something solid is still standing. Three things, specifically.
The buyer behaviour shift is independently documented. This is not a vendor claim. G2's 2026 buyer research found 51% of B2B buyers now begin vendor research on an AI chatbot, up from 29%, and 69% reported switching a vendor decision based on what AI told them. Forrester's 2026 study found 94% use AI somewhere in the buying process. You can dislike survey methodology in general, but these are buyer-side surveys with no product to sell you.
The outcome is binary and checkable by you, for free. This is the strongest argument against the snake-oil charge and it costs nothing to verify. Open ChatGPT. Type the question your best customer would have typed before they knew you existed. You are either in the answer or you are not. No index, no vendor, no proprietary anything. Whatever else is contested, that specific fact is not, and it is the fact the whole category is ultimately about.
The inputs are the same inputs that have always mattered. The things that correlate with being cited are structured content, third-party corroboration, and clear category positioning. SE Ranking found 71% of ChatGPT-cited pages use structured data. Profound's research found brands present on four or more platforms are 2.8x more likely to be cited. None of that is exotic, and none of it is wasted if AI search stalls tomorrow, because it is the same work that makes you findable and credible anywhere. That asymmetry matters: the downside case for doing this work is that you end up with better-structured content and more third-party proof.
What does not survive is the revenue claim in its strong form. Nobody in this category is running holdouts. Nobody can currently hand you a trustworthy conversion rate for AI-sourced leads. Our own honest position on attribution is that AI referral numbers are a floor rather than a total, because zero-click answers, missing referrers and branded-search follow-ups all leak. If someone quotes you a precise ROI multiple for AEO, ask what the control group was. There isn't one.
Six questions to ask any vendor, including us
The category-level question is not answerable. The vendor-level one is, and these six questions separate measurement from theatre quickly. None of them require you to know anything technical.
1. Show me the per-engine breakdown, not the blended score
In our data, presence ranged from 15.1% on Perplexity down to 3.1% on Google AI Overviews across the same companies and questions. Averaging a 5x spread into one number destroys the only actionable part of it. If a tool cannot show you which engines name you and which do not, it cannot tell you what to fix.
2. What is the sample size, and how many times is each question asked?
Ask directly whether each question is run once or several times, and what the tool does when runs disagree. Given that ten identical calls can produce zero companies in common, one-pass sampling reports noise as signal. Any vendor who has thought about this will have a ready answer. Any vendor who has not will change the subject.
3. Are brand-name questions inside my visibility score?
They should not be. Brand-name questions came back at 74.9% in our data against 4.6% for real buying questions. Including them inflates the number by an enormous margin while measuring almost nothing commercially. This single question separates a lot of tools.
4. Who computed the leaderboard you are on?
If a vendor's own index ranks that vendor first, it is marketing, not evidence. This is not a reason to dismiss the tool, plenty of good tools do this, but it is a reason to discount that specific number to zero when comparing options.
5. What is the causal claim, exactly?
There is a large gap between "companies we work with are more often named in AI answers", which is checkable, and "we increased revenue 40%", which without a control group is a story. Push until you find out which one is being claimed. The honest version is still a good enough reason to buy.
6. What happens if I do nothing for 90 days?
A vendor who tells you your score will collapse is overselling. A vendor who says it will mostly drift with model updates and competitor activity, and that the fixable part is structural, is describing the actual mechanism. Guaranteed timelines are the clearest red flag in the category, because search-grounded engines and model memory move on completely different clocks.
So: is it snake oil? The mechanism is real, cheap to verify yourself, and the work it asks for is work that pays regardless. The certainty is frequently manufactured, and the manufacturing happens in three specific places: the denominator, the sample size, and the causal claim. Ask about those three and the category sorts itself out in about ten minutes.
My own company scored zero, and the zero was correct. That is the part I keep coming back to. A metric that is willing to tell the person who built it that he is invisible is at least measuring something real. The question worth asking is whether the vendor in front of you has built a number that can still do that.
Frequently asked questions
Is AEO and GEO snake oil?
The underlying mechanism is real and measurable: buyers ask AI engines for vendor recommendations, engines name some companies and not others, and which companies get named can be changed. What is frequently sold as snake oil is the certainty layer on top: single blended scores presented as precise, proprietary rankings that cannot be independently checked, and causal claims about revenue that the data does not support. Judge the vendor, not the category. A tool that shows you per-engine, per-query results with sample sizes is measuring something. A tool that gives you one number and a proprietary grade is selling confidence.
Why do AI visibility scores look better than reality?
Usually because brand-name queries are in the denominator. In our own data across 4,463 query-engine observations, brands were named 74.9 percent of the time when the engine was asked about them by name, but only 4.6 percent of the time on category buying questions and 0 percent on general discovery questions. Asking an engine what it thinks of your company mostly measures whether the engine can look you up, not whether it recommends you. Any score that blends those question types together will read far healthier than your actual commercial visibility.
How reliable is a single AI visibility scan?
On its own, not very. We ran the same prompt through the same model ten times via API. On ChatGPT, 61 different companies were named across the ten runs and not one company appeared in all ten. On Claude it was 122 companies with none appearing in all ten. Mean overlap between any two runs was 0.20 on ChatGPT and 0.09 on Claude. A single run is a sample from a noisy distribution, so a score built on one pass per question is measuring run-to-run variance as much as it is measuring you. Repeated sampling and per-engine reporting are what make the number mean anything.
Does AEO actually drive revenue, or is it a vanity metric?
It becomes a vanity metric the moment you measure the wrong questions. Being named more often on questions nobody buys from is a number going up with no commercial event attached. Being named on the specific buying-intent questions your buyers ask is a real commercial position. Nobody can currently give you a trustworthy conversion benchmark for AI-sourced leads, and any vendor quoting one should be asked for the denominator. The honest claim is narrower: AI recommendations demonstrably influence which vendors get shortlisted, and you can measure whether you are in that answer.
What are the real red flags in an AEO or GEO vendor?
Six of them. A single blended score with no per-engine breakdown. No sample size or sampling method disclosed. A proprietary index that ranks the vendor itself at the top. Brand-name queries counted inside your visibility score. Causal revenue claims without a control. And guaranteed timelines, since search-grounded engines and model memory move on completely different clocks. None of these prove bad faith, but each one hides variance that would otherwise be visible to you.
Is the AI search shift itself real, or is it hype?
The behaviour shift is well documented by independent research. G2's 2026 buyer report found 51 percent of B2B buyers now start research on an AI chatbot, up from 29 percent, and 69 percent said they switched a vendor decision based on AI input. Forrester's 2026 study found 94 percent use AI somewhere in the buying process. That part is not a vendor claim, it is buyer survey data. What remains genuinely uncertain is the size and attribution of the revenue effect for any individual company, and that uncertainty is exactly where overselling happens.
How should I evaluate AEO results without fooling myself?
Freeze your question set before you start, so the denominator cannot drift. Report per engine, never blended, because engine-to-engine spread is large. Sample each question multiple times rather than once. Exclude questions that contain your brand name from the visibility number. Track the specific questions you care about commercially rather than an aggregate. And treat any single scan as one noisy observation, with the trend across several scans as the actual signal.
Sources and further reading
- TofuBofu first-party scan data: 75 completed scans across 44 domains, 4,463 query-engine observations, June 24 to July 27, 2026. Self-selected sample skewed to B2B services and SaaS
- TofuBofu first-party variance test: one fixed prompt, 10 identical API calls each to ChatGPT, Claude and Gemini, July 2026, default parameters
- G2 2026 AI Search Insight Report: 51% of B2B buyers start research on an AI chatbot, up from 29%, and 69% switched a vendor decision based on AI
- Forrester B2B Buying Study 2026: 94% of buyers use AI somewhere in the buying process
- SE Ranking AI search study: 71% of ChatGPT-cited pages use structured data
- Profound research: brands present on four or more platforms are 2.8x more likely to be cited