Measurement
The engine that shortlists you most is the one your analytics calls worthless
By Arnav Mukherjee, founder of TofuBofu · August 12, 2026
Corrected 14 August 2026
The engine rates first published here were computed against our development database rather than the production one. Every per-engine figure on this page has been replaced with the production set, and the ranking of engines survived the change: Perplexity still leads the buying question by a wide margin, and it still sends almost no traffic.
Second correction, on ChatGPT 2026-08-18, extended to Claude on 2026-08-28, and it bears on the argument this page makes. Every ChatGPT figure below, and every Claude figure, was measured through the API in the mode that answers from training weights and does not search the web. That was the default, and it is why the number is low. We began forcing web search on ChatGPT on 2026-08-18, and in a first probe on a single brand the engine went from searching on none of ten questions to all ten. Claude followed on 2026-08-28, and a forced probe returned an answer assembled from a directory listing and three round-up articles. So the 1.7 percent and the 1.2 percent describe the instruments we used to run, not what either engine finds today. We have not replaced either, because no corpus exists on the other side of those changes yet and inventing one would be worse than leaving a number we can explain. Read the ChatGPT and Claude rows as history. The traffic half of this page is unaffected: it comes from someone else's referral data and never depended on our probe.
Two things were removed rather than restated. A sole-find analysis, which said what share of wins came from Perplexity alone, and a comparison against non-buying questions. Both were real calculations on the wrong corpus, and rather than repeat them from memory against the right one we have taken them out until they are recomputed. What was wrong here was the corpus, not the method, and the honest response to that is a smaller page.
Earlier this month we came close to dropping Perplexity from our own product. The case for it was tidy. It answers fewer questions than the other engines, it is another vendor relationship to keep alive, and the best public data on AI referral traffic says Perplexity sends almost nothing to B2B sites and that the visitors it does send convert worse than the ones from ChatGPT. Fewer engines, less to maintain, no measurable loss.
Before doing it I counted what each engine actually contributes on the only questions that matter commercially: somebody asking for the best provider of something, with no brand handed to them. Not questions about a category, not questions that name a company and ask whether it is any good. Just the shortlist moment.
On that population Perplexity is the most valuable engine we run, and it is not close. The engine the traffic data says to cut is the engine building the shortlists, and the engine the traffic data says to worship is the one that almost never names anybody.
Everything turns on which questions you count
A scan asks an engine many kinds of question, and they are not interchangeable. Some describe a category. Some probe a named company. Some ask what to consider when choosing. Exactly one kind is the buying moment, and in our product that is a defined thing rather than a judgement call: a question is a buying question if it sits in the category-level or long-tail vendor sections, which is what the scoring uses.
That definition is doing more work than any single number below it. Count a brand's mentions across every question a scan asks and you sweep in the probes that contain the company's own name, which every engine answers with that company between 82 and 100 percent of the time. We published inflated figures for exactly that reason once, and the correction cost us more than the caution would have. The rates here count buying questions only.
Here is every engine on buying questions, counted across 46 completed reports covering 34 brands, 419 distinct buying questions and 2,421 answered buying cells, on firms that arrived already suspecting they were missing from AI answers.
The spread is the finding. Perplexity names the brand at more than four times ChatGPT's rate and more than six times Claude's, on the same questions, on the same days, about the same companies. A check run on one engine is not a smaller version of a check run on six. It is a different measurement.
That ordering has a mechanism behind it, and this next part is our read rather than a measurement. Claude and ChatGPT, in the mode these figures were measured in, answer from what they already absorbed. They are fluent about a company they have encountered and noticeably careful about turning that into a recommendation. Both can now be made to search, which is why we expect this ordering to move, and is not a reason to restate it before we have measured it. Perplexity searches first and answers second, so on a question that demands current vendors it goes and finds some. The shortlist moment is precisely the moment retrieval beats recall. We think that predicts where the ordering goes as the engines change, but it is reasoning from how they work, not a second measurement.
Two scoreboards, pointing in opposite directions
The traffic scoreboard says the opposite, and it is also right
Here is the other side, measured properly by somebody else. Orbit Media analysed 97 B2B GA4 accounts covering 28.9 million sessions from July 2025 through June 2026. AI sources were about 0.5 percent of all traffic, roughly one visitor in two hundred, and those visitors converted at around three times the rate of other organic sources, seven times on the per-site median. ChatGPT converted 2.1 percent of visitors into leads against 0.5 percent for Google Search.
And the distribution: ChatGPT accounted for 82.3 percent of AI-referred visits. Everything else split what was left. Their read on Perplexity is blunt, that it is not sending much traffic to B2B sites and that the visitors it does send are less likely to convert.
Both findings are sound and they are counting different events. Whether an engine appears in your analytics depends on two things unrelated to how you are doing inside it: somebody has to click, and the engine has to identify itself when they do. An engine that answers so completely that nobody leaves contributes nothing to a referral report while shaping who gets called. Referral data is not a ranking of engines by importance. It is a ranking of engines by how legible they are to your instruments.
Which produces the specific error worth avoiding. A team opens GA4, sees one engine at eighty-something percent and the rest near zero, and concludes the rest do not matter. What they have measured is the edge of their own reporting. On our numbers the engine that would be cut first is the engine most likely to name you, and the engine that survives the cut is one of the two least likely to.
Your analytics can only report engines that send an identifiable click. Which engines name you, per engine, on the buying questions your buyers actually ask, takes a scan.
Run a free AI visibility scanWhat this page does not yet tell you
The most useful version of this analysis is not the average rate per engine, it is the sole find: a question where exactly one engine named the brand and every other engine that answered did not. An average tells you how an engine performs. Sole finds tell you what a single-engine check would have missed, which is the actual decision anyone is making when they simplify a stack.
We ran that analysis, published it, and then found it had been computed against the wrong database. It is not restated here, because a number recomputed from memory is a number nobody should trust, and repeating the shape of a finding while its evidence is under review is how the first error happened. It goes back up when it comes off the production corpus. The one thing worth saying without it is the check anybody can run on their own report: ask how many engines answered each question, and how many of your wins rest on exactly one of them.
What to do when the scoreboards disagree
Keep them apart and say what each is blind to. Mention share tells you whether an engine puts you on the list, covers every engine, and knows nothing about revenue. Traffic share tells you who clicked, covers only engines that pass a referrer, and cannot see the answer that ended the search. Conversion share is the only one touching money and has the smallest sample of the three. One blended figure hides all three blind spots at once, which is why our own headline number opens onto the per-question detail instead of standing alone.
Then check the definition before you trust any rate, including ours. Ask which questions were counted as buying questions, because that single choice moves every number on this page, and it moves them by more than the difference between the engines. Ask what the denominator was, and specifically whether questions the engine did not answer were counted as the engine failing to name you. Ask which database the figures came from, which is a question we would not have thought to publish two weeks ago and now put first. If a report cannot answer those three, its ranking is decoration.
And do not let an engine leave quietly. The practical hazard in all of this is that engines fail silently: a missing engine lowers an average instead of raising an alarm, so the day one stops answering looks identical to the day your visibility dipped. Whatever you use, know how many engines answered each question, and notice when the number changes.
Frequently asked questions
Which AI engine names brands most often on buying questions?
In our production database it is Perplexity, and the ordering is what matters rather than the absolute level. Across 46 completed reports covering 34 brands, 419 distinct buying questions and 2,421 answered buying cells, Perplexity named the scanned brand on 7.4 percent of the buying questions it answered, Google AI Mode on 4.0 percent, Microsoft Copilot on 3.1 percent, ChatGPT on 1.7 percent, Claude on 1.2 percent and Gemini on 0.4 percent. A buying question here means the product's own definition, somebody asking for the best provider of something with no brand supplied. Every brand in that corpus ran a scan because somebody suspected they were already missing from AI answers, so this is a sample selected on the problem it measures. One caveat on ChatGPT and Claude: both were measured through the API in the mode that answers from training weights without searching the web. We began forcing web search on ChatGPT on 2026-08-18 and on Claude on 2026-08-28, so those two figures describe the probes we used to run rather than what either engine finds today, and neither post-cutover rate is measured yet.
Should I stop tracking Perplexity if it sends no traffic?
In our data that would be an expensive simplification. Perplexity names the scanned brand on buying questions at more than four times ChatGPT's rate and more than six times Claude's, while sending a small fraction of the referral traffic. Orbit Media's referral study is right that Perplexity sends little traffic to B2B sites and converts worse than ChatGPT, and both things are true at once: an engine can be putting you on shortlists while your analytics records nothing, because being named and being clicked are different events.
Which AI engine sends the most traffic to B2B websites?
ChatGPT, overwhelmingly. Orbit Media analysed 97 B2B GA4 accounts covering 28.9 million sessions through June 2026 and found ChatGPT accounted for 82.3 percent of AI-referred visits, with every other engine far behind. The same study put AI sources at about 0.5 percent of total traffic while converting at roughly three times other organic sources, seven times on the per-site median, and reported ChatGPT turning 2.1 percent of visitors into leads against 0.5 percent for Google Search. Small, high intent, and concentrated on one engine.
Why do mention rates on buying questions look so low?
Two reasons, and the first is the sample. These brands came to us suspecting they were invisible, and most of them were right, so the absolute level describes that population rather than B2B generally. The second is the denominator, which is the part most published figures get wrong including ours once. Counting a company's mention rate across every question a scan asks includes probes that contain the company's own name, and every engine answers those with that company 82 to 100 percent of the time. Including them inflates the rate roughly two and a half times. These figures count buying questions only.
Is a high mention rate on one engine automatically good news?
No, and our own research is the cautionary case. In our 15-region study of managed service providers, Claude named plenty of companies and corroborated the fewest of any engine: six firms it recommended have no trace of existing anywhere we could check, and all six were Claude's. A mention rate measures an engine's behaviour as much as your standing, which is the argument against blending engines into a single score. Being named by an engine that invents companies is not the same evidence as being named by one that retrieves them.
Why is Gemini's rate described as a floor?
Because two of the faults behind it are ours, not Gemini's. On Vertex, gemini-3.5-flash spends most of its token budget on reasoning before it writes anything, so at our production limit it truncated mid-answer and any brand named late in an answer was lost. Separately, our brand-matching guard could force a mention to absent but could not rescue one until we made it symmetric in August 2026. Both depress Gemini specifically, so 0.4 percent is a lower bound rather than a clean measurement, and we publish it as one.
How should I report AI visibility to a board or a client?
Report three numbers separately and say what each cannot see. Mention share: on what percentage of buying questions does each engine name you, per engine, never blended, and state which questions counted as buying. Traffic share: what analytics recorded, covering only engines that pass an identifiable referrer when somebody clicks. Conversion share: what those visitors did, the only one touching revenue and the smallest sample. Three numbers with three stated blind spots beats one confident number, and it survives the follow-up question.
Sources and further reading
- Orbit Media: AI search conversion rates. 97 B2B GA4 accounts, 28.9 million sessions to June 2026, the 82.3 percent ChatGPT share and the conversion multiples.
- Our 15-region managed services study. Per-engine corroboration rates, and the six recommended firms with no trace of existing.
- G2 2026 AI Search Insight Report. 51 percent of B2B software buyers begin vendor research on an AI chatbot more often than Google, and 69 percent chose a different vendor than planned based on an AI answer.
- Forrester B2B Buying Study 2026. 94 percent of buyers use AI somewhere in the buying process.
- Our production scan database, read on 12 August 2026: 46 completed reports, 34 brands, 419 distinct buying questions, 2,421 answered buying cells, using the product's own buying-question definition.