Measurement
We asked the same buying question as four different buyers. The vendor lists barely overlapped.
By Arnav Mukherjee, founder of TofuBofu · August 11, 2026
Every AI visibility tool, ours included, reports a mention rate. You appear in some percentage of the answers to your buying questions. It is a reasonable number and it hides something I had not measured properly until this week.
On 11 August I asked four of the six engines we track, ChatGPT, Claude, Bing Copilot and Google AI Mode, the same buying question in two categories, framed four ways. Neutral. As a fifteen-person startup. As a large enterprise. And asking for the cheapest option. Same category, same day, same engines, only the buyer changed. Across 48 calls the engines named 97 distinct firms, and almost none of them survived all four framings.
In managed IT services, 2 of 63 firms were named under all four. In sales engagement software, 5 of 34. Which means a mention rate is not really a measure of your visibility. It is a measure of your visibility to whichever buyer your question set happened to describe.
First, the objection: is this just AI being random?
This was my first reaction and it is the right one. We have published before that these answers move on their own, so a difference between two answers proves nothing until you know how much they differ when nothing changes at all.
So the probe includes a control. The neutral question was asked three separate times per engine per category, with nothing changed. Comparing those runs to each other gives the noise floor: how much the vendor list moves for no reason. Then the persona framings are compared against the same baseline.
Asking again gets you about 55 percent of the same names. Asking as a different buyer gets you about 25 percent. Changing who you say you are moves the answer roughly twice as much as the engine's own randomness does, and the effect held on every engine individually, which is what makes it worth taking seriously on a probe this small.
Per engine, with a second finding hiding in it
The persona column behaves the same everywhere, which is the finding. The other column is the surprise: these engines are not equally stable. Ask ChatGPT the same question twice and roughly three quarters of the names come back. Ask Claude and roughly a quarter do.
That has a direct practical consequence. A single unsampled check is far more misleading on some engines than others, and a competitor who appears or vanishes on Claude between two monthly scans may be telling you nothing whatsoever. It is also why we sample rather than asking once, and why per-engine reporting beats a blended score.
What a single mention rate averages over
See which buyer you are visible to
Run a free scan across six AI engines, then pin your own buying questions so you can measure the buyers you actually want rather than a generic average.
Get your free auditWhat to do with this
The fix is not a more sophisticated score. It is refusing to average over buyers who were never the same population.
Name three personas from deals you have actually won
Not a marketing persona document. The size or maturity of company you sell to, the budget-led buyer who opens with price, and the buyer with a specific constraint such as an industry, a certification or a region. If you cannot name real deals in a segment, it does not get a persona, because each one multiplies what you have to measure.
Write each buying question three times, once per persona
The same question with the qualifier attached is enough. Best provider for a company our client's size. Most affordable provider. Best provider for a regulated industry. These are close to the sentences buyers actually type, and they are what produce the divergent lists in the first place.
Report a mention rate per persona and per engine, never blended
This is the whole point. A blended number moves for reasons you cannot act on. A per-persona number tells you which buyer you are losing, and a per-engine number tells you whether a change is real or is Claude being Claude. Both are only useful if the questions stay fixed between scans.
Aim at the persona you can win rather than the one you want
The neutral category question is the most contested and the least winnable, because it is the one every competitor is also optimising for. A qualified question with a specific buyer attached has a shorter list of firms genuinely positioned for it. That is the same pattern we keep finding: specificity is where a smaller name still gets named.
Say who you are right for, in writing, with specifics
This probe does not prove you can move the persona association, so treat it as a mechanism rather than a promise. But an engine describes you from what it can retrieve, and a site that never states client size, budget range, industry fit or constraints leaves it inferring your fit from whatever else it finds. Stating it plainly at least gives it something accurate to work from.
What this probe does not show
Two categories, four of our six engines, four framings, 48 calls in a single day. That is a probe, not a study, and I would rather state the limits than have you discover them. Each persona cell was measured once, so individual cells still carry the noise the control measured. The two categories were chosen because we know them well, and nothing here says the effect is the same size in yours.
It also does not show that you can change which persona an engine associates you with. That is the obvious next question and it needs a before-and-after on a real site, not a one-day probe. I have a hypothesis, stated above, and no measurement behind it yet.
What survives all of that is the direction and its size relative to the control. Persona overlap sat below the noise floor on all four of the engines probed, independently, and the gap between 0.55 and 0.25 is not a subtle one. G2's 2026 research found 51 percent of B2B buyers now begin vendor research on an AI chatbot, up from 29 percent, and Forrester's 2026 study found 94 percent use AI somewhere in the buying process. Those buyers are not one population, and they are not typing the same sentence.
Frequently asked questions
Does the buyer persona change which vendors AI recommends?
Substantially, and by more than run-to-run randomness explains. We asked four of the six engines we track (ChatGPT, Claude, Bing Copilot and Google AI Mode) the same buying question in two categories, framed four ways: neutral, as a fifteen-person startup, as a large enterprise, and asking for the most affordable option. We also asked the neutral version three separate times to measure how much the answer moves on its own. Re-asking the identical question returned a vendor list overlapping about 55 percent. Changing the framing dropped that to about 25 percent. Almost every firm the engines named was named for some buyers and not others.
How do you separate a persona effect from AI randomness?
By measuring the randomness first and using it as the baseline, which is the step usually missing. We ran the identical neutral question three times per engine per category and compared the vendor sets. That is the noise floor: a median overlap of about 0.55 on a Jaccard measure. Then we compared the neutral run against each persona framing, giving about 0.25. Persona framing roughly halved the overlap relative to simply asking again, and the direction held on all four engines individually, which is what makes the effect credible on a probe this small.
What does this do to my AI visibility score?
It makes a single blended number a persona-blind average of populations that barely share vendors. If your question set skews to one framing, your score describes your standing with that buyer and quietly says nothing about the others. A firm can be well positioned for enterprise buyers and absent for price-led ones, and one number will show a mediocre middle that matches neither. The useful reading is a mention rate per persona per engine, which also tells you which segment is winnable rather than only whether you are winning overall.
Which engine is most stable when you ask the same thing twice?
In this probe, ChatGPT and Google AI Mode were the steadiest, with same-question overlap around 0.74 and 0.67. Bing Copilot and Claude were far less stable at about 0.36 and 0.28. That matters practically: a single unsampled check is much more misleading on some engines than others, and a competitor appearing or vanishing on Claude between two scans may be telling you nothing at all. Every engine still showed persona overlap well below its own noise floor, so the effect was not an artifact of the unstable ones.
How many personas should I actually track?
Three is usually enough and it should map to how you already segment deals. A practical set for most B2B firms is the size or maturity you sell to, the budget-led buyer who leads with cost, and the buyer with a specific constraint such as compliance, an industry or a region. Add a persona only when you can name deals that fit it, because each one multiplies the number of questions you measure and the cost of measuring them. The point is coverage of the buyers you actually want, not completeness.
Can I influence which persona an engine associates with me?
This probe does not test that, so treat it as a mechanism rather than a proven result. What the engine can retrieve about you is the description it works from, so a site that never states client size, budget range, industries or constraints leaves the engine to infer your fit from whatever else it finds. Firms that state plainly who they are right for, with specifics, are giving an engine the material to match them to a qualified question. That is consistent with the rest of our data, where specificity is what separates a named firm from an unnamed one.
How reliable is a probe of this size?
It is a small controlled probe and should be read as one. Two categories, four of the six engines we track, four framings, 48 calls, 197 firm mentions across 97 distinct firms, run on 11 August 2026. Each persona cell was measured once, so individual cells still carry the noise the control measured. What supports the headline is that the gap between 0.55 and 0.25 is large, and that every engine showed persona overlap below its own noise floor independently. It does not establish the size of the effect precisely, and it does not generalise beyond these two categories.
Sources and further reading
- TofuBofu persona probe, 11 August 2026: 2 categories, 4 engines (ChatGPT, Claude, Bing Copilot, Google AI Mode), 4 buyer framings, 48 calls, 197 firm mentions across 97 distinct firms, with a three-run same-question control per engine per category.
- Engine disagreement across five US metros: 85 percent of firms named by exactly one of four engines, none by all four.
- The blended AI visibility score trap: why one number across engines hides the movement you need to act on.
- G2 2026 B2B buyer research: 51 percent begin vendor research on an AI chatbot, up from 29 percent, and 69 percent switched vendor based on AI.
- Forrester B2B Buying Study 2026: 94 percent of buyers use AI somewhere in the buying process.