Measurement
Only four of every ten brands survive a second run. We weight those questions at 50%.
By Arnav Mukherjee, founder of TofuBofu · August 20, 2026
On 11 August I ran a probe to decide whether persona tracking was worth building. Ask a buying question as a fifteen-person startup, as a large enterprise, as a buyer who opens with price, and see whether the engine names different vendors. It did. The vendor lists held about a quarter of their names in common.
Then I ran the control, which is the only reason I did not ship a feature that reports noise as insight. Same questions, same engines, no persona at all, asked again three times. Those runs disagreed with each other too, just half as much: re-asking the identical question held about 55 percent of the names, reframing it held about 25. That gap was enough to keep the finding and the noise floor underneath it was more than enough to scare me. The feature survived on its control, and I published the control alongside the result rather than the headline on its own.
What that probe could not tell me is whether the wobble is spread evenly across the questions we ask. It is not. Someone has now measured it properly at scale, and the answer is worse than a general warning about non-determinism, because the instability concentrates exactly where the money is.
The wobble is old news. The shape of it is not.
That AI answers move on their own is not a revelation and we have written about it more than once. Our own control gave a median overlap of about 55 percent between identical runs, and the average was the least interesting number in it. Per engine, the noise floor ran from 0.74 on ChatGPT down to 0.28 on Claude. The same single-sample method is roughly three times noisier on one engine than on another, which means an unsampled check is not uniformly unreliable, it is unreliable in a pattern.
Ronald Sielinski's paper on measurement uncertainty, submitted to arXiv on 9 March 2026 and revised on 9 June, is the first serious statistical treatment of this I have read. Its argument is one sentence long and it indicts most of the category: citation visibility metrics "should be treated as sample estimators of an underlying response distribution rather than fixed values". He samples Perplexity Search, OpenAI SearchGPT and Google Gemini across three consumer product topics, daily over nine days and again at ten-minute intervals, then runs bootstrap confidence intervals over the results. The finding: "many apparent differences between domains fall within the noise floor of the measurement process", and single-run metrics give "a misleadingly precise picture".
Two things about that paper deserve to be stated rather than borrowed. It measures which domains get cited, not which brands get named, and those are related but separate questions. And three consumer product topics is a narrow base, which the paper does not hide. What it establishes is not a rate you can quote at your board, it is a method: report the interval, not the point.
The quiet line in it is the one that closes the most popular escape hatch. Rank instability shows up "not only among top-ranked domains but throughout the frequently cited domain set". The comfortable response to variance is to say the leaders are stable and only the tail churns. Measured, they are not.
Conductor put 14,000 calls through it. Purchase intent came last on the axis that decides who you see.
Conductor ran 14,000 API calls: ten industries, seven intent types, four engines, fifty runs per cell, each run stateless with no conversation history so every call arrives as a stranger. The four engines were ChatGPT, Perplexity, Claude and Gemini. Then they asked how much of each answer survives a repeat, using pairwise Jaccard similarity across all 1,225 possible run pairs per cell. Their definition, in their words, is "the percentage of brands shared between any two runs. A brand overlap of 63% means that out of every 10 unique brands that appeared across any two runs combined, six appeared in both."
Purchase prompts came out at 40 percent brand overlap, the worst of the seven on that axis, with 59 percent lead-brand stability, which is mid-table: education prompts hold their lead brand only 30 percent of the time. Out of every ten brands appearing across two runs of the same purchase question, four showed up in both. Comparison prompts came out at 63 percent overlap and 91 percent lead-brand stability, the best of the seven on both counts.
Most write-ups of this study report the blue bars and stop. The amber bars are where it gets interesting, because the two axes do not rank together. Education prompts hold the second-highest brand overlap at 60 percent and the lowest lead-brand stability of all seven at 30 percent. The list of names repeats, and the name at the top of it changes in seven runs out of ten. Pricing prompts sit at 58 percent overlap with 59 percent lead stability, the same lead stability as purchase intent despite a much steadier list.
Our read on the education result, and it is a read rather than something Conductor asserts: it is downstream of those prompts mostly declining to name anyone. Conductor found that "In 45% to 72% of Education prompts, depending on the LLM, [AI] responded without naming a single brand", and 20 percent of pricing prompts returned no brands either. When only two or three names surface at all, whichever one lands first is close to a coin toss. High overlap over a tiny set is not stability, it is a small sample wearing a large number.
Read the two ends of the chart together, because the gap between them is the useful part. The question a buyer asks when they are ready to choose is the one engines answer least consistently, on both measures at once. The question they ask when they are weighing two names they already have is the one engines answer most consistently, on both measures at once.
That is not a quirk, it is the mechanism. A comparison prompt hands the engine its own answer set, so there is little left to vary. A purchase prompt hands it an open field and asks it to pick, and an open field is where a non-deterministic system has room to move. The more freedom you give the engine, the more the answer is a sample rather than a fact.
This argues against our own scoring, so let me argue it
Our scoring weights decision-stage questions at 0.5, consideration at 0.3 and awareness at 0.2. Half the score rides on the intent type Conductor measured as the least reproducible of the seven. Anyone reading this in good faith should push on that, so here is the defence and the concession, separately.
The defence. Stability and importance are different properties, and a metric that chases the first ends up measuring the wrong thing to three decimal places. The way to see it is to follow the opposite rule all the way down. If reproducibility set the weights, the highest-weighted question would be the one carrying your own company name, because every engine answers those with your company 82 to 100 percent of the time. The name is sitting in the question. It is the most stable measurement available and it tells you nothing at all, since the engine is reading your name back to you rather than choosing you.
I am not arguing that hypothetically. We shipped it. Our published engine rates were once computed across all intents, brand-name probes included, and they read about 10 percent for the field and 19 percent for the leader. Strip the probes out and the leader is 7.4 percent. The same corpus, the same engines, the same day, inflated roughly two and a half times by a denominator that had quietly filled up with the most reproducible and least meaningful question we ask. That is what optimising for a steady number looks like from the inside, and every tool reporting one flattering aggregate is somewhere on that road.
The concession. A noisy high-value signal earns more samples, not the same single sample as everything else. Our free tier samples each question once, which tells you whether you exist on that question at all and is not enough to date a trend. Paid sampling is adaptive: two runs per question, and a third only when the first two disagree, which spends the extra calls on precisely the cells where the wobble surfaced. That design was right for the general case, and Conductor's result says it should be pointed harder at buying questions specifically, because that is where the variance concentrates. Comparison questions at 91 percent lead-brand stability do not need the third call. Spending the same depth on both is not neutrality, it is buying precision where we already had it.
And the part I would argue with our own code about. Our alert logic calls a change worth emailing when the score moves five points or more, which is a real threshold and I stand by it. It also flags any single query that flipped between named and absent. At one or two samples a question, a single flip is exactly what noise looks like. That rule is defensible as an alert, since a lost query is a specific thing you can go and look at, and it is indefensible as evidence that something changed. Those two readings should not share one word.
See the per-question spread, not just the headline number
A free scan reports every engine separately and shows you the full verbatim answer behind each result, so you can check the reading by hand instead of trusting the average.
Run your free scanIntent is one axis. There are two more, and they multiply.
Run-to-run variance within an engine is the first axis. Engine-to-engine disagreement is the second, and it is larger. Fractl published an AI visibility index in August, built from 96 sector prompts run 15 times each across GPT-4o, Gemini 2.5 Flash and Claude Sonnet 4.6, producing 4,320 responses and more than 8,500 distinct brands. Only 900 of those brands, 11 percent, were named by all three models. Seventy-seven percent, 6,264 brands, were named by exactly one.
We found the same shape in a different vertical, and we found it first. In five legal metros across four engines we logged 136 firm-and-market entries, of which 116 were named by exactly one engine and not a single one by all four. Different denominators, different professions, same conclusion: for most brands, being named is a property of one engine rather than a property of the market. We wrote that up on 5 August.
The third axis is the surface. Ten identical API calls to ChatGPT on one fixed buying question named 61 different companies with none appearing in all ten; Claude named 122, also none in all ten. That test is published in full, and it is worth noting that its mean overlap figures sit well below the 55 percent noise floor from the persona control. Different prompts, different months, mean against median: the two probes bound the wobble, they do not pin it. Which is Sielinski's whole point restated in our own data. A variance number without its method attached is not a number.
The objection to our own positioning, which the same study raises
Fractl did one thing we cannot replicate, because it needs a backlink index we do not run. They mapped every brand against Ahrefs authority data. Five percent of brands, 471 of them, were underrepresented: strong traditional SEO metrics, minimal mentions from the models. Four percent, 377 brands, were the reverse, modest SEO and high model recall.
Here is the number a careful reader will hit me with, so I will put it in myself: 91 percent of brands were aligned. Authority and AI mentions mostly track each other. Taken alone that reads as an argument that AI visibility is a repackaging of SEO, which is the objection our whole category deserves and mostly ducks.
It does not overturn the positioning, and the reason is what "aligned" measures. It is a correlation between two rankings, not evidence that domain rating causes a citation. Both are downstream of a company being real, established, and written about by other people. A brand with traffic, links and coverage tends to appear in AI answers for the same reason it appears in search results, and neither one is producing the other.
Which is exactly what a floor looks like. Necessary, mostly already satisfied, and not sufficient. Our line has never been that SEO does not matter; it is that SEO is the floor and AI visibility is a separate layer on top of it. The 91 percent is the floor doing its job. The 5 percent is the group for whom the floor is met and the layer above is not, and no additional domain rating will move them, because the gap they have is not an authority gap. If you are in that 5 percent, the 91 percent is not consolation, it is the proof that your problem is somewhere your SEO tooling cannot see.
What a visibility number has to carry before you act on it
Four things, and you can demand all four from any vendor including us.
1. The question set, by intent. A score computed over questions containing your own company name is not the same measurement as a score computed over questions a stranger would ask, and the first runs far higher because the answer is sitting in the question. We got this wrong ourselves and shipped inflated rates on eight engine guides before catching it, which is why the correction is now published on the pages that carry the numbers. Ask which intents are in the denominator.
2. The sample count per question. One run per question is a presence check, and it is a perfectly honest thing to sell as long as it is not dressed as a trend line. If a vendor cannot tell you how many times each question was asked, they are reporting a single draw from a distribution with a 40 percent floor on the questions that matter most.
3. Per-engine results, never a blend. Averaging six engines into one figure hides the only fact you can act on, which is which engine is failing you. It also averages together engines with noise floors of 0.74 and 0.28, which are not the same instrument. A brand named on one engine and invisible on five can post a comfortable-looking composite, and the average is where that gets buried.
4. A stated threshold for movement. Below some size, the correct report is no change detected. A vendor who draws a confident line through every monthly reading is selling you a picture of their own sampling error, and it will look like progress roughly half the time.
None of this makes the measurement worthless, and the fashionable conclusion that AI visibility cannot be measured is lazy. Zero is stable. A brand that is absent on every buying question, on every engine, across repeated samples, is not the victim of variance. Across 46 reports and 2,421 answered buying cells, our best-performing engine named the scanned brand on 7.4 percent of buying questions and every other engine sat lower, on a corpus that is not a random sample because nearly every firm in it ran a scan already suspecting it was missing. Wide error bars do not rescue a company that never appears in any draw. They tell you that the difference between 12 percent and 15 percent is not a story, and that the difference between 0 percent and 12 percent is.
Frequently asked questions
Are AI visibility scores accurate?
They are estimates, and most are reported as though they were readings. Ask an engine the same buying question twice and the list of companies it names changes substantially. In our own control probe, run on 11 August 2026 across two categories and four of the six engines we track, re-asking the identical question returned a vendor list overlapping about 55 percent by median Jaccard similarity. That average hides the more useful result: the per-engine noise floor ran from 0.74 on ChatGPT down to 0.28 on Claude, so the same single-sample method is roughly three times noisier on one engine than another. A single-sample score is one draw from that spread. It is not wrong, it is imprecise, and the imprecision is almost never printed next to the number.
What is a good AI visibility score?
The question is malformed without two extra pieces of information, and any vendor who answers it with a single benchmark number is guessing. First, which questions were asked, because a score computed over questions containing your own company name runs several times higher than the same score computed over questions a stranger would ask. Second, how many samples produced it, because a score built from one run of each question carries a margin of error wide enough to swallow most month-to-month movement. Get those two facts, then judge the number. Without them there is no comparable scale to be good on.
Why does AI give different answers to the same question?
Three reasons stack. The models are non-deterministic by design, so identical inputs can produce different outputs. Engines that retrieve live pages may pull a slightly different set each time, which changes the supporting cast in the answer. And the models themselves are updated, sometimes without announcement. Ronald Sielinski's arXiv paper on measurement uncertainty puts the consequence plainly: citation distributions follow a power-law form and exhibit substantial variability across repeated samples, and bootstrap confidence intervals show that many apparent differences between domains fall within the noise floor of the measurement process.
Which kinds of AI questions are the most reliable to track?
Comparison questions, and they win on both axes. Conductor ran 14,000 API calls across ten industries, seven intent types and four engines at fifty runs per cell, and comparison prompts came out most stable at 63 percent brand overlap between runs and 91 percent lead-brand stability. Purchase prompts were the least stable at 40 percent overlap and 59 percent lead-brand stability. The two axes do not rank together, which is the part most summaries miss: education prompts held 60 percent brand overlap but only 30 percent lead-brand stability, the worst of the seven, so the list repeats while the name at the top of it does not. If you want a number that moves for real reasons, watch the comparison set, and treat the purchase set as the one that needs more samples before you read anything into it.
If buying questions are the least stable, why weight them the highest?
Because stability and importance are different properties, and optimising for the first would mean measuring the wrong thing precisely. Follow the opposite rule to its conclusion and you would weight hardest the questions that carry your own company name, which every engine answers with your company 82 to 100 percent of the time because the name is sitting in the question. That is a perfectly stable measurement of nothing. We know, because we shipped it: rates computed across all intents ran about 10 percent for the field and 19 percent for the leader, against 7.4 percent for the leader once brand-name probes came out of the denominator. Our scoring weights decision questions at 0.5, consideration at 0.3 and awareness at 0.2, and that stays. The correct response to a noisy high-value signal is more samples on that signal and a stated margin, not a quieter signal with a cleaner chart.
How many samples does an AI visibility measurement need?
More than one, and more on buying questions than on comparison questions. Sielinski's paper argues visibility metrics should be treated as sample estimators of an underlying response distribution rather than fixed values, and provides practical guidance on the sample sizes required for interpretable confidence intervals. In practice, the free tier of most tools including ours samples each question once, which is enough to establish presence or absence at all and not enough to date a trend. Paid sampling in our product is adaptive: two samples per question, and a third only when the first two disagree, which concentrates the extra calls on exactly the cells where the wobble showed up. Conductor's result says that depth should be pointed at buying questions specifically rather than spread evenly, because comparison intent at 91 percent lead-brand stability does not need the third call.
My score moved three points this month. Is that real?
Probably not, on its own. A three-point move on a single-sampled score sits inside the noise for most brands, and the honest read is no change detected rather than a small decline. What is worth reading is a specific query flipping from named to absent and staying flipped across two consecutive scans, or a move large enough to clear a stated threshold. Our own alert logic treats a score change of five points or more as significant, and separately flags any single query that flipped, which is the part I would argue about: at one or two samples a question, one query flipping is exactly what noise looks like.
Sources and further reading
- Conductor, AI Brand Recommendation Study: Why Intent Type Predicts AI Output Consistency. 14,000 API calls across 10 industries, 7 intent types and 4 engines at 50 stateless runs per cell, published 17 June 2026 and last updated 17 July 2026. Overlap computed as pairwise Jaccard similarity across all 1,225 run pairs per cell. Source of every intent-level figure in this post.
- Ronald Sielinski, Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement, arXiv:2603.08924. Submitted 9 March 2026, revised 9 June 2026. Perplexity Search, OpenAI SearchGPT and Google Gemini across three consumer product topics, sampled daily over nine days and at ten-minute intervals, with bootstrap confidence intervals. The statistical case for reporting visibility with uncertainty estimates.
- Fractl, AI Visibility Index. 96 sector prompts run 15 times each across GPT-4o, Gemini 2.5 Flash and Claude Sonnet 4.6, 4,320 responses, 8,500+ brands, with an Ahrefs authority cross-map. Published August 2026.
- Our persona probe and its control, 11 August 2026. Source of the 55 percent noise floor and the per-engine spread from 0.74 to 0.28.
- Ten identical API calls per engine. Our first-party test of how far a vendor list moves when nothing at all changes.