Measurement
Where do the questions in your AI visibility report come from?
By Arnav Mukherjee, founder of TofuBofu · August 10, 2026
I was on a demo in July with the marketing team of a large consumer brand, walking through a report our product had produced for them. It was going fine. Then one of the brand managers stopped me on the page listing the buying questions we had scanned, and asked the sharpest question anyone has asked me about this product.
"How do you know these are the right questions?"
I gave a decent verbal answer about deriving them from their category and their site. It was true and it was not an answer. The honest version is that a language model had read their website and written the questions, and there was nothing behind them except that model's judgement. So a few weeks later we built a test designed to tell us whether that judgement was any good, and the test came back worse than I expected.
The test, and why it was built to be able to say no
The idea we wanted to try was simple: before writing a brand's buying questions, go and look up how people actually phrase things in that category, and use those real phrasings to seed the writing. Cheap to build, easy to sell, obviously sensible. Which is exactly the kind of idea that should be tested before it is built, because obviously sensible ideas are the ones nobody checks.
So we wrote the falsification first. Take the 10 most recent completed scans in the database. For each one, reconstruct what that scan knew about the brand, run the demand lookup it would have run, and count how many of the observed phrasings the questions we already generated had covered. If the answer was "most of them", the model was already finding what buyers ask, the feature was unnecessary, and we would have learned that for the price of about thirty searches.
The 10 scans were a real spread rather than a curated set: four managed IT providers, an IT services group, a kitchen and bathroom remodeler in Pennsylvania, an HVAC firm in Pittsburgh, a B2B SaaS in automated QA testing, and two Indian packaged food brands. Between them the generated sets contained 122 buying questions. The lookups returned 90 observed phrasings.
Replay against 10 stored scans, 7 August 2026
Zero. Not a low overlap, not a partial match on a few. In ten scans across seven categories, not one phrasing that a real person had typed into Google appeared in the set of questions we had written for that brand. When a result comes back that clean, my first assumption is that I have measured something other than what I meant to measure. That assumption turned out to be half right, and the half that was wrong is the part worth reading.
What a 100 percent miss rate actually means
Put the two lists side by side for one brand and the reason jumps out. This is a national managed IT provider. On the left, what our planner wrote from their website. On the right, what people actually search in that category.
These are not the same kind of object. The left column is a set of qualified buying questions, each one carrying a constraint the model inferred from the company's own pages: healthcare, multi-site, cloud optimization, enterprise application management. The right column is what people type when they are still assembling a list, and it is blunt. Who are the top providers. What should this cost. Which one is best. Give me the list.
The measurement backs that up. Across the whole sample the generated questions averaged 7.8 words, the observed phrasings 5.5. Seventy five percent of the generated questions contained a choosing word such as best, top or which. Only 56 percent of the observed ones did, because a lot of observed demand is not a question at all: "managed service provider companies", "list of MSP companies in USA", "top MSPs".
So the honest reading of the zero is not that our questions were wrong. It is that they were unfalsifiable. The model had written a plausible set, we had shipped it in a report to a buyer, and there was no line in the report, and no artifact behind the report, connecting any single question to a human being who had ever phrased it that way. That is a provenance problem, not an accuracy problem, and it happens to be the exact problem the brand manager put her finger on in about four seconds.
Two registers, one buyer, and the gap in the middle
The eight questions we had genuinely missed
A miss rate of 100 percent is a headline, not a decision. The decision needed a second and much stricter count: of those 90 phrasings, how many were real buying questions in the brand's own category and geography that the generated set had failed to ask? Not fragments, not competitor brand names, not other people's cities. Questions we should have been embarrassed to miss.
Eight of the ten scans had at least one. The remodeler's set, written entirely around Glassport and Pittsburgh in specialist language, had never asked the plainest commercial question in the trade: what does this cost. The HVAC firm's set had eight variants of Pittsburgh service questions and had missed the People Also Ask question sitting on Google for its own metro, phrased as a buyer would phrase it: which company offers the best heating and cooling services in Pittsburgh. Three of the four managed IT sets missed "top 10 best managed it services", which is the shape of the thing a buyer actually wants when they open a chat window.
One genuine miss per scan is a modest result compared to ninety. It is also the result that justified building the thing, and reporting it as ninety would have been the dishonest version. Eight real misses out of ten scans meant the lookup was worth the search credits. Ninety was mostly two vocabularies passing each other.
Why nobody can hand you the real prompt list
There is a tempting fix on the market for all of this, which is to buy a tool with a prompt database. Ahrefs sells Brand Radar against a database of hundreds of millions of prompts, and it is a genuine asset that we cannot match on scale. But it is worth being precise about what any such database can be, because the precision is where the buying decision lives.
Nobody has a feed of what people type into ChatGPT. The prompts are private. What vendors have is a proxy: clickstream panels, search data, submitted prompts, inference from behaviour. Every one of those proxies is a sample of a different surface than the one you care about, in exactly the way our 90 phrasings were a sample of the search box rather than of an assistant. That is not a reason to dismiss the data. It is a reason to ask which surface it came from, and to be suspicious of any tool that will not say.
The same applies to our own lookup, which is why we describe it the way we do. It tells us a human phrases this need this way. It does not tell us how many people do. We checked whether we could get volume from the search infrastructure we already pay for and the answer was no: the trends endpoint returns a relative index rather than counts, and the autocomplete endpoint returns an ordinal score that is not comparable across different seed queries. So we anchor to phrasing and we say phrasing. A claim about volume would need a second vendor and a second bill, and until we pay for one we should not imply it.
One detail from the data is worth pulling out on its own, because it is directly actionable. Two of the managed IT lookups surfaced "best managed it services reddit" without anything in our seeds pointing that way. Buyers are appending reddit to vendor searches because they want an opinion nobody paid for. If your firm has no presence in those threads, you are absent from both halves of one behaviour.
See your own questions, and where each one came from
A free scan runs your buying questions across all six engines and shows which ones are anchored to phrasings people actually search.
Get your free auditThe limits of this test, stated plainly
Ten scans is a small sample and I am not presenting it as a study of the category. It is one run, on one day, against one demand source, with at most three seed lookups per scan. Two of those seed lookups failed outright and returned nothing, though every scan stayed measurable through its other seeds.
There is a specific weakness worth naming because it shaped the result. The lookups were seeded from each brand's category, not from its city, so a remodeler in western Pennsylvania got back phrasings about Tampa and Orlando, and a Pittsburgh HVAC firm got "best HVAC companies in Houston". Those were correctly counted as missed and correctly excluded from the strict count, but they inflate the headline number. A geography-aware seed would produce a lower miss rate and a more useful one.
And the deepest limit is the one in the diagram above. Observed search demand is the closest public evidence we can get to how buyers phrase a need, and it is still not the surface where the answer gets given. Anyone who tells you they have solved that, including us, is selling you a proxy. The difference between vendors is whether they say which proxy.
What we changed, and what you should ask your own vendor
Four things came out of this, and the fourth is the one that applies to you whichever tool you use.
Observed phrasings now seed the writing, they do not replace it
Before the questions are written, the scan looks up how people phrase things in that category and hands those phrasings to the planner as raw material. The generated set is still the scored set. We deliberately did not swap brand-specific questions for generic category fragments, because that would move a customer's score for a reason they cannot act on, and a score that moves for methodology reasons is worse than no score.
At most one genuinely missing observed question gets added per scan
Inside the existing question budget, so nobody's plan quietly changes. It has to clear a gauntlet: a real choosing word, the brand's own category, the brand's own geography, not a competitor's brand name, and not already covered. That is the gate that turned 90 into 8, and it runs on every new scan rather than only in the replay.
Anchored questions carry their source phrase, and unanchored ones carry nothing
No badge of shame on the questions a model wrote to cover a gap, because covering a gap is legitimate work. Search demand does not record the buying situation of a 40-person firm whose IT lead just resigned, and that buyer is still going to ask an engine. What changed is that where a question does have a human phrasing behind it, you can see the phrasing.
Ask any vendor where the questions came from
Including us. It is one sentence and the answer separates a measurement product from a plausible artifact. Then ask what happens to those questions next month, because a tool that regenerates the set on every scan cannot show you a trend, it can only show you a new test with a new score. We reuse the previous scan's questions on repeat scans for exactly that reason, and we only top up when the budget grows.
Question provenance is one item on a longer list of things in a report you can check rather than trust. The rest of that list, eleven checks with a pass and fail condition on each, is in how to check whether your AI visibility report is telling you the truth.
This matters more than it would have two years ago because of where the decision now starts. G2's 2026 buyer research found 51 percent of B2B buyers begin vendor research on an AI chatbot, up from 29 percent, and 69 percent changed their vendor choice based on what an AI told them. If a report claims to tell you how you appear in that moment, the questions it asked are the entire experiment. A wrong question does not produce a wrong answer. It produces a confident answer to something nobody asked, which is harder to spot and more expensive to act on.
None of this changes the underlying position, which is that being findable on Google is the floor and being named by an engine is a separate layer you can win even when your rank does not move. It sharpens the measurement of that layer. You cannot claim to measure whether buyers can find you if you cannot say where the buyer's question came from.
Frequently asked questions
Where do the questions in an AI visibility report come from?
In most tools, one of three places, and the difference matters. Either you type them in yourself, or the tool writes them for you from your website using a language model, or the tool draws them from a database of prompts it has collected. Ours were written from the website by a model, which is the most common approach and the one with the weakest provenance story, because nothing in it is evidence that a human ever phrased a question that way. Ask any vendor which of the three they do. It is a fair question and it has a short answer.
Are AI-generated prompts in visibility tools accurate?
Accurate is the wrong test, because a generated question is not a claim about the world, it is a probe. The right test is whether it is the question your buyer would ask. We ran that test on 10 stored scans by pulling the phrasings people actually search in each brand's category and checking them against the questions our planner had written. Across 90 observed phrasings, the generated sets covered none of them. That does not mean the generated questions were wrong. It means they had no independent corroboration, which is a different and more fixable problem.
Why did zero out of 90 real search phrasings match the generated questions?
Mostly because the two are written in different registers. In our sample the generated questions averaged 7.8 words and 75 percent contained a choosing word such as best, top or which. The observed phrasings averaged 5.5 words and only 56 percent contained one. Search demand skews to short, list-shaped fragments like 'list of MSP companies' or 'top MSPs'. A model writing from a company website skews to long qualified questions like 'best remote IT support services for multi-site operations'. Both are things buyers want. Only one of them had evidence attached.
Can Google search data validate the prompts people type into ChatGPT?
Only partly, and it is important to be honest about the limit. Google's People Also Ask and related searches record how people phrase things in a search box. An AI prompt is a different surface with a different grammar: longer, conversational, carrying context about company size, industry and the situation that triggered the search. Observed search demand is real evidence that a human phrases a need this way, which is more than a generated question has on its own. It is not proof that the same wording is what anyone types into an assistant. Any vendor selling a prompt database should be asked which of the two surfaces it actually sampled.
How many questions should an AI visibility scan track?
Between 10 and 25 is enough for a focused firm, and depth beats breadth because each question needs sampling across engines to survive run-to-run variance. The count is not the interesting number anyway. A hundred questions with no stated origin is a weaker artifact than fifteen where you can see which ones came from observed demand and which were written to cover a gap. If a report cannot tell you which is which, its question count is decoration.
Why do buyers add 'reddit' to vendor searches?
Because they are looking for an opinion that nobody paid for. Two of the managed IT scans in our sample surfaced 'best managed it services reddit' as an observed phrasing, unprompted by anything we fed the lookup. If buyers append reddit to the query and engines lean on Reddit for corroboration, an absent brand is missing from both halves of the same behaviour.
What should I ask a vendor about the prompts in their report?
Four questions, and none of them require technical knowledge. Where did each question come from, and can you show me the source. How many of them are anchored to something observed rather than written. What happens to the questions when I scan again next month, because a set that regenerates each time destroys your trend line. And which questions were excluded and why. A vendor who can answer all four is measuring. A vendor who cannot is generating a plausible artifact, which is exactly what we were doing until we tested it.
Sources and further reading
- TofuBofu, AI visibility for managed IT providers, 2026: the first-party research programme this measurement discipline came out of, with full method and per-region data.
- G2 2026 B2B buyer research: 51 percent begin vendor research on an AI chatbot, up from 29 percent, and 69 percent switched vendor based on AI.
- Ahrefs Brand Radar: the prompt-database approach, and the published scale claim discussed above.
- Google Keyword Planner documentation: what search-box volume data does and does not report, which is the reason a conversational prompt shows as zero.