Measurement
Why an AI engine names you on one question and skips you on the next
By Arnav Mukherjee, founder of TofuBofu · September 3, 2026
An engineering consultancy we scan, selling design and compliance work into two regions, reads zero on every scan we hold of them. Not low. Not slipping. Zero, on every buying question, on every engine, every time we've run it. A firm with real offices, real clients and a real website, and our own product has never once found their name inside an answer.
So in the middle of August I stopped asking the question our report was built on and asked a duller one. Not whether the engines named them. Whether the engine had gone and looked. Five buying questions written for the two markets this firm actually sells into, put through the ChatGPT API twice with no search tool attached, then twice more with web search forced. Twenty calls, one firm, one afternoon.
From memory it searched on none of the ten calls and named them on none of the ten. Those two zeros are one fact wearing two hats. Then we forced it to look, and the answer changed in a way that turns out to be a content instruction rather than a bug report.
An engine that isn't retrieving can't find you at all
OpenAI puts the switch in one sentence, and with the code formatting flattened it reads: "With tool_choice: auto, search is optional. Use tool_choice: required or a specific web search tool choice when search must run." Optional is carrying a lot of weight there. Our probe wasn't even on auto. It called the older completions endpoint with no tools attached, so the switch wasn't merely left unset, it didn't exist.
That half of the probe isn't a sample and it isn't subject to luck. A call carries a web search step or it doesn't, and one with no tools attached can't produce one however many times you run it. Zero searches, zero mentions, and the second follows from the first.
Here's why that lands on you rather than on us. Every fix this whole category sells, the comparison page, the schema block, the third-party listing, the review profile, arrives on the live web long after a model's training data stopped being collected. A model answering from its weights can't see any of it. So a firm can execute an AEO plan flawlessly and stay pinned at zero forever on that reading, while the report keeps telling them to publish harder. The advice and the instrument were pointed at two different worlds.
A memory answer doesn't look broken either, which is the reason nobody catches it. It arrives fluent and structured, full of real company names with plausible reasons attached to each one, and it reads exactly like an answer built from a live search. Before you argue about anyone's visibility number, including ours, make them tell you whether the call that produced it ran a search. The one-call test for that is published and it costs about a cent.
Forced to look, it named them three times. The split underneath is the finding.
With search forced, the engine searched on all ten calls and named the firm on three of them. Before anyone builds anything on that second figure, here is the sample it came from: one brand, five questions, two runs per question per mode, twenty calls in total, against an engine that returns different answers to identical prompts. Three of ten is a direction. It's not a rate, we won't publish it as one, and one of those three questions named the firm on the first run and skipped it on the second.
The split inside those ten is what's worth your afternoon.
Broad category questions named them on none of four attempts. Best providers of their service, asked once for each region they sell into. The question their entire website is written to answer, the one their homepage headline is a direct reply to. Retrieval on, engine out reading the live web, and they didn't appear.
Questions carrying their own differentiator named them on three of four. Every naming in the entire probe came from a question containing the specific thing this firm says it is, and more exactly the thing it says it isn't: independent of the contractors who normally bundle this work into a larger deal. That clause is their positioning. It's the sentence a founder writes on the homepage and then privately wonders whether a single human being reads.
The mechanism isn't mysterious once you say it out loud. A broad category question has thousands of pages competing to answer it, written by firms with more reviews, more press and more domain authority than you, and a retrieval engine picks a handful of them. You lose that auction on the day you enter it, and you lose it again next quarter. A question carrying a specific constraint has almost nobody competing to answer it, because almost nobody has bothered to write the sentence that answers it. If you're the firm that says out loud what you aren't, you're the only match on the page.
Twenty calls on one brand can't carry a rule for everybody, and I'm not asking it to. What it can do is name a mechanism you can go and falsify on your own domain this week, which is more than a leaderboard will ever give you.
Two of the three namings didn't come from their website
Retrieved answers arrive with their sources attached, which is the first time this engine has ever told us where a name came from. The ten retrieved answers cited 53 URLs between them. Thirty nine of those were Google search wrappers, links pointing at a search the model composed rather than at any publisher, and we throw them away instead of storing them, because recording a search engine as the citing source would poison every channel analysis built on top of it.
That leaves fourteen real citations across ten answers, which tells you how thin the evidence trail is even when an engine is trying. Of the three answers that named this firm, one cited their own domain. The other two reached them through Google local listings: an office address in one region, a maps link in the other.
Three data points is an anecdote and I won't dress it up as anything else. It's a checkable anecdote though, and the thing worth checking is uncomfortable. This firm thinks about its website constantly. Nobody there has looked at the local listing in a year, and on the day we ran this, the listing did two thirds of the work. Google publishes what drives those listings, and its own help page says local results rest mainly on "relevance, distance, and popularity", none of which your homepage copy controls.
So the asset carrying you into an AI answer may not be the asset you've been paying for. Worth ten minutes before your next content invoice.
Which questions are you invisible on?
Run a free scan and read the actual buying questions, the actual answers, and which engines named you on which ones. One scan a month, six engines, no card.
Run a free AI visibility scanRun this on your own firm before you believe me
Write the sentence. One line saying what you are and what you're explicitly not, the kind a competitor couldn't honestly copy. Independent of the contractors who normally bundle this. Audit-only, no implementation arm. Migrations exclusively, never greenfield builds. If your line survives being pasted onto a rival's homepage without anyone noticing, you haven't written it yet.
Ask two questions on the same day. The broad category question your homepage answers, then a question containing your sentence. Same engine, same session, ten minutes apart. If the broad one skips you and the narrow one names you, you've reproduced the split on your own domain and you now know where your next four articles go.
Search your own site for that sentence. Not the idea, the words. If it appears once inside a hero banner image and nowhere in a page title, a heading or a body paragraph, no retrieval engine can match a question against it, and you've been asking an engine to infer a claim you never made in text.
Most B2B sites fail the third step, and they fail it for a reason that sounds like good taste. Specific claims feel narrowing. A founder writes "independent" once, decides it sounds defensive, and spends the rest of the site on the same category language every competitor uses. Vague positioning has always been weak marketing. Under retrieval, it's an asset an engine can't index.
What this probe doesn't show
One firm, one engine, twenty calls, one afternoon. I'd be embarrassed to call that a study. Every number above is a count on a named sample, and the only one that doesn't depend on sampling luck is the search count, because a call with no tools attached can't produce a search.
There's no before and after here either, and there won't be one. Every completed scan of this firm reads zero, so there's no second number to line up against the first. We also moved our own instrument three times in the six weeks around these readings, which means any line drawn across those dates would measure us rather than them. We stamped those dates in code so no chart of ours can draw it, including one I asked for and was refused.
And the recommendations we produced for this firm are still open. Diagnosed, drafted, nothing published, which is where this loop breaks for almost everyone. So take the mechanism, not the outcome. The mechanism is the part you can act on without us: an engine that isn't retrieving can't find you at all, and once it is, it finds you where you're specific and loses you where you're broad.
Frequently asked questions
Why does an AI engine name my company on one question and skip me on the next?
Two separate things decide it, and they stack. First, whether the engine went and looked at the live web at all. An engine answering from training memory can't see anything you've published recently, so it can't name you for reasons that have nothing to do with your site. Second, once it is looking, question shape decides the result. In our probe of one firm, broad category questions named them on none of four attempts and questions carrying the firm's own differentiator named them on three of four. A broad category question has thousands of pages competing to answer it and the engine picks a handful. A question carrying a specific constraint has almost nobody competing, because almost nobody has written the sentence that answers it.
Does ChatGPT search the web before it answers a buying question?
In the consumer app it frequently does. Through the API it doesn't unless the caller asks, and OpenAI documents that as a setting: with the tool choice left on auto, search is optional, and you use a forced tool choice when search must run. Our probe wasn't even on auto. It called the completions endpoint with no tools attached, so it ran a web search on none of ten calls. Ask any visibility vendor whether their calls force retrieval, and ask them to prove it on the exact configuration they run your scans on.
Should I write for broad category questions or specific ones?
Write the specific ones first, and treat the category question as a long campaign rather than this quarter's work. The broad question is answered from whatever the web says about your category in general, which favours the firms with the most reviews, press and domain authority. You don't beat that by existing. The narrow question, the one that names a constraint or an exclusion you can honestly claim, has almost no competition and it is the one where a retrieval engine found the firm in our probe. Pages that state what you are, what you are explicitly not, and where you sit are the ones an engine can match against a specific ask.
Is three namings out of ten calls a real mention rate?
No, and we won't publish it as one. That figure comes from one brand, five questions, two runs per question in each mode, twenty calls in total, against an engine that returns different answers to identical prompts. One of the three questions named the firm on the first run and skipped it on the second. Treat three of ten as direction, never as a rate. The number in the same probe that does not depend on sampling is the search count: with no tool attached the engine searched on none of ten calls, and with search forced it searched on all ten. That half is a configuration fact rather than a measurement.
Why would a Google local listing get me named when my website doesn't?
Because a retrieval engine reaches your name through whatever the search index will hand it, and a business listing is structured, third-party and easy to match against a place-shaped question. Of the three answers that named the firm in our probe, one cited their own domain and two arrived through Google local listings, an office address in one region and a maps link in the other. Three data points is an anecdote rather than a rate. It is checkable on your own firm in ten minutes though, and Google publishes what drives those listings: local results rest mainly on relevance, distance and popularity. Most B2B firms think about their website constantly and have not looked at their listing in a year.
What is the fastest way to test this on my own firm?
Three steps, one afternoon. Write down the single sentence that says what you are and what you are explicitly not, the one a competitor could not honestly copy. Then ask an AI engine two questions on the same day: the broad category question your homepage is written to answer, and a question that contains your sentence. Read which one names you. Last, search your own site for that sentence. If it appears once in a hero banner and nowhere in a page title, a heading or a body paragraph, no engine can retrieve it, and the broad question was never going to carry you on its own.
Does this mean my visibility score is wrong?
It means you should ask what produced it before you act on it. A zero from an engine that never ran a search is a fact about the instrument, not about your firm, and no amount of publishing will move it. A zero from an engine that did search is real information, and the next question is which questions it was asked. A report showing you absent across a set of broad category questions is telling you something different from a report showing you absent on questions carrying your own positioning. The second is a much more serious finding, and most reports never separate the two.
Sources and further reading
- OpenAI, Web search guide. The vendor's own statement that search is optional under an auto tool choice and has to be forced when it must run. Read 3 September 2026.
- Google, Improve your local ranking on Google. Google's own statement that local results rest mainly on relevance, distance and popularity. Read 3 September 2026.
- Google Search Central, local business structured data. The markup that describes a physical location to a crawler, which is separate from the listing itself and worth having on both.
- TofuBofu first-party probe, 17 August 2026. Five buying questions for one engineering consultancy, each run twice with no tools attached and twice with web search forced, twenty calls in total on gpt-4o. Search behaviour, brand naming and every citation URL recorded per call, and every figure in this post recomputed from those twenty rows.
- TofuBofu production database, read 1 September 2026. Every completed scan of this firm, all of which read a visibility score of zero, and the recommendation set generated from the most recent one.