Measurement
We forced ChatGPT to search the web. The category question still returned nothing.
By Arnav Mukherjee, founder of TofuBofu · August 18, 2026
In the middle of August, a customer of ours, a B2B engineering consultancy selling into two regions, won a serious inbound enquiry from a buyer who told them, unprompted, that he had found them through ChatGPT. That same week, on that same engine, our own report said what it had said on every run since we started scanning them. Absent. Not low, not slipping. Absent, on every buying question, every time.
Both of those things were true, and that combination should worry anyone paying for a visibility number. So before writing another report telling this customer what to fix, I checked something we should have checked a year earlier. Not whether ChatGPT named them. Whether ChatGPT had gone and looked.
It had not. Ten calls, five buying questions written for the two markets this firm actually sells into, put through the exact code path production was running. Zero web searches. Zero mentions. Those two zeroes are one fact wearing two hats, and the second is a consequence of the first, which means the number sitting in our own report was never a reading of that engine's opinion. It was a reading of what a model happened to remember.
The API does not search unless you make it
There are two ChatGPTs, and the distance between them decides whether AEO work is measurable at all. There is the consumer app your buyer opens, which decides for itself when to go and look something up. And there is the API, which is where a scanner talks to it. Every tool in this category picks one of those surfaces, and almost none of them tell you which, a choice we have argued through in public before.
On the API, retrieval is a switch you throw. OpenAI's own web search guide puts it in one line, and with the code formatting flattened it reads: "With tool_choice: auto, search is optional. Use tool_choice: required or a specific web search tool choice when search must run." Optional is carrying a great deal of weight in that sentence. In an earlier probe of ours, left on auto, the model chose to search on one run out of three. Same question, same configuration, minutes apart, reaching a retrieval engine once and a memory engine twice, with nothing in the answer text to tell the two apart.
Our own probe was not even on auto. It called the older completions endpoint with no tools attached at all, so the switch was not merely left unset. It did not exist. That is the shape of the simplest integration anyone would write against this API, which is exactly why it is worth saying out loud rather than quietly correcting.
Here is why that mistake survives contact with a working product. A memory answer does not look broken. It arrives fluent, structured, full of real company names with plausible reasons attached to each one. It is indistinguishable by inspection from an answer built on a live search, and it is equally indistinguishable to the customer reading the report you built on top of it.
I wanted the second reason to be money, and our own numbers do not support it. At OpenAI's published rates, read on 17 August and checked again on 20 August 2026, the web search tool costs $10.00 per 1,000 calls on top of tokens, with gpt-4o at $2.50 per million input tokens and $10.00 per million output. Against our own median token counts that is $0.0036 per call without search and $0.0149 with it, about 4.1 times. The multiple looks alarming and the absolute number is trivial. A free scan's ChatGPT calls cost about seven cents with search forced, against under two cents without. A whole month of them for a customer on our $99 plan costs us under $1.80, which is under two percent of what that customer pays. Nobody skipped this to protect a margin. The cheap path and the default path were the same path, and it returns confident prose either way. Defaults do not require a decision, which is precisely why they survive audits.
The test costs one call and you should run it against anything you pay for. Ask the engine what today's date is, and to cite the source it used. From memory, ours answered word for word: "I'm unable to provide real-time information, including today's date, as I don't have access to current data or the internet. My responses are based on the information available up to October 2023." With search forced, the same call returned "Aug 17, 2026, 12:57:36 AM". Parse the date, hold it against your own clock, and there is nothing left to argue about.
We tried a better-sounding control first and it failed, which is worth knowing before you design your own. Asking for a news story published in the last seven days, with a URL and a publication date, sounds like a stronger freshness test. It false-alarms. Even with search forced, the model replied that it had no access to specific news articles. A control that reports a healthy instrument as broken gets switched off within a fortnight, and then you have no control at all. The date probe survives because its answer is checkable against a clock you already own.
Now put that memory answer next to what this whole category sells. We tell firms to publish comparison pages, fix their schema, earn third-party listings. Not one of those things can enter a model's training weights. A customer could execute an AEO plan flawlessly and stay permanently invisible on that instrument, and the report would keep telling them to try harder. The advice and the measurement were pointed at two different worlds.
What forcing it changed, and what it did not
Two numbers came out of that probe and they do not deserve equal trust. Running them together is how a twenty-call sample gets promoted into a claim, so I am going to separate them.
Whether a search ran: zero of ten before, ten of ten after. That is barely a measurement. It is a configuration fact. A call carries a web search step or it does not, our code now asserts the presence of that step on every single call, and a call with no tools attached cannot produce one however many times you run it. This is the number the argument rests on, and it is not subject to sampling.
Whether the engine named the brand: zero of ten before, three of ten after. That one is a sample, drawn from a non-deterministic engine, on one company, on five questions run twice each. One of those questions named them on one run and not the other. Three of ten is a direction. It is not a rate, and I am not publishing it as one.
The split inside those ten is the finding, and it is the reason this post exists rather than a changelog entry.
The broad category questions named them zero of four. Best providers of their service, asked once for each of their two regions. The question their entire website is written to answer. Retrieval on, the engine out reading the live web, and they did not appear.
Questions carrying their own differentiator named them three of four. Every naming in the whole probe came from a question containing the specific thing this firm says it is, and more exactly the thing it says it is not: independent of the contractors who normally bundle this service. That clause is their positioning. It is the sentence a founder writes on the homepage and then privately wonders whether anybody reads.
That accounts for eight of the ten, and the last two are the sharpest thing in the set. The fifth question also used the word independent, and it named them zero of two. What separates it from the two that worked is where the geography sits. The winning questions put the supplier somewhere: a consultant based in these countries. The failing one put the customer somewhere: a firm serving these markets. Our read, and it is a read rather than a measurement, is that a retrieval engine matches a supplier-shaped phrase against supplier-shaped pages, and the market you serve is a claim almost every competitor also makes. A separate probe two days earlier landed the same way: a long buyer-style prompt placing the project in the region returned nothing, while a shorter one placing the consultant in the region returned this firm first.
Read all of that as a content instruction, because that is what it is. Turning retrieval on does not buy you the category. The category is crowded, it is answered from whatever the web says about that category in general, and a small independent firm does not win it by existing. What retrieval buys is the ability to be found on the language that is uniquely yours: what you are, what you are explicitly not, and where you sit. Vague positioning has always been weak marketing. It is now an unindexable asset.
Where it found them, and why we had to stop counting URLs
Retrieved answers arrive with their sources attached, which is the first time this engine has ever told us where a name came from. The ten retrieved answers cited 53 URLs between them. Thirty nine of those were Google search wrappers, links to a search the model composed rather than to any publisher, and we drop them rather than store them, because recording google.com as the citing source would be false and would quietly poison every channel analysis built on top of it.
Of the three runs that named this customer, one cited their own website. The other two reached them through Google local listings: an office address in one region, and a maps link in the other. Three data points is an anecdote and I will not dress it as anything more. But it is a checkable anecdote, and it points at something worth testing on your own firm. The asset carrying you into an AI answer may not be the asset you have been investing in. This company has a website they think about constantly and a local listing they almost certainly have not looked at in a year, and the listing did two thirds of the work.
Retrieval also broke our own counting, quietly, on the day we turned it on. It will break the same way in any tool that adds this, so it is worth walking through properly.
Before that change, no engine we query returned an answer with links in the body text. Retrieved ChatGPT answers do. So the customer's own domain now appears inside answers that never say their name in a sentence. Our verification is a literal name check against the answer text, which means that link would have satisfied the check whose entire job is to confirm the engine named them. A footnote would have been promoted to a recommendation. By us. Automatically. In our customer's favour, on the report they pay for.
It cuts the other way too, and that half is worse, because it corrupts a different surface. The same check decides which competitors an answer named, and that list feeds a leaderboard the customer reads to work out who is beating them. A rival whose URL appeared in a citation, in an answer that recommended somebody else entirely, would have been recorded as a rival being recommended.
So every name check now runs against a copy of the answer with URLs and markdown link targets blanked out, character offsets preserved so quoted excerpts still line up with the real text. Three places, and it has to be all three: the check that can veto a mention, the check that can rescue one the analysis model missed, and the filter that decides which competitors were named. The visible label survives, the target does not. A company named in a sentence still counts. A company present only as a link does not.
Two consequences of that, and I would rather state them than let a sharp reader find them.
The masking applies to ChatGPT and to no other engine. Two of our other five can also return links, so right now two of our six engines are counted under slightly different mention rules. That is deliberate, and it is not laziness. Tightening a rule on a second engine inside the same release would move that engine's numbers on the same day for a second reason, and then nobody, including us, could read what the re-baseline meant. One change, one stamped date, one thing to explain. The second engine gets its own date.
The error runs one way, and it runs the safe way. Masking only removes text from the copy used to verify a name. It can turn a mention that would have counted into an absence. It cannot invent one. So this rule can under-count our customers and it cannot over-count them, which is the direction I want any residual error in a paid measurement to point. A tool that rounds in the direction of its own invoice is not measuring anything.
Which ChatGPT is your report measuring?
Run a free scan and read the actual questions, the actual answers, and which engines named you. One scan a month, six engines, no card.
Run a free AI visibility scanThe three objections worth making
"Three of ten is noise." Correct, and I will push it harder than the objection usually does. We have published a first-party test of exactly this instability: one fixed question sent ten times through the ChatGPT API with nothing changed named 61 different companies, and not one of them appeared in all ten runs. An engine that unstable cannot support a three-in-ten rate on any sample this small, which is why the argument does not rest there. It rests on the search counts, which are configuration rather than sampling. What the naming counts contribute is a direction and a shape, and the shape is what reproduced: on a separate probe two days earlier, with differently worded questions, the category phrasing missed and the positioning phrasings hit. Two probes on one brand is still one brand. It is a hypothesis with a mechanism, offered as one.
"You were also asking about the wrong countries." True, and since it is a genuine confound, here it is. The scans that produced those absent readings carried no geography at all, because of a separate defect in the classifier that decides which markets a brand's questions get pointed at. Two faults, either one sufficient to make a real firm look invisible. We fixed the geography one on 17 August in its own release with its own stamped date, and retrieval on 18 August in another, because shipping two instrument changes together makes the result unreadable. The probe in this post sidesteps the confound by using questions written for the two markets this customer actually sells into, so retrieval is the only thing varying between its two halves. Note what that implies for the parametric half: it had the corrected markets too, and still named them zero times out of ten.
"This is a bug report wearing a blog post." It is a bug report. We shipped a scanner whose most-used engine could not see anything our customers published, and it took a customer's inbound enquiry to make me go and check. Publishing it is not penance and it is not humility. It is that the same default sits underneath any tool built the obvious way against this API, the test that exposes it costs one call, and you can now run that test on us exactly as easily as on anyone else. A category that only ever publishes its wins has no way left to tell a measurement company from a dashboard company.
The trend line we are refusing to draw
We shipped forced retrieval on 18 August 2026 and stamped that date into the code, because there is an obvious and very attractive chart available to us now, and it would be a lie.
The customer above reads absent on ChatGPT before that date. Rescan them and some of those cells will read as named. Put the two scans on one line and you have a beautiful before and after, the kind that becomes a case study on a homepage. It would measure our repair, not their progress. Nothing about their business changed between those two points. We changed the question we were asking.
That was not a hypothetical temptation. I asked for that chart. I wanted it as our first case study, and I was refused in writing, in the spec, before a line of the code was written. The stamp shipped in the same release as the change specifically so that no surface of ours can draw a continuous line across it, including one I might later ask for.
There is a register of those dates in our codebase and it currently holds four. Three of them shipped inside six days this month: 12 August, when we made our brand matcher symmetric and widened how much of an answer it reads, which lifted measured rates by roughly 3 percent for reasons that have nothing to do with any customer. 17 August, when the market classifier stopped pointing questions at the wrong countries, which changes which questions get asked at all. 18 August, retrieval. A fourth date is already stamped for a change still in flight. Every affected customer's trend chart breaks at those dates and says why.
The rule underneath all of it is one line: relabel, never recompute. We do not re-run old scans against the new instrument and we do not restate old numbers. Those numbers were not wrong, they were mislabelled, and a tool that silently recomputes its own back catalogue every time it improves is not correcting the record. It is deleting the evidence that it changed, and handing you a smooth line that has never once been interrupted by the truth.
Which brings me to our own published figure. Our ChatGPT mention rate, 1.7 percent of the buying questions it answered across 46 reports covering 34 brands, every one of which came to us already suspecting it was missing from AI answers, describes an instrument we have retired. It is a fact about what a model remembered. We have not replaced it, because no corpus exists on the other side of 18 August yet, and inventing one would be a far worse offence than leaving a gap we can explain. We will recompute after roughly forty post-cutover scans and not a day earlier. Perplexity's 7.4 percent, measured on that same corpus of brands that suspected they were invisible, still means what it always meant: nothing about how we query Perplexity changed that week. If you have read our own post arguing that mention share and traffic share are different scoreboards, the ChatGPT row in it is now history and carries a correction saying so.
This is the part I would press hardest on with any vendor in this category, ours included. Instruments in AI visibility move constantly, because the engines move constantly, and that is fine. What is not fine is a supplier quietly improving their probe and letting the resulting step land in your dashboard as your improvement. Three questions. Does your ChatGPT call run a live search. When did that last change. Does the chart break at that date. The third one separates a measurement company from a dashboard company, and almost nobody volunteers the answer.
The reframe worth keeping is smaller than it sounds. This work has always had two halves: be retrievable, and be worth retrieving. The category has spent two years selling the second while measuring neither. Retrieval is the floor, exactly as being indexed by Google is the floor, and a floor is not an achievement. What gets you named standing on it is whether the web says something specific enough about you to answer a specific question, which is a positioning problem long before it is ever a schema problem.
Frequently asked questions
Does ChatGPT search the web before it recommends a company?
In the consumer app it often does. Through an API it does not unless the caller asks for it, and OpenAI documents that as a setting rather than a default: with the tool choice left on auto, search is optional, and you use a forced tool choice when search must run. We measured what optional means in practice. Our production probe ran a web search on zero of ten calls, and an earlier probe of ours left on auto searched on one of three runs. So the same question, asked the same way, sometimes reaches a retrieval engine and sometimes reaches a memory engine, and nothing in the answer text tells you which one you got.
How can I tell whether a tool is measuring memory or live search?
Ask the engine what today's date is and to cite the source it used. From memory ours answered, word for word: 'I'm unable to provide real-time information, including today's date, as I don't have access to current data or the internet. My responses are based on the information available up to October 2023.' With search forced, the same call returned 'Aug 17, 2026, 12:57:36 AM'. Parse the date, compare it against your own clock, and there is nothing left to argue about. We tried a better-sounding control first, asking for a news story published in the last seven days, and rejected it: even with search forced the model replied that it had no access to specific news articles, so it reports a healthy instrument as broken. Ask any vendor to run the date probe on the exact configuration they use for your scans.
Why would a visibility tool not turn web search on?
Cost is the obvious answer and our own numbers do not support it. At OpenAI's published rates the web search tool is $10.00 per 1,000 calls on top of tokens, with gpt-4o at $2.50 per million input tokens and $10.00 per million output, which works out against our measured token counts at $0.0036 per call without search and $0.0149 with it, about 4.1 times. The multiple is large and the absolute number is trivial: a free scan's ChatGPT calls cost about seven cents with search forced, against under two cents without. The real reason is duller and harder to fix. Attaching no tool is the default, a memory answer reads exactly like a retrieved one, and nothing in the output looks broken. Defaults do not require a decision, which is why they survive.
If ChatGPT answers from memory, can new content ever help?
Not on that call, and that is the whole problem. A model answering from training weights cannot see a page published after its training data was collected. So a customer who does exactly what an AEO report told them to do, publishes the comparison page, fixes the schema, earns the listings, cannot move on that measurement however good the work is. The advice and the instrument were pointed at two different worlds. A number produced that way tells you what a model absorbed during training, which is a fact about the past rather than about whether your buyer will be shown your name.
Did forcing search actually make the brand visible?
Partly, and the half that failed is the more useful half. Across five buying questions run twice in each mode for one customer, the memory probe named them zero of ten and the forced-search probe named them three of ten. The split matters more than the total. Broad category questions named them zero of four, and every naming came from a question carrying their own differentiator, the thing they say they are not. A fifth question that used the word independent but attached the geography to the buyer's market rather than to where the supplier sits named them zero of two. Retrieval is necessary and it is not sufficient. Treat three of ten as direction from one brand's probe, never as a rate.
Is being cited by ChatGPT the same as being recommended?
No, and forcing retrieval is exactly when that distinction starts to cost you. Retrieved answers carry inline links, so a plain check for your brand inside the answer text will match a link to your own domain and record a win where the engine only footnoted you. It cuts the other way too, and that half is worse: a competitor named only inside a link target counts as that competitor being recommended, and then feeds a leaderboard. We now run every name check against a copy of the answer with URLs and link targets blanked out, in three places: the check that can veto a mention, the check that can rescue one, and the filter that decides which competitors were named. When you read anyone's report, ask which of those two events it is counting.
Our scan improved after our vendor changed something. Is that real?
Ask what changed and when. A vendor who alters how an engine is queried has moved the instrument, and a rise across that date measures the vendor rather than you. We shipped this retrieval change on 18 August 2026 and stamped the date in code so no chart of ours can draw a continuous line across it. That was not a hypothetical temptation: I asked for exactly that before-and-after as our first case study, and it was refused in writing before the code shipped, because nothing about the customer's business had changed between the two readings. Demand the date, then judge the movement inside one regime rather than across two.
Sources and further reading
- OpenAI, Web search guide. The vendor's own statement that search is optional under an auto tool choice and must be forced when it has to run. Read 20 August 2026.
- OpenAI, API pricing. Web search at $10.00 per 1,000 calls plus search content tokens billed at model rates, gpt-4o at $2.50 per million input tokens and $10.00 per million output. Read 17 August 2026 and checked again 20 August 2026.
- TofuBofu first-party probe, 17 August 2026. Five buying questions for one live customer, each run twice with no tools attached and twice with web search forced, twenty calls in total. Search behaviour, brand naming and every citation URL recorded per call. Token medians and latency from a companion run the same day.
- TofuBofu first-party probe, 15 August 2026. The same engine under an auto tool choice, which searched on one run of three, plus four differently shaped prompts used to separate the effect of question wording from the effect of retrieval.
- TofuBofu production database, recomputed 12 August 2026. 46 completed reports, 34 brands, 419 distinct buying questions, 2,421 answered buying cells. Every brand in it ran a scan because somebody already suspected it was missing, which is a sample selected on the problem being measured.