Measurement
Does Gemini search the web? Not unless you call it the right way, and it won't tell you.
By Arnav Mukherjee, founder of TofuBofu · October 1, 2026
TL;DR
- Don't trust an HTTP 200. We asked Vertex to search, it agreed, and returned zero grounding metadata six times out of six.
- You can't spot it by reading. The grounded answer was shorter than the ungrounded one on two of three questions.
- It changes the shortlist. Four grounded and four ungrounded runs shared 2 of the 14 firms they named.
- Ungrounded scoring punishes the specialist. MSP Corp was named in 4 of 4 grounded runs and 0 of 4 ungrounded.
- Ask your vendor for the retrieval date per engine. Ours are 18 August, 28 August, and Gemini last.
Gemini was the last engine on our roster answering from memory. Every other engine we query was retrieving, and we knew the fix for Gemini was a search tool in the request, so this looked like an afternoon's work. We passed the flag, the API returned 200, and the answers came back long and confident and full of named vendors.
Nothing had searched. The endpoint had accepted an instruction it doesn't implement, thrown it away, and returned a perfectly healthy response. That's the worst failure mode a measurement tool can have, because there is no error to find and the output looks exactly like success. Here's how we caught it, and what it did to the shortlist underneath.
How we measured this
- Model and date.
gemini-3.5-flashon Google Vertex AI, locationglobal, temperature 0.7, 3,000 output tokens. Every call in this post was run on 1 October 2026. - Three call shapes. The OpenAI-compatible endpoint with no search instruction. The same endpoint passing
google_search. And the nativegenerateContentroute passingtools: [{"googleSearch": {}}]. - Three questions for the metadata test. Unbranded B2B buying questions, one run each per shape, nine calls: Canadian healthcare managed IT, UK fintech cybersecurity consulting, and B2B SaaS revenue attribution tools.
- Four runs each way for the shortlist test. One question, the Canadian one, run four times ungrounded and four times grounded, so that run-to-run variance could be separated from the effect of retrieval. Firm names were read out of the answers by hand rather than pattern-matched, after a first pass at pattern-matching returned section headings instead of companies.
- The honest limit. One question, one model, one day. The metadata result is binary and reproducible. The shortlist divergence is a shape measured on a single question, and it should be read as one, not as a rate.
What does a silent yes look like?
Like a success. The compatibility endpoint took the search flag without complaint and returned 200 on every attempt, and the response carried no grounding chunks, no list of searches performed, and no citation metadata of any kind. Identical, in every respect we can inspect, to not having asked.
| Call shape | HTTP | Searches run | Grounded sources |
|---|---|---|---|
| Compat endpoint, no search asked | 200, 200, 200 | 0, 0, 0 | 0, 0, 0 |
| Compat endpoint, search asked for | 200, 200, 200 | 0, 0, 0 | 0, 0, 0 |
| Native route, googleSearch tool | 200, 200, 200 | 6, 8, 5 | 14, 18, 19 |
As a sentence, because a table can't be quoted: across three buying questions, the compatibility endpoint ran 0 searches and returned 0 grounded sources whether or not we asked it to search, while the native route ran 19 searches and returned 51 grounded sources on the same three questions. Asking politely through the wrong door achieves exactly as much as not asking.
Now the part that makes this dangerous rather than merely annoying. You cannot see it in the answer. All nine responses ran between 5,121 and 6,654 characters, and on two of the three questions the grounded answer came back shorter than the ungrounded one. Both named five vendors. Both cited the correct privacy legislation for the market. Both read like an analyst who had done the work.
Does it change who gets recommended?
We ran one question eight times to find out, four ungrounded and four grounded, because a single pair proves nothing when the model is non-deterministic. The question: the best managed IT services providers for a mid-sized healthcare company in Canada.
The ungrounded answers were remarkably consistent with each other. Four firms appeared in all four runs, with only the fifth slot rotating. So run-to-run variance is real but small, and it isn't what produces the result below.
| Firm | Ungrounded runs | Grounded runs |
|---|---|---|
| TELUS Business | 4 of 4 | 0 of 4 |
| SysGen Solutions Group | 4 of 4 | 0 of 4 |
| Calian | 4 of 4 | 2 of 4 |
| Compugen | 4 of 4 | 1 of 4 |
| MSP Corp | 0 of 4 | 4 of 4 |
| BlueBird iT | 0 of 4 | 3 of 4 |
| F12.net | 0 of 4 | 3 of 4 |
| GAM Tech | 0 of 4 | 3 of 4 |
| ProServeIT | 0 of 4 | 2 of 4 |
| Supra ITS | 2 of 4 | 0 of 4 |
| IT Weapons | 0 of 4 | 1 of 4 |
| Microserve | 0 of 4 | 1 of 4 |
| WBM Technologies | 1 of 4 | 0 of 4 |
| Nucleus IT | 1 of 4 | 0 of 4 |
Stated as sentences, because these are the numbers worth lifting. Fourteen distinct firms were named across the eight runs. Two of the fourteen were named by both call shapes. Seven firms appeared only when the engine searched, including the two it reached for most often, MSP Corp in four runs out of four and BlueBird iT in three. Five firms appeared only when it didn't search, including TELUS and SysGen, which every single ungrounded run named and no grounded run mentioned at all.
Read the two columns as two different questions, because that's what they are. Ungrounded, the engine returned the largest and most written-about IT companies in Canada. Grounded, it returned smaller providers with named healthcare practices, the kind of firm with a decent website and almost no press. An ungrounded answer is closer to a measure of fame than a measure of visibility, and the entire gap between those two things lands on the specialist.
Which column is your firm in? A free TofuBofu scan puts your real buying questions to all five engines and shows you the verbatim answer, the sources each engine pulled, and who got named instead of you.
Run a free scanWhy does this break an AI visibility number?
Because a low number has two opposite causes and cannot tell them apart. The engine looked and picked somebody else, or the engine was never allowed to look. The first is a content problem you can fix. The second is a fact about the tool measuring you, and no amount of writing will move it.
Three consequences follow, and all three are checkable rather than theoretical.
- A silent engine is not an absent brand. When a forced search doesn't fire, the honest record is no signal, not a zero. A zero goes into a denominator and drags a rate down with it, which turns our own plumbing failure into your bad news.
- Turning retrieval on moves the number upward, and that movement isn't yours. Grounding an engine raises its measured mention rate. If a tool switches retrieval on and your chart steps up, the step is the tool improving. Anyone presenting that as progress is selling you your own instrument.
- So the cutover date is part of the data. ChatGPT crossed that line for us on 18 August 2026 and Claude on 28 August 2026, both recorded, and Gemini is the third and last. Across any of those dates we publish counts and say why, rather than drawing a trend through a change in the instrument.
We fixed ours by moving Gemini onto the native route with the googleSearch tool, which is the shape that grounded in our testing, and by logging it loudly when a forced search comes back with nothing. That change is merged and goes out with our next release, so as I write this our own production scans are still on the ungrounded path. I'd rather say that plainly than describe a fix in the present tense before it has reached anybody.
What should you ask your vendor?
Four questions. None needs a technical answer, and the fourth is the one that separates a measurement product from a dashboard.
- Which engines do you force retrieval on, named one by one? "All of them" is not an answer, because the shapes differ per vendor and two of ours needed a different mechanism each.
- What do you record when a forced search doesn't fire? You want no signal. If the answer is a zero, their rates include their own failures.
- Can you show me the sources an engine returned for one specific answer about my company? A tool that reads grounding metadata can produce that list immediately. A tool that never looked has nothing to show.
- What date did retrieval switch on for each engine? Every one of those dates is a seam in your own history. A vendor who can't name them has a chart with two different measurements on it and hasn't noticed.
SEO is still the floor underneath all of this. A page nothing can crawl can't be retrieved by anything, and no API flag fixes that. But nothing about those fourteen firms changed between the two sets of runs. Same question, same model, same morning, and the only variable was whether the engine went and looked. That's the layer AEO lives on.
Frequently asked questions
Does Gemini search the web when you ask it a question?
In the consumer Gemini app it can. Through the API it only searches if the calling code asks it to, in the one specific way that works, and a request that asks in any other way can come back looking perfectly healthy. We tested three call shapes on the same three buying questions. The Vertex OpenAI-compatible endpoint returned HTTP 200 whether or not we passed a search flag, and in six out of six calls it returned no grounding metadata of any kind. The native generateContent route with the googleSearch tool returned 14, 18 and 19 grounded sources on the same three questions, from 6, 8 and 5 real searches.
How can you tell whether an AI answer was grounded in a web search?
Not by reading it, which is the problem. Across our nine calls the answers ran between 5,121 and 6,654 characters and the grounded answer was shorter than the ungrounded one on two of the three questions. Both kinds named five vendors, cited the right privacy legislation and read with equal confidence. The only reliable tell is metadata the API either returns or does not: grounding chunks and the list of searches the model actually ran. If a tool cannot show you those, it cannot tell you whether the answer was retrieved, and neither can you.
Does it actually change which companies get recommended?
On our test question it changed almost all of them. We asked the same question, the best managed IT providers for a mid-sized healthcare company in Canada, four times ungrounded and four times grounded. The ungrounded runs named seven distinct firms and the grounded runs named nine. Only two firms appeared in both sets. TELUS and SysGen were named in all four ungrounded runs and in none of the grounded ones. MSP Corp was named in all four grounded runs and in none of the ungrounded ones. This is one question on one day, so read it as a shape rather than a rate.
Why does an ungrounded engine name the big brands?
Because without retrieval it answers from what it absorbed in training, and what it absorbed is correlated with how much has been written about you. Our ungrounded runs returned large, widely covered Canadian IT companies. The grounded runs returned smaller providers with healthcare practices, the sort of firm that has a good site and little press. So an ungrounded measurement is closer to a measure of fame than a measure of visibility, and the gap falls entirely on the smaller specialist. That is the opposite of what an AI visibility tool is for.
What should I ask an AI visibility vendor about this?
Four questions, and the last one is the one nobody expects. Which engines do you force retrieval on, named one by one? What do you record when a forced search does not fire, a zero or no signal? Can you show me the sources an engine returned for a specific answer about my company? And what date did retrieval switch on for each engine, so I know which part of my own trend line is your instrument changing rather than my visibility moving? A vendor who cannot answer the fourth has a chart that mixes two different measurements and has not noticed.
Is a low Gemini number evidence that Gemini ignores my company?
Not on its own, and this is the specific error worth guarding against. A brand can come out low because the engine looked and chose somebody else, or because the engine was never allowed to look. Those are opposite diagnoses with opposite fixes, and they are indistinguishable from the number alone. Any figure for an engine that was answering from training memory is a measurement of the measuring tool. We record a forced search that fails to fire as no signal rather than as an absence, and we date every change to our own instrument, because a brand reading an unexplained jump would write down a number that means something else.
Did you fix it, and what does that do to past numbers?
We moved Gemini onto the native route with the googleSearch tool, which is the shape that grounded in our testing. That change is merged and ships in our next release, so at the time of writing production is still running the ungrounded path. It also raises Gemini's measured mention rate across the board, and that rise is our instrument improving rather than any brand becoming more visible, so we record the cutover date and tell customers to read counts rather than direction across it. ChatGPT crossed the same line on 18 August 2026 and Claude on 28 August 2026. Gemini was the last engine on our roster still answering from memory.
Sources and further reading
- TofuBofu first-party measurement, 1 October 2026. Every figure on this page comes from the nine metadata calls and eight shortlist runs described in the method block above, against
gemini-3.5-flashon Vertex AI. Our retrieval cutover dates are recorded in our own measurement-shift log, which is what the customer-facing charts read. - Google Cloud, Vertex AI generative AI inference reference. The documentation for the two routes compared here, including the
toolsfield the native route accepts. - Atil et al., "Non-Determinism of 'Deterministic' LLM Settings", arXiv:2408.04667. Why the same prompt twice is not the same test, and the reason we ran four runs per call shape rather than one.