Now live across the AI ecosystem: ChatGPT GPT Store · MCP Registry · mcp.so

Measurement

Does Gemini search the web? Not unless you call it the right way, and it won't tell you.

By Arnav Mukherjee, founder of TofuBofu · October 1, 2026

TL;DR

  • Don't trust an HTTP 200. We asked Vertex to search, it agreed, and returned zero grounding metadata six times out of six.
  • You can't spot it by reading. The grounded answer was shorter than the ungrounded one on two of three questions.
  • It changes the shortlist. Four grounded and four ungrounded runs shared 2 of the 14 firms they named.
  • Ungrounded scoring punishes the specialist. MSP Corp was named in 4 of 4 grounded runs and 0 of 4 ungrounded.
  • Ask your vendor for the retrieval date per engine. Ours are 18 August, 28 August, and Gemini last.

Gemini was the last engine on our roster answering from memory. Every other engine we query was retrieving, and we knew the fix for Gemini was a search tool in the request, so this looked like an afternoon's work. We passed the flag, the API returned 200, and the answers came back long and confident and full of named vendors.

Nothing had searched. The endpoint had accepted an instruction it doesn't implement, thrown it away, and returned a perfectly healthy response. That's the worst failure mode a measurement tool can have, because there is no error to find and the output looks exactly like success. Here's how we caught it, and what it did to the shortlist underneath.

How we measured this

  • Model and date. gemini-3.5-flash on Google Vertex AI, location global, temperature 0.7, 3,000 output tokens. Every call in this post was run on 1 October 2026.
  • Three call shapes. The OpenAI-compatible endpoint with no search instruction. The same endpoint passing google_search. And the native generateContent route passing tools: [{"googleSearch": {}}].
  • Three questions for the metadata test. Unbranded B2B buying questions, one run each per shape, nine calls: Canadian healthcare managed IT, UK fintech cybersecurity consulting, and B2B SaaS revenue attribution tools.
  • Four runs each way for the shortlist test. One question, the Canadian one, run four times ungrounded and four times grounded, so that run-to-run variance could be separated from the effect of retrieval. Firm names were read out of the answers by hand rather than pattern-matched, after a first pass at pattern-matching returned section headings instead of companies.
  • The honest limit. One question, one model, one day. The metadata result is binary and reproducible. The shortlist divergence is a shape measured on a single question, and it should be read as one, not as a rate.

What does a silent yes look like?

Like a success. The compatibility endpoint took the search flag without complaint and returned 200 on every attempt, and the response carried no grounding chunks, no list of searches performed, and no citation metadata of any kind. Identical, in every respect we can inspect, to not having asked.

Three unbranded B2B buying questions, one run per call shape, gemini-3.5-flash on Vertex, 1 October 2026.
Call shape HTTP Searches run Grounded sources
Compat endpoint, no search asked200, 200, 2000, 0, 00, 0, 0
Compat endpoint, search asked for200, 200, 2000, 0, 00, 0, 0
Native route, googleSearch tool200, 200, 2006, 8, 514, 18, 19

As a sentence, because a table can't be quoted: across three buying questions, the compatibility endpoint ran 0 searches and returned 0 grounded sources whether or not we asked it to search, while the native route ran 19 searches and returned 51 grounded sources on the same three questions. Asking politely through the wrong door achieves exactly as much as not asking.

Now the part that makes this dangerous rather than merely annoying. You cannot see it in the answer. All nine responses ran between 5,121 and 6,654 characters, and on two of the three questions the grounded answer came back shorter than the ungrounded one. Both named five vendors. Both cited the correct privacy legislation for the market. Both read like an analyst who had done the work.

WHAT YOU READ WHAT THE RESPONSE CARRIES Asked to search, through the wrong door Five named vendors. Correct legislation. Confident. 5,688 characters. Asked to search, through the right door Five named vendors. Correct legislation. Confident. 5,796 characters. HTTP 200 0 searches, 0 sources, no grounding metadata at all HTTP 200 6 searches, 14 sources, every one of them listed Same model, same question, same minute. The left column is all you get by reading the answer.

We ran one question eight times to find out, four ungrounded and four grounded, because a single pair proves nothing when the model is non-deterministic. The question: the best managed IT services providers for a mid-sized healthcare company in Canada.

The ungrounded answers were remarkably consistent with each other. Four firms appeared in all four runs, with only the fifth slot rotating. So run-to-run variance is real but small, and it isn't what produces the result below.

Firms named for one Canadian healthcare MSP question, gemini-3.5-flash, 4 runs per call shape, 1 October 2026.
Firm Ungrounded runs Grounded runs
TELUS Business4 of 40 of 4
SysGen Solutions Group4 of 40 of 4
Calian4 of 42 of 4
Compugen4 of 41 of 4
MSP Corp0 of 44 of 4
BlueBird iT0 of 43 of 4
F12.net0 of 43 of 4
GAM Tech0 of 43 of 4
ProServeIT0 of 42 of 4
Supra ITS2 of 40 of 4
IT Weapons0 of 41 of 4
Microserve0 of 41 of 4
WBM Technologies1 of 40 of 4
Nucleus IT1 of 40 of 4

Stated as sentences, because these are the numbers worth lifting. Fourteen distinct firms were named across the eight runs. Two of the fourteen were named by both call shapes. Seven firms appeared only when the engine searched, including the two it reached for most often, MSP Corp in four runs out of four and BlueBird iT in three. Five firms appeared only when it didn't search, including TELUS and SysGen, which every single ungrounded run named and no grounded run mentioned at all.

Read the two columns as two different questions, because that's what they are. Ungrounded, the engine returned the largest and most written-about IT companies in Canada. Grounded, it returned smaller providers with named healthcare practices, the kind of firm with a decent website and almost no press. An ungrounded answer is closer to a measure of fame than a measure of visibility, and the entire gap between those two things lands on the specialist.

Which column is your firm in? A free TofuBofu scan puts your real buying questions to all five engines and shows you the verbatim answer, the sources each engine pulled, and who got named instead of you.

Run a free scan

Why does this break an AI visibility number?

Because a low number has two opposite causes and cannot tell them apart. The engine looked and picked somebody else, or the engine was never allowed to look. The first is a content problem you can fix. The second is a fact about the tool measuring you, and no amount of writing will move it.

Three consequences follow, and all three are checkable rather than theoretical.

We fixed ours by moving Gemini onto the native route with the googleSearch tool, which is the shape that grounded in our testing, and by logging it loudly when a forced search comes back with nothing. That change is merged and goes out with our next release, so as I write this our own production scans are still on the ungrounded path. I'd rather say that plainly than describe a fix in the present tense before it has reached anybody.

What should you ask your vendor?

Four questions. None needs a technical answer, and the fourth is the one that separates a measurement product from a dashboard.

SEO is still the floor underneath all of this. A page nothing can crawl can't be retrieved by anything, and no API flag fixes that. But nothing about those fourteen firms changed between the two sets of runs. Same question, same model, same morning, and the only variable was whether the engine went and looked. That's the layer AEO lives on.

Frequently asked questions

Does Gemini search the web when you ask it a question?

In the consumer Gemini app it can. Through the API it only searches if the calling code asks it to, in the one specific way that works, and a request that asks in any other way can come back looking perfectly healthy. We tested three call shapes on the same three buying questions. The Vertex OpenAI-compatible endpoint returned HTTP 200 whether or not we passed a search flag, and in six out of six calls it returned no grounding metadata of any kind. The native generateContent route with the googleSearch tool returned 14, 18 and 19 grounded sources on the same three questions, from 6, 8 and 5 real searches.

How can you tell whether an AI answer was grounded in a web search?

Not by reading it, which is the problem. Across our nine calls the answers ran between 5,121 and 6,654 characters and the grounded answer was shorter than the ungrounded one on two of the three questions. Both kinds named five vendors, cited the right privacy legislation and read with equal confidence. The only reliable tell is metadata the API either returns or does not: grounding chunks and the list of searches the model actually ran. If a tool cannot show you those, it cannot tell you whether the answer was retrieved, and neither can you.

Does it actually change which companies get recommended?

On our test question it changed almost all of them. We asked the same question, the best managed IT providers for a mid-sized healthcare company in Canada, four times ungrounded and four times grounded. The ungrounded runs named seven distinct firms and the grounded runs named nine. Only two firms appeared in both sets. TELUS and SysGen were named in all four ungrounded runs and in none of the grounded ones. MSP Corp was named in all four grounded runs and in none of the ungrounded ones. This is one question on one day, so read it as a shape rather than a rate.

Why does an ungrounded engine name the big brands?

Because without retrieval it answers from what it absorbed in training, and what it absorbed is correlated with how much has been written about you. Our ungrounded runs returned large, widely covered Canadian IT companies. The grounded runs returned smaller providers with healthcare practices, the sort of firm that has a good site and little press. So an ungrounded measurement is closer to a measure of fame than a measure of visibility, and the gap falls entirely on the smaller specialist. That is the opposite of what an AI visibility tool is for.

What should I ask an AI visibility vendor about this?

Four questions, and the last one is the one nobody expects. Which engines do you force retrieval on, named one by one? What do you record when a forced search does not fire, a zero or no signal? Can you show me the sources an engine returned for a specific answer about my company? And what date did retrieval switch on for each engine, so I know which part of my own trend line is your instrument changing rather than my visibility moving? A vendor who cannot answer the fourth has a chart that mixes two different measurements and has not noticed.

Is a low Gemini number evidence that Gemini ignores my company?

Not on its own, and this is the specific error worth guarding against. A brand can come out low because the engine looked and chose somebody else, or because the engine was never allowed to look. Those are opposite diagnoses with opposite fixes, and they are indistinguishable from the number alone. Any figure for an engine that was answering from training memory is a measurement of the measuring tool. We record a forced search that fails to fire as no signal rather than as an absence, and we date every change to our own instrument, because a brand reading an unexplained jump would write down a number that means something else.

Did you fix it, and what does that do to past numbers?

We moved Gemini onto the native route with the googleSearch tool, which is the shape that grounded in our testing. That change is merged and ships in our next release, so at the time of writing production is still running the ungrounded path. It also raises Gemini's measured mention rate across the board, and that rise is our instrument improving rather than any brand becoming more visible, so we record the cutover date and tell customers to read counts rather than direction across it. ChatGPT crossed the same line on 18 August 2026 and Claude on 28 August 2026. Gemini was the last engine on our roster still answering from memory.

Sources and further reading

Arnav Mukherjee

Founder, TofuBofu