Now live across the AI ecosystem: ChatGPT GPT Store · MCP Registry · mcp.so

Measurement

Every engine gave a different shortlist. Not one firm appeared on all of them.

By Arnav Mukherjee, founder of TofuBofu · August 5, 2026

On 3 August we ran a batch of discovery queries to find firms for our own research pipeline. Nine questions, put to the six AI engines we track, of which four returned answers that day. They were the sort of thing a buyer types when they have a problem and no shortlist: who are the best personal injury lawyers in Columbus, who are the best legal marketing agencies, and so on across five metros and four professional categories.

I was reading the stored answers for something else entirely when I noticed the four Columbus shortlists had almost nothing in common. Not different orderings of the same firms. Different firms. I checked Tampa, and found the same thing: ChatGPT named seven firms, Claude named six, Bing Copilot named fourteen, Google AI Mode named eight, and exactly three names appeared on as many as three of those four lists.

That is a measurement problem before it is anything else. If you run a firm in Tampa and you check one engine to see whether AI recommends you, what exactly have you learned? So I counted properly across all nine queries, and the answer turned out to depend almost entirely on how specific the question was.

The local result: 136 firms, and not one consensus pick

Five metros, one buying question each, four engines answering. That produced 136 firm-and-market entries, covering 134 distinct firms, since one national brand turned up in three separate markets. Of those 136 entries, 116 were named by exactly one engine. Fifteen were named by two. Five were named by three. Not a single firm, in any of the five metros, was named by all four.

That last number is the one that stopped me. Across five separate markets and twenty separate answers, the four engines never once agreed on a single firm.

Named by exactly 1 engine
116
Named by 2 engines
15
Named by 3 engines
5
Named by all 4 engines
0

136 firm-and-market entries across Charlotte, Columbus, Denver, Phoenix and Tampa, named by the four engines that answered, August 2026. Counted per market, so a firm named in two markets counts once in each.

A method note, because this is exactly where analyses like this fall apart. Engines write the same firm several ways. Bing Copilot returned "Gunn Law Group P.A." and Google AI Mode returned "Gunn Law Group, P.A.", which is one firm named twice, not two firms named once. We have been bitten by this before: an earlier regional study of ours counted shortened names as separate companies and published firms as invisible when an engine had in fact named them. So the figures above are computed after stripping punctuation and common suffixes such as PA, PLLC, LLC and Injury Attorneys. Before normalisation the raw numbers were 144 firms and 90 percent named by exactly one engine. After normalisation, 136 and 85 percent. The conservative figure is the one in the table.

One more caveat worth stating plainly, because it is the difference between blocked and zero. Six engines were queried; four returned answers. Perplexity and Gemini returned nothing in this collection run for reasons on our side, not because they declined the question. Everything here describes four engines, and none of it should be read as a finding about the other two.

The same engines agree far more on a broader question

The four remaining queries in the batch asked the same engines for the best marketing agencies in four categories: legal marketing, SaaS marketing, B2B SEO, and agencies in Austin. Same engines, same day, same extraction, same normalisation. The behaviour is visibly different.

Those four questions produced 138 distinct agencies, of which 108, or 78 percent, were named by exactly one engine. Still a scattered field, but with a meaningful head: 25 agencies were named by two engines, three by three engines, and two were named by all four. Scorpion was on every engine's legal marketing list. Kalungi was on every engine's SaaS marketing list. In the local queries, the equivalent count was zero.

Why the same four engines converge on one question and scatter on the other

Category question "best SaaS marketing agency" Engine A Engine B Engine C Engine D Shared names well documented, every engine finds them Local question "best injury lawyer in Columbus" Engine A Engine B Engine C Engine D Separate shortlists thin, uneven local evidence Conceptual. The narrower and more local the question, the less evidence each engine shares with the others.

The mechanism is not mysterious. A category question has a small set of heavily documented answers. Large agencies get written about in directories, roundups, review sites and press, so each engine retrieves overlapping evidence and arrives at overlapping names. Ask about personal injury firms in one metro and the candidate pool is hundreds of small practices, each with a modest and unevenly distributed footprint. Small differences in what each engine indexed produce large differences in what it says.

There is a second effect worth noting. Exactly one firm in the entire local set appeared in more than one metro: a national personal injury brand that turned up in three of the five. Every other name was local to its own market. National advertisers can occupy a slot in many metros at once, but the rest of the field stayed genuinely local, which is a more hopeful reading than the usual one. Most of the slots are still winnable.

Some of that disagreement is not disagreement

Four different shortlists could be four defensible opinions. In a metro with two hundred capable firms, reasonable sources will pick differently, and that would be fine. So I took the single Columbus answer that had first caught my eye and checked all six names by hand.

Four of the six I could find no trace of: no website, no bar listing, no directory entry, nothing in a general web search. One, Schottenstein Zox & Dunn, was a real and well-known Columbus firm, and it stopped existing as an independent firm when it combined with Ice Miller LLP effective 1 January 2012. It was recommended in 2026 as a current option. The sixth, Gallagher Sharp, is a real firm founded in 1912 with a Columbus office, but its practice is insurance and civil defence litigation, which is the opposite side of the docket from the plaintiff-side injury representation the question asked for.

So of six confident, well-formatted recommendations, not one was a currently operating Columbus plaintiff personal injury firm. I want to be careful about the strength of that claim: absence from a web search is not proof that a firm does not exist, and a small practice can have a thin footprint. But four unfindable names, one dissolved fourteen years ago and one on the wrong side of the case is not a near miss.

We have seen this pattern before in a regional managed IT study, where one engine's entire shortlist proved unverifiable. Finding it again in an unrelated profession, with a different question, suggests it is a property of thin evidence rather than a quirk of one industry. The practical consequence for you is simple and slightly bleak: from inside a single answer, a buyer cannot tell an omission from an invention, and neither can you.

See all six engines separately, not blended

Run a free scan on your own buying question and read each engine's answer verbatim, including the ones that never mention you.

Get your free audit

What this changes about measuring your own visibility

Four things follow from the numbers above, and they are worth acting on in this order.

1

Treat one engine as one data point, not an answer

If 85 percent of firms in a local answer are named by exactly one engine out of four, then the engine you happen to check is close to a coin flip on your own name. Absence from it is weak evidence of absence generally, and presence in it is weak evidence of health. This is the specific reason we report every engine separately rather than averaging them into one reassuring percentage.

2

Ask the question your buyer asks, with the city in it

The gap between a category question and a local one was the largest single effect in this data. If your firm is bound to a metro, measuring yourself on the national category question tells you about a race you are not in. Localise the query to the market you actually serve, then judge the result on that.

3

Re-run the same question rather than a new one

Because the shortlists are this unstable, a changed question and a changed result are indistinguishable. We rebuild each repeat scan from the previous scan's exact questions for this reason, after watching two consecutive scans share no questions and turn an ordinary result into an apparent collapse. Whatever you use, hold the question fixed.

4

Fix the corroboration, not the individual answer

There is no lever that moves one engine's shortlist. What every engine draws on is the same underlying evidence: consistent naming across directories, a site that states plainly which city and which practice you serve, and third-party listings that agree with each other. Profound's research found that appearing across four or more platforms makes a brand about 2.8 times more likely to be cited. Breadth of corroboration is the thing that converts a name one engine happens to know into a name several of them can verify.

None of this makes AI recommendations less consequential. G2's 2026 research found 51 percent of B2B buyers now begin vendor research on an AI chatbot, up from 29 percent, and 69 percent had switched their vendor choice based on what AI told them. A buyer in Columbus gets one of these shortlists, not four, and has no way of knowing how little the other three would have resembled it. The instability is not a reason to dismiss the channel. It is the reason to measure it properly.

The honest limits of this dataset

Nine queries is a small study and I would rather say so than dress it up. It is one question per metro, run once, on one day, by four engines. Engine behaviour shifts with model updates, and a rerun next month would produce different names. Two of our six engines returned nothing in this run, so this is not a six-engine finding.

What I do think survives all of that is the direction and the size of the gap. Two unanimous picks on the category questions against zero on the local ones, from the same engines on the same day with the same processing, is a difference too large to come from sampling noise alone. If you sell into a metro rather than a category, assume your visibility is four different numbers until you have checked.

Frequently asked questions

Do different AI engines give the same recommendations?

It depends entirely on how specific the question is. In our data, four engines asked for the best personal injury firms in five US metros produced 136 firm-and-market entries between them, covering 134 distinct firms, and 116 of those entries, or 85 percent, were named by exactly one engine. Not a single firm was named by all four in any of the five metros. On broader category questions the same four engines agreed far more often, converging unanimously on two agencies out of 138 named. Agreement is a function of query specificity, not a fixed property of the engines.

Is checking one AI engine enough to know if I am visible?

For a locally bound firm, no, and the gap is much larger than most people assume. If 85 percent of firms in our local sample were named by exactly one engine out of four, then a founder who checks a single engine is sampling a shortlist that the other three engines largely do not share. Being absent from the one you happen to check tells you little, and being present tells you even less. The only reading that survives is a per-engine one, which is why we report every engine separately rather than blending them into one number.

Why do AI engines disagree more on local questions?

A national category question has a small set of heavily documented answers. Large agencies and vendors are written about in directories, roundups, review sites and press, so every engine retrieves overlapping evidence and lands on overlapping names. A metro-level question about small professional firms has a long, thin tail: hundreds of plausible candidates, each with a modest and unevenly distributed web footprint. Small differences in what each engine indexed or retrieved produce large differences in output, so the shortlists diverge.

Does engine disagreement mean the engines are wrong?

Not necessarily, and this is the important distinction. Four different valid shortlists can exist in a metro with two hundred capable firms. But some of the divergence in our sample was not a difference of opinion. Checking one engine's six Columbus recommendations by hand, four were firms we could find no trace of, one had merged into another firm effective January 2012, and the remaining one was a real firm whose practice is insurance and civil defence rather than plaintiff-side personal injury. Disagreement and invention look identical from inside a single answer.

Which AI engine names the most companies?

In this sample the retrieval-heavier engines were more generous. Across the five metros, Bing Copilot named 52 firms and Google AI Mode named 45, while ChatGPT named 33 and Claude 31. A longer list is not automatically a better one, since more names also means more room for weakly grounded ones, but it does mean your odds of appearing differ by engine before anything about your own site is considered. Two engines returned no data in this collection run for reasons on our side, so this ranking covers four engines, not six.

What should a local firm do about AI visibility given this?

Stop treating visibility as one score and start treating it as four to six separate results. Measure each engine on the buying question your clients actually ask, with your city in it, and track which engines name you over time rather than what a blended average says this month. Then work on the corroboration that every engine draws from: consistent naming across directories, a site that states plainly which city and practice area you serve, and third-party listings that agree with each other. Broad corroboration is what turns a name that one engine happens to know into a name several of them can verify.

Do national brands dominate local AI answers?

Partially, and the pattern is visible even in a small sample. Exactly one firm in our five-metro set appeared in more than one metro, a national personal injury brand that appeared in three. Every other name was local to its own market. So the crowding effect is real but narrow: a national advertiser can occupy a slot in many metros at once, while the remaining slots stay genuinely local and, on this evidence, genuinely winnable by firms that are easy to verify.

How do you count company names across engines without double counting?

Carefully, because this is where analyses of this kind quietly break. Engines write the same firm several ways, so Gunn Law Group P.A. and Gunn Law Group, P.A. are one firm reported twice unless you normalise. We learned this the hard way on an earlier study where a matching bug counted shortened names as separate companies. The figures here are computed after stripping punctuation and common suffixes such as PA, PLLC and Injury Attorneys. Before normalisation the raw counts were 144 entries and 90 percent named by exactly one engine; after normalisation, 136 and 85 percent. We report the conservative figure.

Sources and further reading

  • TofuBofu Research: our regional multi-engine studies, including the per-engine raw answers behind the counts above.
  • Ice Miller LLP and Schottenstein Zox and Dunn combine: the 2011 announcement of the combination effective 1 January 2012, for the firm still being recommended in 2026.
  • Gallagher Sharp LLP: the firm's own site, describing an insurance and civil defence practice.
  • G2 2026 B2B buyer research: 51 percent begin vendor research on an AI chatbot, up from 29 percent, and 69 percent switched vendor based on AI.
  • Profound: brands appearing across four or more platforms are about 2.8 times more likely to be cited.

Keep reading: Why one blended visibility score hides the truth · AI visibility for local businesses · AI keeps recommending the big names