Fundamentals
AI keeps recommending the big names. Can a smaller firm break in?
By Arnav Mukherjee, founder of TofuBofu · August 3, 2026
We ran the same experiment across 15 US states and Canadian provinces. One open question per region, put to four AI engines: who are the best managed IT service providers here. Then we took every company name the engines volunteered and checked it against two independent references, the region's own Clutch directory and a live web-search pass. The full per-region data sits on our research pages.
I went in expecting to write about which companies win. The pattern that actually came out of the data was about which companies keep winning everywhere. Symmetrio appeared in the answers for 8 of the 15 regions, Integris in 7, and Ntiva in 5, while Telus Business, Bell Business, CDW, Thrive and Dataprise each showed up in four. A firm genuinely rooted in one province cannot also be the best provider in seven others, which makes these lists a measure of name recall wearing a regional label rather than a measure of regional expertise.
Meanwhile, in British Columbia, firms sitting on the province's own directory with active review profiles were absent from every answer. Adastra, rated 4.9 across 15 reviews, went unmentioned by all four engines. Dyrand Systems, 4.8 across 9 reviews, the same. If you run a smaller firm, this is the experience you already have: the answer lists companies you have never once met in a competitive deal.
The size of the gap, in numbers
Aggregate the 15 regions and the picture sharpens. Across all of them, the four engines named 595 company entries. Of those, 98 turned up in at least one independent check. The remaining 497, or 84 percent, appeared in neither the region's directory nor the web-search results for the same question.
Run the comparison the other way and it gets worse for the smaller firm. Those same 15 regional directories listed 228 managed IT providers between them. Real companies, actively marketing, most with published reviews. Of the 228, exactly 150 were named by none of the four engines. Two thirds of the firms a buyer would find by opening the obvious directory are missing from the answer that buyer now gets instead.
Those two numbers describe one system. Names that recur across regions crowd the top of these answers, and the region's actual supplier base thins out toward the bottom. G2's 2026 research found 51 percent of B2B buyers now begin vendor research on an AI chatbot, up from 29 percent, and that 69 percent had changed their vendor choice based on what AI told them. A shortlist with that shape reaches a lot of pipelines.
Founders describe the symptom long before they find a name for it. A small app studio owner posted this after testing the question their own buyers ask:
"Across ChatGPT, Gemini and Claude, same result every time. Same competitors. We weren't mentioned once. Which is a bit of a problem, because if this becomes how people search, you basically don't exist unless you're one of those answers."
A small business owner on r/smallbusiness
Why breadth beats quality inside a retrieval system
The mechanism explains the fix, so it is worth being precise about it. An AI engine does not rank vendors against criteria. It assembles a plausible answer from text it can retrieve and reconcile, an approach formalised by Lewis and colleagues in their 2020 work on retrieval-augmented generation and now standard across search-grounded systems. Ask for the best providers in a region and the model is looking for names it can state with confidence, where confidence comes from having encountered the same entity in enough independent contexts to treat it as real and relevant.
A national brand clears that bar without trying. Trade press covers it, national directories list it, review sites compare it, partner pages cite it, job ads name it and local news mentions it across every market it touches. A ten-person specialist has a website, perhaps one directory profile, and a handful of reviews. Both firms might serve a 40-person law firm equally well. Only one of them has left a trail wide enough for a model to follow.
You cannot transmit service quality to a model. You can transmit footprint, which is the useful half of an otherwise bleak finding. Everything that looks like unfairness here reduces to a measurable, buildable property of your public presence.
Three different problems that look identical from your desk
When you read an AI answer that omits you, the names occupying your slot fall into three groups. They feel the same when you are the one left out, and they call for different responses.
The national incumbent
A large firm with genuine multi-market presence, such as a telecom's business division or a national reseller. It earned its corroboration honestly, across thousands of sources you cannot match on volume.
What to do: Do not contest the broad query. Compete on qualified questions where its generic material has nothing specific to say and yours does.
The multi-region roll-up
A private-equity-backed group that acquired local providers and now appears under one brand in many markets. In our data these were the names recurring in five to eight regions at once.
What to do: This one is beatable on specificity. Roll-ups tend to publish one national services page per line of business, which leaves every vertical and city question thinly covered.
The name that checks out nowhere
An entry that appears in the answer and in neither the directory nor the search results. Across our 15 regions, 497 of 595 named entries fell into this bucket.
What to do: Nothing to out-compete, which is oddly the hardest case. The slot is occupied by something a buyer cannot verify, and the only counter is being the concrete, checkable option in the same answer.
The third category deserves a caveat, because I do not want to overstate it. An entry missing from a regional directory is not proof of fabrication. A national firm has no reason to appear in a provincial MSP listing, and a real specialist with a thin web presence can slip past a search check. What the 84 percent figure supports is narrower and still useful: most of what these engines volunteer in response to a regional buying question cannot be confirmed from the two sources a diligent buyer would consult first.
Where the incumbent's advantage holds, and where it thins out
The engines are not one weather system
Treating AI visibility as a single condition hides the most actionable finding in our data. The four engines behaved differently enough that they may as well be different products. Of the companies each named across the 15 regions, Perplexity had 87 of 257 clear an independent check. Gemini had 18 of 137. ChatGPT had 8 of 114. Claude had none of its 129.
The ordering tracks how each engine resolves the question. Perplexity retrieves live pages and summarises what it finds, so its answers stay tethered to material currently on the web. The others lean harder on what the model already holds, which is why their lists drift further from anything a buyer can look up.
For a smaller firm this is close to a strategy. The retrieval-heavy engines are the ones where new published work can change an answer inside weeks, so they are where effort converts to evidence fastest. The training-heavy engines move on model release cycles and reward patience rather than campaigns. Working on both at once, and expecting them to respond on the same timeline, is how founders conclude that none of this works.
Two things that will not rescue you
The instinctive response to invisibility is to collect more reviews. The effect is real and small. Kevin Indig's analysis of G2 data found roughly 10 percent more reviews associates with about 2 percent more AI citations, and G2 holds around 22.4 percent share of voice in software answers, so where you gather them matters as much as how many you gather. Our regional data showed the ceiling without ambiguity: Adastra held a 4.9 rating and 15 reviews on the exact directory a buyer researching British Columbia would open, and four engines named it zero times.
The second false hope is your Google ranking. A founder on r/SaaS described running 50 question variations across ChatGPT, Perplexity and Claude, and reported that a competitor ranking below them on Google was recommended in most of the responses while their own product appeared in none. Their numbers are their own observation rather than anything we have verified, but the shape matches what our regional study found from the other direction: 47 companies appeared in web search results for these questions and in no AI answer at all.
This is the practical meaning of a position we hold consistently. SEO is the floor, since an engine that cannot crawl or parse your pages will never cite them. AEO is a separate layer built on agreement across sources, and it can move while your rankings sit still.
See which names AI gives instead of yours
Run a free scan across six AI engines on your real buying questions and read the answers your buyers are getting, verbatim.
Get your free auditWhere a smaller firm has the structural advantage
Stop contesting the question the incumbent already owns. On a broad category query the model has thousands of sources to draw from, and the ones it trusts most are the ones repeated most, which is a contest decided before you enter it.
Buyers ask qualified questions too, and every qualifier they add strips candidates out of the pool. Consider what happens to the available sourcing as a question narrows. "Best IT provider in Texas" has more published material behind it than any model can weigh. "Managed IT for a 40-person law firm in Austin that needs help maintaining SOC 2" has almost none, because few firms have written that page. The national brand has a services page describing managed IT in general terms. You have done that exact engagement eleven times and can describe the compliance evidence, the response commitments and the two things that usually go wrong in month one.
Publish that clearly and you can be the strongest available source for the question. This is not a consolation prize for firms too small to compete, and it is worth resisting the instinct to treat it as one. A qualified question carries far higher buying intent than the category question, because the person asking it has already scoped their problem. Placing fourth on a query that produces browsers is worth less than owning the handful of questions that produce meetings.
The break-in sequence
Write down the questions you should already own
List the qualified questions your best-fit buyers ask out loud: vertical, company size, compliance requirement, city, technology stack, migration type. Ten is plenty. These are the questions where the incumbent has published something generic and you have done the work repeatedly. Anything you cannot describe in specific detail comes off the list.
Check what the engines answer today, per question
Ask each question on the engines your buyers use and save the response text, not a score. You are looking for which names occupy your slot and which of the three categories above they fall into, because that determines whether the question is winnable this quarter or at all.
Publish one specific page per question
State the vertical, the constraint and the geography in plain sentences near the top, then answer the question the way you would answer it on a call. Add FAQ schema, which SE Ranking found on 71 percent of pages ChatGPT cites. A page that could be about any firm in any city gives a model nothing to distinguish you with.
Build corroboration across four or more sources
One review platform will not carry it. Aim for the directory your industry actually uses, genuine participation on Reddit, association and partner listings, and coverage somewhere you did not pay for. Profound's finding that brands on four or more platforms are roughly 2.8 times more likely to be cited is the most useful number to organise this around.
Make your entity unambiguous everywhere
Same legal name, same description, same locations across your site, your schema and every listing. Engines resolve entities by matching features, and inconsistent features are how a firm gets confused with another or quietly dropped. Our data included cases where engines disagreed with each other about a single company's name.
Re-ask the same questions monthly and compare answers
Not to watch a score move, but to see whether a specific answer changed and which engine changed first. Grounded engines should move before the others. If nothing has shifted after two months on a question, the honest conclusion is usually that the page you published was not specific enough to beat what was already there.
What this sequence will not do
It will not put you on the broad category query against a national brand within a year, and any tool promising otherwise is selling something. It will not fix a training-based answer quickly, since those move when models do. It will not help if engines cannot crawl or render your site, which is the floor problem you have to clear first.
What it does is convert an unfair-feeling outcome into a set of tasks with observable results. One number is worth holding on to while you work through them. G2's 2026 research found that one in three B2B buyers ended up choosing a vendor they had not heard of before AI recommended it. The crowding at the top of these answers is real, and it is measurable, and it has a documented rate of leakage that runs directly in your favour.
Frequently asked questions
Why does ChatGPT only recommend big companies?
Being named is a retrieval outcome rather than a quality ranking. An engine assembles its answer from text it can find and cross-check, and a national brand has been described, listed, reviewed and linked in far more places than a ten-person specialist. Breadth of corroboration is what the model responds to. In our 15-region study the effect was stark: engines named 595 company entries in total, and 497 of them appeared in neither the region's own industry directory nor our web-search check.
Is AI deliberately biased toward large brands?
There is no size preference written into these systems. The bias is a side effect of how they are trained and how they retrieve. Large brands occupy more of the public text these models learn from and search across, so they surface more often and with more confidence. The practical result resembles bias closely enough that the distinction only matters for one reason: footprint is something a smaller firm can build, whereas a deliberate preference would be something you could do nothing about.
Can a small company realistically get recommended by AI?
Yes, through corroboration rather than scale. Profound's research found brands present across four or more platforms are about 2.8 times more likely to be cited, which means breadth of independent sources matters more than depth on any one of them. G2's 2026 research also found that one in three B2B buyers ended up choosing a vendor they had not known before AI recommended it, an outcome that requires smaller names to get in. The realistic route is to win narrow, qualified questions before attempting the broad category question.
Do more reviews make AI recommend me?
Reviews help, slowly, and they will not carry the load on their own. Kevin Indig's analysis using G2 data found roughly 10 percent more reviews associates with about 2 percent more AI citations. Our regional study showed the ceiling plainly: a British Columbia firm rated 4.9 with 15 reviews on the exact directory a buyer would consult was named by none of the four engines we asked. One strong profile on one platform is a single source, and single sources move AI answers weakly.
Why do Perplexity and ChatGPT give such different vendor lists?
They resolve the question differently. Perplexity retrieves live pages and summarises them, so its answers stay closer to what is currently published. ChatGPT, Claude and Gemini lean more on what the model already holds. Our study measured the gap directly: of the companies each engine named across 15 regions, 87 of Perplexity's 257 cleared an independent check, against 18 of 137 for Gemini, 8 of 114 for ChatGPT and none of Claude's 129. For a smaller firm this ordering is useful, because the retrieval-heavy engines respond to new published work first.
Which queries can a smaller firm actually win in AI answers?
The qualified ones. A broad category question such as best IT provider in Texas is where generic corroboration is thickest and national names cluster. Add a vertical, a company size, a compliance requirement or a specific city, and the pool of confidently answerable sources collapses, often to whoever has published something specific. A specialist that has done the work eleven times and written about it clearly can be the strongest available source for that exact question, and those questions carry higher buying intent than the broad one.
How long does it take to break into an AI answer?
It varies by engine, which is why sequencing matters. Search-grounded engines such as Perplexity, Copilot and Google AI Overviews re-read the live web, so new pages, listings and reviews can change their answers within weeks. Answers drawn primarily from training data shift on model release cycles, measured in months. Working on the sources the grounded engines retrieve gives you observable movement while the slower half compounds in the background.
Is this just SEO under a new name?
The two overlap without being the same game. SEO is the floor: an engine that cannot crawl or parse your pages will never cite them. AEO is a distinct layer, and you can win it while your Google rank stays flat, because an AI answer is assembled from agreement across many sources rather than from one ranked list. Our regional data showed the split directly, with 47 companies appearing in web search results but no AI answer, and hundreds appearing in AI answers with no web search presence at all.
Sources and further reading
- TofuBofu Research: Top MSPs in British Columbia: per-engine answers, the directory cross-check, and the methodology behind the figures in this post.
- TofuBofu Research index: the same open-question study across 15 US states and Canadian provinces.
- r/smallbusiness: tested if ChatGPT recommends my business: the founder quoted above.
- Profound research: brands present on four or more platforms are roughly 2.8 times more likely to be cited.
- G2 2026 B2B buyer research: 51 percent start on an AI chatbot, 69 percent switched vendor based on AI, one in three chose a vendor they did not previously know.
- Lewis et al., Retrieval-Augmented Generation (arXiv:2005.11401): how retrieval grounds a generated answer in source documents.
- SE Ranking 2026: 71 percent of ChatGPT-cited pages use structured data.
- Am I too early for AEO?: the same concentration effect measured across our wider scan data, applied to the question of timing.