Reference guide
AI visibility: what it is, how to measure it, and how to improve it
A TofuBofu reference guide, maintained from our own measurement.
The short version
AI visibility is whether AI assistants name your company when a buyer asks for a recommendation. It is measured as a mention rate per engine per question, not as a rank and not as traffic. Measuring it has become close to free. Changing it has not.
The definition, precisely
A buyer opens ChatGPT, Claude, Perplexity, Gemini, Google AI Mode or Microsoft Copilot and asks something like "best managed IT provider for a 40 person law firm". The engine replies with three to five named companies. AI visibility is the answer to one question: are you one of them, and on how many of the questions that matter.
Three things it is not, each of which people confuse it with:
The discipline of improving it is called answer engine optimization, or AEO. AI visibility is the measurement. AEO is the work.
What a good score looks like, and why the question is a trap
Everyone wants a benchmark number. There is not an honest one, and here is the data that shows why.
We went through every completed scan in our database and pulled 3,538 question-and-engine cells: one cell per time we put one question to one engine for one real company and recorded whether that company was named. Mention rate depends almost entirely on what kind of question was asked.
Read the fourth row again, because that is the one that costs money. On the question a buyer actually asks when they are choosing a vendor, the companies we scanned appeared 5.4 percent of the time. Ninety-five times in a hundred, the answer named somebody else. A high score on questions that already contain your name is not reassurance, it is the engine repeating what you told it. The only comparison worth making is against the companies named on your own buying questions, and your scan already produces that list. Ignore any vendor quoting an industry-average score, because the average depends on a question mix nobody has standardised. Why the buying question carries the weight, and why it is now answered at the top of the funnel rather than at the bottom, is set out in the TOFU, MOFU, BOFU guide.
The four measurement traps
The blended score
Averaging six engines into one number destroys the only actionable signal in the report. Engines disagree hard on specific questions: across five metros we collected 136 firm-and-market entries from four answering engines and 85 percent were named by exactly one engine, with not one firm appearing on all four. A hard zero on a single engine is a specific, fixable finding. Averaged, it becomes a slightly lower number nobody acts on.
Silence scored as absence
Engines time out, refuse and return empty responses. That is no evidence about you in either direction, and it must be recorded as its own state and dropped from the denominator. This matters most for Google AI Overviews, which frequently renders nothing at all on commercial queries: in our data its median stored answer length on those queries is zero characters. Scoring that as 'not mentioned' invents a visibility problem you do not have.
A question set that changes
If the questions regenerate each month, this month and last month are two different experiments and the line between them is decoration. We watched two consecutive scans of one brand share zero questions, turning an 8 percent result into a zero. It looked like a collapse. It was two unrelated measurements.
One sample per question
Answers are non-deterministic. Put the identical question to one engine twice and the shortlist can come back different, with nothing altered at your end. A single run reported to two decimal places is a coin flip wearing a percentage sign.
Measurement, and the two places it breaks
What the tooling costs, and why that is the wrong question now
Tracking has commoditised faster than almost anyone expected. Paid tools start around fifty dollars a month, several vendors give away a free checker, and in March 2026 G2 created a dedicated software category for AI search visibility. A capability gets its own category page and a free tier in the same year when its price is heading toward zero.
That is good for buyers and it moves the real question. If measurement is nearly free, what you are actually paying for is everything after the number: whether the tool explains which of forty gaps is worth an afternoon, whether it produces the asset that closes one, and whether it can re-run the same questions and prove the answer changed.
Before paying anyone, run the checks in how to check whether your AI visibility report is telling you the truth. It takes twenty minutes and it separates a tool that measured something from one that generated something.
Measure yours, free, across all six engines
Every engine's verbatim answer to every buying question, including who got named instead of you. No card, and you keep it monthly.
Get your free auditDoes it actually produce revenue?
The best public dataset is Orbit Media's analysis of 97 B2B GA4 accounts, 28.9 million sessions from July 2025 to June 2026. AI-referred visitors were roughly 3 times more likely to convert into leads than other organic traffic, and 7 times on the per-site median. ChatGPT drove 82.3 percent of that AI traffic and converted 2.08 percent of visitors, against 0.5 percent for Google Search.
And the volume is tiny: 0.5 percent of all traffic, about one visit in 200. The authors are explicit that the true figure is higher, because Google AI Mode and AI Overviews are recorded as organic search in GA4 and assistant apps land in direct.
The honest position: today this is a small, exceptionally high-quality channel that is systematically undercounted, sitting in front of a buying behaviour that is not small at all. G2's 2026 research found 51 percent of B2B buyers now begin vendor research on an AI chatbot, up from 29 percent, and one in three ended up with a vendor they had not heard of before the AI named it. The mention matters even when the click never happens. More on that in cited by AI but getting no clicks.
How to improve it
Ordered by leverage per hour. The full version, with the evidence for each, is in the AEO guide.
Confirm the crawlers can read you at all
Binary, and more often broken than anyone expects. Of 28 distinct domains we tried to crawl, 11 returned zero pages, and their websites were fine in a browser. If a crawler cannot fetch you, nothing else on this list matters and nothing tells you.
Get corroborated by sources that are not you
The biggest lever by a distance. Engines repeat what independent sources agree about you. Directories, review sites, industry rankings, roundups.
Describe yourself identically everywhere
Category, geography, buyer size and specialism, consistent across your site, your directory profiles and your review-site descriptions. When sources contradict each other, an engine has no confident description to repeat, and it names somebody else. An afternoon of work.
Put the facts an engine needs into readable text
Across 64 real answers we categorised, engines described the vendors they named using awards and reviews 95 percent of the time, geography 88, compliance 73, vertical specialism 72, size-fit 62 and price 38. A fact that lives only in a map embed or a badge image does not exist.
Answer each buying decision on its own page
What it costs, who it is for, how it compares to the obvious alternative. One strong page per decision rather than three thin pages for three phrasings of the same one.
One thing worth saying plainly, because the category avoids it: being findable on Google is the floor, not the goal. Of 228 firms listed in their own regional directories, our engines named 89 and skipped 139. Rank is necessary and nowhere near sufficient, because an answer is decided on corroboration.
Frequently asked questions
What is AI visibility?
AI visibility is whether AI assistants name your company when someone asks them a question your business could answer. It is measured as a mention rate: of the buying questions you care about, across the engines your buyers use, on what share does the answer include you. It is not a ranking, because generated answers have no positions, and it is not traffic, because most AI answers produce no click at all. A company can have excellent SEO, a strong brand and zero AI visibility, and the only way to know is to ask the engines and read what comes back.
How do you measure AI visibility?
Put a fixed set of buying questions to each engine, sample each question several times because answers are non-deterministic, store the verbatim answers, and count the share where your brand is named. Four rules keep it honest: keep the question set stable between checks or your trend line is meaningless, record an engine returning nothing as its own state rather than as an absence, report per engine rather than blending, and always keep the raw answer so any verdict can be checked against its own evidence.
What is a good AI visibility score?
There is no universal good number, and any vendor quoting one is inventing it. Mention rates depend enormously on question type. In our own data across 3,538 measured question-and-engine cells, brands were named on 56.9 percent of questions that asked about them by name, 5.4 percent of category buying questions, and 0 percent of awareness and informational questions. The number that matters is the middle one: on the question a buyer asks when choosing a vendor, the answer named somebody else 95 times in 100. A strong score on questions containing your own name is the engine repeating what you told it, not evidence you are being recommended. Compare yourself against the competitors named on your own buying questions, not against a benchmark.
Why should I not trust a single AI visibility score?
Because it averages away the only actionable signal. Engines disagree violently on specific questions: across five metros we collected 136 firm-and-market entries from four answering engines and 85 percent were named by exactly one engine, with no firm appearing on all four. A blended number turns a hard zero on one engine, which is a specific and fixable finding, into a slightly lower average. Always get the per-engine breakdown before you interpret the headline.
How much do AI visibility tools cost?
Tracking has commoditised fast. Paid tools start around fifty dollars a month, several vendors offer a free checker, and in March 2026 G2 created a dedicated software category for AI search visibility. Prices climb into the hundreds for higher prompt volumes and daily rather than weekly refreshes. The important buying question is no longer what a tool measures, since measurement is close to free, but what it does after the measurement: whether it explains the result, whether it produces the fix, and whether it can prove the answer changed.
Does AI visibility actually drive revenue?
The traffic it produces is small and unusually high quality. Orbit Media analysed 97 B2B GA4 accounts covering 28.9 million sessions and found AI-referred visitors roughly 3 times more likely to convert into leads than other organic traffic, 7 times on the per-site median, while AI accounted for just 0.5 percent of all traffic. So it is currently a small, excellent channel that is undercounted, because Google AI Mode and AI Overviews are recorded as organic search and assistant apps land in direct. Mentions that never produce a click still matter, since the buyer who reads a recommendation and searches your name later arrives as direct traffic.
How do I improve my AI visibility?
In order of leverage: make sure the crawlers can read you at all, get corroborated by independent sources, describe yourself consistently everywhere, put the facts an engine needs into plain text, and answer each buying decision on its own page. Corroboration is the big one, because engines repeat what other sources agree about you rather than what you claim. Structured data is a real floor and a day of work, not a strategy.
Sources and further reading
- TofuBofu, the AEO guide: the discipline behind the measurement, with the full evidence base.
- TofuBofu, AI visibility for managed IT providers, 2026: the 15-region study behind the 228 directory firms and the 139 unnamed.
- TofuBofu, five-metro engine disagreement study: 136 entries, 85 percent named by exactly one engine.
- TofuBofu, what engines say about the vendors they name: descriptive language across 64 answers.
- Orbit Media, AI traffic conversion rates: 97 B2B GA4 accounts, 28.9M sessions, 3x and 7x conversion, 0.5 percent of traffic.
- G2 2026 B2B buyer research: 51 percent begin on an AI chatbot, up from 29 percent; one in three chose a vendor unknown before AI named it.