Now live across the AI ecosystem: ChatGPT GPT Store · MCP Registry · mcp.so
Solutions / AI Visibility Scan

The measurement layer

See exactly where AI recommends you, and where it does not

You cannot fix what you cannot see. The scan asks all six engines the questions your buyers ask, stores every answer word for word, and records whether an engine named you or named somebody else. Everything else TofuBofu does starts from that record.

In every plan, including free ChatGPT Claude Perplexity Gemini Google AI Mode Microsoft Copilot

Every plan queries all 6. See how each engine picks who it names

What we measured

38 of 39

sites we crawled during a scan serve no llms.txt file, and those 39 are not a random sample: each firm ran a scan because someone suspected it was missing from AI answers.

Engines queried

6

ChatGPT, Claude, Perplexity, Gemini, Google AI Mode, Microsoft Copilot. Every plan, including the free one.

Answers kept

100%

Every answer is stored word for word, so any number on your report can be checked against the text it came from.

The gap this closes

  • No rank to check and no console to log into: an engine gives one answer, three names, and no page two.
  • You cannot tell an absence from a bad day, because one check on one engine is an anecdote.
  • The cause is usually plumbing, not marketing: 16 of 39 readable crawls carried no schema an engine could parse, on firms who already suspect they are missing.
  • Nothing tells you any of this is missing. There is no error message for being left out.

What you end up with

  • Every buying question run across all six engines, with each answer stored word for word.
  • A named list of who is being recommended in your place, per question, per engine.
  • The specific technical gaps behind the absence, ranked by what moves first.
  • A repeatable baseline, so next month's scan is a comparison and not a fresh guess.

How it works

Three steps, and you only do the first one.

1

You confirm the questions

Research reads your site and drafts the buying questions your market asks, localized when your market is local. You edit them before anything runs. The same confirmed set is re-asked on every later scan, so your trend line compares like with like instead of quietly comparing two different tests.

2

Six engines answer, and we read the text

ChatGPT, Claude, Perplexity, Gemini, Google AI Mode and Microsoft Copilot each get the questions. Every answer is stored in full and checked for your brand, including domain-style and spaced variants. An engine that returns nothing is recorded as no signal, not as a brand that went unnamed, because those two results are not the same result.

3

The gaps get ranked, then drafted

Buying questions carry the weight, so who-should-I-hire counts for more than what-is-this-category. What you end up holding is an ordered list of the questions where competitors are named and you are not. On a paid plan that list is the brief the content engine writes against.

What you get on each plan

A scan is a delivered artefact, not a dashboard you have to operate. This is what lands, per plan, with the limits read live from the plan configuration and not typed into this page.

Track

Free
5 questions 1 scan/mo
  • One scan a month, no card
  • The same report a paying account gets
  • Every engine's answer in full, plus who was named instead
  • Ranked list of the questions worth chasing first
  • Sentiment off, each question asked once
  • Shareable link that outlives the month

Fix

$99/mo
10 questions 4 scans/mo 10 drafts/mo
  • Weekly scans and a wider question set
  • Sentiment on, plus adaptive sampling
  • Every losing question becomes a finished draft
  • Publishes through your own CMS
  • Alerts when a score moves or a win is lost
  • Technical fixes listed separately from the writing

Dominate

$499/mo
25 questions 4 scans/mo 25 drafts/mo 10 scripts 4 newsletters
  • The full question budget, same six engines
  • Video scripts and newsletters off the same scan
  • Head-to-head detail per competitor
  • A person reads your scan, not a queue position

Rule

$1999/mo
25 questions 4 scans/mo 25 drafts/mo 10 scripts 4 newsletters
  • Portfolio level, several brands at once
  • Human strategy over the same agent production
  • Cadence and question budget set per brand
  • One view of which brands are getting named

The proof, including the parts that flatter nobody

We publish our own scan numbers, which is a habit worth being suspicious of until you can see the method. Across the 46 completed reports behind this page, Perplexity named the brand in 7.4% of the buying questions we put to it and every other engine landed between 0.4% and 4.0%, on a set of firms that suspected they had a problem before they ran anything.

The engines are also less settled than a single check suggests. In July we asked four of them to name managed IT providers across 15 North American regions, resolved every name they produced to a real website, and could corroborate 115 of 566. That is one profession over one fortnight, not a law of nature, and the full method, including what we could not verify, sits on the research page.

Read the MSP AI Visibility Report 2026 →

A worked example: one row of the record

A report is a table of questions by engines, and this is one row of it. The brand and the answer text below are illustrative, written to show the shape. The fields are the fields the report stores.

Question: best managed IT provider for a healthcare clinic in Ohio

EngineVerdictNamed insteadEvidence
ChatGPTAbsentVendor A, Vendor B, Vendor CFull answer stored, 1,840 characters
ClaudeAbsentVendor A, Vendor DFull answer stored, 1,210 characters
PerplexityNamed, briefVendor A, Vendor BNamed in a list of five, no detail
GeminiAbsentVendor BFull answer stored, 990 characters
Google AI ModeAbsentVendor A, Vendor EAnswer plus 19 reference links
Microsoft CopilotNo signalEngine returned nothing, dropped from the denominator

What the row hands to the rest of the product

  • A buying question you lose on five of the six engines that answered, which ranks it high on the fix list.
  • Vendor A, named by four engines, which promotes it to a tracked competitor whether or not you declared it.
  • The verbatim answers, so a claim about your positioning can be quoted back with the sentence it came from.
  • A content brief aimed at this exact phrasing, drafted for you on a paid plan and listed as an idea on the free one.
  • A row in the next scan's comparison, because the question is held still and re-asked.

Multiply that by the questions in your plan and the six engines, and you have the record the score is computed from. The score is the summary. The row is the evidence, and every claim on the report traces back to one.

How it actually works

The mechanism, for anyone who wants it. Open a card to read the detail.

What an engine is doing while it answers a buying question

Two different machines wear the same chat window. One answers from parametric memory, the weights it learned during training, which is why it can produce a confident vendor name with no page behind it. The other retrieves documents first and writes its answer from what it just fetched. Most production systems blend the two, and the blend is what decides whether your website has any say in the answer.

Read the detail

The retrieval half is the half you can win. It runs a search, pulls a handful of candidate documents, and grounds the response in their text. A page that states a buyer's question and answers it in the next sentence is cheap for that step to use. A page that opens with three paragraphs of positioning before it says anything checkable is expensive, so the system reaches past it and quotes somebody else.

The memory half explains the strangest thing our own research turned up. Asked to recommend managed IT providers across 15 North American regions, one engine produced 129 firm names and 2 of 129 resolved to a website we could find, and six of the names in that managed IT study left no trace of existing anywhere. Being fluent and being grounded are separate properties, and a scan tells you which one you are up against on each engine.

Source: Lewis et al. 2020, Retrieval-Augmented Generation (arXiv:2005.11401). The foundational paper describing how models retrieve from an external, non-parametric memory to generate factual, sourced answers.

Why one check is an anecdote and six are a measurement

Ask the same question twice and a model can give you two different vendor lists. That variance is the reason a screenshot proves nothing, and it is the reason the paid tiers ask each question twice and pay for a third call when the first two disagree. A verdict that survives disagreement is worth recording. A verdict from a single lucky draw is not.

Read the detail

The other half of the discipline is what we do with silence. An engine can return no answer at all, and a system that scores silence as absence quietly invents bad news. Every engine that returns nothing leaves the denominator instead of counting against you, and a question with too few answers to judge says so on the report instead of showing a confident zero.

The last piece is holding the questions still. Regenerating the question set between scans is the fastest way to destroy a trend line: two consecutive scans of the same firm once shared no questions at all, and an 8% reading followed by a 0% reading looked like a collapse when it was two different tests. Repeat scans re-ask the set you confirmed, add to it when your budget grows, and trim from the end when it shrinks.

What the crawl sees that no answer can tell you

Every scan also fetches your site, walks your sitemap, and records what a machine reading it would find: which schema types you serve, whether an llms.txt file exists, which topics you have already published. That crawl is how a gap in the answers turns into an instruction you can act on this afternoon.

Read the detail

It is also the part that has to be honest about its own failures. A site that refuses our crawler used to be recorded as a site with nothing on it, which produced advice to add markup the firm already served. One domain answered our crawler with an HTTP 444 and a browser with a 200, so a report would have told a real prospect they published nothing while their sitemap listed 676 URLs. The crawler now identifies itself with a contact address, and a site we could not read is reported as unread instead of empty.

What it costs

The prices and limits below read from the live plan configuration, so this page cannot drift from what your account gets.

Plan Price Questions Scans
Track Free 5 1/mo
Fix $99/mo 10 4/mo
Dominate $499/mo 25 4/mo

Rule is portfolio level and sales-assisted. Talk to us and we will scope it against the number of brands you carry.

See where you stand first

Run a free scan and find out which of these gaps you actually have.

Get your free audit

Frequently asked questions

Which engines does the scan cover?

All six: ChatGPT, Claude, Perplexity, Gemini, Google AI Mode and Microsoft Copilot. Engines are not plan-gated, so a free scan asks the same six a Dominate scan asks. What the plan changes is how many questions you get, whether sentiment runs, and how many times each question is sampled.

Why does my scan differ from what I see when I ask ChatGPT myself?

Because your chat is one sample from a system that varies, and it knows you. Your history and your custom instructions both shape the answer you personally get. The scan asks neutral buying questions from a clean context, across engines, and verifies the answer text server side, so it reflects what a stranger sees.

How is the TofuBofu AI Visibility Index calculated?

It is the share of buying questions, across every engine that answered, where your brand appears in the answer text. Every question in a scan is a buying question, so there is nothing else in the denominator, and an engine that returned nothing leaves the denominator instead of counting against you. Questions the engines answer without naming any vendor at all leave it too, because an answer that recommends nobody is not an absence you could have prevented.

Does running a scan change my AI visibility?

No. The scan observes. It asks questions the way a user would and records what comes back, and it never touches your site or signals an engine. The Index moves when the underlying reality moves, which is the only reason it is worth trusting.

Does a low score mean my SEO is bad?

Not on its own. Firms that hold page one on Google go missing from ChatGPT all the time, because the retrieval layer rewards structure and corroboration, not rank. The scan measures the AI layer on its own terms and tells you which of the two you have a problem with.

What do I do with the result?

Start with the buying questions where competitors are named and you are not, in the order the report ranks them. Each one maps to a page we can draft, and the plumbing gaps the crawl found are an afternoon of work, not a project.

Sources and further reading

Related solutions

Blog Content FAQ Schema Pages Comparison Pages