Tools
Cloudflare's AEO dashboard does not ask about you. It asks about your category.
By Arnav Mukherjee, founder of TofuBofu · August 14, 2026
On 17 August 2026 I opened our own scan screen and read a line that said the brand in front of me sold into one country, followed by a list of seven. Saudi Arabia, the UAE, Iraq, Indonesia, Vietnam, Jordan, the Philippines. One country, seven names, printed side by side.
The cause was category inference. A regex matched the firm's industry, the industry mapped to a market scope, and the code that set the scope never looked at the office list sitting next to it. The label was the small half of the problem. The large half was that market scope decides which buying questions we generate, so the firm had been scored on questions for a market it half sells into. We shipped the fix that day. What it cost while it was live was not a wrong word on a screen. It was the wrong questions.
Eleven days earlier, Cloudflare had shipped a product whose entire design rests on getting that same inference right, once, and then reusing the answer across everybody in the category. This post is a read of what they built, from their own documents, because they published more real methodology in one engineering post than most of this category has published in a year.
What Cloudflare actually shipped
The AEO Visibility Dashboard went out on 6 August 2026, in early access, requested from the Overview tab of the Cloudflare dashboard. It reports four numbers. Citation Rate, defined by Cloudflare as "the share of answers in your category that cite your site as a source". Prominence, how much of an answer is yours and how early it lands. Mention Rate, how often assistants name you whether or not they cite you. Share of Voice, your slice of citations against your competitors. It sits beside Agent Readiness, the existing technical check, in what Cloudflare calls its AEO Suite: one tool asks whether an agent can read your site, the other asks what the assistants say once it can.
Give them the asset first, because it is real and no competitor can rebuild it. Cloudflare sits between AI platforms and the sites they fetch. Their AI Operator Activity view shows real crawl and referral traffic per operator, including the errors the crawlers hit, and their blog names the pattern worth acting on: the operator that crawls thousands of your pages and refers nobody back. That is first-party infrastructure data. It is better than anything the rest of us can assemble, and it is the reason their own count that fewer than half of all HTML page requests now come from a human is worth more than an estimate from someone reading their own access logs.
Their framing of the problem is also correct, and they said it better than we usually do: there is no impression count and no missed-click report, so when a competitor gets named instead of you, the sale is gone and nothing tells you it happened. That is the whole reason this category exists.
Everybody samples prompts. Judge the sampling.
The press release leads on network data, and it is more careful than a skim of it suggests. It says most tools fill the gap by "only sending test prompts to AI chatbots and sampling the responses", and that Cloudflare's data creates "deeper, more transparent insights than sampling test prompts alone". The words only and alone are load-bearing. Cloudflare is not claiming it avoids prompt sampling. It is claiming it pairs prompt sampling with data nobody else has, which is true, and which is a fair thing to lead a launch with.
The engineering blog, published the same day by Matthew Conroy and Jack Galilee, says the rest of it plainly: "we probe the leading assistants (today, Anthropic's Claude and OpenAI's GPT) with likely customer prompts to see how they respond". There is no other way. Crawl data tells you an engine came to your door. Referral data tells you somebody arrived after clicking. Neither tells you what the assistant said about you to the buyer who never clicked, and that conversation is where the shortlist forms. G2's 2026 buyer research puts 51 percent of B2B buyers starting vendor research on an AI chatbot, up from 29 percent, with 69 percent changing a vendor choice on what AI told them. Most of that influence leaves no trace in a log file.
So the interesting question was never whether a vendor asks the models. Everyone asks. The question is which models, on whose questions, how many times, and whether the same questions come back next month. Cloudflare answers three of those four in public, which is three more than most of this category manages, and the answers are worth reading closely.
One panel per category, reused across every account
Here is the design, in their words. Cloudflare infers your industry and category from your site, using health and fitness as the industry example and sports apparel as the category. It then builds a benchmark for that category by querying the assistants with likely prompts, and this is the phrase that should change how you read the number: "without specifying your brand". Then: "Rather than re-querying models every time a site owner runs a scan, we run this panel once per category and reuse the baseline across all accounts in that domain."
They give three reasons and all three are good ones. Zero latency, because results load from a snapshot. Lower compute, because aggregating at category level avoids redundant model calls across thousands of scans. And an Industry Fit score, which falls out of reusing one corpus: mapping which brands keep appearing together tells you "whether an AI assistant views your site alongside your actual competitors". That last one is the sharpest idea in the announcement and it is close to the finding we consider our own most useful, that the competitor set AI frames you against is often not the one you declared.
Now hold that design against how the same launch is described elsewhere. The press release says Share of Voice shows how a brand stacks up "across the specific questions its customers are asking AI assistants". The engineering blog says "Because these numbers are specific to your site, you can experiment, re-run the scan, and measure the impact on the exact questions that bring you business."
Both sentences are defensible about the output. Your Citation Rate genuinely is your number, computed for your domain. Neither is true of the input. The questions are not the specific ones your customers are asking, because they were never asked about you and were not chosen for you. They are the questions Cloudflare considers likely for a category it assigned you to, asked once, reused across you and every rival in that category. Those are two different products. A number that tells you where you stand inside a category is useful. It is not the same as knowing what came back when a buyer asked the question that actually closes your deals.
Two designs, two failure modes. A shared panel breaks when the category is inferred wrongly, and every number downstream inherits the error. A per-brand set breaks when the questions are not the ones buyers ask, which is why ours are confirmed with the founder before the first scan runs.
The strongest case for the shared panel, and where it stops working
Take the counter-argument seriously, because it is a good one. A category baseline is cheap, it returns instantly, it costs the user nothing at the point of use, and it buys something a per-brand measurement genuinely cannot: real comparability. If your Share of Voice is computed from the same answers as your rival's, the gap between you is a clean gap. Run two brands through two separately generated question sets and part of the difference between them is the question sets. Cloudflare picked comparability and speed. That is a defensible choice, not a shortcut, and for a sports apparel brand in a dense, well-populated category with a stable vocabulary, I think it is close to the right one.
It also happens to get one thing right that most of this category gets wrong. Cloudflare queries the assistants "without specifying your brand". That is the correct discipline and it is not a small detail. We had to learn it the hard way: our own published engine rates were inflated by roughly 2.6 times for a period because the denominator included probes that contained the company's own name, and every engine names a brand back to you when the name is sitting in the question. A panel that never names you cannot make that mistake.
So here is the honest boundary. The shared panel works where the category is dense, the vocabulary is stable, and the buying question does not carry geography. It degrades exactly where B2B services live. Two managed IT firms can sit in the same inferred category while one sells to hospitals across a country and the other sells to dentists in one city, and the question that decides each of their deals contains a word the panel does not know about. The panel will still return a confident number for both.
And category boundaries in this domain are not a hypothetical failure. In our fifteen-region managed IT study, run over two days in July 2026 with every name resolved to a real website, the most frequently named provider was Symmetrio, named in eight of fifteen regions by one engine and no other, with a different head office each time. Symmetrio is a real company and a good one. It is a staffing and recruiting firm, not a managed service provider. The same study returned a telecom carrier, a teleconferencing hardware maker and a national technology reseller as answers to a managed IT question.
Be precise about what that shows. Those are facts about how an engine answered, not accusations against any of those firms, and they are not evidence about Cloudflare's classifier either. They are models drawing category lines from vocabulary rather than from what a company sells. But it is the same job Cloudflare's inference step has to do, on the same messy vocabulary, and it has to do it once for everyone in the bucket. When our own inference broke on 17 August, one brand got the wrong questions. When a shared panel's inference breaks, everyone assigned to that panel gets them.
Two assistant families, and the one they left out
Get the count right first, because the wording invites a misread. Cloudflare probes two assistant vendors, Claude and GPT, and prompts "each assistant multiple times across different models" through Cloudflare AI Gateway. That is two families across several models, not two models, and the word today in their sentence signals the list is expected to grow. Not named: Perplexity, Gemini, Google AI Mode, Microsoft Copilot.
Perplexity is the omission that matters, and I want to make that case from the right evidence rather than the convenient one. Across 46 completed reports and 2,421 answered buying-question cells measured on 12 August 2026, on sites whose owners already suspected they were missing from AI answers and which are therefore not a random sample of the web, Perplexity named the brand on 7.4 percent of buying questions, Google AI Mode on 4.0, Copilot on 3.1, ChatGPT on 1.7, Claude on 1.2 and Gemini on 0.4.
Two of those six cannot carry an argument, and I would rather say so than let a reader lean on them. ChatGPT's 1.7 percent measures an instrument we retired on 18 August 2026: until then our ChatGPT probe answered from training weights with no web search, so that figure describes what a model remembered, not what ChatGPT finds when it retrieves. It is left in place because it is what we measured and we do not restate history we have not re-run, but no claim about ChatGPT's behaviour today may rest on it, including mine. Gemini's 0.4 is a floor rather than a reading, because our own token budget truncated its answers for part of that window.
Strip both out and the argument is still there, because it never needed them. Perplexity leads the engines we can read cleanly, and the reason is mechanical: it retrieves before it answers instead of recalling. In that same fifteen-region managed IT study in July 2026, 101 of the 257 firm names Perplexity produced could be resolved to a real company and corroborated independently, against 2 of 129 for Claude. A retrieval-first engine names firms that exist, and it names the ones that published something findable. That is the engine most likely to reward a site that just fixed its content, which makes it the engine you most want in the panel when you are trying to prove a fix worked.
None of that makes Cloudflare's number wrong. It makes it narrow in a checkable direction, and narrow in the direction that most understates a site that has done the work recently.
See your own questions, on six engines, before you pick a tool
Confirm the buying questions your buyers actually ask, then read every engine's verbatim answer and the rival it named instead of you. Free, no card.
Get your free auditWhat they fix, and the half nobody fixes
It would be easy and wrong to say Cloudflare only measures. Their Agent Readiness side ships remediation that we cannot match and should not pretend to: every failed check comes with a next step, and where a Cloudflare setting solves it there is a link straight into that setting, so managed robots.txt or serving clean Markdown to agents is one click on infrastructure they already run. Everything else gets a button that hands the job to your coding agent. Make the change, re-scan, watch the check go green. If your problem is that agents cannot read your site, the company that terminates your traffic is genuinely the best-placed vendor on earth to fix it.
That is the plumbing half, and it is a solved problem now. The other half is not, and this is where three heavyweights have now stopped at the same line. Adobe completed its acquisition of Semrush on 28 April 2026 and said out loud it was buying brand visibility, generative engine optimization and agentic search optimization. Ahrefs shipped Brand Radar. Cloudflare shipped this. Three serious companies have concluded the category is real. Not one of them writes the comparison page that answers the question you are losing, puts it live on your CMS, and re-asks that same question a month later to show the answer changed.
Which matters more once you notice what the pre-computed panel does to proof. Cloudflare's blog says you can re-run the scan and measure the impact. It does not say how often the category panel is refreshed, and that is the one methodological fact missing from an otherwise unusually open post. If the baseline is a snapshot shared by everyone in your category, then publishing something excellent on Tuesday cannot move your number until that snapshot is rebuilt, whenever that is. A per-brand design has the opposite property and the opposite cost: it is slower and it cannot compare you cleanly to a rival, but the questions you were losing last month are the questions asked again this month, so movement means something.
Measurement is commoditising and it will keep commoditising, because it is a few API calls away from anyone with an engineering team, and the three companies above just proved how few. Knowing your Citation Rate is 4 percent tells you nothing you can act on by Friday. The work is the fix and the proof. A dashboard is the receipt.
The unanswered question is the price
Nobody outside Cloudflare knows what this will cost or which plan it lands on. I checked on 20 August 2026: the public plans page lists Free at zero, Pro at 20 dollars a month billed annually, Business at 200 dollars a month billed annually, and Enterprise on request, and it does not mention AEO or Agent Readiness anywhere. The AEO tag on their blog still holds exactly one post, the 6 August one. Their developer changelog carries no AEO entry since. Two weeks in, nothing has moved, which is normal for early access and worth knowing if you were planning around it.
Their own press release is the most precise source on this, in the small print rather than the headline: the forward-looking statement lists the timing of general availability for the AEO Suite and its features as an open uncertainty. So if you are budgeting for AI visibility measurement over the next two quarters, treat this one as priced unknown rather than as bundled free. That is not a criticism. It is the difference between a plan and a hope.
Four questions to ask any AI visibility vendor, ours included
Which engines, by name, and how many did you leave out. How many times was each question asked, because a single run of a non-deterministic system is an anecdote rather than a measurement. Whose questions were these, mine or my category's, and who chose them. And are the same questions asked again next month, because a score computed on a fresh question set every time cannot produce a trend line, only the appearance of one.
Cloudflare answers the first three in public and the fourth is the one their post leaves open. Two assistant families across several models, prompted repeatedly through AI Gateway, on a category's questions chosen by Cloudflare, against a baseline whose refresh cadence is unstated. Ours are six engines, adaptive sampling on paid plans, your own buying questions confirmed with you before the first scan runs, and the same questions re-asked every time so the line means something. We answer the fourth and we lose on the first three counts of scale, comparability and instant results.
Pick whichever fits the decision in front of you. If you want to know how your category behaves and where you sit in it, Cloudflare is about to make that nearly free, and I would take it. If you need to know which specific question lost you a deal, read the answer the buyer read, and prove a month later that the fix moved it, a shared category baseline structurally cannot tell you. Just never accept a number from anyone who will not answer all four.
Frequently asked questions
What is Cloudflare's AEO Visibility Dashboard?
It is a measurement product Cloudflare released on 6 August 2026, in early access, requested from the Overview tab of the Cloudflare dashboard. It reports four metrics: Citation Rate, which Cloudflare defines as the share of answers in your category that cite your site as a source, plus Mention Rate, Prominence and Share of Voice. It joins Agent Readiness, an existing technical check, in what Cloudflare calls its AEO Suite. Agent Readiness asks whether an agent can reach and read your site at all. AEO Visibility asks what the assistants say once it can.
Does Cloudflare measure AI visibility from network data or by asking the assistants?
Both, and its press release is more careful about it than a skim suggests. It says that most tools fill the gap by only sending test prompts and sampling the responses, and that Cloudflare data gives deeper insight than sampling test prompts alone. The words only and alone are doing real work there. Cloudflare is not claiming it avoids prompt sampling. Its engineering blog says plainly that it probes the leading assistants with likely customer prompts. Crawl and referral data are genuinely Cloudflare's to see and nobody else can assemble that. What an assistant recommends is only knowable by asking it, so everyone in this category asks, and the question worth arguing about is how.
Which AI assistants does Cloudflare probe?
Two families, per its own blog: Anthropic's Claude and OpenAI's GPT, and it prompts each of them multiple times across different models through Cloudflare AI Gateway. So it is two assistant vendors rather than two individual models, and the word today in that sentence signals the list is expected to grow. Perplexity, Google Gemini, Google AI Mode and Microsoft Copilot are not named. Perplexity is the omission that matters most, because it retrieves before it answers rather than recalling from training weights, and retrieval is what separates a named firm that exists from a named firm that does not.
What does it mean that the panel is pre-computed per category?
Cloudflare states that rather than re-querying models every time a site owner runs a scan, it runs the panel once per category and reuses the baseline across all accounts in that domain. It gives three reasons: results load instantly from a snapshot, aggregating at category level avoids redundant AI calls across thousands of scans, and reusing one corpus lets it derive an Industry Fit score from which brands keep appearing together. The trade is real and it is well argued. It also means the questions behind your number were chosen for a category you were assigned to, not for your business, and that two rivals in the same inferred category are scored against the same underlying answers.
How does Cloudflare decide what category my site is in?
It infers it. The blog says Cloudflare infers your industry and category from your site, and offers health and fitness as an industry example with sports apparel as a category. Inference is the single point of failure in any category-level design, because everything downstream inherits it and a wrong category still produces a confident-looking number. We know the failure mode from the inside: our own scan screen spent a period telling multi-market firms they sold into one country, because an industry regex set the market scope and never read the office list. We shipped the fix on 17 August 2026.
Which Cloudflare plan will include the AEO Visibility Dashboard?
Cloudflare has not said, and as of 20 August 2026 its public plans page does not mention AEO or Agent Readiness at all. That page lists Free at zero, Pro at 20 dollars a month billed annually, Business at 200 dollars a month billed annually, and Enterprise on request. The press release places AEO Visibility in early access with no general availability date, and its own forward-looking statement lists the timing of general availability as an open uncertainty. If you are budgeting for AI visibility measurement in the next two quarters, treat the price of this one as unknown rather than as free.
Does Cloudflare account for AI answers changing between runs?
Yes, and it is the most credible paragraph in the whole announcement. Cloudflare states that AI assistants rarely answer the same question the exact same way twice, and that it uses Cloudflare AI Gateway to prompt each assistant multiple times across different models to account for that variance. It also says it uses exact text analysis rather than a model grading its own output. Both are the correct instincts and most of this category quietly skips them while presenting one confident score. What the post does not say is how often the category panel is refreshed, which is the one methodological fact missing from an otherwise unusually open piece of writing.
Sources and further reading
- Cloudflare engineering blog, 6 August 2026: Matthew Conroy and Jack Galilee on the methodology in Cloudflare's own words, including probing Claude and GPT, inferring industry and category from the site, querying without specifying your brand, running the panel once per category, Industry Fit, and prompting each assistant multiple times through AI Gateway. Every quotation in this post is from this page or the press release below. Read on 20 August 2026.
- Cloudflare press release, 6 August 2026: the four metrics, early access from the Overview tab, the network-layer framing with its only and alone qualifiers, and the forward-looking statement naming general availability timing as an uncertainty.
- Cloudflare plans and pricing: checked 20 August 2026 for AEO Suite availability by tier. Free, Pro, Business and Enterprise are listed; AEO and Agent Readiness are not mentioned.
- Adobe completes its Semrush acquisition, 28 April 2026: Adobe's own framing of the deal around brand visibility, generative engine optimization and agentic search optimization.
- TofuBofu managed IT visibility study, July 2026: fifteen North American regions, four engines, every name resolved to a real website. Per-engine grounding from 101 of 257 for Perplexity down to 2 of 129 for Claude, and the Symmetrio case.
- G2 2026 AI Search Insight Report: 51 percent of B2B buyers begin vendor research on an AI chatbot, up from 29 percent, and 69 percent changed a vendor choice based on what AI told them.