Now live across the AI ecosystem: ChatGPT GPT Store · MCP Registry · mcp.so

Measurement

How many queries should you track for AEO? Is there a right number?

By Arnav Mukherjee, founder of TofuBofu · September 30, 2026

TL;DR

  • Run a set of at least ten on every engine. Below that, one answer flipping moves you 4 points.
  • Multiply before you compare. 50 prompts on one engine is not 50 on five.
  • Adding queries doesn't raise your score. It's a rate, so new queries move it either way.
  • Freeze the set. Changing it mid-programme breaks the only trend line you have.
  • Stop around 25. A category has a finite number of ways to ask the hiring query.

We watched two consecutive scans on the same company go from 8% to 0% and spent an afternoon looking for what broke. Nothing had. The two scans shared no queries at all, because we were regenerating the set each time, so we'd run two different tests and compared the results as though they were one.

That bug is fixed. Repeat scans now reuse the previous set by default, so a trend line survives. But the afternoon taught me the thing this whole page is about, which is that how many queries you track matters far less than whether the resulting number is capable of telling you anything at all, and that almost every piece of advice in this category treats the question as one of coverage when the thing that actually decides it is arithmetic.

What counts as one query?

One thing asked of one engine. Ask five engines the same thing and that's five queries, not one.

This sounds pedantic until you price it. Every visibility score you will ever be shown is a fraction, and the thing we're defining here is its denominator, which means a vendor can double the headline number they quote you without changing anything about what they measure. So when someone quotes a prompt count, there's exactly one useful follow-up. Across how many engines?

Two of those quote a bigger headline number than the third, and one of them measures a single engine, which means a buyer comparing the three on prompt count alone would rank them in almost exactly the wrong order for the thing they're trying to find out. Multiply first. Then compare.

Why is ten the floor, and 25 roughly the ceiling?

Because of what one answer changing does to your number. Your index is times recommended over queries asked, so the smallest possible move is one divided by the total. Here's that step size at each of our own plan sizes, on our current roster:

What one answer flipping does to your score, by set size. Plan figures read from the TofuBofu pricing page, 30 September 2026.
Queries per scan Across 5 engines One flip moves you
5 (our free tier)25 queries4.0 points
10 (Fix, $99)50 queries2.0 points
25 (Dominate, $499)125 queries0.8 points

Stated as a sentence: with a set of five on five engines, one answer changing moves your score four points, so a single engine having an off day looks exactly like a trend. At a set of 25 the same flip moves it 0.8 points and you can see real movement through the noise.

That's the honest case for the floor, and it applies to our own free tier as much as anyone's. Under ten you have a snapshot rather than a measurement. A free monthly baseline is genuinely useful for the one job of finding out whether you have a problem worth spending money on, and it is the wrong instrument to run a quarterly programme against, report to a board with, or hold an agency to. Use it for the first question. Buy resolution for the second.

More queries buy resolution, then stop buying anything queries in the tracked set noise a set of 5: one flip is a big move 10: usable floor 25: diminishing returns start past here, mostly paraphrases Conceptual. The step sizes are in the table above.

Why doesn't tracking more queries raise your score?

Because it's a rate, and a rate has two ends. Add queries you already win and it rises. Add queries you lose and it falls. Neither tells you anything about the world.

This has a consequence people discover too late: changing the set breaks the trend. If you expand from ten to twenty-five in month three, the month-four number isn't comparable to month two, and a drop could be you, the engines, or the seven new queries you happened to be losing. Freeze the set at the start of a programme. When you do change it, write down the date and treat the comparison as broken across that line.

See what five well-chosen queries tell you. A free TofuBofu scan puts real buying queries to all five engines and shows you the verbatim answers, including who got named instead of you.

Run a free scan

What belongs in the denominator, and what doesn't?

Three exclusions we make deliberately, and each one would flatter the number if we didn't:

Ask any vendor how they handle these three. The answers will tell you more about the quality of the number than the size of the tracked set will.

Where this advice breaks

Three cases, stated plainly, because a rule with no edges is a slogan.

Frequently asked questions

How many queries should I track for AEO?

Enough that one answer changing doesn't look like a trend, which in practice means a set of at least ten buying queries put to every engine you care about. The arithmetic is the whole answer. A query is one thing asked of one engine, so a set of ten run on five engines is fifty queries and your score moves in steps of 2 percentage points. A set of five on five engines is twenty-five queries and every step is 4 points. At that resolution a single engine having an off day reads as a collapse. A set of twenty-five on five engines is 125 queries and a step is 0.8 points, which is fine enough to see a real trend. Every figure here assumes the five-engine roster we sold on 30 September 2026; change the engine count and the arithmetic moves with it.

What counts as one query?

One thing asked of one engine. The same thing put to five engines is five queries, not one, and this matters because it's the denominator of every visibility number you'll be quoted. If a vendor sells you 50 prompts, ask across how many engines: 50 prompts on one engine and 50 prompts on six are different products at the same headline number. Profound's own pricing page describes its Starter tier as ChatGPT tracking only with 50 prompts tracked, and its Growth tier as three answer engines with 100 prompts, which is the distinction made visible in a vendor's own words.

Does tracking more queries improve my score?

No, and expecting it to is the most common misunderstanding of this metric. Your visibility index is a rate: times recommended divided by queries asked. Adding queries you already win pushes it up, adding queries you lose pushes it down, and neither movement means anything changed in the world. That's why the set has to be stable. If you add ten this month your trend line is broken and you won't know whether the drop is you or the measurement. Pick the set, then leave it alone, and add only with a clear note in the record that the comparison isn't clean.

Is there a point where more queries stop helping?

Yes, and it arrives sooner than most people expect, because the number of genuinely distinct buying queries in one category is finite. A managed IT provider in one city has maybe fifteen to twenty-five ways a buyer phrases the hiring query, and after that you're generating paraphrases that score the same pages against the same answers. Paraphrases cost money and add no information. The better use of budget past that point is more engines, more brands, or actually fixing what the existing set already told you. Our own top tier sells 25 queries per scan, and that is not a technical ceiling.

Should informational queries be in the tracked set?

Some, and not many, and they should be scored separately from the buying ones. A query like 'what is managed IT' tells you whether an engine will surface you during education. A query like 'best managed IT provider in Vancouver' tells you whether you're on the shortlist. Mixing them into one average hides the second inside the first, which is why we weight buying-intent queries at half the score on their own. Keep a few discovery queries for context and never let them dilute the number you're actually managing.

What about queries that name my own brand?

Keep them, score them apart, and never count them in your visibility rate. Asking an engine 'what is Acme Corp like to work with' will almost always produce your name, because you put it in the question. Including those in the same average inflates the number and hides the absences that matter. They're worth running because they tell you about sentiment and about factual errors an engine is repeating about you. They're just answering a different question, so we exclude brand-naming queries from the visibility calculation rather than letting them flatter it.

How many queries do the tools actually give you?

It varies more by engine count than by set size, which is why the headline numbers mislead. Profound's Starter is 50 prompts on ChatGPT only and its Growth tier is 100 prompts across three answer engines, both read from their own pricing page in August and re-verified in September 2026. TofuBofu sells 5 queries per scan on the free tier, 10 on Fix at $99 a month and 25 on Dominate at $499, each run on the same roster on every plan, read from our own pricing page on 30 September 2026. Multiply before you compare: the same prompt count on one engine and on five is a five-fold difference in what you actually learn.

Sources and further reading

Arnav Mukherjee

Founder, TofuBofu