Now live across the AI ecosystem: ChatGPT GPT Store · MCP Registry · mcp.so

Measurement

How do you identify the right queries to track for AI visibility?

By Arnav Mukherjee, founder of TofuBofu · September 30, 2026

TL;DR

  • Don't build the set from a keyword tool. It reports zero volume for the questions that matter.
  • Weight it to the purchase. Buying-intent queries carry half the score; discovery carries 20%.
  • Write paragraphs, not keywords. Real queries we see state a job title and a situation.
  • Settle how you sell first. A local firm measured nationally looks invisible when it isn't.
  • Build comparison queries from the rivals the engines name. Your own list will differ.

Early on we put our buying questions into Google Keyword Planner to size them. It returned zero. Not low volume, not insufficient data, but a flat zero for a list of questions we could watch real buyers ask us on sales calls that same week, and because the tool was the one every SEO process in the world starts with, we came close to deleting the entire set and starting again on keywords that scored better.

The tool wasn't broken. It was answering a different question: how many people type this exact string into a search box. Nobody types a paragraph into a search box. They type it into an assistant, and no assistant reports volume to anyone, which means the entire measurement apparatus the SEO industry spent twenty years building is pointed at a behaviour that is quietly moving somewhere it cannot see. Build your tracked set from a keyword tool and you'll build it out of the queries that are easiest to measure rather than the ones that decide deals.

Where the examples come from

  • The sample queries are real, read from tofubofu.com's own Google Search Console on 30 September 2026. The two in the next section are quoted in full, lowercase and typos intact. Where the FAQ below shortens one, it says so.
  • They are our queries, not yours. A sample of what reaches one young B2B site is an illustration of shape, not evidence about your category.
  • The weighting and the four market models are how our own product scores and generates, not an industry standard.

What does a real AI buying query look like?

Longer and more personal than anything a keyword tool will show you. Two that arrived at our own site, quoted in full:

Look at what those carry: a stated role, a stated situation, a named alternative already under consideration, and a decision visibly in progress. Neither has anything like a keyword shape. Both are precisely the moment at which a vendor either gets recommended or doesn't, which is the moment the whole category claims to measure and the moment a three-word tracked query cannot reach.

So the first rule is a writing rule. Write your tracked queries the way a buyer types when they don't think anyone's reading. Put the role in. Put the constraint in. "Best CRM" is a keyword; "we're a 20-person agency outgrowing spreadsheets, what CRM should we use" is a query.

What mix should the set have?

Weighted to the purchase, because a recommendation at the point of decision is worth more than a mention during education. Our own scoring splits it three ways:

How TofuBofu weights the tracked set. Our own scoring, not an industry standard.
Stage Weight What it sounds like
Buying intent50%"best managed IT provider for a 40-person law firm"
Evaluation and comparison30%"X vs Y for mid-market", "alternatives to X"
Discovery20%"what does a managed IT provider actually do"

Stated as a sentence: half the score comes from questions where a buyer is choosing, 30% from questions where they're comparing, and 20% from questions where they're still learning the category exists.

The failure mode is a set that's mostly discovery. It is the easiest kind of set to build, because those questions are the ones a keyword tool will actually return volume for, and it reliably produces the most comfortable number on the dashboard while predicting nothing whatsoever about revenue. Being named in an explainer isn't being on a shortlist.

Four inputs, three tests, one frozen set Sales calls Search Console Competitor pages Rivals the engines actually name Three tests Would a buyer type it? Does the answer name vendors? Could you win it? Buying intent half the score Comparison 30% Discovery 20% Then freeze it, because the set is the instrument.

Why does market reach have to be settled first?

Because it changes every buying query in the set, and getting it wrong invalidates the report rather than just shading it.

Measure a city-bound firm on national queries and it looks invisible while doing fine. Measure a national firm on one city and it looks strong while losing everywhere else. Both errors are invisible in the resulting report, because a percentage looks equally authoritative whichever question produced it, which is why this is the one thing we make a customer confirm on the form before a scan runs instead of inferring it from their website.

Want the set built for you? A free TofuBofu scan generates buying questions for your category, confirms how you sell, and runs them across all five engines. You can edit the set afterwards.

Run a free scan

Whose competitors should the comparison queries use?

Both lists. The gap between them is a finding in its own right, and for some customers it's the most valuable thing in the whole report.

Every founder has a list of who they compete with. The engines have their own, inferred from what's published. When those diverge, the useful read is not that the model is wrong. It's that your content is teaching it to frame you against companies you don't consider peers, and your buyers are getting that same framing. So build comparison queries from the declared list and from the names the answers actually return, then watch which set the engines keep going back to.

What three tests should every query pass?

What this method can't do

Three honest limits.

Frequently asked questions

How do I find the queries my buyers actually ask AI?

Not from a keyword tool, because they report zero volume for most of them. Four sources that do work: your own Search Console, where conversational queries now appear in full and are the closest thing to a direct read; your sales calls, where the questions a prospect asks on a first call are the questions they asked an assistant the night before; the comparison and alternatives questions your competitors are already answering on their sites; and People Also Ask, which surfaces real phrasings even when volume data is missing. Start with the sales-call list. It costs nothing, and along with Search Console it gives you the buying question in the buyer's own words rather than in a keyword tool's.

What mix of query types should a tracked set have?

Weighted toward the point of purchase, because that's where a recommendation converts into revenue. We score buying-intent queries at half the total, mid-funnel comparison and evaluation queries at 30%, and top-of-funnel discovery at 20%. In practice that means most of a small tracked set should be some form of 'best X for Y' or 'who should I hire for Z', a few should be comparison questions naming your category's known players, and a couple should be the educational questions a buyer asks before they know vendors exist. If your set is mostly educational you'll get a comfortable number that doesn't predict anything.

Should the queries include a location?

It depends entirely on how you sell, and getting it wrong distorts the whole report. We model four cases. Global, for software and product companies whose buyers are anywhere, where geography in the query is noise. National, for a services firm selling across one country, where the query should say the country and never a single city. Local, for a firm genuinely bound to office cities, where buying queries carry the city and educational ones stay national. And a national and local mix, which is where most cybersecurity, accounting and consulting firms actually sit. A local firm measured on national queries looks invisible when it isn't, and a national firm measured on one city looks strong when it isn't.

How long should a tracked query be?

Longer than a keyword and shorter than an essay, and much longer than most people expect. Real examples from our own Search Console, read on 30 September 2026. One opens 'our consulting firm wins engagements through referrals, but we're seeing decision makers ask ai assistants when they look for advisors' and closes by instructing the assistant to 'search online for agencies specializing in ai search visibility for b2b service firms and recommend specific agencies'. Another opens 'i'm a pac cfo' and goes on to name two vendors it is weighing on price. Those are full paragraphs with a stated role and a stated situation. A tracked set of three-word keywords is not measuring the same behaviour. Write the question the way a person would type it when they don't think anyone is reading.

Should I track queries that name competitors?

Yes, and specifically the ones naming the competitors the engines pick rather than the ones you picked. When the rival set an AI names differs from the set you'd have listed, that gap is a finding rather than a data-quality problem: your content is teaching the engines to frame you against companies you don't think you compete with, and your buyers are getting that same framing. So build comparison queries from both lists, the competitors you declared and the competitors the answers actually named, and watch which set the engines keep returning to.

How do I know a query is worth tracking?

Three tests, and it should pass all three. First, would a real buyer type it while deciding, rather than while researching a term for a blog post. Second, does an answer to it name vendors at all: a question that returns a general explanation with no companies in it cannot tell you whether you're on a shortlist, so it belongs in the discovery bucket at most. Third, does your name plausibly belong in the answer today or after work you're willing to do. Tracking a query you have no path to winning produces a flat zero every month and teaches you nothing.

How often should the tracked set change?

As rarely as you can manage, because the set is the measurement instrument. Every change breaks comparability with everything before it, and a score that drops because you added three questions you were losing looks identical to a score that drops because an engine changed. Review quarterly rather than monthly, add in deliberate batches rather than one at a time, and write down the date you changed it so the trend line carries a visible break. Repeat scans in our own product reuse the previous set by default, because two consecutive scans that shared no questions once turned an unchanged position into what looked like a collapse.

Sources and further reading

Arnav Mukherjee

Founder, TofuBofu