Getting started
How to evaluate an AEO/GEO agency and compare two proposals
By Arnav Mukherjee, founder of TofuBofu · September 30, 2026
TL;DR
- Insist on a baseline in week one, on a written query set. Without a before, nothing is provable.
- Get the engines named in the contract. Profound's entry tier is "ChatGPT tracking only".
- Reject four proofs: one screenshot, a post count, schema added, backlinks acquired.
- Don't let the agency own the measurement. They'd control all three inputs to their own grade.
- Ask for a client whose number fell. No such story means they aren't measuring.
A founder showed me two AEO proposals last month and asked which was better. One listed 12 blog posts, schema implementation and a monthly report. The other listed a content audit, entity optimisation and quarterly strategy reviews. Same ballpark price. No overlapping line items at all.
Neither document said which AI engines would be checked, which questions would be asked, or what number was supposed to move. They weren't comparable because neither was measurable. That's the actual problem with buying in this category, and it's fixable with about ten minutes and the right nine questions.
How do you compare two proposals that share no line items?
Stop comparing deliverables. Rewrite both onto one page under three headings, and score each one out of three.
| Axis | What a strong proposal states | The weak version |
|---|---|---|
| Measurement | Named engines, a fixed query count, a written query set, a baseline before work starts | "We track AI visibility" |
| The work | What gets written, against which losing questions, who publishes and where | "12 blog posts a month" |
| The proof | Which number moves, by when, re-scanned on the same frozen set | "Monthly reporting" |
Stated as a sentence: a proposal that names its engines, freezes its query set, records a baseline and commits to a re-scan is comparable to any other proposal that does the same, whatever the deliverables in between. One that doesn't isn't comparable to anything, including itself six months later.
Why is the baseline clause the one to fight for?
Because a visibility score is a rate over a tracked set, and a rate only means something against a fixed denominator.
If the query set changes during the engagement, the closing number and the opening number are measurements of different things. Add questions you were losing and the score falls even if the work succeeded. Add questions you were winning and it rises even if nothing happened. An agency that won't write the query set into the contract is an agency whose results you can't check, and that's true whether or not anyone intends it.
Three specific things to get in writing before signing:
- The set itself, as a list of the actual questions, approved by you.
- The baseline reading, per engine, taken before any work starts.
- The re-scan commitment, on that same set, with the raw per-engine answers delivered rather than a summary chart.
Which four deliverables look like proof and aren't?
All four are legitimate work. None of them is evidence that more AI answers now recommend you.
- A screenshot of one answer naming you. Assistants vary between runs, accounts and days. One screenshot is an anecdote, and a screenshot is also the easiest thing in this category to produce on demand.
- A count of published posts. Output, not outcome. Twelve posts that target nothing a buyer asks will move nothing.
- Schema markup added. Worth doing and not a result. Google's own guidance states you don't need new machine-readable files, AI text files or markup to appear in its generative features, because they run on the core Search systems.
- Backlinks acquired. A proxy for authority at best, and entirely upstream of the thing you bought.
Take your own baseline before either agency starts. A free TofuBofu scan puts real buying questions to all five engines and gives you a dated position you own. Then their first report has something to be checked against.
Run a free scanWhat nine questions should you ask both vendors, including us?
- Which engines, exactly, at this price? Get the list in the contract. Coverage is where the tiers differ and where the surprises live.
- How many queries, and across how many engines? The same query run on five engines counts as five. Multiply before comparing two quotes.
- Who writes the query set, and can I see it? If you can't approve it, you can't audit the result.
- Will you record a baseline before any work? No baseline, no proof, whatever follows.
- Is the set frozen for the engagement? And what's the process if it changes.
- What happens when an engine returns no answer? A silence is not an absence. A firm that hasn't considered this hasn't looked at the raw data.
- Whose competitors do you track? The ones we name, or the ones the engines name. The gap between those lists is itself a finding.
- Who publishes, and to whose site? Content that sits in a shared drive awaiting review isn't published.
- Show me a client whose number went down. And what you did. This is the most revealing question on the list.
Ask us the same nine. If we duck one, that's information too.
Should the agency own the measurement?
Preferably not, and we'll state our interest plainly: we sell measurement, so weigh this accordingly. The structural point stands whoever you buy from.
If one firm picks the questions, runs the scans and writes the report, it controls all three inputs to its own grade. That's not an accusation of bad faith, it's an arrangement you'd flag in any other category and it's the default here. The cleanest fix is that you own the measurement account and the agency works against it, so the history survives the relationship ending. The second best is a query set fixed in the contract with raw per-engine answers delivered to you.
What no honest proposal can promise
Three things, and a vendor offering any of them is telling you something useful about themselves.
- A specific score by a specific date. Nobody controls the engines. A credible commitment is to the work and the measurement, not to the outcome.
- Query volume for AI assistants. The assistants don't publish it. Any volume figure in a proposal is a model, and it should say so.
- A guaranteed position in a named answer. There's no submission process and no ranking to buy. Anyone implying otherwise is selling something that doesn't exist.
And the concession we owe on our own side: a specialist label is not a qualification. A competent SEO team that has read how retrieval works often beats a new agency with an AEO badge and a template, and Google's own position is that optimising for its generative features is still SEO, rooted in the same ranking systems. Judge the method, not the category name on the invoice.
Frequently asked questions
How do I compare two AEO proposals that list different things?
Stop comparing the deliverables and force both onto the same three axes. First, the measurement: which engines, how many queries, whose query set, and is the baseline recorded before work starts. Second, the work: what gets written, who publishes it, and to whose site. Third, the proof: what specific number is expected to move, by when, and what happens if it doesn't. Most proposals in this category are thick on the second axis and silent on the first and third, which is exactly why two of them look incomparable. Rewrite both onto one page under those three headings and the gaps become obvious in about ten minutes.
What is the single most important clause in an AEO proposal?
A baseline measured and recorded before any work begins, on a query set that's written down and frozen. Without it nothing that follows can be proved, because there's no before. And it has to be their measurement of your position today, in writing, not a promise to measure later. The reason this clause matters more than the deliverables is arithmetic: a visibility score is a rate over a tracked set, so if the set changes during the engagement the end number isn't comparable with the start number. An agency that will not write down the query set is an agency whose results you will not be able to check.
What do agencies present as proof that isn't proof?
Four things, and all of them are real work rather than dishonest work. A screenshot of one ChatGPT answer naming you: assistants vary between runs and between accounts, so a single screenshot is an anecdote. A count of blog posts published: output, not outcome. Schema markup added: necessary plumbing, and Google states plainly that you don't need special AI files or markup to appear in its AI features. Backlinks acquired: a proxy at best. None of the four is a measurement of whether more AI answers now recommend you, which is the only thing you're buying.
How many engines should the proposal cover?
As many as your buyers use, and the number is less important than whether it's named in the contract. The trap is that engine coverage is the main thing vendor tiers differ on, and it's usually buried. Profound's own pricing page describes its entry tier as ChatGPT tracking only. Peec's tiers below Enterprise say 'Choose 3 models', with Claude on Enterprise alone. Scrunch lists four engines on Core and nine on Enterprise. So 'we track AI visibility' can honestly mean one engine. Ask which engines, at your price, in writing, and ask what happens when a new one matters.
Should the agency own the measurement?
Preferably not, and this is the conflict nobody puts in a proposal. If the same firm chooses the queries, runs the scans and reports the results, they control all three inputs to their own grade. That isn't an accusation, it's a structure you'd flag in any other category. The cleanest arrangement is that you own the measurement account and the agency works against it, so the data survives the relationship ending. Second best is a query set fixed in the contract with the raw per-engine answers delivered to you, not just a summary chart. We sell measurement, so weigh that as you like: the structural point holds whoever you buy it from.
What should an AEO engagement actually deliver in 90 days?
A recorded baseline in week one, a frozen query set you've approved, published content against the specific buying questions you're losing, and a re-scan on the same set that shows movement or doesn't. That last clause is the one to insist on, because a re-scan on a changed set can show anything. Be realistic about the timeline too: content has to be published, crawled and then retrieved before an answer can change, so a 90-day window is a reasonable first read rather than a verdict. Any proposal promising a specific score by a specific date is selling certainty that nobody in this category can supply.
Is a specialist AEO agency better than a general SEO agency?
Not automatically, and the label tells you less than the method. Google's own guidance says that from Search's perspective, optimizing for generative AI search is optimizing for the search experience and thus still SEO, and that its AI features are rooted in its core ranking and quality systems. A competent SEO team that has read how retrieval works is often better than a new agency with an AEO label and a template. What separates them is measurement: does the proposal specify engines, a frozen query set and a before-and-after, or does it just rename the same deliverables. Ask for a specimen report from an existing client with the numbers redacted.
What should I ask that agencies don't expect?
Three questions. What does your measurement do when an engine returns no answer at all, because scoring a silence as an absence quietly understates every client's position and a firm that hasn't thought about it hasn't looked closely. Which competitors will you track, the ones we name or the ones the engines name, because when those two lists differ the gap is itself a finding. And show me a client whose number went down and tell me what you did about it. The third one is the most revealing: an agency with no such story either has very few clients or isn't measuring.
Sources and further reading
- Google Search Central, "Optimizing your website for generative AI features on Google Search". Source of the "still SEO", the machine-readable-files and the core-ranking-systems statements. Last updated 10 July 2026.
- Profound pricing. "ChatGPT tracking only" entry tier. Read 28 August 2026, re-verified 7 September.
- Peec AI pricing. "Choose 3 models", Claude on Enterprise only. Read 28 August 2026, re-read 8 September.
- Scrunch pricing. Four engines on Core, nine on Enterprise. Read 8 September 2026.