Now live across the AI ecosystem: ChatGPT GPT Store · MCP Registry · mcp.so

Getting started

You sold a client on AEO. Here is what you actually deliver.

By Arnav Mukherjee, founder of TofuBofu · August 20, 2026

I pulled the public archive of the AEO subreddit, every post from October 2021 to 17 August 2026, and counted what people are actually asking. The place went from 20 posts in October 2025 to 298 in May, 307 in June and 287 in July. In the 90 days ending 17 August there are 914 posts, and 162 of them, 18 percent, carry the vocabulary of agency work.

The most-replied question in that window, with 50 replies, is titled "Just sold a client on GEO/AEO, now how do I deliver?" That is a purchase order. Somebody has taken money and does not yet have a delivery motion. The top comment, at 22 points, tells him this was something to work out before charging, which is satisfying and useless.

One day before that window opens, on 19 May, sits the most-replied question about measurement in the entire five-year archive, with 54 replies: "How are you measuring brand visibility in ChatGPT and Perplexity for clients?" Its author states the symptom in one line: his clients rank well on Google but barely get mentioned when someone asks ChatGPT the same question. Two threads, one problem. Demand for this service arrived before the playbook for it, and the people selling it are improvising in public.

So here is the playbook, sequenced by which measurements are stable enough to prove a move inside a quarter rather than by what is easiest to put on a proposal. It includes the promise you should refuse, and the argument for what your retainer is actually for once the technical work is done.

A definition, before the number gets quoted back at me

That 18 percent depends entirely on what you count, so here is the count. The window is 20 May to 17 August 2026 inclusive. A post is in the agency bucket if its title or body contains any of agency, agencies, client, clients, retainer, retainers, deliverable, deliverables, consultant or consultants, whole word, case insensitive. That gives 162 of 914.

One word is doing most of the work. Drop agency and agencies and the same window returns 98 posts, 11 percent. I am publishing both because a corpus statistic without its term list is not a finding, it is a vibe with a decimal point, and I have watched this exact number get quoted internally at two different values computed two different ways. The percentage is at least robust to where you put the window edge: every start date I tested between 18 and 20 May lands between 17.6 and 17.8 percent.

Scope the plumbing as a week, and invoice it that way

Every AEO guide of the last two years opens identically. Add Organization schema, ship an llms.txt, mark up the FAQs. Do it. It is real work, a competent developer clears it in days, and your client has probably not done it: across 39 sites we crawled during real scans, 38 served no llms.txt and 26 carried no FAQ schema, and those firms are not a random sample because nearly all of them ran a scan already suspecting they were invisible.

What that checklist will not do is separate your client from the firm beating them. We crawled 14 firms our scans measured as absent alongside the 8 distinct rivals the engines named instead, and the technical rows came out level or backwards: Organization schema on 86 percent of the invisible firms against 88 percent of the named ones. Only one axis moved in the right direction, and it was FAQ markup, on 5 of the 8 named firms against 5 of the 14 invisible ones. Eight firms is a small, selected sample and one of them swings any percentage by more than 12 points. The full table and its caveats are in the study itself, and I am not going to re-run them here.

The commercial instruction is what matters for you. This work has a finish line and it can be verified from outside, which makes it a fixed-fee project, not a recurring line. Quote it as one, deliver it in the first fortnight, and let the client watch it close. An agency whose quarter is a march down that checklist is charging monthly for something that stopped changing in week two, and the first competent buyer to notice will say so in the renewal meeting.

Lead with comparison content, for a measurement reason

This is the most useful paragraph in the post if you are three weeks from a client review, and I have not seen anyone else make the argument.

AI answers do not repeat. Ask an engine the same buying question twice and the list of companies changes. Conductor measured exactly how much, putting 14,000 API calls through ten industries, seven intent types and four engines at fifty runs per prompt, and published the result on 17 June 2026. The instability is not spread evenly across intents, and that unevenness is the whole opportunity.

Comparison prompts were the steadiest of the seven, at 63 percent brand overlap between two runs and 91 percent lead-brand stability. Purchase prompts had the lowest overlap of any intent type, 40 percent, which is to say six of every ten brands appearing across a pair of runs appeared in only one of them.

Read what that does to your client review. The open buying question, the one your client cares about most and the one you almost certainly wrote into the proposal, is the noisiest cell available. You can genuinely improve your client's standing on it and still be unable to demonstrate that you did, because the second reading disagrees with the first for reasons that have nothing to do with your work. Meanwhile the comparison question, which converts at least as well and which almost nobody targets deliberately, is the cell where a real move clears the noise.

A delivery sequence you can defend in a client review Month 1: instrument Freeze the question set Baseline + fixed-fee plumbing Months 2 to 3: publish Comparison and alternatives Named case studies, profiles Month 4 on: hold Re-read vs the baseline Same questions, same engines Where a move is provable, and where it is not Comparison questions Steadiest of seven intent types measured. A real move clears the noise. Report it. Open buying questions Lowest run-to-run overlap of the seven. Worth the most, proves the least once. Intent stability from Conductor, 14,000 API calls, published June 2026. The sequence is ours.

So work both and report them differently. Build the comparison and alternatives pages first, aimed at the specific rivals the engines are naming instead of your client, which you take from the scan rather than from the client's own opinion of who they compete with. Those two lists are frequently different, and the gap between them is itself a finding worth putting on slide one. Keep the open buying questions in the report as the strategic line, sampled more than once, and say plainly that they move slowly and noisily.

Then the layer almost every agency under-scopes, because it is the only one that is not a page on a domain you control: corroboration. Named case studies with the client organisation in them, complete and correctly categorised profiles wherever your client's category gets compared, and third-party coverage written by somebody other than your client. It is slower, it is harder to invoice, and it is the part a competitor selling schema audits cannot copy by Friday.

Run the baseline before the kickoff call

A free scan gives you the question set, the per-engine result and the rivals the engines name instead of your client, which is the list your comparison pages should target.

Run a free scan

The objection: I just argued you out of your own retainer

Take the argument seriously, because it is the strongest thing anyone will say back to me. You came here to sell a monthly fee and I have told you the technical work is a fixed-fee fortnight. If the plumbing finishes, the pages get written and the profiles get filled in, what exactly is the client buying in month seven?

Concede the part that is true first. If your retainer is a monthly schema audit and a screenshot of a dashboard, it deserves to be cancelled, and it will be. That model survived in SEO for a decade because ranking reports looked like progress. It will not survive here, because the client can now open ChatGPT themselves and check.

Now the case for the fee, and it comes straight out of the Conductor numbers rather than from anything I would like to be true. Even comparison prompts, the steadiest cell of the seven, hold the same lead brand only 91 percent of the time. Purchase prompts hold theirs 59 percent of the time. A position in an AI answer is a distribution, not a rank. You do not win it and keep it. You hold it, or you do not, and the only way to know which is to keep reading it.

That relocates the retainer. It is not maintenance of the plumbing, it is four things that are genuinely continuous:

The measurement is a repeated measure, and one reading is not a measurement. This is not a philosophical point, it is what fifty runs per prompt demonstrates. The practitioner behind the most-replied post in the whole window, an AMA from someone running local AEO at a stated average retainer of $12,500, reached the same conclusion from the other end: he tracks at least fifty prompts because below that, in his words, you are reacting to noise. That is his judgement rather than a measurement, and it happens to agree with the study.

The rival set moves underneath you. Your comparison pages target the firms the engines name instead of your client. That list is an output of the engines, not a fact about the market, and it changes. A comparison page aimed at a company the engines stopped recommending four months ago is a page working on a question nobody is being asked.

Corroboration is other people's property. Nothing about profiles, coverage or named case studies is a thing you finish. It accrues, it goes stale, and it depends on human beings outside the engagement doing something. That is a cadence, not a task.

The question set has to be defended. Your client will launch a service line, enter a city, drop a vertical. Every one of those is pressure to change the questions, and changing them silently destroys the comparison you have spent two quarters building. Somebody has to own that decision, in writing, and explain the cost of it. That is a job.

Sell it that way and the retainer gets easier to defend, not harder, because every line of it is something the client can see is still happening. The fixed-fee fortnight is the part that makes the rest credible: an agency that voluntarily prices the easy work as a project is visibly not padding.

What goes in the report, and what you refuse to put in the contract

The 54-reply measurement thread is mostly people naming tools at each other. A tool is not a reporting method. Four rules, written from the supplier's side, will keep you out of trouble with a client who is paying attention.

Freeze the question set on day one. If the questions change between readings you have not measured a change, you have run a different test. This is the most common way an AEO report misleads without anyone intending it, and it almost always resolves in the agency's favour, which is precisely why the client will eventually stop believing it.

Report per engine, and separate presence from position. A composite hides the one fact worth acting on: a client named on one engine and absent on five looks like partial success in an average. The Conductor table adds a second reason. Education prompts had the second-highest brand overlap of the seven, 60 percent, and the lowest lead-brand stability of all of them, 30 percent. The same brands keep showing up and the order reshuffles almost every run. A report that collapses those into one score can tell a client they are visible and stable when they are visible and churning.

Sample more than once, and publish a movement threshold. One run per question tells you whether your client exists on that question at all. It does not date a trend. Decide before the first reading how large a move has to be before you call it one, put the threshold in the deck, and write no change detected the rest of the time. A consultant who says that in month two is the one still there in month twelve.

Commit to the work on a date, and to the movement never. You can put the pages, the markup, the profiles and the report on a calendar. Nobody on earth can put a named position in an AI answer on a calendar, and the people who do it anyway are why this category has a credibility problem your prospect is already carrying into the first call.

Then the scope line to write in explicitly: reviews. You cannot get them and neither can your client, because a review comes from your client's customer and that is the whole point of it. What you can legitimately sell is guidance and process. Make sure the profiles exist, are claimed and sit in the right category. Help the client build the request into the moment a job is delivered, so the ask is theirs and it is timed well. Manufacturing the reviews is off the table: it breaks the platforms' terms, it is illegal advertising in most markets, and it gets wiped when detected. Say this in the pitch rather than after the contract. Every competitor who quietly implies otherwise has handed you the differentiator.

Why the work is worth selling at all

Underneath the retainer, the client's actual fear is that the shortlist is being assembled somewhere they cannot watch. That is what the author of the measurement thread was describing when he wrote that his clients rank well on Google and barely get mentioned when someone asks ChatGPT the same question. He was not asking for a tool. He was reporting that a position he had already paid for stopped covering the ground it used to cover.

That gap is the entire business. A first-page Google rank buys no seat in an assembled answer, because the answer is built from what the wider web says about a company, weighted by which sources the engine leans on. Search position is the floor you still have to hold. This is a separate layer on top of it, and a firm can win the layer while its Google position sits exactly where it is. Sell that, deliver the four things in that order, price the finishable part as a project, report it honestly, and you have a service. Sell a dashboard and you have resold somebody else's subscription with your logo on the invoice.

Frequently asked questions

What are the actual deliverables in an AEO engagement?

Four, in this order. A machine-readable site, meaning schema and an answer-shaped page structure, which is the cheapest work and the least differentiating. Comparison and alternatives pages aimed at the specific rivals the engines already name instead of your client. Named, specific case studies plus corroboration on platforms your client does not own. And a measurement report showing named or absent per engine on a question set you freeze at the start. Everything else is either a variant of those four or a tool subscription with your logo on the invoice.

Should I start with schema and llms.txt?

Do it in week one and do not sell it as the outcome. We crawled 14 firms the engines never name alongside the 8 distinct rivals the same engines recommended instead, and the standard technical checklist did not separate the two groups: Organization schema sat on 86 percent of the never-named firms against 88 percent of the named ones, and three of the eight things we measured were more common among the invisible firms. That is 8 named firms, so one firm moves any percentage by more than 12 points, and every firm in it was selected because somebody already suspected it was invisible. It is a small, selected sample. It is still enough to stop you building a quarter's plan out of a checklist.

Should AEO be sold as a retainer or a one-off project?

Both, split honestly. The technical work is a project and should be a fixed fee, because it finishes and it can be verified. The retainer is not for maintaining that work, it is for holding a position that does not stay held. Conductor asked the same prompts 50 times each and found that even comparison prompts, the steadiest of the seven intent types, keep the same lead brand only 91 percent of the time, and purchase prompts only 59 percent. A position in an AI answer is a distribution, not a rank you win once. Publishing cadence, a rival set that changes underneath you, corroboration on platforms your client does not control, and a repeated measurement are all genuinely continuous. A retainer that is a monthly schema audit deserves to be cancelled.

Which deliverable should I lead with if I need to show a client progress fast?

Comparison content, for a measurement reason rather than a content reason. Conductor ran 14,000 API calls across ten industries, seven intent types and four engines at fifty runs per prompt, and comparison prompts were the most reproducible of the seven at 63 percent brand overlap between runs and 91 percent lead-brand stability. Purchase prompts had the lowest overlap of any intent type at 40 percent. Work the cell where the measurement is steady enough that a real move clears the noise, or you will win something inside a quarter and be unable to show it.

How do I report AI visibility to a client without overpromising?

Report named or absent, per engine, on a question set you freeze at the start and do not change. Sample each question more than once, because a single run is a presence check and not a trend. State a movement threshold below which you will write no change detected. And never blend engines into one composite, because a client named on one engine and invisible on five can be shown a comfortable average that hides the only fact worth acting on. Separate presence from position too: the Conductor data shows an intent type can have high brand overlap and still reshuffle its leader nearly every run.

Can I promise a client I will get them reviews on G2 or Google?

No, and say so out loud in the pitch, because it separates you from the agencies that imply otherwise. A review comes from your client's customer. Neither you nor your client can manufacture one. What an agency can legitimately do is guidance and process: make sure the profiles exist, are claimed and are categorised correctly, and help the client build the request into the moment a job is delivered. Manufacturing reviews breaks the platforms' own terms, breaks advertising law in most markets, and gets wiped when detected. Scope it as guidance, write the client's own obligation into the contract, and refuse the work if what they want is the other thing.

How long before an AEO engagement shows results?

Long enough that month one should be sold as instrumentation rather than improvement. The honest sequence is a baseline reading and a frozen question set in month one, published assets in months two and three, and a second reading compared against that baseline after it. Anyone quoting a fixed number of weeks to a named position is quoting a number nobody has measured. What you can commit to on a date is the work: the pages, the markup, the profiles and the report. Promise the work, report the movement, and never invert those two.

Sources and further reading

Related reading