Now live across the AI ecosystem: ChatGPT GPT Store · MCP Registry · mcp.so

Getting started

Your agency's six months are up. Can you tell if it moved your AI visibility?

By Arnav Mukherjee, founder of TofuBofu · August 11, 2026

I was on a call on 10 August with the owner of a managed IT firm in British Columbia. His six-month agency agreement was ending and he had asked the obvious question: what happens after the six months? The answer he got back was, in his retelling, that you just stop paying us. He did not want to stop. He got a website out of it, social accounts, a presence that did not exist before. What he could not answer was whether any of it had changed anything, and he said so plainly: for everything we think we know about marketing, we know nothing.

Later the same day, a founder building a video product put the same problem from the supplier's side. The video, he said, is just a paintbrush. It leads to a result but it is not the result, and what he wants is many answers to the question every client asks after delivery: now what?

Both of them are circling the same gap. Agency work is measured in deliverables, because deliverables are what can be counted and invoiced. The thing the client is actually buying is whether a buyer hears their name. Nothing in a standard monthly report measures that, and four of the proof points most commonly offered instead do not survive contact with data we have already published.

Four things that look like proof and are not

None of these is a scam. Each is real work that a competent agency genuinely did. The problem is only that each is an input being presented where an outcome belongs, and in every case we have a measurement that shows the gap.

1

We added schema markup and an llms.txt file

We crawled 14 managed IT firms that AI engines never named, against the specific rivals those same engines did name in the same regions. Organization markup was present on 86 percent of the invisible firms and 88 percent of the named ones. More of the invisible firms served an llms.txt than the named ones did. This is plumbing worth having, and on our evidence it is not what separates a firm that gets recommended from one that does not. Price it as plumbing.

2

Look, ChatGPT mentions you now

Two problems, both measured. Engines disagree: across five US metros, 85 percent of firms were named by exactly one of the four answering engines and none appeared on all four, so one screenshot tells you close to nothing about the other five. And a chat answer is not necessarily a live lookup. Across 25 vendor-recommendation queries, ChatGPT opened with a memory caveat 15 times and Claude flagged its cutoff 11 times, while the four retrieving engines did it zero times. The answer may predate the work.

3

We published twelve blog posts

Across 3,538 question-and-engine cells from real scans, the brands were named on 5.4 percent of buying questions and on zero percent of the 602 awareness and informational cells. Zero, not merely low. Industry-trends thought leadership is the default content deliverable and it sits exactly in the band where nobody gets named. The posts may be good. They are aimed at the part of the funnel where AI engines do not hand out names.

4

Impressions and rankings are up

Genuinely good news, and a different scoreboard. Search rank is the floor and being named in an AI answer is a distinct layer built on top of it. Our regional research found 139 of 228 directory-listed firms named by no engine at all, and a separate control group of 88 ranked MSP 501 firms in those regions had 66 named by no engine. Firms can be findable and still absent from the shortlist. Both numbers matter and neither substitutes for the other.

Put those together and the picture is uncomfortable but useful. A firm can receive every deliverable on the list, pay for six months of it, and have no way of knowing whether the thing it was buying happened. Not because anyone lied, but because nobody in the arrangement was measuring the outcome.

What gets reported, and what you are actually buying

REPORTED MONTHLY Countable, invoiceable, and all of it upstream of the decision. Pages shipped Posts published Schema added Impressions assumed to produce WHAT YOU ARE BUYING A buyer asks who to call, and an engine says your name Measured per question, per engine, dated, before and after the work. The arrow is the assumption. Nobody in the arrangement is usually testing it.

Take the baseline before the renewal conversation

Run a free scan across six AI engines on your real buying questions and get a dated record of who the engines name today, before anyone tells you what has improved.

Get your free audit

The test to run before you renew

This takes an afternoon by hand and it is worth doing at least once even if you later automate it. The point is not to catch anyone out. It is to give the engagement a scoreboard it currently does not have.

1

Write down five questions your buyers actually ask

Not what you would like to rank for. The literal sentences a prospect says on a first call, in the market you actually sell to. If you serve one province or one metro, the question has that place in it, because a firm bound to a region is not competing for the national version and should not be measured on it.

2

Ask them across several engines and save the answers

Not one engine, and not just the score. Save the text, because the wording is where the useful information is. When an answer names a competitor, it usually says something specific about why: a vertical, a certification, a named client, a response commitment. That sentence is the gap you are actually paying to close.

3

Get the change list from your agency and map it to questions

Ask which specific change is meant to move which specific question. A good agency will have an answer and it will be a reasonable one. A vague answer here is the single most informative moment in the whole review, and it is a fair question rather than a hostile one.

4

Separate the floor from the layer explicitly

Track ranking and AI mentions as two lines, not one blended number. They move at different speeds and can move in opposite directions. A blended score lets a genuine ranking improvement disguise no movement at all on the thing you were worried about, and it lets real citation progress get killed because the ranking line is flat.

5

Agree what counts as moved, before the next period starts

Name the questions, the engines and the date of the next check while everyone is still happy. Corroboration takes time to build and it is unreasonable to judge it on one screenshot, but it is equally unreasonable for the target to be defined after the results are in. Write it down now and the next renewal conversation takes ten minutes.

This is not an argument against agencies

The constraint most owner-led firms hit is volume, and that is exactly what an agency solves. The MSP owner on that call is not short of judgement. He is short of hours, and every hour he spends approving a social post is an hour not spent on a client. Handing that to someone else is a good decision. Nothing in this post argues otherwise.

The narrower claim is that whoever does the work should not be the only party holding the scoreboard. That is not a trust problem, it is a structural one, and it exists in every service arrangement where the deliverable and the outcome are different objects. A monthly report full of completed tasks is a perfectly honest document that simply cannot answer the question the client is asking.

Good agencies gain from this. An independent measurement turns the renewal conversation from a defence of how busy the last six months were into a discussion of what moved, and that is a much easier conversation to win when the work is genuinely good. It also settles arguments in the other direction: when the citations start arriving, there is a dated record proving they were not there before.

Why this suddenly matters more

For most of the last decade, a marketing retainer could be judged on rankings and traffic. Both were imperfect and both were at least visible in a dashboard the client could open. The share of the buying journey those numbers cover is now shrinking.

G2's 2026 research found 51 percent of B2B buyers now begin vendor research on an AI chatbot, up from 29 percent, and 69 percent have switched vendor based on what AI told them. Forrester's 2026 study found 94 percent use AI somewhere in the buying process. The buyers who never make the shortlist never arrive on your site, so they never appear in the analytics your agency reports from. The retainer can be doing excellent work on a shrinking pool and every chart in the deck will look fine. That is the specific failure this test is designed to catch.

Frequently asked questions

How do I know if my marketing agency is helping my AI visibility?

Measure the same buying questions on the same engines before and after, and keep that measurement outside the agency's reporting. Everything in a normal monthly report is an input: pages shipped, posts published, schema added, impressions. None of those is the outcome you are buying, which is whether an engine names your firm when a buyer asks who to call. The two are genuinely separable, and in our own matched crawls the firms carrying the most checklist items were frequently the ones no engine named.

My agency added schema and an llms.txt file. Is that proof of progress?

It is proof of work done, and our own data says it is weak proof of outcome. We crawled 14 managed IT firms that AI engines never named against the rivals those engines did name in the same regions. Organization markup was on 86 percent of the invisible firms and 88 percent of the named ones. More of the invisible firms served an llms.txt than the named ones did. These are worth having and they are not a differentiator, so an invoice line for them should be priced as plumbing rather than presented as a result.

The agency showed me a ChatGPT screenshot naming us. Does that count?

Only weakly, for two measurable reasons. First, engines disagree: across five US metros we found 85 percent of firms were named by exactly one of the four answering engines and none appeared on every list, so one engine's answer tells you almost nothing about the other five. Second, a chat answer is not always a live lookup. Across 25 vendor-recommendation queries, ChatGPT opened with a memory caveat 15 times and Claude flagged its cutoff or asked the buyer to verify 11 times, while the four retrieving engines did it zero times. A screenshot can be reporting training data rather than anything the agency changed.

We published twelve blog posts. Why has nothing changed in AI answers?

Very likely because of what they were about. Across 3,538 question-and-engine cells from real scans, the brands we scanned were named on 5.4 percent of buying questions and on zero percent of the 602 awareness and informational cells. Not a low rate, zero. Thought-leadership posts about industry trends are the default agency content deliverable and they sit precisely in the band where nobody gets named. The same effort spent on the questions buyers ask when choosing a vendor lands in the band where mentions actually happen.

What should I ask my agency for before renewing?

Three things, and all three are reasonable asks. A written list of the buying questions your customers actually ask. A before-and-after measurement of those exact questions across several engines, dated, taken at the start of the engagement and again now. And a plain statement of which changes they expect to move which question. If no baseline was taken at the start, that is not a reason to walk away, it is a reason to take one today and agree what moved counts as from here.

Does this mean hiring an agency is a mistake?

No. Volume is a real constraint and most owner-led firms cannot produce and publish consistently on their own, which is exactly what an agency solves. The argument here is narrower: whoever does the work should not also be the only party holding the scoreboard. A good agency benefits most from an independent measurement, because it converts a renewal conversation about how busy they have been into one about what actually moved, and that is a far easier conversation to win when the work is good.

Can I run this check myself without buying another tool?

Yes, and you should at least once, because doing it by hand is what makes the numbers real to you. Write down five questions a buyer would ask, ask each one in several engines, and record which firms are named in every answer along with the date. Save the answers themselves, not just a score. It is tedious to repeat monthly across six engines, which is why tools exist, but a single manual round is enough to tell you whether the gap between the deliverable list and the outcome is real at your firm.

Sources and further reading

Keep reading: Can you do AEO and GEO yourself, or do you need an agency? · How to check whether your AI visibility report is telling you the truth · How to measure AI visibility