Measurement
How often should I check my AI visibility?
By Arnav Mukherjee, founder of TofuBofu · July 6, 2026
A founder messaged me, genuinely rattled. He had asked ChatGPT for the best providers in his niche, seen his company named third, and felt great. The next morning he asked again to show a colleague, and he was gone. Nothing on his site had changed overnight. He wanted to know what he had done wrong. The honest answer was: nothing. He had just watched non-determinism in action, and drawn exactly the wrong conclusion from it.
This is the trap at the center of the "how often should I check" question. AI answers wobble on their own. If you check too often, you measure the wobble instead of your actual standing, and you will make yourself miserable reacting to noise. The right cadence is built around that reality.
Why the answer moves when you did not touch anything
Three things make AI answers unstable, and none of them are about you. First, the models are non-deterministic: ask the same question twice and you can get two different answers, even with settings dialed to their most repeatable. Researchers ran five LLMs configured to be deterministic across eight tasks and ten runs and reported accuracy variations of up to 15 percent between runs that should have been identical. Second, retrieval rotates: engines that pull live sources may grab slightly different pages each time, so the supporting cast in the answer shifts. Third, the models themselves get updated, sometimes quietly, and an update can reshuffle who gets named.
There is a fourth cause, and it is the one nobody in this category volunteers: the tool measuring you changes too. Ours moved four times in nine days last month. On 12 August we widened the analysis window and made the brand matcher symmetric, which lifted measured rates by roughly 3 percent. On 17 August the geo classifier stopped truncating a brand's list of markets, which changes which questions get asked in the first place. On 18 August our ChatGPT probe started searching the web instead of answering from training weights. On 20 August repeat scans stopped re-asking retired awareness questions. Three of those four carry no published size, because we never measured one, and saying so is more useful than inventing a number.
We weight buying questions at half the Index, deliberately, because those are the ones a purchase actually turns on. Our read is that they are also the least repeatable of the lot, since they are the questions where an engine has the widest field of vendors to choose between. That is why the paid plans sample a question more than once and only spend a third call when the first two disagree. A single unsampled reading of a buying question is close to worthless, which is why sampling within a check matters more than checking more often.
Put together, that means a day-to-day change in whether you appear is usually noise. It is the equivalent of judging your weight by stepping on the scale every hour. The number jumps around for reasons that have nothing to do with the trend, and if you react to each jump you will chase phantoms.
Daily noise vs the monthly trend
The instrument moves too. Ours moved four times in nine days.
Everything above is about the engines wobbling. There is a second source of movement that almost nobody in this category will discuss, because admitting it is expensive: the tool measuring you changes, and when it does, your chart moves for reasons that have nothing to do with your market or your work.
Here are ours, with dates, because we publish them in the source code the product reads. 12 August 2026: we widened the text window the analyser reads and made the brand matcher symmetric, so it could rule a mention back in as well as out. That lifted measured rates by roughly 3 percent. 17 August: the geo classifier stopped collapsing a brand selling into seven markets down to one country, which changes which buying questions get generated at all. 18 August: our ChatGPT probe started searching the web instead of answering from training weights. 20 August: repeat scans stopped re-asking awareness questions that generation had stopped writing weeks earlier, which changes the denominator.
Three of those four carry no published number, and that is deliberate rather than coy. We measured the size of the first one. We did not measure the other three, and inventing a figure to look rigorous is the exact failure this whole page is arguing against. What we do instead is relabel and never recompute: no stored result is rewritten, no scan is re-run, and a customer whose trend line crosses one of those dates gets told on the chart that the step may be ours.
Now put that beside a daily cadence. In the nine days those four changes landed, a daily chart would have shown you nine readings and four genuine discontinuities, with no way to tell which was which. A monthly cadence would have shown you one step and a date to ask about. Frequency does not fix this. Disclosure does.
And the outside world moves during your measurement window
One worked example from the same fortnight, and it is not about us. Google began rolling out its August 2026 spam update at 16:27 UTC on 18 August and declared it finished at 08:49 UTC on 21 August, in its own words on its own status dashboard: "The rollout was complete as of August 21, 2026."
Anyone who sampled Google's AI surfaces on the 19th and compared it to the 22nd measured a Google rollout, whatever their dashboard told them they were measuring. That confounder is invisible at any cadence, but a weekly or daily reader is far more likely to act on it, because a monthly reader would have averaged straight through it. Before you read a step in any chart as your own doing, check whether one of the engines was mid-rollout when it happened.
The right cadence: monthly, sampled, trend-based
For almost every B2B company, monthly is the sweet spot. AI citation patterns move more slowly than paid ad impressions but faster than organic search rankings, so a month is long enough for real change to show and short enough that you catch problems early. Inside each monthly check, sample: run every query at least three times per engine and average whether you appear. That smooths out the built-in wobble so you are comparing like with like from one month to the next.
Then judge the trend, not the dot. One month at 14 percent share of voice and the next at 12 is probably flat, within the noise. Three months running 9, 13, 17 is a real climb. The discipline is to make decisions from the line, and to treat any single reading as one data point rather than a verdict.
When to check sooner
Monthly is the baseline rhythm, not a rule against ever looking in between. There are three good reasons to run an extra check:
You shipped a fix. You added FAQ schema, rewrote a page answer-first, or earned a batch of reviews and mentions. A follow-up check confirms whether it landed. Give it a few weeks first, though, because engines need to re-crawl and recall lags.
A competitor moved. If a rival suddenly starts showing up everywhere, it is worth a look to see what changed and whether it pushed you down.
A model launched or updated. A major new model can reshuffle answers across a whole category. When one lands, a check tells you where you now stand. These are event-driven, not a reason to switch to daily monitoring.
What to do
1. Set a monthly check
Same query set, same engines, same day each month. Consistency is what makes the trend readable. Put it on the calendar so it actually happens.
2. Sample within each check
Run every query at least three times per engine and average. This is the single most important habit for not fooling yourself with noise.
3. Judge the trend, not the reading
Compare months, look at the direction over a quarter, and treat any one number as a dot on a line rather than a result to react to.
4. Add event-driven checks
After you ship a fix, when a competitor moves, or when a model updates. Wait a few weeks after a fix before expecting it to show.
5. Resist daily monitoring
Unless you are confirming a just-made change, daily checking of a non-deterministic signal costs attention and returns mostly noise.
6. Ask your vendor when its own measurement last changed
Any tool that has improved its matching, its question generation or its engine probes has moved your history. Ask for the dates and ask whether old scans were recomputed or relabelled. A vendor that cannot answer has either never improved or never told you.
Get a sampled monthly reading, automatically
Run a free scan now, then track the trend over time instead of chasing daily noise.
Get your free auditFrequently asked questions
How often should I check my AI visibility?
Monthly for most companies. AI answers vary run to run even when nothing on your side has changed, so checking daily mostly measures noise and invites panic. A monthly cadence, where each check samples several runs per query and you watch the trend over months, gives you signal instead of noise. Check sooner only when there is a real reason: you shipped a fix, a competitor moved, or a new model launched.
Why does my AI visibility change when I have not changed anything?
Because the answer was never stable to begin with. Large language models are non-deterministic even when you configure them not to be: researchers ran five of them across eight tasks and ten runs each, with settings dialled to be repeatable, and reported accuracy variations of up to 15 percent between runs that should have been identical. On top of that the sources an engine retrieves rotate and the model itself gets updated. And the tool measuring you moves too: ours changed four times in nine days in August 2026, and we publish those dates. So day-to-day movement is usually the instrument, not a shift in your standing.
Should I check my AI visibility every day?
No. Daily checking of a noisy, non-deterministic signal is a recipe for chasing ghosts. You will see yourself appear and disappear from one day to the next and be tempted to react to changes that are not real. Unless you have just made a specific change and want to confirm it, daily monitoring costs attention and gives back mostly noise.
How many times should I run each query when I check?
At least three, and average the result. Because a single run can name you or miss you by chance, one reading is unreliable. Sampling several runs per query on each engine and averaging whether you appear is what turns a coin flip into a measurement you can compare month to month.
When should I check more frequently than monthly?
When something real has changed. If you just published a fix, added schema, or earned a batch of reviews and mentions, a follow-up check confirms whether it landed. If a competitor makes a visible move, or a major model updates, a check tells you where you now stand. These are event-driven checks on top of your regular monthly cadence, not a reason to switch to daily monitoring.
How long before my AEO changes show up in AI answers?
Usually weeks, not days. AI engines need to re-crawl your updated pages and, for training-based recall, the change propagates even more slowly. That lag is another reason monthly is the right rhythm: checking the day after you publish will often show nothing, not because the fix failed but because the engines have not caught up yet. Give it a few weeks before judging.
How do I know a change in my score is real and not my tool changing?
You ask, and a serious vendor publishes the dates. Our own measurement moved four times in nine days in August 2026: on 12 August we widened the analysis window and made the brand matcher symmetric, which lifted measured rates by roughly 3 percent; on 17 August the geo classifier stopped truncating a brand's markets, which changes which questions get asked; on 18 August our ChatGPT probe started searching the web instead of answering from training weights; and on 20 August repeat scans stopped re-asking retired awareness questions. Three of those four have no published size because we never measured one, and saying so is the point. We relabel history rather than recompute it, and a customer whose trend line crosses one of those dates is told the step may be ours. If your vendor cannot name a single date on which its own measurement changed, that is not a sign it never did.
Should I watch a single number or a trend?
A trend. Any single reading sits on top of real volatility, so it can mislead in either direction. The reliable signal is the direction over several months: is your mention rate and share of voice climbing, flat, or slipping. Treat one month's number as a data point, not a verdict, and make decisions from the line, not the dot.
Sources and further reading
- Non-Determinism of "Deterministic" LLM Settings (arXiv, August 2024): five LLMs, eight tasks, ten runs, with accuracy variations up to 15 percent between runs configured to be identical. General-purpose tasks rather than vendor recommendations, cited here for the mechanism only.
- Google Search Status Dashboard, August 2026 spam update: began 18 August 2026 at 16:27 UTC, complete 21 August at 08:49 UTC, in Google's own record.
- Our own measurement-shift dates are published in the module the product itself reads, and every one of them is annotated on the customer trend line that crosses it.