Methodology
Most AI visibility scores are noise with a number attached.
Ask an assistant the same question twice and you can get different brands. A panel of twenty-five prompts sampled once a day produces a figure whose margin of error is wider than the range it is meant to move within. It will show a change every month, and none of those changes will mean anything.
How noisy, specifically
Three figures set the floor for any honest measurement design:
- 65% of AI-cited sources change from one day to the next
- ±44 points of error on a single prompt asked once
- 57.8% of ChatGPT runs return zero citations at all
We watched this happen while building our own tooling. Running one commercial query twice, forty minutes apart, returned two differently worded answers naming a different number of vendors. The underlying citation pattern held; the prose did not. A design that cannot survive that is not a measurement.
The design
| Choice | Why |
|---|---|
| At least 120 prompts, frozen and approved in writing | Precision comes from the number of distinct prompts, not from re-asking one. Freezing the set is what makes month two comparable to month one. |
| Five repetitions per prompt per engine | Enough to characterise the noise. Beyond roughly ten, repetition buys almost nothing. |
| Reported per engine, never blended | Engines retrieve differently and move at different times. A blended figure hides the only actionable information — which engine changed. |
| A confidence interval on every figure | So you can tell real movement from sampling noise. |
| The instrument declared on every number | API sampling, consumer interface and search-result capture are three different measurements. We label which produced the figure. |
| Noise floor measured at onboarding | On your category and an approved competitor set, before any movement is reported. The floor differs by market. |
Prompts are written for your positioning, not your category
A prompt set built around a category label rather than what you actually sell will under-report you, and the resulting "you are invisible" finding will be an artefact of the prompt rather than a fact about your brand.
We have measured this on live sites. One vendor appeared third with a generic description on a broad category prompt. On a prompt written from their own positioning, they appeared second — and the answer misquoted their pricing by roughly twenty dollars a month, sourced from a third-party video. Same company, same day, different question.
Capture is done in a rendered browser
AI Overviews are frequently absent from fetched HTML and appear only when a page is actually rendered. Every measurement here is captured from a real browser in a logged-out session — logged-in sessions personalise results and invalidate comparison between periods.
Any tool that reports AI Overview presence from a plain HTTP fetch will under-report it, and will do so silently. That produces confident false negatives, which are worse than no data because they look like good news.
What we can prove, and what we cannot
| Confidence | Covers |
|---|---|
| Provable | Indexation coverage; crawler behaviour verified from your server logs by IP rather than user-agent; Search Console impressions, clicks and position on a frozen query set; conversions on target pages. |
| Sampled | Movement in AI answers. Reported with intervals, per engine, against the noise floor — never as a single score. |
| Not claimed | Revenue attributed to an AI citation. The industry cannot currently do this cleanly, and we will not present a number that implies otherwise. |
Flat months are reported as flat
Most months the honest finding is that nothing measurable changed. That is what the report will say. A retainer that only ever reports gains cannot be checked, and a client who cannot check the reporting has no way to know whether the work is real.
Your baseline starts with the audit
The prompt panel is built and frozen during the diagnostic audit. Every later reading is measured against it.