A Confidence Checklist Before Reporting an AI Visibility Lift
Before reporting a lift, answer seven questions: was the prompt set unchanged across both periods; were both periods sampled the same number of times; was the gap between your treatment and control changes larger than either set normally swings (not merely "did the control stay flat" — a control that rose can still sit well below a treatment that rose further); did your change exceed competitors' change by more than normal variation (not merely "were competitors flat"); is the change larger than this metric's normal period-to-period variation; did shipped work precede it by a plausible interval; and can someone else reproduce the number from the recorded rows? A lift that fails the control question or the normal-variation question is not a lift you can defend — and the fastest way to lose credibility on this channel is to claim one that later turns out to have been a model update.
Reporting a lift is a claim about causation made from observational data on a system you do not control. The checklist is not bureaucracy; it is what stops the claim from being withdrawn later.
The seven questions
| Question | If the answer is no |
|---|---|
| Same prompt set both periods? | The denominator moved — not comparable |
| Same sampling depth both periods? | You may be comparing 1 run with 5 |
| Control gap bigger than normal variation? | A shared trend could explain it — or could be masking your effect |
| Your change larger than competitors' change? | A shared move can still contain your effect — compare sizes, and claim causation neither way |
| Bigger than normal variation? | It is probably noise |
| Shipped work precedes it plausibly? | Correlation without a mechanism |
| Reproducible from the rows? | It is an assertion, not a measurement |
The two that usually fail
The control comparison, and normal variation. Most teams have no control set at all, and most have never computed their own period-to-period variation, so they have no idea whether a six-point move is remarkable or Tuesday. Note what the control question is actually asking: not whether the control stayed still, but whether the gap between the two changes is bigger than the noise. Treating a risen control as automatically disqualifying throws away real results. Both are cheap to fix and neither can be fixed retrospectively, which is the argument for setting them up before you need them.
How to report when the checklist fails
Say what you observed and what you cannot rule out. 'Presence rose from 40% to 55% across the same 20 prompts; our control set rose by a similar amount over the same period, so we cannot separate our work from a platform-wide shift' is a more useful sentence than a lift claim, and it costs nothing except the pleasure of announcing a win. Teams that report this way are believed when they do claim one.
For agencies and vendors
The incentive runs the other way here: a lift is the thing a client or a prospect wants to hear. We hold ourselves to this list and would rather report a flat period honestly, because the alternative is being the vendor whose numbers stop being trusted the first time a model update is mistaken for a result. If you are evaluating tools, ask which of the seven their reporting actually supports.
FAQ
What counts as normal variation? — Compute it from your own history: the spread of your presence rate across periods when nothing changed. Until you have three or four periods you do not know, and should say so rather than guessing.
Is a lift ever reportable without a control? — With explicit caveats, yes — say that a platform change cannot be ruled out. What is not defensible is reporting it as a result of your work while knowing you cannot separate the two. The mirror error is also worth naming: discarding a result because the control moved at all, when the treatment moved considerably further.
How long should we wait before reporting? — At least two periods after the work shipped, and longer where the metric is volatile. The pressure to report quickly is exactly the pressure that produces withdrawn claims.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.