A Normal-Variation Worksheet for AI Visibility Rates
Normal variation is how much your visibility rate moves between periods when you changed nothing. Compute it from your own history: for each assistant, record the rate in every quiet period (no shipped changes to the mapped pages, no prompt-set change, no known platform change), then take the range and the pooled mean. A later move is worth investigating only when it falls outside that range, is clearly larger than sampling noise, and holds for a second period. For the noise part, the binomial standard error √(p(1−p)/n) is a useful floor: a 30% rate from 12 runs has a standard error of about 13 points. Until you have three or four quiet periods, you don't know your normal variation, and your report should say so.
Several guides on this site end with the same instruction: compare a change against your normal variation. This is the worksheet that computes it. It needs nothing but the rows you're already recording and a spreadsheet.
The worksheet, filled with illustrative data
Four quiet periods, one prompt set of 25 prompts, three runs per prompt per assistant, so 75 answers per assistant per period. Each cell is mentions out of answers. Every number is synthetic.
| Assistant | Period 1 | Period 2 | Period 3 | Period 4 | Quiet range | Pooled mean |
|---|---|---|---|---|---|---|
| A | 12/75 (16%) | 16/75 (21%) | 10/75 (13%) | 14/75 (19%) | 13% to 21% | 17% |
| B | 6/75 (8%) | 4/75 (5%) | 7/75 (9%) | 5/75 (7%) | 5% to 9% | 7% |
| C | 23/75 (31%) | 19/75 (25%) | 26/75 (35%) | 22/75 (29%) | 25% to 35% | 30% |
| D | 0/75 (0%) | 2/75 (3%) | 1/75 (1%) | 0/75 (0%) | 0% to 3% | 1% |
The calculation
Per assistant, over quiet periods k:
rate_k = mentions_k / answers_k
quiet_range = min(rate_k) .. max(rate_k)
pooled_mean = sum(mentions_k) / sum(answers_k)
se = sqrt(pooled_mean * (1 - pooled_mean) / answers_per_period)
New period:
flag if rate_new is outside quiet_range
AND |rate_new - pooled_mean| > about 2 x se
confirm only if it holds in the next period
and the control set moved by lessReading a new period
Say period 5 follows a batch of page changes. Assistant A comes in at 20/75, about 27%. That's roughly nine points above its pooled mean and five above its highest quiet period; its standard error at 17% on 75 answers is about 4.4 points, so the move is a little over two of them. Worth investigating. Assistant C comes in at 27/75, 36%. That's six points above its mean, but only one above its own highest quiet period, and C's standard error at 30% is about 5.3 points. That's barely outside C's quiet range and only about one standard error from its mean: not enough to flag. Similar-looking moves, different verdicts, and the difference is entirely in each assistant's own history.
Why the formula understates the noise
The binomial standard error assumes independent draws. Repeated runs of the same prompt in the same window probably aren't fully independent, since they share a model version and an index state, and prompts differ from one another. The formula also knows nothing about platform changes. So treat it as a floor on the noise, and trust the observed quiet-period range more than the formula whenever the two disagree.
What counts as a quiet period
One with no shipped changes to the pages your prompts map to, no change to the prompt set or the sampling surfaces, and no platform change you know of. Perfectly quiet periods are rare. Use the quietest you have, label the worksheet with what did happen in each, and accept that a noisy history produces a wide range, which is itself an honest finding.
FAQ
Can I compute this per prompt? — Only with more runs than most teams have. With three runs a period, a single prompt's rate can only be 0%, 33%, 67% or 100%, so per-prompt variation is mostly arithmetic. Work per assistant first.
What if we've already shipped changes everywhere? — Then this prompt set has no quiet periods, and normal variation is unknown. Say that, and set up a control set before the next round.
Does a bigger sample fix it? — It narrows sampling noise, which falls with the square root of the number of runs: four times the runs halves the standard error. It doesn't touch platform change, which no sample size removes.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.