Why the Same AI Search Prompt Can Produce Different Citations
Five things vary between two runs of the same prompt: sampling inside the model itself, which pages the retrieval step happens to fetch that time, personalisation and location attached to the session, the model or system-prompt version you were routed to, and the live index, which changes underneath everyone continuously. Only the last is something your work changes. The practical consequence is that a single check of a single prompt is an anecdote, and a visibility number is only meaningful with its sample size and date attached — which is why a report that says 'we are cited for X' without saying how many times it was checked is not reporting a measurement.
This surprises people who arrive from rank tracking, where the same query on the same day from the same location returns a stable answer. Generated answers are not that, and treating variance as a bug leads to chasing changes that were never real.
The five causes
| Cause | What it looks like | Under your control? |
|---|---|---|
| Model sampling | same sources, different wording and order | No |
| Retrieval variance | different sources cited for the same question | Indirectly — via retrievability |
| Personalisation / location | different results per user or region | No |
| Model or prompt version | step change on a date, then stable | No |
| Index freshness | new sources appear over days or weeks | Yes — this is where your work lands |
Telling them apart
Wording changes with the same sources is sampling, and it is noise. Different sources for the same question is retrieval variance — worth watching, because a page that appears in some runs and not others is a page on the edge of being retrievable rather than a page that is invisible. A step change on a specific date that then holds is usually a model or system change, not your content; check whether competitors moved at the same moment, because if everyone shifted, you did not cause it. A gradual appearance over days after you published is index freshness, which is the one pattern that can legitimately be attributed to work.
What this means for measurement
- Sample repeatedly — a prompt checked once is an anecdote, and three to five runs per prompt per period is the minimum that distinguishes presence from luck
- Report presence as a rate, not a binary: cited in 4 of 5 runs is a different fact from cited once
- Always attach sample size, date and engine to any number you publish internally, or it cannot be compared to the next one
- Treat single-run changes as noise until they persist across a period
- Watch competitors in the same sample, since a shift that moves everyone is a platform change rather than a result
The honest implication for vendors, us included
Any tool reporting a clean binary — you are cited, you are not — is either sampling repeatedly and hiding the distribution, or checking once and presenting an anecdote as a measurement. Ask which. It is a fair question to ask us too, and the answer should be a methodology rather than a reassurance.
FAQ
Does this mean AI visibility cannot be measured? — It means it is measured as a rate over repeated samples rather than as a position. That is a familiar shape from survey and experiment work, and it is perfectly rigorous once you stop expecting rank-tracker determinism.
How many samples are enough? — Enough that a one-run change does not move your headline number. Three to five per prompt per period is a practical floor for most teams; more if you are making a decision on a small difference.
Should I report a range instead of a number? — Reporting the rate and the sample size gives your reader the same information with less ceremony. What matters is that the denominator is visible.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.