Skip to content
Fenn

All posts

Why the Same AI Search Prompt Can Produce Different Citations

6 min readAEOMeasurement

Five things vary between two runs of the same prompt: sampling inside the model itself, which pages the retrieval step happens to fetch that time, personalisation and location attached to the session, the model or system-prompt version you were routed to, and the live index, which changes underneath everyone continuously. Only the last is something your work changes. The practical consequence is that a single check of a single prompt is an anecdote, and a visibility number is only meaningful with its sample size and date attached — which is why a report that says 'we are cited for X' without saying how many times it was checked is not reporting a measurement.

This surprises people who arrive from rank tracking, where the same query on the same day from the same location returns a stable answer. Generated answers are not that, and treating variance as a bug leads to chasing changes that were never real.

The five causes

CauseWhat it looks likeUnder your control?
Model samplingsame sources, different wording and orderNo
Retrieval variancedifferent sources cited for the same questionIndirectly — via retrievability
Personalisation / locationdifferent results per user or regionNo
Model or prompt versionstep change on a date, then stableNo
Index freshnessnew sources appear over days or weeksYes — this is where your work lands

Telling them apart

Wording changes with the same sources is sampling, and it is noise. Different sources for the same question is retrieval variance — worth watching, because a page that appears in some runs and not others is a page on the edge of being retrievable rather than a page that is invisible. A step change on a specific date that then holds is usually a model or system change, not your content; check whether competitors moved at the same moment, because if everyone shifted, you did not cause it. A gradual appearance over days after you published is index freshness, which is the one pattern that can legitimately be attributed to work.

What this means for measurement

  1. Sample repeatedly — a prompt checked once is an anecdote, and three to five runs per prompt per period is the minimum that distinguishes presence from luck
  2. Report presence as a rate, not a binary: cited in 4 of 5 runs is a different fact from cited once
  3. Always attach sample size, date and engine to any number you publish internally, or it cannot be compared to the next one
  4. Treat single-run changes as noise until they persist across a period
  5. Watch competitors in the same sample, since a shift that moves everyone is a platform change rather than a result

The honest implication for vendors, us included

Any tool reporting a clean binary — you are cited, you are not — is either sampling repeatedly and hiding the distribution, or checking once and presenting an anecdote as a measurement. Ask which. It is a fair question to ask us too, and the answer should be a methodology rather than a reassurance.

FAQ

Does this mean AI visibility cannot be measured? — It means it is measured as a rate over repeated samples rather than as a position. That is a familiar shape from survey and experiment work, and it is perfectly rigorous once you stop expecting rank-tracker determinism.

How many samples are enough? — Enough that a one-run change does not move your headline number. Three to five per prompt per period is a practical floor for most teams; more if you are making a decision on a small difference.

Should I report a range instead of a number? — Reporting the rate and the sample size gives your reader the same information with less ceremony. What matters is that the denominator is visible.

Put this to work on your own website.

Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.