A Repeated-Sampling Protocol for AI Visibility Experiments
Fix five parameters and do not change them mid-experiment: the prompt set, the engines, the number of runs per prompt per period (three to five is the practical floor), the cadence (weekly is usually right — daily mostly samples noise, monthly misses the window in which you could act), and the counting rule for mention versus citation. Record every run as a row — prompt, engine, timestamp, run number, outcome, and the cited URL where there is one — rather than recording a summary. Summaries cannot be re-analysed and cannot answer a question you did not think of at the time; rows can. The change that quietly invalidates everything is editing the prompt set mid-experiment, because your denominator moved and the two periods are no longer comparable.
This is the boring infrastructure under every honest AI-visibility claim. Without it you have impressions of a trend; with it you have something two people can disagree about productively.
The five fixed parameters
| Parameter | Practical default | Why it matters |
|---|---|---|
| Prompt set | 20 buyer-intent prompts | the denominator — changing it breaks comparison |
| Engines | the ones your buyers use | each behaves differently; mixing them hides that |
| Runs per prompt | 3-5 per period | distinguishes presence from luck |
| Cadence | weekly | daily samples noise, monthly misses the window |
| Counting rule | mention vs citation, written down | the most common source of disagreement |
Record rows, not summaries
One row per run: prompt id, engine, timestamp, run number, outcome (cited, mentioned, absent), and the URL if cited. A summary of '60% presence this week' cannot answer which prompts moved, whether one engine drove it, or whether a competitor appeared — and those are precisely the questions the summary prompts. Rows cost nothing extra to collect and preserve every future question.
What to compute
- Presence rate per prompt: runs cited divided by runs sampled
- Presence rate overall: the mean across prompts, reported with the sample size
- Stability: how much each prompt's rate moves between periods, which tells you how much movement is normal for you
- Citation share: where your URL is cited rather than merely mentioned
- Competitor presence in the same runs, which is your control for platform-wide shifts
The invalidating change
Editing the prompt set mid-experiment. Adding three prompts you expect to win on lifts the average with no change in the world, and it is very easy to do accidentally while tidying. If the set must change, start a new baseline and say so — a visible discontinuity is honest, an invisible one is not.
FAQ
Is this worth it for a small site? — The protocol is the same at any size; the prompt set is just shorter. Ten prompts sampled properly beats fifty sampled once, because the first can be compared over time and the second cannot.
Can I automate it? — Yes, and most teams should once the manual version has taught them what the numbers mean. Whatever you automate, keep the row-level output — a tool that only exposes summaries removes the analysis you will eventually want.
How long before the data is useful? — Three to four periods before trends mean much. One period is a baseline, two is a line between two points, and neither is a trend.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.