When to Stop an AEO Experiment
Write four stopping rules before starting: a date, a result threshold that would count as success, a futility threshold at which you stop regardless, and a validity condition whose failure ends the experiment immediately — a changed prompt set, a platform update mid-run, a shipped change to the control topics. Decided in advance, these make stopping a rule rather than a judgement made by whoever is most invested. The rule teams omit is futility, and its absence is why AEO experiments run for quarters producing a result described as 'still early' every time it is reviewed.
The hardest part of measuring this channel is that a null result looks identical to an experiment that has not finished. Without a pre-committed stopping point, the ambiguity always resolves in favour of continuing.
The four rules
| Rule | Set it to | Stops the experiment when |
|---|---|---|
| Date | A specific review date | That date arrives, regardless of result |
| Success threshold | A movement larger than normal variation | Observed and reproduced across runs |
| Futility threshold | A point where no plausible effect remains | Reached — stop and say so |
| Validity condition | Prompt set, control set and method unchanged | Any of them changes mid-run |
Futility is the rule that gets omitted
It is the only one that can conclude 'this did not work', which is why it is left out. Set it concretely: if after N periods the treatment has not diverged from the control by more than the normal period-to-period variation, stop. That is a real finding and it frees the effort for something else.
The validity condition matters more here than elsewhere
You do not control the platform. A model update mid-experiment invalidates the comparison, and continuing past it produces a number that mixes two different systems. When the validity condition fails, the experiment ends — it does not pause and resume, because the periods are no longer comparable.
What to do with an inconclusive result
Record it as inconclusive, with what would have been needed to conclude — more runs, more prompts, a longer period, a larger effect. That sentence is what makes the next experiment better designed rather than the same experiment run again hoping for a clearer answer.
Our position
A measurement vendor benefits from experiments running longer, and we would rather say plainly that many AEO experiments should be stopped earlier than they are. A programme that ends three experiments honestly learns more than one that keeps six alive and describes all of them as promising.
FAQ
Can we extend an experiment? — Once, decided before you see the current result, with a stated reason. Extending after seeing the data is choosing the stopping point to suit the answer, which is the specific thing the rules exist to prevent.
What if the result is positive but small? — Check it against your own normal period-to-period variation before calling it anything. Small and within noise is not a result; small and outside noise, reproduced, is.
Is a null result worth reporting? — Yes, and more so in a channel this young. Most published AEO claims are positive, which tells you about publication rather than about what works.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.