Skip to content
Fenn

All posts

When to Stop an AEO Experiment

5 min readAEOMeasurement

Write four stopping rules before starting: a date, a result threshold that would count as success, a futility threshold at which you stop regardless, and a validity condition whose failure ends the experiment immediately — a changed prompt set, a platform update mid-run, a shipped change to the control topics. Decided in advance, these make stopping a rule rather than a judgement made by whoever is most invested. The rule teams omit is futility, and its absence is why AEO experiments run for quarters producing a result described as 'still early' every time it is reviewed.

The hardest part of measuring this channel is that a null result looks identical to an experiment that has not finished. Without a pre-committed stopping point, the ambiguity always resolves in favour of continuing.

The four rules

RuleSet it toStops the experiment when
DateA specific review dateThat date arrives, regardless of result
Success thresholdA movement larger than normal variationObserved and reproduced across runs
Futility thresholdA point where no plausible effect remainsReached — stop and say so
Validity conditionPrompt set, control set and method unchangedAny of them changes mid-run

Futility is the rule that gets omitted

It is the only one that can conclude 'this did not work', which is why it is left out. Set it concretely: if after N periods the treatment has not diverged from the control by more than the normal period-to-period variation, stop. That is a real finding and it frees the effort for something else.

The validity condition matters more here than elsewhere

You do not control the platform. A model update mid-experiment invalidates the comparison, and continuing past it produces a number that mixes two different systems. When the validity condition fails, the experiment ends — it does not pause and resume, because the periods are no longer comparable.

What to do with an inconclusive result

Record it as inconclusive, with what would have been needed to conclude — more runs, more prompts, a longer period, a larger effect. That sentence is what makes the next experiment better designed rather than the same experiment run again hoping for a clearer answer.

Our position

A measurement vendor benefits from experiments running longer, and we would rather say plainly that many AEO experiments should be stopped earlier than they are. A programme that ends three experiments honestly learns more than one that keeps six alive and describes all of them as promising.

FAQ

Can we extend an experiment? — Once, decided before you see the current result, with a stated reason. Extending after seeing the data is choosing the stopping point to suit the answer, which is the specific thing the rules exist to prevent.

What if the result is positive but small? — Check it against your own normal period-to-period variation before calling it anything. Small and within noise is not a result; small and outside noise, reproduced, is.

Is a null result worth reporting? — Yes, and more so in a channel this young. Most published AEO claims are positive, which tells you about publication rather than about what works.

Put this to work on your own website.

Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.