How to Separate an AI Model Change From a Website Improvement
Use a control group you never touch. Split your prompt set in two before making changes: a treatment set covering topics whose pages you are about to improve, and a control set covering topics you will deliberately leave alone. Sample both on the same schedule, then compare the *size* of each change rather than its direction — the difference between them is the quantity that carries information. Treatment up 30 points while control is up 5 is a 25-point gap worth investigating; treatment and control both up 20 is not. Direction alone tells you almost nothing, because a shared trend and a real effect can happen at once. Without a control set every visibility change has two explanations and no way to weigh them, which is why a reported lift with no control is a story rather than evidence.
This is the single most useful discipline in AEO measurement and the most commonly skipped, because it requires deciding in advance not to improve some pages for a while.
Setting it up
- Split the prompt set before you change anything — retrofitting a control after you have shipped is not a control
- Match the sets roughly on competitiveness and topic, so you are not comparing your easiest questions with your hardest
- Freeze the control topics: no content changes, no new internal links, no metadata edits for the measurement period
- Sample both sets on the same days, on the same engines, with the same counting rule
- Write down the expected effect and the period before starting, so you are not choosing the comparison after seeing the data
Reading the result: compare sizes, not directions
The instinct is to read the two sets as up or down and conclude from the pattern. That is wrong often enough to be dangerous. Treatment rising from 20% to 50% while control rises from 20% to 25% is two ups — and a 25-point gap between them that is very much worth investigating. A shared trend and a real effect are not mutually exclusive. What carries information is the difference between the two changes, read against how much each set normally moves.
| Gap between the changes | What it supports | What to check first |
|---|---|---|
| Treatment change much larger than control | An effect worth investigating | Is the gap bigger than each set's normal period-to-period swing? |
| Treatment change much smaller than control | Your work may be underperforming, or the sets are mismatched | Were the sets comparable to begin with? |
| Gap close to zero, both moved | A shared trend; no evidence either way about your work | Did the sets move for the same reason, or coincidentally? |
| Gap close to zero, neither moved | No signal in this period | Was the period long enough to show anything? |
| Treatment fell further than control | Possible harm — the case to act on quickly | Reproduce before reverting; check sampling depth |
The assumptions this comparison rests on
Two of them, and they are worth stating in any report that uses this method. **Parallel trends**: the two sets would have moved together had you changed nothing — which you cannot prove, but you can make more plausible by checking that they tracked each other for a period before the intervention. **No spillover**: improving the treatment pages did not also help the control topics, which is a real risk when the pages share a site, internal links or a domain-level reputation signal. Spillover makes a genuine effect look smaller than it is. The approach is a rough difference-in-differences read, and it inherits that method's assumptions along with its usefulness.
The case that saves the most work
Both sets falling by a similar amount. Without a control, a broad decline reads as a content failure and triggers a quarter of rewriting that changes nothing, because the cause was a model update. Note the caveat though: a similar-sized fall in both is consistent with no effect *and* with a real effect that a platform decline is masking. It is grounds not to panic, not proof that nothing happened.
What a control cannot do
It cannot give you statistical significance in the strict sense — you do not control assignment, the engines are not a stable population, and your sample sizes are small. It cannot prove causality at all. What it gives you is more modest and still valuable: a way to weigh the most common alternative explanation against your own, with the size of the gap and your known variation as the weights. State it that way in reports. Claiming causality from a two-set comparison is the overreach that makes the next honest claim harder to believe.
FAQ
How big should the control set be? — Similar in size to the treatment set, and never so small that one prompt dominates it. If you only have twenty prompts total, ten and ten is better than eighteen and two.
Can I ever improve the control pages? — Yes, after the measurement period ends. Freezing them permanently would mean deliberately underserving part of your market to preserve an experiment, which is the wrong trade.
What if I have already shipped changes everywhere? — Then you cannot attribute this period, and saying so is the honest reporting move. Set up the split before the next round rather than constructing a retrospective control.
Both sets went up. Is that definitely not my work? — No, and reading it that way discards real results. Compare the sizes: if treatment rose far more than control, and the gap is larger than either set normally swings between periods, there is something to investigate. A shared platform trend can carry a genuine effect along with it.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.