How to Run a Fair Two-Week Evaluation of an AEO Platform
Fix the inputs before the trial starts and change nothing during it: the same 20 buyer-intent prompts for every vendor, the same engines, the same definition of mention versus citation, and a hand-checked baseline recorded before day one. Then run each platform for two weeks on those inputs and record three things daily — what it reported, what it proposed or shipped, and what you could independently verify. Judge on reconcilability and on work that reached production, not on dashboard polish. The trap is treating a stochastic system as deterministic: a single prompt checked once proves nothing, which is why the protocol samples repeatedly and compares distributions rather than snapshots.
A trial run on vendor-supplied prompts, vendor-chosen engines and vendor framing measures the vendor. That is a demo with a longer runtime. A fair evaluation holds the inputs still and lets the products differ.
Before day one
- Write the prompt set: about 20 questions a real buyer would ask, in their words, covering problem, comparison and pricing intent. This is yours and it goes to every vendor unchanged
- Pick the engines that match your buyers and freeze the list — adding one mid-trial invalidates the comparison
- Write down the counting rule: what a mention is, what a citation is, and how a partial or hedged reference is scored
- Hand-check the full prompt set once and record the result. Without this baseline you cannot tell a platform's reading from reality
- Decide the decision criteria now, in writing, including what would make you buy nothing
During the two weeks
| Record daily | Why it matters |
|---|---|
| What the platform reported | the claim under test |
| What it proposed or shipped | separates measurement from execution |
| What you verified by hand | the only number you fully trust |
| Anything that needed support | support quality is part of the product |
| Time you spent | a tool that saves reporting time but costs review time is not a saving |
The stochastic trap
These engines do not return the same answer twice, and a fair share of apparent movement inside two weeks is resampling rather than anything either party did. So do not judge on a snapshot: compare the shape of results across repeated samples, treat single-prompt changes as anecdote, and be suspicious of any vendor whose trial narrative attributes a two-week improvement to their product. We would be making that up too.
What two weeks can and cannot prove
It can prove: whether the platform's numbers reconcile with your hand-checks, whether proposed changes are sane and reviewable, whether anything actually shipped, and how the product behaves when you disagree with it. It cannot prove ranking or citation outcomes — those move on the engines' schedule, not yours. Anyone promising demonstrated visibility gains inside a fortnight is describing noise, and treating that promise as disqualifying will save you a year.
Scoring at the end
Score three things: reconcilability (could you reproduce their numbers by hand?), shipped work (did anything reach production, with an approval trail?), and friction (how much of your team's time did it consume?). A platform that wins on dashboards and loses on all three is a reporting subscription. Run Fenn through exactly this protocol — our own evidence-log template is designed to be the artifact you keep afterwards, whichever way the decision goes.
FAQ
Is two weeks long enough? — For process and reconcilability, yes. For outcomes, no, and no protocol fixes that. Buy on process quality and verify outcomes over the following quarter.
Should I trial two platforms at once? — Yes, on identical inputs, provided each has its own approval path so shipped changes do not interleave. Sequential trials drift because the market moves between them.
What if the numbers disagree with my hand-check? — Ask for the sampling methodology in writing. A defensible gap has an explanation (different engines, different refresh, different citation rule). An unexplained gap is the finding.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.