Skip to content
Fenn

All posts

How to Run a Fair Two-Week Evaluation of an AEO Platform

6 min readAEOBuyer evaluation

Fix the inputs before the trial starts and change nothing during it: the same 20 buyer-intent prompts for every vendor, the same engines, the same definition of mention versus citation, and a hand-checked baseline recorded before day one. Then run each platform for two weeks on those inputs and record three things daily — what it reported, what it proposed or shipped, and what you could independently verify. Judge on reconcilability and on work that reached production, not on dashboard polish. The trap is treating a stochastic system as deterministic: a single prompt checked once proves nothing, which is why the protocol samples repeatedly and compares distributions rather than snapshots.

A trial run on vendor-supplied prompts, vendor-chosen engines and vendor framing measures the vendor. That is a demo with a longer runtime. A fair evaluation holds the inputs still and lets the products differ.

Before day one

  1. Write the prompt set: about 20 questions a real buyer would ask, in their words, covering problem, comparison and pricing intent. This is yours and it goes to every vendor unchanged
  2. Pick the engines that match your buyers and freeze the list — adding one mid-trial invalidates the comparison
  3. Write down the counting rule: what a mention is, what a citation is, and how a partial or hedged reference is scored
  4. Hand-check the full prompt set once and record the result. Without this baseline you cannot tell a platform's reading from reality
  5. Decide the decision criteria now, in writing, including what would make you buy nothing

During the two weeks

Record dailyWhy it matters
What the platform reportedthe claim under test
What it proposed or shippedseparates measurement from execution
What you verified by handthe only number you fully trust
Anything that needed supportsupport quality is part of the product
Time you spenta tool that saves reporting time but costs review time is not a saving

The stochastic trap

These engines do not return the same answer twice, and a fair share of apparent movement inside two weeks is resampling rather than anything either party did. So do not judge on a snapshot: compare the shape of results across repeated samples, treat single-prompt changes as anecdote, and be suspicious of any vendor whose trial narrative attributes a two-week improvement to their product. We would be making that up too.

What two weeks can and cannot prove

It can prove: whether the platform's numbers reconcile with your hand-checks, whether proposed changes are sane and reviewable, whether anything actually shipped, and how the product behaves when you disagree with it. It cannot prove ranking or citation outcomes — those move on the engines' schedule, not yours. Anyone promising demonstrated visibility gains inside a fortnight is describing noise, and treating that promise as disqualifying will save you a year.

Scoring at the end

Score three things: reconcilability (could you reproduce their numbers by hand?), shipped work (did anything reach production, with an approval trail?), and friction (how much of your team's time did it consume?). A platform that wins on dashboards and loses on all three is a reporting subscription. Run Fenn through exactly this protocol — our own evidence-log template is designed to be the artifact you keep afterwards, whichever way the decision goes.

FAQ

Is two weeks long enough? — For process and reconcilability, yes. For outcomes, no, and no protocol fixes that. Buy on process quality and verify outcomes over the following quarter.

Should I trial two platforms at once? — Yes, on identical inputs, provided each has its own approval path so shipped changes do not interleave. Sequential trials drift because the market moves between them.

What if the numbers disagree with my hand-check? — Ask for the sampling methodology in writing. A defensible gap has an explanation (different engines, different refresh, different citation rule). An unexplained gap is the finding.

Put this to work on your own website.

Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.