Skip to content
Fenn

All posts

An AI Visibility Baseline Worksheet for a New SaaS Website

6 min readAEOMeasurement

A baseline is a fixed set of prompts, run against named engines in two dated sampling windows, with every mention and cited URL logged before you change a single page. Without one, any later 'improvement' is unfalsifiable. The worksheet below has the exact fields.

The most common AI-visibility mistake isn't a bad tactic — it's starting tactics before measuring anything. If you don't know how often engines name you today, nothing you observe next month means anything. A baseline turns 'we feel more visible' into a comparison between two dated snapshots.

Scope: what a baseline does and doesn't claim

A baseline measures appearance in sampled answers: whether your brand is mentioned, and whether any of your URLs are cited, for a fixed prompt set on named engines during a dated window. It does not claim market share, rank, or user behaviour — answers vary between runs, so a baseline is a sample, not a census. Record enough runs to see the variance rather than pretending it away.

The worksheet fields

  • Prompt set: the fixed list (25 is a workable size) with a version number — changing a prompt means a new version, never an edit in place
  • Engines: each engine and mode you sample (with or without web browsing), named per row
  • Window: start and end date-time for each sampling pass; two windows minimum, a week apart
  • Runs per prompt: how many times each prompt was asked per engine per window
  • Mention: brand named in the answer text (yes/no per run)
  • Citation: one of your URLs linked or referenced as a source (record the exact URL)
  • Owner and evidence pointer: who ran it, and where the raw answer text is stored

A filled example — fictional, for shape only

Lumina Desk is a fictional B2B scheduling SaaS we use to show the worksheet's shape; these numbers are illustrative, not observed results. Prompt set v1 (25 prompts), two engines with browsing enabled, three runs per prompt per window. Window A (Sept 1–2): mentioned in 4 of 150 runs, zero URL citations. Window B (Sept 9–10): mentioned in 6 of 150 runs, one citation of its pricing page. The value isn't the numbers — it's that every one of them is attached to a prompt id, an engine, a window and a stored answer.

A failure worth checking

The classic self-inflicted wound: someone rewrites three prompts between windows because the answers 'weren't relevant'. The comparison is now meaningless, and worse, it looks like a visibility change. If a prompt is wrong, version the whole set and restart both windows. Prompt drift is the baseline killer.

What would change the decision

If your category answers rarely cite any vendor at all, a mention-heavy baseline with zero citations is normal, not a crisis — the decision shifts from 'get cited' to 'get named accurately'. And if answers vary wildly between runs of the same prompt, increase runs per window before drawing any conclusion.

Put this to work on your own website.

Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.