Skip to content
Fenn

All posts

How to Chart AI Visibility Without Overstating It

5 min readAEOReportingMeasurement

Chart the rate with its denominator in the label, show the spread across repeated runs instead of a single line, mark every prompt-set version change and every known platform change on its date, start the y-axis at zero for rates, never join points across a method change, and plot each assistant separately rather than as one blended 'AI' line. Keep first-party counts from Google or Bing on their own panel with their own unit. Each rule removes one way a correct number turns into a misleading picture. The test is simple: someone who sees only the chart, with nobody narrating, shouldn't come away believing more than the data supports.

Most overstatement in AI visibility reporting isn't in the numbers. It's in the chart. A line that climbs straight through a prompt-set change. An axis starting at 10% that turns a three-point move into a cliff. One blended series hiding the single assistant that did all the moving.

The seven rules

RuleHowWhat it prevents
Show the denominatorAxis label or footnote: 'share of 300 sampled answers'A sample rate read as a population figure
Show the spreadEach run as a faint point, the period mean as the lineRun-to-run noise read as a trend
Mark method changesA vertical rule at each prompt-set version changeA redefinition read as a movement
Mark platform changesAnnotate dated vendor announcementsA vendor update read as your result
Zero-based axis for rates0% to a sensible ceiling, same on every panelThree points drawn as a cliff
No line across breaksStart a new series after a method changeTwo instruments joined into one trend
One series per assistantSmall multiples on a shared scaleOne assistant's move hidden in an average

Showing spread without a statistics lecture

The simplest honest version: plot every run's result as a faint point and draw the period mean through them. A reader sees at once whether the points sit within a couple of points of the line or scatter across twenty. If you want an interval instead, a standard binomial interval for a proportion, such as the Wilson interval, gives a reasonable width. State its limit, though: repeated runs of one prompt in one window share the same model and index state, so they're probably not fully independent, and the true uncertainty is likely wider than the interval shows.

An illustrative before and after

Before: one line labelled 'AI visibility' rising from 12% to 21% over six weeks, y-axis running 10% to 22%. After, same data: four small panels, one per assistant, each on a 0% to 40% axis, runs as faint points. A vertical rule in week three reads 'prompt set v2 to v3'. The whole rise sits after that rule, in one assistant. The first chart says visibility nearly doubled. The second says something changed when we changed the questions, in one product, and we can't yet tell whether it's real. Both are drawn from identical, illustrative numbers.

Where first-party counts go

Search Console's generative AI report counts impressions; Bing Webmaster Tools' AI Performance report counts citations. Give each its own panel, labelled with its unit and source, below the sampled panels. Never put either on the same y-axis as a sampled rate. They answer different questions, and a shared axis invites the reader to compare heights that aren't comparable.

The chart spec to copy

Chart: [metric] by assistant, [date range]
Unit: share of sampled answers (denominator per point: __)
Y-axis: 0% to __% (never truncated for rates)
Series: one per assistant, shared scale; no blended 'AI' line
Points: each run faint; period mean as the line
Breaks: new series after prompt-set change (v__ on __)
Annotations: platform changes on __ (source: vendor release notes)
First-party counts: separate panel, own unit, own source line
Footnote: prompt set v__, __ prompts, __ runs, sampled __ to __

FAQ

Is a blended all-assistants line ever useful? — As a secondary figure, with the per-assistant panels beside it. On its own it averages products that behave differently and can sit flat while one assistant rises and another falls.

Where do I find platform changes to annotate? — Some are published. OpenAI keeps ChatGPT release notes, Anthropic publishes release notes for its API and for the Claude apps, Google publishes Gemini app release notes and a Search Status Dashboard for ranking updates. Many changes are never announced, so a step change with no annotation is unexplained, not evidence of your own work.

Should competitors be on the chart? — On the same denominator and the same scale, yes. It's the cheapest context there is. A competitor series drawn from a different prompt set is decoration.

Put this to work on your own website.

Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.