Skip to content
Fenn

All posts

A Monthly Operating Checklist for AI Visibility Tracking

5 min readAEOProgrammeMeasurement

Once a month, check the instrument before reading the numbers. Start with data quality: failed or refused runs and whether they stayed in the denominator, prompts or surfaces that changed without a version bump, the surface register's last-verified dates, and the month's platform changes. Then exports: archive the month's first-party AI data from Search Console and Bing Webmaster Tools and your own raw rows, because Google's report has row limits and a preview report's definitions can change. Then the numbers: each assistant's rate against its normal variation, the control set, the prompt-to-page map and open corrections. Last, decide queued prompt changes in one batch. Decisions about the programme itself belong in the quarterly review; the monthly pass is maintenance.

Read the numbers first and you'll spend the meeting explaining a spike that turns out to be forty runs that errored and were dropped. The order in this checklist is the point: it catches problems in the instrument before they turn into a story about the market.

The twelve checks, in order

OrderCheckWhat a failure looks like
1Failed and refused runs counted and handled as declaredA rate rose because errors were quietly dropped
2Prompt set unchanged, or version bumpedA prompt edited in place mid-series
3Surface register re-verifiedA default model changed and the row didn't
4Platform changes logged for the monthA step change with no annotation
5Search Console generative AI data archivedRow limits hid pages; the month was never saved
6Bing AI Performance figures archivedPreview definitions changed; old numbers lost
7Raw rows backed up with the methodology noteA chart nobody can re-derive
8Each assistant's rate read against its quiet-period rangeAn ordinary wobble reported as a trend
9Control set compared with treatment by size of changeA shared move claimed as a result
10Prompt-to-page map states updatedA backlog built on last quarter's states
11Correction tracker re-tests runA closed error quietly back
12Prompt change requests decided in one batchFive small edits, five broken comparisons

Keep your own copy of every source

Google's help page for the generative AI report says the same 1,000-row limit applies as in the standard Performance report, and that the newest data can be preliminary. Bing's AI Performance report launched as a preview, and Microsoft has said it will keep refining its metrics. Neither is a history you control. Archive each month's figures, by export where the report offers one and by a dated screenshot where it doesn't. The same goes for any tool you use, ours included: Fenn exports visibility, prompts, competitors and sources to CSV, and a tool's history is only yours if you keep a copy.

The log to fill in

Monthly operating log: [month]
Runs: __ planned, __ completed, __ failed (kept in denominator as no-mention: yes/no)
Prompt set: v__ (changed this month: yes/no; bridge period __ to __)
Surfaces re-verified on __; changes: __
Platform changes logged: __
Archived: Search Console AI (__ rows), Bing AI Performance (export / screenshot), raw rows (__)
Assistants outside their quiet range: __
Control vs treatment gap: __ points (normal range __)
Corrections: __ re-tested, __ closed, __ reopened
Prompt change requests: __ accepted, __ rejected; next version v__

What the monthly pass doesn't decide

Whether an experiment worked, whether to spend more, whether to change the programme's scope. Those need more than one month of evidence and belong in the quarterly review, where the control set and the corrections list are standing items. A monthly meeting that starts making strategy calls on one month of noisy data ends up reversing them the following month.

FAQ

What if a check fails? — Fix it before the report goes out, or send the report with the failure named. A month reported as 'instrument broken on these dates' is fine. A broken month reported as normal isn't.

Should this be weekly? — The data-quality checks can be, if you sample weekly. Keep the archive and the change-request batch monthly, because batching is what stops the prompt set drifting.

How long should it take? — Short, once it's routine. If it keeps running long, the usual cause is that rows aren't being stored as they're collected, and the fix is upstream of the checklist.

Put this to work on your own website.

Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.