A Monthly Operating Checklist for AI Visibility Tracking
Once a month, check the instrument before reading the numbers. Start with data quality: failed or refused runs and whether they stayed in the denominator, prompts or surfaces that changed without a version bump, the surface register's last-verified dates, and the month's platform changes. Then exports: archive the month's first-party AI data from Search Console and Bing Webmaster Tools and your own raw rows, because Google's report has row limits and a preview report's definitions can change. Then the numbers: each assistant's rate against its normal variation, the control set, the prompt-to-page map and open corrections. Last, decide queued prompt changes in one batch. Decisions about the programme itself belong in the quarterly review; the monthly pass is maintenance.
Read the numbers first and you'll spend the meeting explaining a spike that turns out to be forty runs that errored and were dropped. The order in this checklist is the point: it catches problems in the instrument before they turn into a story about the market.
The twelve checks, in order
| Order | Check | What a failure looks like |
|---|---|---|
| 1 | Failed and refused runs counted and handled as declared | A rate rose because errors were quietly dropped |
| 2 | Prompt set unchanged, or version bumped | A prompt edited in place mid-series |
| 3 | Surface register re-verified | A default model changed and the row didn't |
| 4 | Platform changes logged for the month | A step change with no annotation |
| 5 | Search Console generative AI data archived | Row limits hid pages; the month was never saved |
| 6 | Bing AI Performance figures archived | Preview definitions changed; old numbers lost |
| 7 | Raw rows backed up with the methodology note | A chart nobody can re-derive |
| 8 | Each assistant's rate read against its quiet-period range | An ordinary wobble reported as a trend |
| 9 | Control set compared with treatment by size of change | A shared move claimed as a result |
| 10 | Prompt-to-page map states updated | A backlog built on last quarter's states |
| 11 | Correction tracker re-tests run | A closed error quietly back |
| 12 | Prompt change requests decided in one batch | Five small edits, five broken comparisons |
Keep your own copy of every source
Google's help page for the generative AI report says the same 1,000-row limit applies as in the standard Performance report, and that the newest data can be preliminary. Bing's AI Performance report launched as a preview, and Microsoft has said it will keep refining its metrics. Neither is a history you control. Archive each month's figures, by export where the report offers one and by a dated screenshot where it doesn't. The same goes for any tool you use, ours included: Fenn exports visibility, prompts, competitors and sources to CSV, and a tool's history is only yours if you keep a copy.
The log to fill in
Monthly operating log: [month]
Runs: __ planned, __ completed, __ failed (kept in denominator as no-mention: yes/no)
Prompt set: v__ (changed this month: yes/no; bridge period __ to __)
Surfaces re-verified on __; changes: __
Platform changes logged: __
Archived: Search Console AI (__ rows), Bing AI Performance (export / screenshot), raw rows (__)
Assistants outside their quiet range: __
Control vs treatment gap: __ points (normal range __)
Corrections: __ re-tested, __ closed, __ reopened
Prompt change requests: __ accepted, __ rejected; next version v__What the monthly pass doesn't decide
Whether an experiment worked, whether to spend more, whether to change the programme's scope. Those need more than one month of evidence and belong in the quarterly review, where the control set and the corrections list are standing items. A monthly meeting that starts making strategy calls on one month of noisy data ends up reversing them the following month.
FAQ
What if a check fails? — Fix it before the report goes out, or send the report with the failure named. A month reported as 'instrument broken on these dates' is fine. A broken month reported as normal isn't.
Should this be weekly? — The data-quality checks can be, if you sample weekly. Keep the archive and the change-request batch monthly, because batching is what stops the prompt set drifting.
How long should it take? — Short, once it's routine. If it keeps running long, the usual cause is that rows aren't being stored as they're collected, and the fix is upstream of the checklist.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.