A robots.txt Review Checklist for Search and Answer Engines
Review robots.txt in five passes: inventory every user-agent rule and date when each was added (2023-era wildcard blocks catch 2026 crawlers nobody considered); audit paths per agent against what you actually want retrievable; check for the robots-vs-noindex trap (blocking crawl on a page you want deindexed prevents engines from ever seeing the noindex); fetch the file as each major AI agent to confirm the CDN serves them the same file; and verify the sitemap line points at a sitemap that parses. Most sites' files encode a policy nobody currently at the company chose.
robots.txt is the oldest contract on the web and the least reviewed: written once, edited in emergencies, inherited across replatforms. Meanwhile the population of agents reading it changed completely — a file last reviewed before AI crawlers existed is running your AI visibility policy by accident.
The five passes
- Agent inventory: list every User-agent block and answer, per block, 'who added this and why?' Unattributable blocks are policy debt — especially wildcards that predate AI crawlers and now govern them silently
- Path audit per agent: for each allowed agent, walk the disallow list against your current sitemap. Sections launched since the last review are governed by default rules nobody checked
- The noindex trap: a page you want OUT of indexes must be crawlable so engines can see its noindex — disallowing it in robots.txt preserves stale index entries forever. Blocking and deindexing are opposite tools
- Serve-check as each agent: curl the file with GPTBot/ClaudeBot/PerplexityBot user-agents — edge products sometimes serve bots a different response, and the file your editor shows isn't necessarily the file bots receive
- Sitemap line: present, absolute URL, and the target parses. A robots.txt pointing at a 404 sitemap is a broken handshake at the front door
A worked review
Real pattern from the field: a marketing site's file contained a 2023 'User-agent: * / Disallow: /api/' pair (fine), a CCBot block someone added during a scraping scare (defensible, but it also expresses a training-data policy — keep or reconsider deliberately), and no explicit rules for GPTBot or ClaudeBot — meaning the wildcard governs them, which happened to be permissive. The review's output wasn't edits; it was a one-paragraph written policy ('search and answer crawlers allowed, bulk scrapers blocked, /api/ off-limits') that the file now provably implements. The paragraph is the deliverable; the file is its compilation.
A failure worth checking
The inherited wildcard: a replatform copies robots.txt from the old CMS, including a 'Disallow: /' under some User-agent added for a long-dead reason — or worse, a staging file ships to production. Entire engines quietly lose access, and because robots.txt failures produce no errors anywhere, the symptom is months of unexplained absence. The five-pass review catches it in minutes; nothing else reliably does.
FAQ
Should I block AI training crawlers but allow answer crawlers? — That split is a real policy many sites legitimately choose (block CCBot-class bulk collectors, allow GPTBot/ClaudeBot-class retrieval). The review's job isn't to pick your policy — it's to guarantee the file implements the one you picked, on purpose, currently.
How often should robots.txt be reviewed? — On every replatform or CDN change (the high-risk moments), when a new significant crawler emerges, and annually regardless. Put the written policy paragraph in the repo next to the file; future reviewers thank you.
Does robots.txt affect AI citations directly? — It gates retrieval, which gates everything: a blocked crawler can't cite what it can't read. Beyond that gate, citations depend on content and competition — robots.txt is the door, not the pitch.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.