How to Verify an AI Search Crawler Can Retrieve Your Pages
Verify retrieval in four checks per important URL: fetch it with each AI crawler's user-agent (curl -A 'GPTBot' etc.) and confirm a 200 with full HTML, not a challenge page; confirm robots.txt allows that agent on that path; confirm your answer content exists in the raw server response, not only after JavaScript; and check your CDN/WAF logs for bot-mitigation rules that block silently at the edge before your origin ever sees the request. A page that fails any one of these is invisible to that engine no matter how well-written it is.
AI visibility work has a brutal prerequisite: the crawler has to be able to read the page. Teams polish answer-first content for months while their WAF serves every AI crawler a 403 — a failure mode that produces zero error messages on your side and zero citations on theirs. Retrieval verification is twenty minutes; do it before any content work.
The four checks
- User-agent fetch: curl -A with each crawler's published UA string (GPTBot, ClaudeBot, PerplexityBot, Google-Extended for AI features). Expect 200 + your actual HTML. A 403, a CAPTCHA page, or a 200 with challenge markup all mean blocked
- robots.txt: confirm each agent is allowed on the paths that matter — and that a well-meaning wildcard block from 2024 isn't still there
- Server-rendered content: fetch with JS disabled (curl is exactly this) and search the response for your answer text. Several AI crawlers execute little or no JavaScript — content that only exists post-hydration doesn't exist for them
- Edge interference: CDN bot-mitigation, WAF rules, and rate limits often act before your origin logs anything. Check the edge product's own logs for these UAs specifically
A worked check on one URL
Take your best answer page. curl -A 'GPTBot' -sI the URL: status 200? curl -A 'GPTBot' -s the URL piped through grep for your key answer sentence: present? Repeat for ClaudeBot and PerplexityBot — policies differ per bot, so one passing proves one. Then read robots.txt as each agent would. Then check the CDN dashboard's bot report for those UAs over the last week: served or challenged? Fifteen minutes, and you now know something most sites optimizing for AI answers never checked.
A failure worth checking
The silent edge block: bot mitigation enabled years ago for scraper defense, quietly classifying AI crawlers as threats. Origin logs show nothing (requests die at the edge), analytics show nothing (these bots don't run analytics JS), and the only symptom is absence — no citations, ever, from engines whose crawlers never got a page. If your citation count across all engines is exactly zero despite real content, check the edge first, not the content.
FAQ
Should I allow all AI crawlers? — That's a policy choice, not a technical default: training crawlers and search/answer crawlers serve different purposes, and blocking the former while allowing the latter is a legitimate stance. What's not legitimate is thinking you allow them while your WAF disagrees — verify whatever policy you chose is the one actually running.
How often should retrieval be re-verified? — On every CDN/WAF config change, every robots.txt edit, and quarterly regardless — edge products update their bot classifications without asking you. The four checks script neatly; run them in CI if the stakes justify it.
Does verified retrieval guarantee citations? — No — it's the floor, not the ceiling. Retrieval makes you eligible; content quality, structure and competition decide the rest. But nothing above the floor matters while you're below it.
Put this to work on your own website.
Fenn finds what your customers ask, drafts the articles and site fixes, and measures what ChatGPT, Claude, Gemini, Perplexity and Grok say about you — with every change waiting for your approval.