Free AI crawler policy checker: which AI bots your robots.txt lets through, by role
Paste a domain. We read its robots.txt once and resolve every AI token the way a crawler would, then split them by what they do: retrieval agents that fetch pages for AI answers, and training crawlers or policy tokens that do not. No signup.
A readiness check based on one public file, not a measurement of whether AI answers mention you. A site that lets every crawler through can still go unmentioned; that is what the full grade and the ClarAI product look at.
The roles, in the operators' own words
Robots.txt tokens are not interchangeable. The three operators whose crawlers this tool checks each document a split between the agents that fetch pages for answers and the crawlers that collect content for training, and they say plainly which setting changes what. The checker follows those documents, not a guess.
Retrieval agents: these decide answer visibility
- OAI-SearchBot. OpenAI states that sites which disallow it will not appear in ChatGPT search answers.
- PerplexityBot. Perplexity states that it surfaces sites in Perplexity results and respects robots.txt.
- Claude-SearchBot. Anthropic states that blocking it may reduce a site's visibility in Claude's search responses.
- Claude-User. Anthropic states that blocking it may reduce visibility for user-directed web search.
Training crawlers and policy tokens: a policy choice
- GPTBot. OpenAI's training crawler; its setting is independent of OAI-SearchBot.
- ClaudeBot. Anthropic states that it collects content that may contribute to training.
- Google-Extended. Per Google, a policy token that governs Gemini training and grounding in the Gemini app and Vertex AI, not Search inclusion, and not AI Overviews or AI Mode.
- Googlebot (reported on its own row) is what Google's AI features crawl with; blocking it removes the site from Google Search and from AI Overviews and AI Mode.
The longer guide to every token, with the source quotes, is at the AI crawlers guide. Google's own guidance on what its AI features require is summarised in Google-Extended does not do what you think it does.
What this does not measure
- Whether any AI answer mentions, cites or recommends the site. Access is a readiness gate, not a result.
- Anything beyond robots.txt. Indexing, snippet eligibility, structured data and page speed are separate gates; the free grade covers them.
- Whether a crawler actually obeys the file. A robots.txt rule is a request that well-behaved crawlers honour, not a lock.
- User-triggered fetchers such as ChatGPT-User and Perplexity-User, whose operators say robots.txt may not apply.
Also free: the snippet controls checker (the Google gate), the full readiness grade, and the AI visibility checker explainer.
Questions
What does this checker actually read?
Why split retrieval agents from training crawlers?
Does blocking GPTBot remove my site from ChatGPT?
What about Google-Extended and AI Overviews?
Why are ChatGPT-User and Perplexity-User not in the list?
If every crawler is allowed, will AI answers mention my brand?
Do you store the domains people check?
Sources
Every claim about a crawler or a control on this page is limited to what the operator's own document states.
- OpenAI, Overview of OpenAI crawlers and fetchers: https://developers.openai.com/api/docs/bots
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Perplexity, Perplexity crawlers: https://docs.perplexity.ai/guides/bots
- Google, Overview of Google crawlers and fetchers (Google-Extended): https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
- Google, AI features and your website: https://developers.google.com/search/docs/appearance/ai-features
Access is the first gate. Grade the rest free.
The full readiness grade checks robots.txt together with indexing, snippet eligibility, structured data and the other technical gates Google documents. No signup, about twenty seconds.