Free · AI crawler policy checker

Free AI crawler policy checker: which AI bots your robots.txt lets through, by role

Paste a domain. We read its robots.txt once and resolve every AI token the way a crawler would, then split them by what they do: retrieval agents that fetch pages for AI answers, and training crawlers or policy tokens that do not. No signup.

no signup · one file read · about five seconds
OAI-SearchBotPerplexityBotClaude-SearchBotClaude-UserGPTBotClaudeBotGoogle-ExtendedGooglebot

A readiness check based on one public file, not a measurement of whether AI answers mention you. A site that lets every crawler through can still go unmentioned; that is what the full grade and the ClarAI product look at.

The roles, in the operators' own words

Robots.txt tokens are not interchangeable. The three operators whose crawlers this tool checks each document a split between the agents that fetch pages for answers and the crawlers that collect content for training, and they say plainly which setting changes what. The checker follows those documents, not a guess.

Retrieval agents: these decide answer visibility

  • OAI-SearchBot. OpenAI states that sites which disallow it will not appear in ChatGPT search answers.
  • PerplexityBot. Perplexity states that it surfaces sites in Perplexity results and respects robots.txt.
  • Claude-SearchBot. Anthropic states that blocking it may reduce a site's visibility in Claude's search responses.
  • Claude-User. Anthropic states that blocking it may reduce visibility for user-directed web search.

Training crawlers and policy tokens: a policy choice

  • GPTBot. OpenAI's training crawler; its setting is independent of OAI-SearchBot.
  • ClaudeBot. Anthropic states that it collects content that may contribute to training.
  • Google-Extended. Per Google, a policy token that governs Gemini training and grounding in the Gemini app and Vertex AI, not Search inclusion, and not AI Overviews or AI Mode.
  • Googlebot (reported on its own row) is what Google's AI features crawl with; blocking it removes the site from Google Search and from AI Overviews and AI Mode.

The longer guide to every token, with the source quotes, is at the AI crawlers guide. Google's own guidance on what its AI features require is summarised in Google-Extended does not do what you think it does.

What this does not measure

  • Whether any AI answer mentions, cites or recommends the site. Access is a readiness gate, not a result.
  • Anything beyond robots.txt. Indexing, snippet eligibility, structured data and page speed are separate gates; the free grade covers them.
  • Whether a crawler actually obeys the file. A robots.txt rule is a request that well-behaved crawlers honour, not a lock.
  • User-triggered fetchers such as ChatGPT-User and Perplexity-User, whose operators say robots.txt may not apply.

Also free: the snippet controls checker (the Google gate), the full readiness grade, and the AI visibility checker explainer.

Questions

What does this checker actually read?
One file: https://yourdomain/robots.txt, fetched once with a normal browser user agent and held in memory for 24 hours so a repeat check costs nothing. It resolves each token the way a crawler does: a token's own User-agent group wins, otherwise the wildcard group applies, otherwise the token is allowed by default. It shows the exact line that produced each verdict. It does not crawl any page and it writes nothing to a database.
Why split retrieval agents from training crawlers?
Because the operators document them as different things. OpenAI states that sites which disallow OAI-SearchBot will not appear in ChatGPT search answers, while GPTBot is its training crawler and each setting is independent. Anthropic states that ClaudeBot collects content that may contribute to training, while blocking Claude-SearchBot or Claude-User may reduce visibility in Claude's search responses. Perplexity states that PerplexityBot surfaces sites in its results and respects robots.txt. Blocking a training crawler keeps you out of training; it does not by itself remove you from answers.
Does blocking GPTBot remove my site from ChatGPT?
Not on its own. Per OpenAI's crawler documentation, OAI-SearchBot is the token that decides whether a site appears in ChatGPT search answers, GPTBot is the training crawler, and the two settings are independent. A site that blocks GPTBot and allows OAI-SearchBot reads as ready here.
What about Google-Extended and AI Overviews?
Google-Extended is a robots policy token, not a crawler. Per Google it governs whether content may be used to train Gemini models and to ground answers in the Gemini app and Vertex AI, and it does not affect Search inclusion or ranking. AI Overviews and AI Mode source answers from the Search index, which standard Googlebot crawls, so Googlebot is the token that gates those features and it is reported on its own row.
Why are ChatGPT-User and Perplexity-User not in the list?
They are user-triggered fetchers: a person asked the assistant to open a page. OpenAI states that robots.txt rules may not apply to ChatGPT-User, and Perplexity states that Perplexity-User generally ignores robots.txt because a user requested the fetch. A rule for either cannot be relied on, so neither is scored; the checker says so on every result.
If every crawler is allowed, will AI answers mention my brand?
Not necessarily. Access is a readiness gate, not a result: a site every crawler can read can still go unmentioned, and a site that blocks a training crawler can still be cited. This tool never measures mentions or citations. The free readiness grade covers the other technical gates, and the ClarAI product measures the answers themselves, with the sample size printed on every number.
Do you store the domains people check?
The result is held in the server's memory for 24 hours so a repeat check is free, and it is dropped after that. No account is created, no email is asked for, and nothing is written to a database.

Sources

Every claim about a crawler or a control on this page is limited to what the operator's own document states.

  1. OpenAI, Overview of OpenAI crawlers and fetchers: https://developers.openai.com/api/docs/bots
  2. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
  3. Perplexity, Perplexity crawlers: https://docs.perplexity.ai/guides/bots
  4. Google, Overview of Google crawlers and fetchers (Google-Extended): https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
  5. Google, AI features and your website: https://developers.google.com/search/docs/appearance/ai-features

Access is the first gate. Grade the rest free.

The full readiness grade checks robots.txt together with indexing, snippet eligibility, structured data and the other technical gates Google documents. No signup, about twenty seconds.