Only the retrieval tokens (OAI-SearchBot, PerplexityBot, Claude-SearchBot and Claude-User) decide whether an AI answer can read your pages. GPTBot, ClaudeBot and Google-Extended govern training and grounding policy; blocking them does not, by itself, remove you from ChatGPT search, Claude's answers or Google's AI features.
The seven tokens, by role
| Token | Vendor | Role | What blocking it does, per the vendor | ClarAI's audit treats a block as |
|---|---|---|---|---|
| OAI-SearchBot | OpenAI | Retrieval (ChatGPT search) | Sites that disallow it will not appear in ChatGPT search answers. | A fail |
| PerplexityBot | Perplexity | Retrieval (Perplexity answers) | Removes the site from Perplexity results; the bot respects robots.txt. | A fail |
| Claude-SearchBot | Anthropic | Retrieval (Claude search) | May reduce the site's visibility in Claude's search responses. | A fail |
| Claude-User | Anthropic | Retrieval, user-directed fetch | May reduce the site's visibility for user-directed web search. | A fail |
| GPTBot | OpenAI | Training | Keeps content out of OpenAI model training; independent of OAI-SearchBot. | A policy note, unscored |
| ClaudeBot | Anthropic | Training | Keeps content out of Anthropic training; does not affect Claude-SearchBot. | A policy note, unscored |
| Google-Extended | Training and grounding policy | Stops use for Gemini training and for grounding in Gemini Apps and Vertex AI; does not affect Google Search inclusion, AI Overviews or AI Mode. | A policy note, unscored |
Retrieval or search agents: the four that decide answers
- OAI-SearchBot (OpenAI). The search crawler behind ChatGPT search. OpenAI's crawler page says that sites which disallow it will not appear in ChatGPT search answers, and that its setting is independent of GPTBot. This, not GPTBot, is the OpenAI token that decides ChatGPT search visibility.
- PerplexityBot (Perplexity). Surfaces sites in Perplexity results and respects robots.txt. Perplexity states it is not used for training. Blocking it removes the site from Perplexity answers.
- Claude-SearchBot (Anthropic). Anthropic's search agent, which it says improves search result quality; blocking it may reduce the site's visibility in Claude's search responses. This, not ClaudeBot, is the Anthropic token that decides answer visibility.
- Claude-User (Anthropic). Fetches a page because a Claude user asked for it. Anthropic lists it as a token you can block, and says blocking it may reduce visibility for user-directed web search, which is why it sits in this group rather than with the fetchers below.
If you want to appear in an engine's answers, these four are the ones that must be allowed. A site that blocks all crawlers with a wildcard rule and then adds a GPTBot allowance has not opened the door it thinks it has.
Training and policy tokens: the three that do not
- GPTBot (OpenAI). The training crawler. OpenAI publishes its IP ranges and treats its robots setting as independent of OAI-SearchBot and ChatGPT-User. Blocking it keeps your content out of model training and does not by itself remove you from ChatGPT search answers.
- ClaudeBot (Anthropic). Collects web content that could contribute to training, in Anthropic's words. Blocking it does not affect Claude-SearchBot or Claude-User.
- Google-Extended (Google). Not a crawler but a policy token read by Googlebot. It controls whether crawled content may be used to train Gemini models and for grounding in Gemini Apps and in Grounding with Google Search on Vertex AI. Google states it does not impact a site's inclusion in Google Search and is not a ranking signal (common crawlers page, last updated 14 July 2026).
The user-triggered fetchers robots.txt may not govern
Three tokens exist that a robots.txt policy does not reliably reach, because the fetch happens on a person's request rather than on the vendor's schedule.
- ChatGPT-User (OpenAI). Fetches on behalf of a ChatGPT user's action. OpenAI's page says "robots.txt rules may not apply" to it, and publishes an IP file for it.
- Perplexity-User (Perplexity). Perplexity says it "generally ignores robots.txt rules" because a user requested the fetch, and that it is not used for training.
- Google-Agent and Google's other user-triggered fetchers. Google's fetchers page (last updated 19 August 2026) adds a Google-Agent token for agents hosted on Google infrastructure and states that these fetchers generally ignore robots.txt rules.
What that means in practice: robots.txt is a policy for crawlers, not a wall for agents. If a user pastes your URL into an assistant, the fetch happens. Plan for that page to be read.
What Googlebot and Google-Extended actually control
Google's AI Overviews and AI Mode retrieve from the core Search index and crawl with standard Googlebot; Google's "AI features and your website" page says there are no additional requirements to appear in either. Google-Extended is a separate token that governs Gemini training and grounding in Gemini Apps and Vertex AI, and Google states it does not affect inclusion in Search. So: block Googlebot and you leave Search, AI Overviews and AI Mode; block Google-Extended and you leave none of those, but you do affect the Gemini app.
The off switch for AI Overviews and AI Mode is not in robots.txt at all. It is the Search Console "Search generative AI" control (Settings, Include or Exclude), rolled out to every property worldwide on 31 August 2026. Exclude removes a site's content from AI Overviews, AI Mode and generative AI features in Discover, does not affect training, is not a ranking signal, and leaves no public trace (Search Console Help, answer 16908024). The longer version is in Google-Extended does not do what you think it does.
A robots.txt grouped by role
A starting point, not a recommendation to allow or block anything: the grouping is the point. Put the retrieval agents together so nobody blocks one by accident, keep the training tokens as a separate, deliberate choice, and note that Googlebot carries Google's AI features.
# Retrieval and search agents: these decide whether an AI answer can read your pages. # Block one of these and you leave that engine's answers. User-agent: OAI-SearchBot User-agent: PerplexityBot User-agent: Claude-SearchBot User-agent: Claude-User Allow: / # Training and policy tokens: a policy choice, not an answer-visibility switch. # Swap Allow for Disallow here to stay out of training; the answers above are unaffected. User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended Allow: / # Googlebot crawls for AI Overviews and AI Mode as well as classic Search. User-agent: Googlebot Allow: / # Not listed on purpose: ChatGPT-User, Perplexity-User and Google-Agent are # user-triggered fetchers; their vendors say robots.txt rules may not apply to them.
Two cautions. A blanket User-agent: * block placed above these groups still applies to every token you did not name, so check what your wildcard rules say. And a robots.txt change does nothing about the two controls that live elsewhere: Google's Search generative AI control in Search Console, and the snippet controls (nosnippet, data-nosnippet, max-snippet, noindex) that remove a page from AI Overviews and AI Mode even when Googlebot is allowed. Those are covered in snippet controls are your AI visibility off switch.
Beyond the seven
- OAI-AdsBot (OpenAI). Listed on OpenAI's page for fetching advertising landing pages. Not in the seven because it has no bearing on answers.
- Applebot-Extended (Apple). Apple's page (dated 4 September 2026) says it "does not crawl webpages" and governs whether Applebot's crawled data may be used to train foundation models and to generate contextual answers in Siri and Search. So it is not a pure training opt-out either: disallowing it can remove a site from Siri and Apple Search answers.
- Cloudflare's classes. Cloudflare sorts AI bots into Search, Agent and Training and, from 15 September 2026, blocks Training and Agent bots by default on pages that display ads for all new domains onboarding to Cloudflare, while Search stays allowed (Cloudflare blog, 1 July 2026). If your site sits behind Cloudflare, check that setting before you check robots.txt: it can block a retrieval agent your robots.txt allows.
The honest line on JavaScript
No vendor page we have read states whether its crawler executes JavaScript. Google documents rendering for Googlebot; OpenAI, Anthropic and Perplexity say nothing either way. Treat client-rendered content as at risk: put the text an engine should quote, and the JSON-LD that describes the page, in the raw HTML. ClarAI's site audit flags a page whose raw HTML is an empty shell and a page whose structured data appears only after scripts run, and the readiness checks are described in is your site agent-ready?.
How to check your own robots.txt
The free grader fetches your robots.txt and scores the retrieval group only; the training tokens are listed as an unscored policy row. It measures readiness, not visibility: a site that passes can still go unmentioned. Inside ClarAI, the Tech SEO audit runs the same check on every project and checks Googlebot separately because it gates Google's AI features; both were re-based on the vendors' current pages in September 2026 (see how ClarAI works). The background on the crawlers themselves is in GPTBot, ClaudeBot and PerplexityBot: what each one does, and the three definitions these tokens serve are AI visibility, AEO and GEO.
Questions people search
Does blocking GPTBot remove my site from ChatGPT?
No. OpenAI's crawler page says each token's setting is independent, and that OAI-SearchBot is the one whose disallow removes a site from ChatGPT search answers. Blocking GPTBot keeps your content out of OpenAI model training and nothing else.
Do I need to allow ClaudeBot to appear in Claude's answers?
No. Anthropic's help article describes ClaudeBot as the crawler that collects content that could contribute to training, and Claude-SearchBot and Claude-User as the agents whose blocking may reduce visibility in Claude's search responses. Blocking ClaudeBot does not affect Claude-SearchBot.
Does Google-Extended affect AI Overviews or AI Mode?
No. Those features retrieve from the Search index with standard Googlebot. Google-Extended governs Gemini model training and grounding in Gemini Apps and Vertex AI, and Google states it does not impact inclusion in Google Search. The off switch for AI Overviews and AI Mode is the Search Console Search generative AI control, not robots.txt.
Do AI crawlers execute JavaScript?
No vendor page we have read says. Google documents rendering for Googlebot; OpenAI, Anthropic and Perplexity publish nothing on it either way. Assume they do not, and keep the text an engine should quote, and your structured data, in the raw HTML.
Can robots.txt stop an assistant fetching a page a user asked for?
Not reliably. OpenAI says robots.txt rules may not apply to ChatGPT-User, Perplexity says Perplexity-User generally ignores robots.txt rules, and Google says its user-triggered fetchers, including Google-Agent, generally ignore them. robots.txt is a policy for crawlers, not a wall for agents acting on a person's request.
Check which tokens your robots.txt actually blocks
The free grader reads your robots.txt and scores the retrieval group only; a block on a training token is reported as a policy note, never as a failure. No signup, about twenty seconds.
Sources
- Google Search Central, Google's common crawlers, incl. Googlebot and Google-Extended (last updated 14 July 2026)developers.google.com/search/docs/crawling-indexing/google-common-crawlers
- Google Search Central, Google's user-triggered fetchers, incl. Google-Agent (last updated 19 August 2026)developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers
- Google Search Central, AI features and your website (last updated 10 December 2025)developers.google.com/search/docs/appearance/ai-features
- Google Search Console Help, Search generative AI control (rolled out worldwide 31 August 2026)support.google.com/webmasters/answer/16908024
- OpenAI, Overview of OpenAI crawlers (OAI-SearchBot, GPTBot, ChatGPT-User, OAI-AdsBot)developers.openai.com/api/docs/bots
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler? (ClaudeBot, Claude-User, Claude-SearchBot)support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Perplexity, Perplexity crawlers (PerplexityBot, Perplexity-User)docs.perplexity.ai/guides/bots
- Cloudflare, Content Independence Day: no AI crawl without compensation, 1 July 2026 (Search, Agent and Training classes; defaults from 15 September 2026)blog.cloudflare.com/content-independence-day-ai-options/
- Apple, About Applebot (page dated 4 September 2026; Applebot-Extended scope)support.apple.com/en-us/119829