citelaneBETA

Guide

AI crawlers and robots.txt: which bots to allow

Block AI training without disappearing from AI search. A plain-language table of the OpenAI, Anthropic, Perplexity and Google user agents, what each one does, and a robots.txt you can copy.

Updated October 2, 2026 · Citelane

Short answer

To stay citable while opting out of training, allow the search and user agents (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot, Bingbot) and block only the training agents (GPTBot, ClaudeBot, Google-Extended) if you want to.

The user agents that matter

AI-related crawlers and what blocking them does
User agentCompanyUsed forIf you block it
OAI-SearchBotOpenAISurfacing sites in ChatGPT searchYour pages can’t appear as ChatGPT search results
ChatGPT-UserOpenAIVisiting a page when a user asks ChatGPT toOpenAI notes robots.txt may not apply to user-initiated visits
GPTBotOpenAICollecting data that may train foundation modelsOpt out of training; OpenAI says this is separate from search
Claude-SearchBotAnthropicImproving Claude’s search resultsLower visibility in Claude search answers
Claude-UserAnthropicFetching pages for a Claude user’s questionClaude can’t retrieve your page for users
ClaudeBotAnthropicCollecting data that may train modelsOpt out of future training
PerplexityBotPerplexityIndexing pages for Perplexity answersYour pages are less likely to be cited by Perplexity
Google-ExtendedGoogleGemini training and grounding controlsGoogle says it does not affect Google Search ranking or inclusion
Googlebot / BingbotGoogle / MicrosoftSearch indexing (also AI Overviews, AI Mode, Copilot)You drop out of search, and the AI features built on it

Names and behaviour change. Check the vendors’ own pages (linked under Sources) before relying on this table.

A robots.txt that blocks training but keeps AI search

Rules are grouped by user agent; a crawler follows the most specific group that names it, so give training bots their own group:

  • User-agent: GPTBot / Disallow: /
  • User-agent: ClaudeBot / Disallow: /
  • User-agent: Google-Extended / Disallow: /
  • User-agent: * / Allow: /
  • Sitemap: https://example.com/sitemap.xml

Leave out any group you don’t want to block. If you want the most AI visibility possible, the simplest file is an Allow: / for everyone plus your sitemap.

robots.txt isn’t the only gate

CDN bot protection, firewalls and “block AI bots” toggles can stop crawlers even when robots.txt allows them. If an assistant never cites a page you know is relevant, check your CDN’s bot settings and server logs for the user agents above.

Can AI crawlers read your page?

Check robots.txt rules for GPTBot, OAI-SearchBot, PerplexityBot and ClaudeBot, plus noindex, canonical and sitemap. Free, no model calls.

Check AI crawler access

Questions

Does blocking GPTBot remove me from ChatGPT?

Not from ChatGPT search, according to OpenAI: search uses OAI-SearchBot, which you control separately. Blocking GPTBot only signals that your content shouldn’t be used for training.

How fast do robots.txt changes take effect?

OpenAI says about 24 hours for ChatGPT search. Other crawlers re-read robots.txt on their own schedules.

Should I block Google-Extended?

It’s a training and grounding control for Gemini; Google says it doesn’t affect Search ranking. Whether to block it is a content-licensing decision, not an SEO one.

Sources

Related