Guide
AI crawlers and robots.txt: which bots to allow
Block AI training without disappearing from AI search. A plain-language table of the OpenAI, Anthropic, Perplexity and Google user agents, what each one does, and a robots.txt you can copy.
Updated October 2, 2026 · Citelane
To stay citable while opting out of training, allow the search and user agents (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Googlebot, Bingbot) and block only the training agents (GPTBot, ClaudeBot, Google-Extended) if you want to.
The user agents that matter
| User agent | Company | Used for | If you block it |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfacing sites in ChatGPT search | Your pages can’t appear as ChatGPT search results |
| ChatGPT-User | OpenAI | Visiting a page when a user asks ChatGPT to | OpenAI notes robots.txt may not apply to user-initiated visits |
| GPTBot | OpenAI | Collecting data that may train foundation models | Opt out of training; OpenAI says this is separate from search |
| Claude-SearchBot | Anthropic | Improving Claude’s search results | Lower visibility in Claude search answers |
| Claude-User | Anthropic | Fetching pages for a Claude user’s question | Claude can’t retrieve your page for users |
| ClaudeBot | Anthropic | Collecting data that may train models | Opt out of future training |
| PerplexityBot | Perplexity | Indexing pages for Perplexity answers | Your pages are less likely to be cited by Perplexity |
| Google-Extended | Gemini training and grounding controls | Google says it does not affect Google Search ranking or inclusion | |
| Googlebot / Bingbot | Google / Microsoft | Search indexing (also AI Overviews, AI Mode, Copilot) | You drop out of search, and the AI features built on it |
Names and behaviour change. Check the vendors’ own pages (linked under Sources) before relying on this table.
A robots.txt that blocks training but keeps AI search
Rules are grouped by user agent; a crawler follows the most specific group that names it, so give training bots their own group:
User-agent: GPTBot/Disallow: /User-agent: ClaudeBot/Disallow: /User-agent: Google-Extended/Disallow: /User-agent: */Allow: /Sitemap: https://example.com/sitemap.xml
Leave out any group you don’t want to block. If you want the most AI visibility possible, the simplest file is an Allow: / for everyone plus your sitemap.
robots.txt isn’t the only gate
CDN bot protection, firewalls and “block AI bots” toggles can stop crawlers even when robots.txt allows them. If an assistant never cites a page you know is relevant, check your CDN’s bot settings and server logs for the user agents above.
Check robots.txt rules for GPTBot, OAI-SearchBot, PerplexityBot and ClaudeBot, plus noindex, canonical and sitemap. Free, no model calls.
Check AI crawler accessQuestions
Does blocking GPTBot remove me from ChatGPT?
Not from ChatGPT search, according to OpenAI: search uses OAI-SearchBot, which you control separately. Blocking GPTBot only signals that your content shouldn’t be used for training.
How fast do robots.txt changes take effect?
OpenAI says about 24 hours for ChatGPT search. Other crawlers re-read robots.txt on their own schedules.
Should I block Google-Extended?
It’s a training and grounding control for Gemini; Google says it doesn’t affect Search ranking. Whether to block it is a content-licensing decision, not an SEO one.