Reference · checked 2026-10-10
AI crawler user agents list
20 robots.txt tokens from OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Amazon, Meta, DuckDuckGo, Mistral and Common Crawl, grouped by what blocking them actually does. Each row links to the vendor’s own documentation.
Keep search engines and AI search crawlers allowed if you want to be recommended. Training bots (GPTBot, ClaudeBot, Google-Extended and others) can be blocked without leaving ChatGPT, Claude or Google search. Generate that robots.txt or check your current one.
AI search index (5)
These build the indexes AI assistants search. Block them and you stop being a source in those answers.
| robots.txt token | Company | Used for | If you block it |
|---|---|---|---|
OAI-SearchBot | OpenAI | Surfacing websites in ChatGPT search results | Your pages can’t appear in ChatGPT search answers |
Claude-SearchBot | Anthropic | Improving the quality of Claude’s search results | Less visibility in Claude’s search answers |
PerplexityBot | Perplexity | Indexing pages that Perplexity shows and links in answers | Your pages are less likely to be cited by Perplexity |
Amzn-SearchBot | Amazon | Search experiences such as Alexa; Amazon says not used for training | Less visibility in Alexa answers |
DuckAssistBot | DuckDuckGo | Sources for DuckDuckGo’s AI-assisted answers; not used for training | You stop being a source for DuckAssist answers (takes about 72 hours) |
User-triggered fetch (5)
These open a page because a person asked an assistant to. Some vendors say robots.txt may not apply to them.
| robots.txt token | Company | Used for | If you block it |
|---|---|---|---|
ChatGPT-User | OpenAI | Visiting a page when a ChatGPT user or Custom GPT asks | OpenAI says robots.txt may not apply to these user-initiated visits |
Claude-User | Anthropic | Fetching a page when a Claude user asks a question | Claude can’t open your pages for its users |
Perplexity-User | Perplexity | Visiting a page to answer a Perplexity user’s question | Perplexity can’t open your pages for its users |
Meta-ExternalFetcher | Meta | Fetching a link at a Meta AI user’s request | Meta says it may bypass robots.txt for user-requested fetches |
MistralAI-User | Mistral | Visiting a page when a user asks Mistral’s assistant | Mistral’s assistant can’t open your pages for its users |
Model training (7)
These collect content for model training. Blocking them is a licensing choice; vendors say it doesn’t remove you from their search.
| robots.txt token | Company | Used for | If you block it |
|---|---|---|---|
GPTBot | OpenAI | Collecting content that may train OpenAI models | Opts out of training; OpenAI says it doesn’t affect ChatGPT search |
ClaudeBot | Anthropic | Collecting content that may train Anthropic models | Future content is excluded from Anthropic training |
Google-Extendedtoken only, no own user agent | Control token for Gemini training and grounding (crawling uses normal Google agents) | Google says it doesn’t affect Search ranking or inclusion | |
Applebot-Extendedtoken only, no own user agent | Apple | Control token for training Apple’s generative models | Opts out of Apple model training; Siri and Spotlight still work |
Amazonbot | Amazon | Improving Amazon products; may train Amazon AI models | Opts out of Amazon’s general crawl and training |
Meta-ExternalAgent | Meta | Crawling for training Meta AI models and indexing | Opts out of Meta’s AI crawl |
CCBot | Common Crawl | Open web archive that many AI training datasets are built from | Keeps future pages out of Common Crawl |
Search engine (3)
Classic search crawlers. AI Overviews, AI Mode, Copilot and Siri answers are built on their indexes, so keep them allowed.
| robots.txt token | Company | Used for | If you block it |
|---|---|---|---|
Googlebot | Google Search, also the source for AI Overviews and AI Mode | You drop out of Google Search and its AI features | |
Bingbot | Microsoft | Bing Search, also behind Copilot answers | You drop out of Bing and Copilot |
Applebot | Apple | Siri, Spotlight and Safari suggestions | You drop out of Siri and Spotlight results |
Once crawlers can read you, the question is whether ChatGPT, Perplexity and Gemini actually name you for your buyers’ questions. Check one page for free, about a minute, no sign-in.
Run the free AI visibility checkQuestions
Which AI crawlers should I allow?
To stay citable, allow the AI search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amzn-SearchBot, DuckAssistBot), the user-triggered fetchers and the search engines. Blocking training bots is optional and doesn’t remove you from AI search, according to the vendors.
Is Google-Extended a crawler?
No. Google says Google-Extended has no user agent of its own; Google crawls with its normal agents and reads the Google-Extended token in robots.txt as a control for Gemini training and grounding.
How do I check whether my site blocks one of these?
Paste a page into the free AI crawler checker. It reads your robots.txt and reports each bot, with training bots listed separately.
Vendors rename and add agents often. This list was checked against each linked page on 2026-10-10; trust the vendor page if they disagree.