citelaneBETA

Reference · checked 2026-10-10

AI crawler user agents list

20 robots.txt tokens from OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple, Amazon, Meta, DuckDuckGo, Mistral and Common Crawl, grouped by what blocking them actually does. Each row links to the vendor’s own documentation.

In short

Keep search engines and AI search crawlers allowed if you want to be recommended. Training bots (GPTBot, ClaudeBot, Google-Extended and others) can be blocked without leaving ChatGPT, Claude or Google search. Generate that robots.txt or check your current one.

AI search index (5)

These build the indexes AI assistants search. Block them and you stop being a source in those answers.

robots.txt tokenCompanyUsed forIf you block it
OAI-SearchBotOpenAISurfacing websites in ChatGPT search resultsYour pages can’t appear in ChatGPT search answers
Claude-SearchBotAnthropicImproving the quality of Claude’s search resultsLess visibility in Claude’s search answers
PerplexityBotPerplexityIndexing pages that Perplexity shows and links in answersYour pages are less likely to be cited by Perplexity
Amzn-SearchBotAmazonSearch experiences such as Alexa; Amazon says not used for trainingLess visibility in Alexa answers
DuckAssistBotDuckDuckGoSources for DuckDuckGo’s AI-assisted answers; not used for trainingYou stop being a source for DuckAssist answers (takes about 72 hours)

User-triggered fetch (5)

These open a page because a person asked an assistant to. Some vendors say robots.txt may not apply to them.

robots.txt tokenCompanyUsed forIf you block it
ChatGPT-UserOpenAIVisiting a page when a ChatGPT user or Custom GPT asksOpenAI says robots.txt may not apply to these user-initiated visits
Claude-UserAnthropicFetching a page when a Claude user asks a questionClaude can’t open your pages for its users
Perplexity-UserPerplexityVisiting a page to answer a Perplexity user’s questionPerplexity can’t open your pages for its users
Meta-ExternalFetcherMetaFetching a link at a Meta AI user’s requestMeta says it may bypass robots.txt for user-requested fetches
MistralAI-UserMistralVisiting a page when a user asks Mistral’s assistantMistral’s assistant can’t open your pages for its users

Model training (7)

These collect content for model training. Blocking them is a licensing choice; vendors say it doesn’t remove you from their search.

robots.txt tokenCompanyUsed forIf you block it
GPTBotOpenAICollecting content that may train OpenAI modelsOpts out of training; OpenAI says it doesn’t affect ChatGPT search
ClaudeBotAnthropicCollecting content that may train Anthropic modelsFuture content is excluded from Anthropic training
Google-Extended
token only, no own user agent
GoogleControl token for Gemini training and grounding (crawling uses normal Google agents)Google says it doesn’t affect Search ranking or inclusion
Applebot-Extended
token only, no own user agent
AppleControl token for training Apple’s generative modelsOpts out of Apple model training; Siri and Spotlight still work
AmazonbotAmazonImproving Amazon products; may train Amazon AI modelsOpts out of Amazon’s general crawl and training
Meta-ExternalAgentMetaCrawling for training Meta AI models and indexingOpts out of Meta’s AI crawl
CCBotCommon CrawlOpen web archive that many AI training datasets are built fromKeeps future pages out of Common Crawl

Search engine (3)

Classic search crawlers. AI Overviews, AI Mode, Copilot and Siri answers are built on their indexes, so keep them allowed.

robots.txt tokenCompanyUsed forIf you block it
GooglebotGoogleGoogle Search, also the source for AI Overviews and AI ModeYou drop out of Google Search and its AI features
BingbotMicrosoftBing Search, also behind Copilot answersYou drop out of Bing and Copilot
ApplebotAppleSiri, Spotlight and Safari suggestionsYou drop out of Siri and Spotlight results
Allowed is not the same as recommended

Once crawlers can read you, the question is whether ChatGPT, Perplexity and Gemini actually name you for your buyers’ questions. Check one page for free, about a minute, no sign-in.

Run the free AI visibility check

Questions

Which AI crawlers should I allow?

To stay citable, allow the AI search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Amzn-SearchBot, DuckAssistBot), the user-triggered fetchers and the search engines. Blocking training bots is optional and doesn’t remove you from AI search, according to the vendors.

Is Google-Extended a crawler?

No. Google says Google-Extended has no user agent of its own; Google crawls with its normal agents and reads the Google-Extended token in robots.txt as a control for Gemini training and grounding.

How do I check whether my site blocks one of these?

Paste a page into the free AI crawler checker. It reads your robots.txt and reports each bot, with training bots listed separately.

Vendors rename and add agents often. This list was checked against each linked page on 2026-10-10; trust the vendor page if they disagree.