Free tool
AI robots.txt generator
Pick which AI bots may read your site. The default blocks model training but keeps ChatGPT search, Claude, Perplexity, Google and Bing, so you stay citable. Runs in your browser; nothing is sent.
Ticked bots get Disallow: /. Everything unticked falls under User-agent: *.
# AI bots you chose to block User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Amazonbot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: CCBot Disallow: / # Everyone else, including AI search crawlers User-agent: * Allow: /
Upload it to your site root as /robots.txt, then check that AI crawlers can read your page. Crawlable is step one; the free AI visibility check shows whether assistants actually recommend you.
How the rules work
- A crawler follows the most specific
User-agentgroup that names it, so each blocked bot gets its own group withDisallow: /. - Everyone else falls under
User-agent: *. Your private paths go there; blocked bots don’t need them repeated. - Google-Extended and Applebot-Extended are control tokens, not crawlers. Blocking them opts out of training without touching Google Search or Siri.
What each bot does, with links to the vendors’ docs: AI crawler user-agent list. Which OpenAI bot to allow: OAI-SearchBot vs GPTBot vs ChatGPT-User.
Questions
Will blocking GPTBot remove my site from ChatGPT?
Not from ChatGPT search, according to OpenAI. GPTBot collects training data; ChatGPT search uses OAI-SearchBot, which this generator leaves allowed unless you tick it.
What does the “Allow AI search, block training” preset block?
The training-only agents and tokens: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Amazonbot, Meta-ExternalAgent and CCBot. Search engines, AI search crawlers and user-triggered fetchers stay allowed.
Do all AI bots obey robots.txt?
Not all of them, and not always. OpenAI says robots.txt may not apply to ChatGPT-User visits that a user starts, and Meta says the same for Meta-ExternalFetcher. CDN bot protection is the only hard block.
Where do I put the file?
At the root of each host, for example https://yoursite.com/robots.txt. Subdomains need their own file. Most crawlers re-read it within a day.