citelaneBETA

Free tool

AI robots.txt generator

Pick which AI bots may read your site. The default blocks model training but keeps ChatGPT search, Claude, Perplexity, Google and Bing, so you stay citable. Runs in your browser; nothing is sent.

Ticked bots get Disallow: /. Everything unticked falls under User-agent: *.

Your robots.txt
# AI bots you chose to block
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: CCBot
Disallow: /

# Everyone else, including AI search crawlers
User-agent: *
Allow: /

Upload it to your site root as /robots.txt, then check that AI crawlers can read your page. Crawlable is step one; the free AI visibility check shows whether assistants actually recommend you.

How the rules work

  • A crawler follows the most specific User-agent group that names it, so each blocked bot gets its own group with Disallow: /.
  • Everyone else falls under User-agent: *. Your private paths go there; blocked bots don’t need them repeated.
  • Google-Extended and Applebot-Extended are control tokens, not crawlers. Blocking them opts out of training without touching Google Search or Siri.

What each bot does, with links to the vendors’ docs: AI crawler user-agent list. Which OpenAI bot to allow: OAI-SearchBot vs GPTBot vs ChatGPT-User.

Questions

Will blocking GPTBot remove my site from ChatGPT?

Not from ChatGPT search, according to OpenAI. GPTBot collects training data; ChatGPT search uses OAI-SearchBot, which this generator leaves allowed unless you tick it.

What does the “Allow AI search, block training” preset block?

The training-only agents and tokens: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Amazonbot, Meta-ExternalAgent and CCBot. Search engines, AI search crawlers and user-triggered fetchers stay allowed.

Do all AI bots obey robots.txt?

Not all of them, and not always. OpenAI says robots.txt may not apply to ChatGPT-User visits that a user starts, and Meta says the same for Meta-ExternalFetcher. CDN bot protection is the only hard block.

Where do I put the file?

At the root of each host, for example https://yoursite.com/robots.txt. Subdomains need their own file. Most crawlers re-read it within a day.