How to block AI training without disappearing from AI answers

Every major AI company runs one crawler for training and others for search and answers. Block the nine training-only crawlers, leave the rest open, and you are out of the training sets but still in the answers. Here is the list and the robots.txt.

Two different things, two different crawlers

“Should I block AI?” bundles two decisions that have nothing to do with each other. One is training: whether future models get to learn from your pages. The other is answers: whether an assistant can find your site when someone asks, fetch it, and cite it.

The AI companies split these jobs across separate crawlers, each with its own robots.txt rule. OpenAI has GPTBot for training, OAI-SearchBot for ChatGPT search and ChatGPT-User for fetching a page when a user asks about it. Anthropic has ClaudeBot, Claude-SearchBot and Claude-User. Apple has Applebot for Siri and Spotlight and Applebot-Extended as the training opt-out. So you can say no to training and yes to answers, and the vendors say so themselves. OpenAI's crawler documentation puts it plainly: each setting is independent of the others, so a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot.

Most sites that block AI “a bit” get this wrong in one of two ways. They block GPTBot and think they have opted out of AI, while OAI-SearchBot and ChatGPT-User are still welcome. Or they flip a CDN's “block AI bots” switch and vanish from the answers along with the training sets. Here is how to do it on purpose.

The nine names that only concern training

These are the robots.txt names whose makers document them as being about training data and nothing else. Blocking them does not remove you from any answer being given today.

  • GPTBot (OpenAI): training data for GPT models.
  • ClaudeBot (Anthropic): training data for Claude models.
  • Applebot-Extended (Apple): not a crawler but a token that decides whether pages Applebot already fetched may train Apple's foundation models. Plain Applebot keeps you in Siri and Spotlight.
  • CCBot (Common Crawl): the open crawl that many training sets are built from.
  • Bytespider (ByteDance): training crawler.
  • FacebookBot (Meta): a training crawler, which Meta has described as feeding its speech-recognition models.
  • cohere-training-data-crawler (Cohere): training crawler.
  • AI2Bot and AI2Bot-Dolma (Allen AI): the Dolma training dataset.

This is the same list our readiness check marks as Training only, taken from each vendor's own documentation. It changes a few times a year, so the robots.txt generator keeps the current one.

The crawlers to leave open

Everything that fetches pages for search results or for a live answer. The pattern in the names helps: Search, User and fetcher are the answer side.

  • OAI-SearchBot and ChatGPT-User: ChatGPT search and ChatGPT browsing.
  • Claude-SearchBot and Claude-User: Claude's search index and browsing.
  • PerplexityBot and Perplexity-User: Perplexity, which cites its sources in every answer.
  • Googlebot: Google Search, and with it AI Overviews and AI Mode, which use the ordinary search index.
  • meta-externalfetcher: Meta AI fetching a page a user asked about.
  • The smaller answer engines: DuckAssistBot, Kagibot, YouBot, PhindBot, Amazonbot, MistralAI-User and others.

One footnote on the fetchers. ChatGPT-User and meta-externalfetcher act for a person who asked about a specific page, and OpenAI and Meta both say robots.txt may not apply to them. Allowing them costs nothing; blocking them may not do anything.

Two that do not split cleanly

Google-Extended is the token people reach for to keep Gemini out of their content, and it does that. It is not a crawler: it decides what Google may do with pages Googlebot already fetched. But Google documents that it also controls grounding in the Gemini apps and on Vertex AI: whether Gemini can pull your page in while answering. Block it and you are out of Gemini's training and out of its grounding at the same time. It has no effect on Google Search, AI Overviews or AI Mode, which follow Googlebot and the usual snippet controls.

meta-externalagent is the same shape at Meta: it gathers training data and it indexes content for Meta's own products. One token, both jobs.

There is no clean answer for these two. If keeping your pages out of training matters more than being cited by Gemini or Meta AI, block them. If being cited matters more, leave them open. Just decide it knowingly. Our check counts both as answer crawlers for exactly this reason: blocking them can cost you answers.

The robots.txt

A crawler follows the most specific group that names it and ignores the rest, so the file is short: one group per training crawler, and a catch-all that allows everyone else.

# Training-only crawlers: out.
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: cohere-training-data-crawler
Disallow: /

User-agent: AI2Bot
Disallow: /

User-agent: AI2Bot-Dolma
Disallow: /

# Everyone else, including every search and fetch crawler: in.
User-agent: *
Allow: /
Content-Signal: ai-train=no, search=yes, ai-input=yes

Sitemap: https://yourdomain.com/sitemap.xml

The Content-Signal line is the newer, one-line way to say the same thing: no training, yes to search, yes to being read as input for an answer. It is a stated preference that crawlers honouring the convention read, not an enforcement mechanism, and it belongs inside a User-agent group. Keep it; the explicit groups above it do the work with crawlers that have not adopted it.

If your robots.txt already has rules for other bots, keep them and add these groups. The generator's Recommended balance preset builds this file and merges it into what you have, replacing only the AI-crawler groups so the file never ends up with two conflicting rules for the same bot.

What it costs you, honestly

Nothing today. Every assistant that searches the web or fetches pages on request keeps working, and a model that has already read your site does not unlearn it.

What changes is the future. The next generation of models will learn about your business only from what other sites say about it, not from your own pages. For a company that wants to be the definitive source on itself, that is a real cost, and it is why plenty of businesses choose to allow training on purpose. For a publisher whose content is the product, it is the point. Neither choice is wrong; the mistake is making it by accident.

One more limit: robots.txt is a request. The major vendors document that their training crawlers follow it. Scrapers that ignore it exist, and no robots.txt stops them.

Check that it worked

  1. Run a free readiness check on your site. Since 29 September the score counts only the crawlers that feed AI search and answers; the training-only crawlers are listed under Training only and cost no points. A site with the file above scores as fully readable.
  2. Paste the file into the robots.txt validator if you edited it by hand. A stray line can quietly turn an Allow into nothing.
  3. Check your CDN. A “block AI bots” switch usually blocks the search and fetch crawlers along with the training ones, and it does so at the edge, where robots.txt never sees it. If you use one, look at exactly which bots it covers.

Blocking training is one decision inside generative engine optimization, and it is the only one that does not need to cost you visibility. Make it deliberately, then get back to the part that does move answers: being the clearest source on what you do.

Try Crawl Readiness free

Check whether ChatGPT, Claude, Perplexity, and 50+ other AI crawlers can access your website — no signup for the basics.

Run a free readiness check →