Short answer
It depends which ones. Crawlers that only collect training data, such as GPTBot and ClaudeBot, can be blocked without affecting whether AI search tools can show you today. Crawlers that fetch pages for live answers, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, are different: block them and those tools can't use your site in their answers.Two kinds of AI crawler
Training crawlers collect text that may be used to build future models. OpenAI says disallowing GPTBot “indicates a site's content should not be used in training generative AI foundation models”, and Anthropic says restricting ClaudeBot signals that “the site's future materials should be excluded from our AI model training datasets”.
Answer crawlers fetch pages so a tool can answer people now. OpenAI: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links.” Anthropic says disabling Claude-SearchBot “may reduce your site's visibility and accuracy in user search results”. Perplexity describes PerplexityBot as “designed to surface and link websites in search results on Perplexity”, adding that it “is not used to crawl content for AI foundation models”.
Blocking training crawlers: a fair choice
Some businesses don't want their writing used to train AI, and that is a legitimate decision. It costs you nothing in today's answers, because the companies keep the settings separate; OpenAI says “each setting is independent of the others” ( OpenAI).
Blocking answer crawlers: usually a mistake for a business
If you want customers to find you through ChatGPT, Claude or Perplexity, these are the crawlers that make it possible. Shutting them out is like asking not to be listed in a directory your customers use.
A robots.txt that does both
This keeps your content out of OpenAI's and Anthropic's training and Google's Gemini training, while leaving the answer crawlers free:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /Google-Extended only affects Gemini; Google says it “does not impact a site's inclusion in Google Search” ( Google). And robots.txt only works if your firewall lets the crawler reach it in the first place, so check that too.