Short answer

Look in two places. First your robots.txt file, at yourdomain.com/robots.txt, for rules that name AI crawlers. Then your security or CDN settings, because that is where most of the blocks we find come from, usually without the owner knowing.

1. Read your robots.txt

Open yourdomain.com/robots.txt in a browser. It is a plain text file of rules, each starting with User-agent: followed by a crawler's name. A block looks like this:

User-agent: PerplexityBot
Disallow: /

Disallow: / means “stay out of the whole site”. Also check the rules under User-agent: *, which apply to every crawler without its own section. These are the names to look for:

AI crawlers, who runs them, and what they are for
Name in robots.txtCompanyWhat it is forKind
OAI-SearchBotOpenAIFinds pages to show in ChatGPT searchAnswers
ChatGPT-UserOpenAIOpens a page when a ChatGPT user's request needs itAnswers
GPTBotOpenAICollects content for training OpenAI's modelsTraining
Claude-SearchBotAnthropicIndexes pages to improve Claude's search resultsAnswers
Claude-UserAnthropicFetches a page when a Claude user asks a questionAnswers
ClaudeBotAnthropicCollects content that may be used for trainingTraining
PerplexityBotPerplexityFinds and links pages in Perplexity search (not training)Answers
Perplexity-UserPerplexityVisits a page to answer a user's questionAnswers
Google-ExtendedGoogleA robots.txt setting for Gemini training; does not affect Google SearchTraining

Two caveats from the companies themselves. Perplexity says its user-triggered fetcher “generally ignores robots.txt rules”, and OpenAI says of ChatGPT-User that “robots.txt rules may not apply”, because a person asked for the page. So robots.txt mainly governs the crawlers that build each tool's index.

2. Check your security and CDN settings

This is where most blocks come from. Services like Cloudflare and many security plugins can refuse bots before robots.txt is ever read. Cloudflare, for instance, has a setting under Security Settings called Block AI bots, and a newer one for AI bot policies; here is what each does. If someone else manages your site, ask them directly which AI bots are allowed.

3. Test what a crawler gets

A rough test from your own computer: request your homepage while identifying as the crawler, and see whether you get the page or an error.

curl -I -A "OAI-SearchBot" https://yourdomain.com/

A 200 means the page was served; a 403 means it was refused. Treat it as a hint rather than proof: the real crawler comes from the company's own servers, and some security tools treat it differently from a test on your laptop, in either direction.

Found a block?

Decide crawler by crawler. Blocking the ones that only collect training data is a reasonable choice; blocking the ones that fetch pages for answers keeps you out of those answers. Which to allow, and why.

Sources