Your website might be telling AI crawlers they're welcome, while your server quietly turns them away.

A robots.txt file can allow crawlers, but that doesn't guarantee they can reach your content. Firewalls, bot protection and server-level rules can still get in the way, and none of them show up when you read robots.txt.

In this post I look at why this happens, how to find these hidden blocks on your own site, and why checking robots.txt alone isn't enough for AI visibility. The numbers come from the 1.5k+ business websites I have audited.

Two different doors

Every crawler introduces itself by name. OpenAI's training crawler calls itself GPTBot, Anthropic's calls itself ClaudeBot, Google's calls itself Googlebot. What happens next depends on two separate things.

robots.txt is a sign on the door. It is a plain text file at yourdomain.com/robots.txt that says which crawlers are welcome and where. Well-behaved crawlers read it and do what it says. It is a request, not a lock.

Your server is the lock. When a crawler asks for a page, your server, your hosting firewall, your security plugin or your CDN decides whether to send it. If any of them refuses, the crawler gets an error instead of your page, whatever robots.txt says.

So a site can have a perfectly friendly robots.txt and still be closed. The owner reads the sign, sees “everyone welcome”, and never learns that the lock says otherwise.

What 1.5k+ audits show

For every site I audit, my tool reads robots.txt and then requests a page as each crawler, by its full published name, after a normal browser request has confirmed the site is up.

Take GPTBot. Across the 1.5k+ sites, 52 (3.5%) told it no in robots.txt. Far more, 164 (10.9%), refused it at the server. And 134 of those served the very same request the moment it said “Googlebot” instead.

Across every AI crawler I test, 265 sites turned at least one away by name at the server, against 59 that said no in robots.txt. For every site that blocks an AI crawler where its owner can see it, about 4.5 block one where they can't.

Share of 1.5k+ audited sites refusing each AI crawler, in robots.txt and by name at the server
CrawlerRefused in robots.txtRefused by name at the server
Collect training data
GPTBot3.5%8.9%
ClaudeBot3.5%7.8%
CCBot3.5%8.8%
Build AI search results
OAI-SearchBot0.1%3.1%
Claude-SearchBot0.1%2.8%
PerplexityBot2.3%3.1%
Open a page when someone asks
ChatGPT-User0.3%2.9%
Claude-User2.3%2.0%
Perplexity-User0.1%2.1%

The server column counts only sites that served the same request when it said Googlebot, so it is a floor, not a ceiling. The full crawler table is on the research page.

A name block or an address check?

There is one honest complication. Some firewalls don't trust names at all. They check the address a request comes from against the address list the crawler's owner publishes, and refuse anything that only claims to be GPTBot. A firewall like that would refuse my test, then let the real GPTBot in.

That is why the Googlebot comparison matters. Google checks its own crawler by address too, so a firewall that verifies addresses refuses a fake Googlebot as well. If a server refuses “GPTBot” but serves the identical request labelled “Googlebot”, it isn't checking addresses. It is reacting to the name alone, and the real GPTBot carries that same name, so it gets refused too.

It is strong evidence rather than proof for any single site, which is why the test further down is worth running on your own.

Why it happens without anyone deciding it

Some of these blocks are deliberate. Many arrive with something else, and nobody notices:

  • CDN and firewall settings. Cloudflare has settings that block AI crawlers, and its defaults for new domains have changed more than once since 2025 (more on that here).
  • Security plugins. Some ship with a list of “bad bots” to refuse, and AI crawlers can be on it.
  • Rules copied into server files. A blocklist pasted into .htaccess or an nginx config years ago keeps working long after everyone has forgotten it.
  • Challenge pages. Bot protection that asks every visitor to prove it is a browser, usually by running JavaScript. Crawlers don't do that, so they never get past it.

None of these touch robots.txt, which is exactly why reading robots.txt tells you nothing about them.

Two examples from this week

Both tested by hand on 8 October 2026. I've left the sites unnamed because the pattern is the point.

The site that blocks by name

Its robots.txt allows every crawler. Its server answers “403 Forbidden” to GPTBot, ClaudeBot and PerplexityBot, and to the Ahrefs and Semrush crawlers. The same page, requested as Googlebot, OAI-SearchBot or ChatGPT-User, comes back normally. So Perplexity's own search crawler cannot read this site, and nothing in robots.txt would ever tell the owner.

The site behind a checkpoint

Every request I sent without a real browser got a security checkpoint page with an HTTP 429 status, even a request for robots.txt itself. Its host says this mode lets verified crawlers such as Googlebot through. Anything it can't verify gets the checkpoint, including every test I sent, whatever name it used, and an AI tool's page fetcher.

What each block costs you

Not every block is a mistake. Some owners refuse training crawlers on purpose, and that is a fair choice. What matters is knowing which kind you are refusing:

  • Training crawlers (GPTBot, ClaudeBot, CCBot) collect material that future models learn from. Blocking them doesn't stop AI search tools from showing you today.
  • Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) are how AI search tools find your pages in the first place. Block them and those tools can't use your site in their answers.
  • On-request fetchers (ChatGPT-User, Claude-User, Perplexity-User) open a page when someone asks about it. Block them and the AI can't read your site even while a customer is asking about you.

In the table above, each crawler in the second and third groups is refused by name on 2.0% to 3.1% of sites. That sounds small until it is your site. If you do want to block training, do it in robots.txt, where you can see it and change it (which ones to block, and which not to).

Check your own site in five minutes

You don't need special tools. Run these three commands, or send them to whoever looks after your website. Each one asks for your homepage and prints only the status code.

# 1. As a normal browser (the control)
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/129.0 Safari/537.36" https://yourdomain.com/

# 2. The same request, saying it is Googlebot
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://yourdomain.com/

# 3. The same request, saying it is GPTBot
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.1; +https://openai.com/gptbot)" https://yourdomain.com/

On Windows, type curl.exe instead of curl: in Windows PowerShell, plain curl runs a different command.

How to read the three numbers:

  • 200, 200, 200: nothing at your server is refusing GPTBot. Repeat the third command with the other crawler names you care about.
  • 200, 200, then 403 (or 401, 429, 503): your server is blocking GPTBot by name, so the real one is blocked too. Find the rule and decide whether you meant it.
  • 200, then a refusal for both Googlebot and GPTBot: your firewall is probably checking addresses, and the real crawlers may well get in. Your firewall's bot settings or logs will confirm it.
  • A refusal on the first one: the test itself is being blocked, perhaps by location. Try again from another network before reading anything into the rest.

If something is refusing a crawler you want, look in this order: your CDN's bot or AI crawler settings, your security plugin's firewall rules, your host's firewall, then your server configuration files. And read robots.txt last, for completeness, not first. There is a shorter version of this check in the answers section.

The short version

  • robots.txt is a sign on the door. Your server is the lock.
  • Across 1.5k+ audited sites, 265 block at least one AI crawler by name at the server, against 59 that do it in robots.txt.
  • If your server refuses GPTBot but serves the same request as Googlebot, the real GPTBot is refused too.
  • Many of these blocks come with a setting, a plugin or an old rule rather than a decision.
  • Three commands will tell you which kind of site yours is.

If you'd rather have it checked for you, Ru Visibility tests every major search and AI crawler at robots.txt and at the server as part of every audit, and tells you in plain English what is being refused and why.