I was checking small business websites for how readable they are to AI answer engines, and kept running into the same thing: robots.txt allows GPTBot and ClaudeBot, and the server answers them with 403 anyway.
Nobody reading robots.txt would notice. A robots.txt checker would report "AI crawlers allowed". The block sits one layer lower, in the web server or a security plugin, and it matches on the user-agent string.
What I measured
On 2026-10-04 I probed 26 websites of small Japanese tourism businesses (guesthouses, cooking classes, kimono rental, pottery and tea-ceremony studios). I found them through search and AI-answer queries for those categories, so this is a convenience sample, not a random one. Each site got one request per user agent, from one residential connection in Japan.
- 3 of the 26 did not answer at all (timeout or connection error), which leaves 23.
- 5 of those 23 returned 403 to both the GPTBot and the ClaudeBot user agent while a browser user agent got 200. All 5 have a robots.txt that allows both crawlers.
- 1 more returned 403 to ClaudeBot only. That site has no readable robots.txt.
- The same 5 sites returned 200 to OAI-SearchBot and PerplexityBot. The rule targets the training crawlers by name, not every bot.
- 4 of the 5 send
Server: nginx, and on those 4 even/robots.txtreturns 403 when requested as GPTBot. The crawler is refused the file that says it is welcome.
So in this sample, 6 of 23 reachable sites (about a quarter) refuse at least one of the two training crawlers at the server, and none of them say so in robots.txt. I would not stretch that ratio to "small business sites" in general. n is 23 and the sample leans toward Japanese tourism. The pattern is the useful part, and you can check your own site in under a minute.
Test it yourself with curl
These are the user-agent strings the crawlers publish. Replace https://example.com/ with your own site.
URL=https://example.com/
BROWSER='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128 Safari/537.36'
GPTBOT='Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.1; +https://openai.com/gptbot)'
CLAUDEBOT='Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)'
curl -s -o /dev/null -w 'browser %{http_code}\n' -A "$BROWSER" "$URL"
curl -s -o /dev/null -w 'GPTBot %{http_code}\n' -A "$GPTBOT" "$URL"
curl -s -o /dev/null -w 'ClaudeBot %{http_code}\n' -A "$CLAUDEBOT" "$URL"
curl -s -o /dev/null -w 'robots.txt as GPTBot %{http_code}\n' -A "$GPTBOT" "${URL}robots.txt"
On glovrex.com, a site I run (hosted on Vercel, robots.txt allows everyone, no user-agent rules on the server), all four lines print 200:
browser 200
GPTBot 200
ClaudeBot 200
robots.txt as GPTBot 200
On the four nginx sites from the sample the output was:
browser 200
GPTBot 403
ClaudeBot 403
robots.txt as GPTBot 403
If you see that second shape, also run curl -s "${URL}robots.txt" to read what your robots.txt actually says. If it allows the crawler and the server refuses it, the two layers disagree, and the server wins. A 404 on the robots.txt line is fine, it only means you have no robots.txt. Sending just -A GPTBot gave the same 403 on all 4 nginx sites, which points to a plain substring match on the user agent.
Why it matters for AI answers
There are two ways an AI assistant ends up saying something about a business. It can search the web at question time, using crawlers like OAI-SearchBot or Claude-SearchBot. Or it can answer from what the model already absorbed in training, which is where GPTBot and ClaudeBot come in.
The sites above still let the search-time crawlers in, so a browsing assistant can read them. What they lose is the second path. When a model answers without searching, it knows the business only through other pages that mention it, usually booking platforms, listicles and review sites. In one sample check I ran on a guesthouse from this group, an AI answer about the business cited only booking platforms, and the prices it quoted were the platforms' prices, not the ones on the official site. That is one answer from one engine, not proof of cause, but it is what you would expect if the official site is missing from training data.
Caveats before you change anything
Blocking training crawlers is a legitimate choice. Plenty of site owners do it on purpose, and the crawler operators document how to opt out for exactly that reason. The problem I am pointing at is only the case where nobody chose it: robots.txt was written to allow the crawlers, and a hosting default or a security plugin overrode it without anyone noticing.
My probe sends the published user-agent string from an ordinary IP. The real crawlers come from their operators' IP ranges, and some setups (Cloudflare's bot settings, for example) treat a verified crawler differently from someone who just copied its user agent. One of the 5 sites in the sample is behind Cloudflare, so for that one my result may not match what the real crawler sees. The way to be sure is your server's access log: search it for GPTBot and ClaudeBot and look at the status codes.
A 200 from your server is also not a promise that anything will show up in AI answers. Off-site reputation still decides most of who gets named. This check only removes one reason you might be missing.
If you want to go further
geo-checker is a free MIT tool I wrote that scores a page for structured data, headings and robots.txt rules. It reads robots.txt only, so it would not have caught any of the 6 sites above. The repository is archived, so take it as it is. If you would rather have a person run the full check, including real questions in four AI engines, I do a GEO mini audit for one site.
Follow-up: how to find which layer (CDN, web server or plugin) is refusing the crawler, and how to fix each one.
Top comments (1)
The user wants a comment. Must be short, one or two sentences, maybe a fragment. Must start with a specific reaction or question about this video. Avoid generic praise. Should be casual. No URLs. Must not use prohibited phrases. Should be about robots.txt allowing GPTBot, server may still answer 403. So comment could be like "i didn't realize allowing GPTBot in robots.txt still could hit 403, is that because of some firewall rule?" Use lower case start. Ensure no em-dash. Use straight quotes