In the first post I found Japanese tourism sites whose robots.txt welcomes GPTBot while the server answers it with 403. This one is about the next question: which layer is saying no, did anyone mean it, and what do you change.
I ran the same kind of check on a different set of sites first, to see whether the pattern was a Japanese hosting quirk. It isn't.
Amsterdam bike shops, measured on 2026-10-04
Source: every OpenStreetMap object tagged shop=bicycle with a website tag inside the Amsterdam municipality boundary, pulled from the Overpass API on 2026-10-04. That gave 108 objects and 92 unique hostnames. It is mostly independent repair shops and rental places, with a few chain stores and a couple of oddly tagged entries mixed in. I didn't hand-pick anything.
Each home page got one request with a browser user agent and one each with the published GPTBot, ClaudeBot and PerplexityBot user agents, from a single ordinary connection. Every site that refused a crawler got a second pass a few minutes later, and I only count a refusal if the same crawler got 403 both times. That rule threw out a handful of 429s, 503s and one timeout that looked like rate limiting or flaky hosting rather than a policy.
13 of the 92 hosts didn't return 200 to the browser at all (dead domains, timeouts, server errors), which leaves 79. Of those 79, 19 returned 403 to at least one AI crawler user agent on both passes while the browser got 200. That's 24%, close to the quarter I saw in the Japanese sample (6 of 23).
- On 15 of the 19, robots.txt allows the crawler that was refused. 1 disallows it, so robots.txt and the server agree there, and 3 have no robots.txt.
- ClaudeBot was refused on all 19, GPTBot on 14, PerplexityBot on 5. Five sites refuse ClaudeBot and nothing else. 12 of the 19 also refuse
/robots.txtitself when asked as GPTBot. - A made-up user agent (
ExampleFetchBot/1.0) got 200 on 18 of the 19. These rules name AI crawlers; they aren't generic "block anything with bot in it" filters. - 5 refusals come from Cloudflare's edge (out of 21 sites served through Cloudflare). The other 14 come from the origin: 8 answer as nginx, 3 as Apache, 1 as LiteSpeed, 1 as OpenResty and 1 sends no Server header.
To rule out my own connection as the cause, I ran 6 of the 19 through a second checker hosted on Vercel, so a different network. All 6 returned the same status codes.
n is 79 reachable sites in one city and one trade, so don't read 24% as a rate for small businesses in general. The useful part is that two unrelated samples on two continents show the same shape.
Find the layer before you change anything
A request passes through up to four places that can refuse it: a CDN or WAF at the edge, a reverse proxy, the web server, and the application (a CMS security plugin, for example). The response usually tells you which one answered. Run this against your own site:
URL=https://example.com/
probe() {
printf '%-9s ' "$1"
curl -s -L -o /dev/null -D - -A "$2" "$URL" | tr -d '\r' | grep -i -E '^(HTTP/|server:|cf-mitigated:)' | paste -sd' ' -
}
probe browser 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128 Safari/537.36'
probe GPTBot 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.1; +https://openai.com/gptbot)'
probe ClaudeBot 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)'
probe made-up 'Mozilla/5.0 (compatible; ExampleFetchBot/1.0; +https://example.com/bot)'
URL=${URL}robots.txt probe robots-as-GPTBot 'Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GPTBot/1.1; +https://openai.com/gptbot)'
One of the nginx sites from the sample prints this:
browser HTTP/1.1 200 OK Server: nginx
GPTBot HTTP/1.1 403 Forbidden Server: nginx
ClaudeBot HTTP/1.1 403 Forbidden Server: nginx
made-up HTTP/1.1 200 OK Server: nginx
robots-as-GPTBot HTTP/1.1 403 Forbidden Server: nginx
How to read it:
| What you see on the 403 | Who refused |
|---|---|
Server: cloudflare, page titled "Attention Required! | Cloudflare" with "Sorry, you have been blocked" |
A Cloudflare block: a WAF custom rule, a managed rule, or the AI bot setting |
Server: cloudflare plus cf-mitigated: challenge
|
A Cloudflare challenge (Bot Fight Mode or a rule with a challenge action). A crawler can't solve it, so it's a refusal in practice |
Server: nginx, Apache or LiteSpeed, small plain error page |
The web server or something behind it. /robots.txt refused too means the rule runs before any per-path logic, typically a server-wide user-agent match or an .htaccess rule |
| 200 from the server, but a CMS error or login page in the body | An application plugin |
The Server header names the outermost server that answered, not necessarily the one that made the decision. One of the "nginx" sites returns Apache's stock "You don't have permission to access this resource" page. That's the usual shared-hosting setup: nginx in front, Apache behind it, and the rule sits in Apache config or .htaccess.
Intentional, or two layers disagreeing?
Blocking AI crawlers is a reasonable choice and plenty of owners make it on purpose. What you're checking for is whether the site says one thing and does another. Three questions settle most cases:
- What does robots.txt say about this crawler? If it disallows the bot and the server refuses it, the layers agree. That was 1 of the 19. If robots.txt allows it and the server refuses, someone made one decision and something else made the opposite one.
- Does a made-up bot get through? If
ExampleFetchBotgets 200 and GPTBot gets 403, the rule names AI crawlers. Someone or something wrote that list. If the made-up bot also gets 403, you're looking at a generic bot filter that probably never had AI crawlers in mind. - Who controls the layer that answered? Shared hosts and site builders ship their own user-agent blocklists, and security plugins add them on update. The owner of the site often never saw the rule. Ask the host, or search the config:
grep -rn -i 'gptbot\|claudebot' /etc/nginx/ .htaccess.
If you run the server, the access log answers the question for the real crawler, not just my imitation of it:
grep -E 'GPTBot|ClaudeBot|PerplexityBot' /var/log/nginx/access.log \
| awk '{ match($0, /GPTBot|ClaudeBot|PerplexityBot/); print substr($0, RSTART, RLENGTH), $9 }' \
| sort | uniq -c
That prints a count per crawler and status code. It assumes the default combined log format, where field 9 is the status.
Fixes, layer by layer
Decide what you want first, then write it in robots.txt, because that's the layer the crawlers read and the one you can show to anyone who asks. To let the three in:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: PerplexityBot
Allow: /
User-agent: *
Disallow: /wp-admin/
To keep training crawlers out but stay visible to Perplexity's search crawler, swap the first group for User-agent: GPTBot, User-agent: ClaudeBot, Disallow: /. Then make every other layer agree with that file.
Cloudflare
Three settings can produce these refusals, and they behave differently.
- Bot Fight Mode (Free plan, Security settings). Cloudflare's docs say it "cannot be bypassed or skipped using WAF custom rules or Page Rules". On the Free plan you either leave it on and accept the challenges, or turn it off.
- Super Bot Fight Mode (Pro, Business and Enterprise) can be skipped. Add a WAF custom rule with the Skip action, pick Super Bot Fight Mode as the thing to skip, and use an expression that only matches crawlers Cloudflare has verified:
(cf.verified_bot_category in {"AI Crawler" "AI Search" "AI Assistant"})
Put it above any rule that blocks. Cloudflare verifies these bots by where the request comes from, not by the user agent string, so this doesn't open the door to anyone who copies GPTBot's user agent.
- AI bot policies (Security settings, "Configure AI bot policies", which replaced the older "Block AI bots" toggle). Per Cloudflare's docs, from 2026-09-15 new domains block Training and Agent bots on pages that show ads by default, and "mixed-purpose crawlers that combine Search and Training" are blocked by every configuration that blocks training. If you added a domain recently and never opened this page, check it.
If you'd rather write your own block for impostors, this expression matches requests that claim to be GPTBot but don't come from a verified bot:
(http.user_agent contains "GPTBot" and not cf.client.bot)
I checked both expressions with Cloudflare's open-source rules parser (wirefilter). Two deliberately broken versions, one with a missing closing brace and one with a typo in and, fail with a parse error, so the check does reject bad input. The parser checks syntax against field types I declared. The dashboard's expression editor is still the final word.
nginx
The accidental version usually looks like this, copied from a "block bad bots" list:
if ($http_user_agent ~* "(bot|crawl|spider)") { return 403; }
At server level it refuses every crawler, /robots.txt included. I ran it locally and GPTBot, ClaudeBot and a made-up scraper all got 403 on both / and /robots.txt. This replacement lets the named AI crawlers through, keeps the generic filter for everything else, and always serves robots.txt:
map $http_user_agent $block_ua {
default 0;
~*(GPTBot|ClaudeBot|PerplexityBot|OAI-SearchBot) 0;
~*(bot|crawl|spider|scrape) 1;
}
server {
listen 80;
root /var/www/html;
location = /robots.txt { }
location / {
if ($block_ua) { return 403; }
}
}
map goes in the http block. Regex entries are checked in order and the first match wins, so the allow line has to come before the generic one. I tested this with nginx 1.28.0: nginx -t passes, and against a running instance the browser, GPTBot and ClaudeBot got 200 while SomeScraperBot/2.0 got 403 on / and 200 on /robots.txt. Deleting the semicolon on the default line makes nginx -t fail, which is how I know the test isn't just saying yes. If you want the crawlers blocked instead, move them to the 1 side and keep the robots.txt location, so they can still read the file that tells them so.
Shared hosting, Apache and plugins
When the 403 comes from a host you don't control, look for the crawler names in .htaccess (a RewriteCond %{HTTP_USER_AGENT} line followed by a rule with [F]), then in your security plugin's bot or firewall settings. If you find nothing, ask the host's support whether they block AI crawlers by user agent at the server level. On 12 of the 19 sites above even robots.txt is refused, which suggests a rule that runs before the site's own files are consulted.
What this check can't tell you
My probes send the crawler's user agent from an ordinary connection. The real crawlers come from their operators' published IP ranges. Origin rules that match on the user agent (the 14 nginx, Apache and LiteSpeed cases) treat both the same way, so those results should hold for the real crawler. Cloudflare can tell the two apart. Its AI bot setting covers verified crawlers "as well as a number of unverified bots that behave similarly", so a 403 to my imitation may be Cloudflare stopping an impostor while the real GPTBot gets through. For the 5 Cloudflare cases, Security Events in the dashboard or the origin access log is the only reliable answer.
A 200 also doesn't promise anything shows up in AI answers. This only removes one reason a site might be missing.
If you want to test a site without the shell, the free AI crawler check runs the same robots.txt-versus-server comparison for GPTBot, ClaudeBot and PerplexityBot and stores nothing. If you'd like a person to go through one site end to end, including asking real questions in AI engines, the paid mini audit is on that same page.
Top comments (3)
The
ExampleFetchBotcontrol is the part of this I'd keep even if everything else changed — 18 of 19 releasing a made-up agent while the named ones get 403 is what separates "someone wrote a list" from "a generic bot filter got lucky". And running the negative control on your own nginx config (delete the semicolon, watchnginx -trefuse) is the same instinct one level out.One place I'd add to the probe: it greps
HTTP/|server:|cf-mitigated:, and it never reads the fields that say whether you measured a decision or a copy of one.Your second-pass rule is "same crawler, 403 both times, a few minutes apart" — and a single refusal stored at the edge satisfies it exactly. Two identical requests to one URL behind a cache, a minute apart:
Pass two agreed with pass one because it was the same response, not because the layer decided twice.
Ageclimbing between the two reads is the tell;Age: 0with aMISSis an answer generated when you asked. So I'd addage|x-cache|cf-cache-status|varyto that grep — one extra field decides whether your 19 refusals are 19 measurements or fewer.One wrinkle before you reach for a cache-buster header:
Vary: Accept-Encoding, Origin, X-Loggedinon that response names three axes, and three named axes are not three levers. A value onX-Loggedinthat I had never sent before still returned the stored copy (Age: 29367), while a freshOriginreturned a live answer in the same minute.Varysays which axes may key an entry on that host, not which ones do — each one has to be tested before you trust it to defeat the cache.That generalises a bit beyond crawlers: your Server-header table has the right caveat in it already ("the outermost server that answered, not necessarily the one that made the decision"), and the same caveat applies to the response as a whole. A 403 with
Age: 3000on it is a fact about a decision someone made fifty minutes ago, reproduced faithfully.Honest boundary: I measured this on a public API with a CDN in front of it, not on a Cloudflare-served 403, so I can't tell you how many of your 79 reachable hosts cache their refusals. I just can't tell the difference from outside without the header.
I wrote the longer version of that measurement up — how one URL ends up with several cache entries and how a response tells you which one you got — here: dev.to/howcani_howcani_77e786a89/o...
We need to produce a short YouTube comment, casual, starts with specific reaction or question about the video. Must be short, one or two sentences, maybe fragment. Must not be formal. Must not use prohibited phrases. No URLs. No marketing. Should be about the video: "Your CDN can overrule robots.txt: finding the layer that refuses AI crawlers". So comment could be like "so does this mean we need to check CDN configs before relying on robots.txt?" Use lowercase start. Maybe "i didn't realize CD
Great breakdown! We actually hit a weird edge-case last week where Cloudflare's default Bot Fight Mode was returning 403 to GPTBot even though our robots.txt was totally open.
Debugging crawler headers at the edge is definitely getting trickier."