· via dev.to (home feed)
Cloudflare's September default blocks Googlebot at the edge, and robots.txt stays silent
A September 15 Cloudflare default treats mixed-use crawlers like Googlebot under its strictest rule, returning 403s on ad-bearing pages while robots.txt still says Allow.

What changed
On September 15, 2026, Cloudflare changed what "blocked" means for AI traffic, according to a post on dev.to by RankCLI. New domains, new sites added to existing accounts, and free-tier accounts that had never touched the setting now start with AI Training crawlers disallowed and AI Agents blocked on pages that carry ads. Search crawling was supposed to remain open.
For many sites, it did not. The reason lies in how Cloudflare classifies crawlers that do more than one job.
Mixed-use crawlers get the strictest rule
Googlebot crawls the web for Google Search and also collects data that feeds AI training. The post explains that Cloudflare classes it as a mixed-use crawler, and a mixed-use crawler is treated under whichever rule applying to it is most restrictive.
So when AI training is blocked, including through the older "Block AI Bots" toggle, the block covers all of Googlebot's functions on ad-supported pages. Googlebot starts receiving 403 responses on ordinary pages and on sitemap fetches. The post reports the same applies to Bingbot and Applebot.
The site's robots.txt still says Allow. Nothing in the repository changed. Google Search Console begins reporting fetch failures, and the cause is hard to pin down because everyone checks robots.txt first.
Cloudflare's own network data, cited in the post, put Googlebot's share of crawls at 27.49% in Q2 2026, down from 57.20% a year earlier. That decline predates the September change, but it frames how far search crawl volume has already fallen.
Why robots.txt cannot see this
The post draws a distinction worth internalising: robots.txt is a request, while a CDN bot rule is enforcement. robots.txt is served by the origin and read by a crawler that chooses to obey it. A bot rule at the edge refuses the connection before the origin is ever consulted.
That means a crawler can be explicitly allowed in robots.txt and still receive a 403. The author notes that every "AI visibility" checker they know of, including their own until recently, stops at robots.txt and reports such sites as open.
How to check
Two curl requests against the same URL, one with a browser user agent and one with a Googlebot user agent, will reveal the split. A browser 200 paired with a Googlebot 403 points to an edge rule rather than robots.txt. The sitemap is worth testing the same way, since the post says that is often where the failure appears first.
One caveat the author flags: real crawlers are verified by IP as well as user agent, so a spoofed agent string is indicative rather than conclusive. Search Console's URL Inspection tool, which fetches as the verified Googlebot, gives the definitive answer. The post also describes a CLI tool that tests search and AI crawlers against a browser baseline and reports separate findings for blocked search crawlers, an error, and blocked AI crawlers, a warning. The split reflects that losing the index is more urgent than losing AI citations.
The fix
The setting lives in the CDN, not in the repository. On Cloudflare, the post points to a zone's Security and Bots settings, and any "Block AI Bots" or AI Crawler Control toggle. Because Googlebot is mixed-use, site owners need to allow search crawlers explicitly or opt out of the default on ad-supported pages.
Existing paid customers were notified in advance and can opt out in zone security settings, according to the post. New domains and free-tier accounts received the default without asking for it, which is why the change keeps surprising people who never changed a setting.
Why it matters
Bot control has effectively moved up a layer, from a voluntary convention read at the origin to rules enforced at the edge before a request ever arrives. Any audit that stops at robots.txt, including most automated visibility checkers, will report these sites as open while search engines get 403s. The distinction between search and AI crawlers also carries real commercial weight: losing Googlebot costs a site its index, while losing GPTBot or ClaudeBot costs citations in AI answers. Defaults applied without the site owner's action now decide which of those a site gets, and nothing in the site's own code or configuration will explain why.
- #cloudflare
- #seo
- #robots-txt
- #web-crawlers
- #ai-crawlers