Cloudflare Just Split AI Crawlers Into Three Classes and the Default Could Block Googlebot on Your Site
Cloudflare's new AI crawler taxonomy separates Search, Agent, and Training bots. September 15 defaults block Agent and Training on ad-supported pages. The trap: Googlebot is a mixed-use crawler, so blocking Training blocks Google too.
Cloudflare just reclassified every AI bot that touches your website into three categories: Search, Agent, and Training. On September 15, the defaults flip. Agent and Training crawlers get blocked on any page running ads. That affects more than 20% of all web domains. And the part nobody is talking about: Googlebot does both Search and Training. Block one, you block both.
Three Crawler Classes, One Default That Changes Everything
Until now, Cloudflare gave website owners a single toggle: block AI bots or don't. That toggle now splits into three.
Search crawlers index your content so AI engines can answer questions about it later. These stay allowed by default.
Agent crawlers act in real time on behalf of users. ChatGPT browsing the web for a buyer comparing your product to a competitor's. Blocked by default on ad-supported pages starting September 15.
Training crawlers permanently absorb your content into model weights. Also blocked by default on ad-supported pages starting September 15.
The logic makes sense on the surface. Publishers should control whether their content trains someone else's model. But the implementation creates a problem that most founders will not see coming.
The Mixed-Use Crawler Trap
Googlebot, Applebot, and BingBot are not clean categories. They perform Search and Training in the same crawl session. Cloudflare's rule: when a bot serves multiple purposes, the most restrictive setting applies.
If you block Training crawlers, Googlebot gets blocked. On every page that displays ads.
That is not a hypothetical edge case. Google's own crawler is a mixed-use bot. Googlebot handles Search results and Gemini/AI Overviews training in the same process. BingBot manages Bing indexing and Copilot data pipelines. Applebot feeds Spotlight, Siri, and Apple Intelligence. The three search engines that still drive the majority of organic discovery are dual-purpose crawlers, and Cloudflare's taxonomy treats the Training label as dominant.
The internet's bot-to-human traffic ratio flipped in 2025. More automated requests than human ones. The infrastructure layer had to respond. But the way it responded creates a visibility trap for anyone who does not understand which bots do what.
Early estimates suggest mid-sized publishers blocked for 4 weeks could see 20% to 40% organic session drops, with full recovery taking 4 to 8 weeks after re-enabling Search access. That is not a rounding error. That is a quarter's pipeline.
What This Means If You Run Ads Behind Cloudflare
If your site runs ads and sits behind Cloudflare, here is what September 15 looks like:
- Training and Agent crawlers are blocked by default on ad-supported pages
- Googlebot is classified as a mixed-use crawler (Search + Training)
- Under the most-restrictive-rule principle, Googlebot gets blocked on those pages
- Your content becomes invisible to Google's crawler on the pages generating revenue
You can opt out. Cloudflare is giving website owners explicit controls in their Security settings. But the default is set to block. Founders who do not actively configure their AI crawler settings before September 15 will have their defaults changed for them.
Cloudflare framed this as giving publishers the power to push AI companies to pay for content they use for training. That framing is correct for publishers who want to monetize their content. But for founders who need AI engines to read and cite their content, the same defaults work against them.
I wrote about the initial Cloudflare AI bot blocking shift when the single toggle launched last year. That was a warning shot. This is the structural split that makes it real, because now the categories are granular enough to catch bots you actually need.
Your CDN Is Now Your AI Visibility Gatekeeper
This is the shift I want you to sit with.
We spent two decades optimizing for search engines. Then the game expanded to AI engines that decide who gets cited. Now the game has expanded again. The infrastructure layer sitting between your server and the internet is making visibility decisions on your behalf.
Cloudflare is not wrong for doing this. Publishers deserve control over how their content gets used. But the consequence is concrete: your CDN configuration is now a first-order AI visibility variable. It sits upstream of your content, your search optimization, your earned media strategy. Everything.
Content that AI engines cannot crawl does not get cited. Content that does not get cited does not appear when a buyer asks ChatGPT or Perplexity who to trust. The crawl is the first gate. And Cloudflare just put that gate on a timer.
Three Moves Before September 15
Audit your Cloudflare AI settings now. Go to Security, find the AI traffic controls, and understand which of the three categories you are allowing or blocking. If you have not touched these settings, you are running on defaults that change in eight weeks.
Decide which crawlers serve your business. Search crawlers drive discovery. Agent crawlers power the real-time AI experiences your buyers are already using. Training crawlers build the models that will shape how your brand is understood for years. Each category has a real cost and a real benefit. Make the call intentionally.
Check BotBase for the bots hitting your site. Cloudflare launched BotBase, a searchable database classifying every known bot by type and behavior. Know exactly which bots are crawling your content, what they are classified as, and how your settings affect them.
The founders who will lose here are the ones who treat infrastructure as someone else's problem. Your CDN is no longer just a performance layer. It is the first gate AI engines have to pass to read your content.
If the gate is closed, nothing downstream matters. Not your content quality. Not your citation signals. Not your earned media. The machine never gets to read it.
FAQ
Will Cloudflare's September 15 changes affect all websites?
The new defaults apply to ad-supported pages on domains behind Cloudflare. More than 20% of all web domains use Cloudflare infrastructure. If your site runs ads and uses Cloudflare, Training and Agent crawlers get blocked by default unless you configure otherwise.
Can I still allow Googlebot if I block Training crawlers?
Not under the current rules. Googlebot is classified as a mixed-use crawler performing both Search and Training functions. Cloudflare applies the most restrictive setting to multi-purpose bots. Blocking Training crawlers blocks Googlebot on affected pages unless you explicitly override the default.
What is the difference between Agent and Training crawlers?
Agent crawlers act in real time on behalf of users, like ChatGPT browsing the web during a conversation. Training crawlers permanently absorb content into AI model weights for capability improvement. Both are blocked by default on ad-supported pages starting September 15, but they serve fundamentally different purposes for your brand's AI visibility.