Learning Center

The 411 on Cloudflare's New AI Crawler Rules

July 31, 2026

Show Editorial Policy

shield-icon-2

Editorial Policy

All of our content is generated by subject matter experts with years of ad tech experience and structured by writers and educators for ease of use and digestibility. Learn more about our rigorous interview, content production and review process here.

The 411 on Cloudflare's New AI Crawler Rules
Ready to be powered by Playwire?

Maximize your ad revenue today!

Apply Now

Key Points

  • Cloudflare is replacing its one-click AI bot block with a three-tier system covering Search, Agent, and Training crawlers, effective September 15, 2026.
  • Training and agent crawlers are blocked by default on ad-supported pages for newly onboarded domains, because ad presence signals a human audience is expected.
  • Multi-purpose crawlers like Googlebot face the strictest applicable rule: block training, and you block the entire crawler.
  • According to Similarweb data cited in the source reporting, Google AI Overviews now appear in 43% of searches, and publisher referral traffic is falling as a result.
  • Publishers using the traffic they still have efficiently matters more than ever. Granular crawler control is one piece of the puzzle.

What Cloudflare Changed

Cloudflare announced a significant overhaul of its AI crawler management system. The familiar "block all AI bots" toggle, launched in July 2025, has been retired in favor of three independent controls: Search, Agent, and Training. The new rules take effect September 15, 2026, and they apply to all Cloudflare users, including free-tier accounts.

The three categories aren't arbitrary. Cloudflare is drawing sharp lines around crawler intent:

  • Search crawlers: index content and return referral traffic. Cloudflare's stance is permissive here, provided site owners receive traffic or compensation in return.
  • Agent crawlers: execute real-time tasks on behalf of users. Examples include ChatGPT-User, Gemini, and Claude performing browser-based operations.
  • Training crawlers: scrape content for model training, embedding that data permanently into model architecture. These face the strictest defaults.

For newly onboarded domains, training and agent crawlers are blocked by default on pages displaying ads. Cloudflare's reasoning is direct: an ad on a page signals that a real human is expected to see it. Bots consuming that inventory, or taking that content to train a model, weren't part of the deal.

The Multi-Purpose Crawler Problem

The most consequential piece of this announcement is how Cloudflare handles crawlers that do multiple jobs.

Googlebot, Applebot, and BingBot all serve both traditional search indexing and AI training functions. Under the new rules, these multi-purpose crawlers are governed by the strictest applicable rule. Block training crawlers, and you block the whole crawler. There's no surgical option to let Google index your content while preventing it from feeding your articles into Gemini's training data.

That's a hard trade-off. Publishers who rely on organic search traffic for a significant share of their audience will need to think carefully before pulling the training block lever on major search engine bots. The leverage is real, but so is the cost.

Cloudflare operates infrastructure for over 20% of websites globally. If even a fraction of those domains adopt training blocks on major crawlers, the pressure on Google, Apple, and Microsoft to negotiate AI licensing agreements like the one Reuters pursued increases materially.

Essential Background Reading:

Supporting Tools Worth Knowing About

Cloudflare didn't just change the controls. It shipped infrastructure to make those controls meaningful. Three additions are worth understanding:

  • BotBase: a crawler visibility database for enterprise customers, listing verified bots with classifications and purposes. Publishers can filter, search, and export detection IDs for use in custom security rules.
  • Content Signals expansion: a new "use" signal with three values: immediate (interact, don't store), reference (index, summarize, link back), and full (summarize and reproduce). Cloudflare-hosted robots.txt files will default to "use=reference." Bots that violate the signal lose Verified status.
  • Transfer trust scheme: an experimental mechanism using the RFC 7239 Forwarded header, allowing crawlers to carry identity and usage intent across intermediate layers. Lose trust status, lose access across the network.

The Content Signals work is particularly important for publishers thinking about citation traffic. A bot claiming "reference" intent and then reproducing content in full will lose its Verified status. That's a meaningful enforcement mechanism, assuming AI companies actually comply.

Related Content:

The Traffic Context Driving All of This

None of this happens in a vacuum. The backdrop is a structural shift in how users interact with search.

According to Similarweb data cited in the source reporting, Google AI Overviews appeared in 43% of searches as of May 2026, up from 15% a year earlier. Google AI Mode visits more than doubled year-over-year, reaching 279 million. Average query length has grown significantly as users shift from short keywords to conversational questions, and users are spending more time on Google's platform rather than clicking through to publisher content.

Cloudflare framed this directly: the 30-year implicit agreement between crawlers and site owners, "I crawl you, you send me referral traffic," has broken down. The three-tier system is its attempt to force a renegotiation.

One data point cuts the other way. Similarweb reported that following a May 7, 2026 search update, the proportion of ChatGPT desktop sessions that included landing page visits jumped from 25% in March to nearly 60% by late May. When AI-driven results include prominent links, users do click. The traffic isn't gone. It's just harder to capture.

Next Steps:

What Publishers Should Do Before September 15

The deadline is real and the defaults matter. Here's how to approach the decision:

  • Audit your Cloudflare settings now: If you're already blocking all AI crawlers with the old toggle, understand how that maps to the new three-tier system. The migration behavior should be documented in Cloudflare's official blog.
  • Separate your search dependency from your training tolerance: If Google organic traffic drives a significant portion of your sessions, blocking training crawlers on Googlebot carries real risk under the new strictest-rule logic. That's a business decision, not just a technical toggle.
  • Use Content Signals if you want citation credit: Setting "use=reference" in your Cloudflare-hosted robots.txt signals that indexing and summarization are acceptable but full reproduction is not. It won't stop every bad actor, but it creates an enforcement hook. Understanding what factors drive AI citation rankings helps you make that signal count.
  • Watch BotBase for enterprise insights: If you're managing multiple properties and need visibility into which verified bots are hitting which pages, BotBase is worth evaluating as it rolls out.

See It In Action:

Making the Most of the Traffic You Still Have

Crawler controls determine who can access your content. They don't determine what happens when a real human actually shows up.

Publisher referral traffic is under pressure across the board. Sessions are getting harder to earn. That makes the revenue you generate from each session more critical, and it's where ad monetization strategy becomes the variable you can actually control.

AI bot traffic now accounts for 40% of web traffic, which means the human sessions you're generating are increasingly valuable. If your floor pricing, ad layout, and demand path aren't optimized for the sessions you do have, you're compounding the traffic problem with a yield problem. We work with publishers to make sure that doesn't happen. Our RAMP platform is built to maximize RPS across a full demand stack, so the sessions you earn are doing as much work as possible.

Cloudflare is helping publishers protect their content. We help them make sure that content generates revenue.

Check out our AI Crawler Protection Grader to see where your current crawler configuration stands, and visit our AI crawler resource center for the full picture on protecting your content without leaving revenue on the table.

New call-to-action