Learning Center

Cloudflare's AI Bot Blocker Is Also Blocking Googlebot

August 5, 2026

Show Editorial Policy

shield-icon-2

Editorial Policy

All of our content is generated by subject matter experts with years of ad tech experience and structured by writers and educators for ease of use and digestibility. Learn more about our rigorous interview, content production and review process here.

Cloudflare's AI Bot Blocker Is Also Blocking Googlebot
Ready to be powered by Playwire?

Maximize your ad revenue today!

Apply Now

Key Points

  • Cloudflare's AI bot blocking feature is reportedly issuing 403 errors to legitimate Googlebot and Bingbot crawlers, not just AI training bots.
  • Cloudflare officially plans to block "mixed-purpose crawlers" that combine search indexing and AI training starting September 15, 2026, but some publishers may already be seeing this behavior.
  • The core conflict: Cloudflare now classifies Googlebot and Bingbot as "Search + Training" bots, meaning blocking AI training can block your search visibility at the same time.
  • Publishers using Cloudflare's AI Crawlers & Scrapers feature need to audit their settings now, before the September deadline.
  • Protecting your content from AI scraping is legitimate. Accidentally deindexing your site is not the trade-off you want.

What Happened

Search Engine Journal reports that a Redditor in r/SEO posted an issue that should get every publisher's attention. They were testing Cloudflare's AI Crawlers & Scrapers feature and set "AI Training = Block." The result: both Googlebot and Bingbot started receiving HTTP 403 responses when attempting to fetch the site's sitemap. Disabling the block made the 403s disappear immediately.

Google's John Mueller responded to the post and asked for a direct message to investigate further. The original poster confirmed it wasn't a case of fake bots spoofing Googlebot's user agent. Cloudflare's own dashboard was showing Googlebot and Bingbot as blocked, not impersonators.

This isn't a misconfiguration anecdote. It points to a structural problem with how Cloudflare is classifying bots.

See It In Action:

Why This Matters

Cloudflare's official documentation confirms the mechanic. Starting September 15, 2026, Cloudflare will update default settings for new domains: bots classified as "Training" or "Agent" will be blocked on pages that display ads, while "Search" will remain allowed.

The critical phrase in Cloudflare's announcement: "Mixed-purpose crawlers that combine Search and Training will also be blocked by all configurations to block AI training, including the legacy 'Block AI bots' option."

Cloudflare has classified Googlebot and Bingbot as mixed-purpose. They index for search. They also collect data that feeds AI training. So when you block AI training, you block both functions at once.

That creates a binary choice no publisher should have to make. Allow AI training access to keep your search rankings intact, or block AI training and risk 403 errors wiping out your crawlability. Neither option is clean.

The September default isn't the only concern, either. The Redditor's report suggests some configurations may already be triggering this behavior today, weeks before the scheduled change. Whether that's a fluke, user error, or early rollout behavior isn't clear yet. The safe assumption: check your settings now.

Essential Background Reading:

  • AI Crawler Resource Center for Publishers: Everything publishers need to know about AI crawlers, blocking strategies, and protecting content without sacrificing traffic.
  • AI Content Info: How AI systems use publisher content and what that means for your content strategy and rights.
  • Block AI: An overview of your options for blocking AI access to your site, with guidance on what each approach actually does.
  • AI and Publishers Resource Center: Playwire's full library of publisher-focused AI guidance, covering traffic, revenue, and protection strategy.

What Publishers Should Do

The situation isn't as simple as "turn off AI blocking." Here's a structured way to think through your options:

OptionSearch CrawlabilityAI Training AccessRisk Level
Allow AI Training (no block)PreservedOpenContent scraped freely
Block AI Training (current behavior)Potentially disruptedBlockedSearch visibility at risk
IP allowlist for GooglebotPreservedMixedRequires maintenance
Wait for Cloudflare fixUncertainUncertainHigh if September arrives first

The cleanest interim workaround, based on what's known, is to verify legitimate Googlebot IPs through Google's published IP ranges and create specific allowlist rules in Cloudflare that explicitly permit those crawlers regardless of your AI training block settings.

Here's what to check right now:

  • Cloudflare dashboard review: Log into your Cloudflare account and check the AI Crawlers & Scrapers section. Confirm whether Googlebot and Bingbot appear as blocked.
  • Sitemap fetch test: Use Google Search Console's URL Inspection tool to verify Googlebot can still access your sitemap. A 403 response there will confirm the problem.
  • Bot Fight Mode: If you have Bot Fight Mode enabled alongside AI Training blocking, test whether disabling it resolves the 403s. The original poster's experience suggests Bot Fight Mode may compound the issue.
  • Cloudflare opt-out: Before September 15, Cloudflare says all customers can opt out of the new defaults. Find that setting and decide whether you want to control the transition manually.
  • Monitor crawl errors: Set up crawl error alerts in Search Console if you haven't already. Catching a 403 wave early is much better than noticing a traffic drop three weeks later.

The September 15 deadline is real. Cloudflare has announced it publicly. Publishers who rely on Cloudflare for AI bot protection need to understand that the tool's current architecture treats search indexing and AI training as bundled behaviors for major crawlers. That architecture is the problem, and there's no guarantee it gets resolved before the default changes roll out.

Related Content:

The Larger Trade-Off

This situation shows why AI crawler protection decisions deserve more than a one-click toggle. The publishers most actively protecting their content from AI scraping are, in many cases, the same publishers most dependent on search traffic for their revenue.

Blunt blocking tools don't distinguish between a crawler here to serve your interests and one here to extract your content for free. That distinction matters enormously when search visibility is directly tied to ad revenue, subscription acquisition, or audience growth.

The right approach is selective, configurable, and monitored. Broad switches with undocumented classification logic create exactly the kind of collateral damage this Redditor discovered.

Next Steps:

  • AI Crawler Protection Grader: Assess your current crawler protection setup and identify gaps before the September 15 Cloudflare deadline.
  • Manage Blocking: Practical tools and guidance for managing your blocking configuration with precision rather than blunt toggles.
  • Ad Load and Traffic Stability: Why traffic stability is a revenue variable and how ad load decisions connect to long-term yield performance.
  • AI Info: Playwire's broader perspective on AI developments and how they affect publisher revenue and operations.

How We're Thinking About This

Our AI Crawler Protection Grader and AI crawler resource center exist because this category of decisions is complicated. Protecting your content from AI scraping is a legitimate business objective. Doing it in a way that accidentally tanks your search indexing is not a trade-off worth making.

The publishers we work with rely on us to help them navigate exactly this kind of situation: where the tools designed to solve one problem introduce another. If you're using Cloudflare's AI blocking features and haven't verified your Googlebot access since this report surfaced, that's the first thing to do today.

Search traffic you've earned through years of content investment is worth protecting on both fronts.

New call-to-action