Learning Center

ChatGPT's Fetch Bot Ignores Robots.txt. Now What?

August 17, 2026

Show Editorial Policy

shield-icon-2

Editorial Policy

All of our content is generated by subject matter experts with years of ad tech experience and structured by writers and educators for ease of use and digestibility. Learn more about our rigorous interview, content production and review process here.

ChatGPT's Fetch Bot Ignores Robots.txt. Now What?
Ready to be powered by Playwire?

Maximize your ad revenue today!

Apply Now

Key Points

  • OpenAI says robots.txt may not apply to ChatGPT-User because the fetches are user-initiated, not automated crawling in the traditional sense.
  • TollBit's State of the Bots report found ChatGPT-User reached disallowed pages on nearly half of European sites that had explicitly blocked it.
  • Blocking ChatGPT-User and OAI-SearchBot together removes AI search visibility without guaranteeing protection from page fetching.
  • Cloudflare is moving crawler enforcement to the network layer, with new domains blocked by default for Training and Agent crawlers on ad-serving pages starting September 15.
  • The right response depends on your site's goals. Publishers should audit server logs, not just assume robots.txt is working.

What Happened

TollBit's latest State of the Bots report, covering the first half of 2026, confirmed what a lot of publishers suspected but couldn't prove: ChatGPT-User is reaching pages sites explicitly told it to avoid. Around 15% of identified AI page-fetchers reached disallowed URLs across European sites in the study. ChatGPT-User, Bytespider, and Youbot each hit disallowed pages on nearly half of the European sites that had specifically listed them. ChatGPT-User reached the most sites of any agent in that group.

OpenAI's position: Their crawler documentation states that ChatGPT-User visits a page when a real user asks a question about it. Because a human initiated the request, OpenAI argues robots.txt rules may not apply. Perplexity takes the same stance with Perplexity-User. Anthropic draws a different line and says all three of its bots respect the file.

TollBit counts any request to a disallowed URL as a bypass, regardless of what the operator claims about intent. That's the right framing for publishers trying to understand real-world exposure.

See It In Action:

Why the Blocking Logic Gets Complicated

The agent situation is messier than most publishers realize. OpenAI uses two distinct bots with two different jobs. OAI-SearchBot determines whether your site appears in ChatGPT search results. ChatGPT-User fetches the page content when a user asks about it. Separate systems, separate purposes.

If you block both to shut out all AI traffic, you've given up AI search visibility while still operating under a fetching policy that OpenAI says carries a user-initiated carve-out. That's not necessarily wrong, but it should be a deliberate choice, not an accidental one.

The robots.txt file shows what you asked for. Server logs and CDN records show what happened. Those are not always the same document.

Here's how the major AI companies currently handle robots.txt compliance for their page-fetching agents:

AgentOperatorRespects robots.txt?Notes
ChatGPT-UserOpenAIConditionalUser-initiated requests may bypass
OAI-SearchBotOpenAIYesControls AI search visibility, separate from fetching
Perplexity-UserPerplexityConditionalSame user-initiated rationale as OpenAI
ClaudeBot / Claude-UserAnthropicYesAnthropic states all three bots respect the file
BytespiderByteDanceNo consistent complianceHigh disallow-page reach rate in TollBit data

Essential Background Reading:

What Publishers Should Do

This situation calls for a clear-eyed audit, not a panic-driven block list. The right decisions depend on what your site needs from AI traffic.

Start with your server logs. Check whether ChatGPT-User, Bytespider, or other agents are reaching pages you've marked disallowed. robots.txt is not a firewall. Treating it like one leads to bad assumptions about your actual exposure.

From there, think through the trade-offs clearly:

  • Block both OAI-SearchBot and ChatGPT-User: you remove AI search visibility entirely. The fetching control you've added may or may not hold given OpenAI's documented carve-out.
  • Block ChatGPT-User only: you retain some AI search presence through OAI-SearchBot but still face the user-initiated fetch argument for page content.
  • Block neither: ChatGPT indexes and fetches your content. Potentially useful for referral traffic and citation visibility, depending on your niche.
  • Use network-layer enforcement: Cloudflare's updated controls, active for new domains starting September 15, move the enforcement decision away from voluntary bot compliance and onto the CDN itself. Training and Agent crawlers will be blocked by default on pages with ads, while Search crawlers remain allowed.

Use our AI Crawler Protection Grader to see how your current setup holds up. And if you want the full strategic framework, the AI crawler resource center for publishers covers the decision tree in detail.

Related Content:

The Cloudflare Development Worth Watching

Cloudflare's shift to network-layer enforcement is the most structurally significant development in this story. When a CDN makes a compliance decision at the network edge, the bot's stated policy on robots.txt becomes largely irrelevant. The page never arrives.

The default-block behavior on ad-serving pages for new domains is a direct acknowledgment that AI agent fetching affects publisher monetization, not just content control. Whether the user-initiated loophole argument survives legal or regulatory scrutiny is an open question. Network-layer blocking doesn't wait for that answer.

The broader question for publishers: how much of your current traffic still comes from human users? Whatever AI does to your referral mix, the sessions you retain need to be working as hard as possible.

Next Steps:

Protecting the Revenue You Still Have

AI-driven changes to your traffic mix make yield optimization more important, not less. Fewer sessions with the same ad infrastructure means your revenue per session needs to go up.

We work with publishers across gaming, entertainment, education, and news to maximize RPS from every real human session. Our RAMP platform handles the programmatic complexity so your team doesn't have to. When traffic gets unpredictable, the revenue floor you've built matters more than ever.

Check your actual log data. Update your blocking decisions based on what the bots are doing, not what they say they're doing. And make sure every legitimate session you have is earning what it should.

Talk to our team about what that looks like for your site.

New call-to-action