Learning Center

AmazonBot and Meta Are Eating Your Bandwidth Budget

September 8, 2026

Show Editorial Policy

shield-icon-2

Editorial Policy

All of our content is generated by subject matter experts with years of ad tech experience and structured by writers and educators for ease of use and digestibility. Learn more about our rigorous interview, content production and review process here.

AmazonBot and Meta Are Eating Your Bandwidth Budget
Ready to be powered by Playwire?

Maximize your ad revenue today!

Apply Now

Key Points

  • AmazonBot and Meta-External-Agent generated 63% of all AI bot sessions across three billion website visits analysed by 51Degrees, according to Press Gazette.
  • AI bot traffic climbed 62.5% year over year and now accounts for 13% of web sessions, up from 8%.
  • Publishers block only about 21% of AI scraper bots in their robots.txt files, per Known Agents, even though those bots comply 95.6% of the time.
  • Coverage is the gap: most publishers wrote their robots.txt before half these user agents existed.
  • Every bot session you serve costs bandwidth and returns nothing, which puts the pressure back on the yield of your human sessions.

What the 51Degrees Data Found

Press Gazette reported on new 51Degrees research analysing three billion website visits in the year to 29 May 2026. The sites span telecoms, advertising, media, retail, financial services, and technology.

AmazonBot accounted for more than 4% of all web visits over the period. Meta-External-Agent came in just under 3%.

Together those two bots represented 63% of sessions generated by the 72 AI-related bots 51Degrees identified. The study excluded Google's main search crawler, so this is purely AI-adjacent traffic.

Here is how the top crawlers stacked up on 29 May alone.

BotSessions (29 May)Share of total web sessionsOperator
AmazonBot6.43 million4%+Amazon
Meta-External-Agent4.23 million~3%Meta
AhrefsBot2.17 million2.3%Ahrefs
BingBot1.62 million1.3%Microsoft
ClaudeBot~288,000Not reportedAnthropic
OpenAI Search Bot~129,000Not reportedOpenAI

Those four leaders accounted for 85% of crawl sessions across all 72 identified bots. The AI names that dominate industry conversation, ClaudeBot and OpenAI Search Bot, ranked seventh and ninth.

AmazonBot Outweighs the Crawlers Making Headlines

Most publisher robots.txt files were built around GPTBot, CCBot, and ClaudeBot. Those are the crawlers that generated headlines and lawsuits.

The 51Degrees numbers tell a different story. AmazonBot alone generated more than twenty times the daily sessions of ClaudeBot.

Bot traffic analytics platform Known Agents describes AmazonBot as making "broad, high-volume sweeps that fetch far more pages per visit than a search crawler." That's a materially different load profile from a targeted training crawl.

Your origin servers feel that difference. So does your CDN bill.

Essential Background Reading:

  • AI Info: A primer on how AI is reshaping the publisher landscape, from crawlers to search results.
  • AI Content Info: Background on how AI systems interact with publisher content, before you get into blocking tactics.
  • AI and Publishers Resource Center: A central hub covering the full range of AI issues publishers are navigating right now.
  • Block AI: The basics of what blocking AI crawlers actually involves at the technical level.

Robots.txt Coverage Is the Weak Link

Known Agents found that AI scraper and data providers honor robots.txt requests 95.6% of the time. Publishers, meanwhile, block only around 21% of AI scraper bots in their files.

The bots are following the rules. Publishers just aren't writing enough rules.

Known Agents identified Adweek, The San Diego Tribune, and the Arkansas Democrat Gazette as having 100% coverage of the AI bots it tracks. W magazine, Screenrant, Semafor, CBS News, and Digital Spy sat at the other end, blocking roughly 2% of tracked bots.

That spread has nothing to do with technical sophistication. It reflects how recently someone opened the file and updated it against a current list of user agents.

Related Content:

What Bot Traffic Costs You Right Now

AI bots accounted for 13% of web sessions in May 2026, up from 8% a year earlier. Fastly estimates nearly half of all web traffic comes from bots overall.

51Degrees CEO James Rosewell called the increase a "huge licensing revenue opportunity" for publishers, while warning that unchecked bot traffic is "a growing and unwanted burden, particularly for smaller, independent players who can ill afford the impact of IP theft and increased bandwidth costs."

Rosewell also noted there is "no technical solution that can stop the bots entirely." He's right, and it's the line to remember the next time a vendor pitches you a silver bullet.

Bot sessions consume infrastructure and generate zero ad revenue. Your cost per session goes up while your monetisable session count stays flat.

Next Steps:

What Publishers Should Do This Week

The fix here is unglamorous and mostly a maintenance task. Nobody gets promoted for updating a text file, and that's exactly why the coverage numbers look the way they do.

Work through these in order:

  • Audit your current robots.txt against a live bot list: Pull your file and check it against the user agents hitting you today, rather than the ones that were newsworthy in 2023. AmazonBot and Meta-External-Agent belong on that list.
  • Separate your logs by user agent: Quantify what percentage of your sessions are bots before you decide anything. You cannot make a business decision on an unmeasured cost.
  • Decide bot by bot: BingBot drives search referrals. AmazonBot's high-volume sweeps return considerably less. Blanket blocking costs you distribution you may want.
  • Add edge-level enforcement for repeat offenders: Robots.txt is a polite request. Rate limiting or WAF rules at the CDN handle the crawlers that ignore it.
  • Re-audit quarterly: New user agents launch constantly. A file you updated last year is already stale.
  • Run our AI Crawler Protection Grader: It scores your current coverage in a couple of minutes and shows you exactly which bots are getting through.

Publishers weighing the broader strategic question, block versus licence versus optimise for citation, can work through the tradeoffs in our AI crawler resource center.

See It In Action:

The Human Sessions You Keep Have to Earn More

Blocking bots doesn't add a single human visitor to your site. It stops the bleeding on bandwidth and slows content extraction, which is reason enough to do it.

The revenue side of this equation lives elsewhere. AI Overviews and chat interfaces are compressing referral traffic across the open web, so the sessions you still get carry more weight than they did two years ago.

RPS becomes the metric that matters. If your monetisable session count is flat or declining, revenue per session is the only lever with meaningful room left.

How We Help Publishers Protect and Monetise Their Traffic

We built our AI crawler tooling because publishers kept asking us the same question: which bots are hitting me, and what are they costing me. The Protection Grader answers the first half in minutes.

The second half is what our RAMP platform handles. Managed Service, Self-Service, and Mobile App all run on the same yield infrastructure, with real-time price floor optimisation, multi-variable auction analysis across your full demand stack, and dynamic timeout tuning based on how your users behave.

You see every optimisation and every recommendation, with clear explanations of what changed and why. Quality, Performance, Transparency. No black box, no middleman markup buried in the waterfall.

AI bots are going to keep crawling. Your job is to make sure the humans who show up are worth more to your business than they were last quarter. Talk to our team about what that looks like on your inventory.

New call-to-action