Strategy6 minNewsroom

Cloudflare Blocks Mixed Use AI Crawlers by Default

Portão de bronze semifechado no lobby de marmore de uma antiga sede de jornal ao entardecer, com pilhas de jornais sob luz dourada.

New default policy applies to free customers and new sites: bots combining search, training, and agents are blocked on any ad-monetized page.

Starting Monday, September 15, Cloudflare will block mixed use crawlers by default on any page that displays advertising. This rule automatically affects new customers, new sites registered by existing customers, and the entirety of the free tier. Paying customers with prior configurations retain their existing settings, but any new site created from today onwards will fall under the restrictive regime.


The specific target is the bot that performs three functions within a user-agent: indexing for search, fetching pages on behalf of an agent, and collecting content for training. According to Cloudflare, this mixed-purpose approach allowed large AI companies to capture the commercial value of content under the label of a search bot. From now on, the operator of the crawler must present separate user-agents for each purpose or accept being blocked on the monetized slices of traffic.


What Changes for Publishers and AI Labs


The original announcement from July 1 already indicated the deadline of September 15. Cloudflare's Content Signals Policy, launched in September 2025, categorizes permitted use into three types that the publisher declares via robots.txt: search, AI answers, and AI training. The new default blocks training and agent usage on ad-serving pages while maintaining search as permitted. This choice effectively draws a paid perimeter around the content that supports the advertising model of the open web.


In the same move, Cloudflare has retired the Pay Per Crawl, launched in July 2025, replacing it with Pay Per Use, a marketplace that only charges the AI consumer when the content is actually used in a generated response. Ceramic.ai and You.com are present as initial partners. The logic aligns more with royalties on usage rather than access fees, addressing a frequent complaint from labs: paying for pages that never appear in a response.


How This Resonates in Three Markets


In the United States, where OpenAI, Anthropic, Google, and Perplexity concentrate the majority of affected crawlers, the immediate effect is operational. Bots that currently navigate a wide range of free-tier sites will need separate user-agents by deadline, and each bot exposed individually becomes a target for selective blocking. In a quarterly call transcript in August, Sam Altman told analysts that original quality content continues to be more valuable, not less, as models mature. This rhetoric aligns with the new asymmetry.


For European publishers, Le Monde, Der Spiegel, and Financial Times have been operating for months under their own licensing agreements with OpenAI and News Corp. For them, the new standard caps negotiations and empowers smaller publishers that do not have individual contracts. The German Media Publishers Association has already signaled intentions to use Content Signals as a sector standard in collective negotiations for the next cycle.


In Brazil, where major news organizations still rely predominantly on bilateral commercial agreements with Google and Meta and where PL 2,630 remains stalled, the Pay Per Use mechanism offers a technical alternative that bypasses individual negotiation with each AI lab. This option is more appealing for medium-sized outlets without dedicated legal teams and challenges the notion that only large newsrooms can monetize their own content in AI.


What This Move Does Not Resolve


There are blind spots. Crawlers hosted on residential, rotating IPs and boutique scrapers operate in a gray area that the default network standard does not reach. Matthew Prince, CEO of Cloudflare, admitted in the original July announcement that the policy applies to declared traffic and is a line of defense, not a wall. The battle for training data continues, but now in a darker part of the web.


There is also a rarely discussed strategic collateral effect. By pushing labs towards separate user-agents, Cloudflare also exposes previously obscured behavior metrics: volume by purpose, recrawl frequency, industry coverage. Publishers will begin to see who is genuinely training on their archives and who is merely responding to agent queries. Nathan Benaich from Air Street Capital wrote in the August State of AI newsletter that this transparency favors negotiations for differentiated licensing tiers by type of use, a model that the music industry adopted after two decades of conflict with Napster.


The financial cap of the exercise also deserves attention. Cloudflare has not disclosed transaction volumes from Pay Per Crawl during the pilot and there are no public projections for Pay Per Use. Small publishers gain technical control; whether they realize relevant revenue depends on the appetite of the labs to populate the marketplace. The response will come in the next two quarters, and it matters more than the default blocking itself: it will indicate whether the AI data market becomes a priceable infrastructure or continues to operate on opportunistic capture.

The week's analysis, by email

One weekly edition with what matters to people who decide. No ads, no sponsorship.

One-click cancellation, at any time.

Strategy