AIToday

Cloudflare lets site owners choose which AI bots can access their content

Hacker News16h agoSend on LINE
Cloudflare lets site owners choose which AI bots can access their content

Key takeaway

Cloudflare is giving website owners more nuanced control over which AI bots can access their content, replacing an all-or-nothing blocking option with three separate categories: Search (allowed by default), Agent, and Training (both blocked by default on ad-supported pages starting September 15, 2026). The shift addresses a year of feedback that sites were unfairly forced to choose between search visibility and protection from AI model training, and introduces a "content use" framework—immediate, reference, or full—that lets owners specify how bots can store and reuse their material.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Cloudflare launched new controls allowing website owners to manage AI traffic by three use cases—Search (indexing for search results), Agent (real-time tasks like ChatGPT fetches), and Training (model fine-tuning)—replacing the previous binary "Block AI Bots" option. Starting September 15, 2026, new domains will block Training and Agent bots by default on ad-supported pages, while allowing Search. The company also introduced BotBase, an Enterprise feature providing a searchable database of all known bots with their classifications and content-use behaviors.

  • Why it matters

    Website owners have long faced a choice between search visibility and protecting their content from AI training—an unfair trade-off that advantages large incumbents. Cloudflare's taxonomy lets small sites allow beneficial search traffic while blocking or restricting bots that train models or scrape content without compensation. The "content use" levels (immediate, reference, full) give owners granular control over how their material is reused, addressing a year of feedback that the original "block all AI" approach was too blunt.

  • What to watch

    The September 15, 2026 default change will affect multi-purpose crawlers like Googlebot, Applebot, and BingBot—they will be blocked on Training if an owner has selected that restriction, since Cloudflare enforces the most restrictive rule when a crawler serves multiple purposes. Site owners can opt out of the new defaults anytime before that date via Security settings. Later in 2026, BotBase will expand from visibility to direct control, letting Enterprise customers enforce rules on specific bots in real time.

In Depth

One year after launching its "Content Independence Day" initiative, Cloudflare is rethinking how website owners should govern AI traffic to their sites. The company's original "Block AI Bots" feature offered a single-purpose solution: stop crawlers trained on model data. But conversations with customers and observation of the market revealed that approach was inadequate. Website owners wanted the ability to distinguish between different kinds of AI automation—some beneficial, some extractive—rather than blocking everything. Cloudflare's new framework defines three use cases. Search encompasses any crawling that indexes or collects content to answer questions later; the assumption is that site owners should receive referral traffic or compensation. Agent covers real-time automated behavior acting on a person's behalf, such as ChatGPT User-Agent fetches or browser-use agents like Gemini or Claude driving Chrome; these bots visit a web application to complete a task immediately. Training refers to crawlers that absorb content into a model's underlying architecture. Each of these three categories can now be managed independently on Cloudflare, available to all customers including the Free tier. Beyond these three, Cloudflare's taxonomy also tracks other bot behaviors—ads verification, feed fetching, SEO crawling, security testing, data collection, and monitoring—a full accounting that reflects the evolved state of automated web traffic. The company is now calling on bot operators to separate their crawlers by purpose, arguing that transparency benefits site owners by clarifying why a given crawler is visiting and enabling finer-grained access control. Starting September 15, 2026, Cloudflare will apply new defaults to all new domains onboarding to the platform. On pages that display ads, Training and Agent bots will be blocked by default, while Search will remain allowed. The reasoning is straightforward: an ad signals that a website owner intended human attention to land there and monetize it; bots that prevent that attention undermine the site owner's business model, while Search traffic typically funnels visitors back to the site. Multi-purpose crawlers—those that combine Search with Training, including Google, Apple, and Bing's bots—will be subject to the most restrictive rule a customer has applied; if a site owner has selected to block Training, those multi-purpose crawlers will be blocked entirely. Cloudflare will notify customers of the upcoming change and allow them to opt out by changing their settings anytime before September 15. For Enterprise customers, Cloudflare is also launching BotBase, a new dashboard feature that provides a searchable database of all known bots, their classifications across the taxonomy, and their content-use behaviors. This visibility tool will later expand to direct control, allowing Enterprise customers to manage known bots in real time. Cloudflare is also introducing a new optional field in robots.txt, part of its Content Signals standard, called "use" which allows site owners to express one of three content-reuse preferences: immediate (interact but store nothing), reference (index, excerpt, and link back), or full (summarize and reproduce). All customers already using Cloudflare's managed robots.txt will now have use=reference added automatically alongside the existing search=yes,ai-train=no signals. Bots that abuse these signals will lose their Verified status and cease being allowed on the platform, creating an incentive for operators to honor site owner preferences and maintain compliance with the emerging standard.

Context & Analysis

Cloudflare's update reflects a fundamental shift in how the web community views AI bot traffic. One year ago, the company launched a stark binary choice—"Block AI Bots"—in response to what it called an existential threat: AI systems scraping content for training without compensating creators. But a year of customer feedback revealed the flaw: that approach harmed small websites most, forcing them to choose between discoverability via search engines and protection from AI training, a Faustian bargain that entrenched large incumbents who control both search and AI infrastructure. The new three-category taxonomy (Search, Agent, Training) and the "content use" framework acknowledge that not all AI automation is predatory—real-time agent behavior and search indexing can drive value back to site owners—while still giving owners the tools to prevent their content from being absorbed wholesale into proprietary models without compensation. By setting Search as allowed-by-default and Training/Agent as blocked-by-default on ad-supported pages, Cloudflare is making a bet that most site owners benefit from search visibility more than they risk from training bots. The September 15, 2026 date is notable because it will affect major multi-purpose crawlers like Googlebot, Applebot, and BingBot—bots that crawl for both search and training—forcing those operators to either separate their crawlers (as Cloudflare encourages) or accept blocking on sites that opt for the restrictive defaults.

FAQ

When do the new defaults take effect?
On September 15, 2026, Cloudflare will set new defaults for all new domains: Training and Agent bots will be blocked by default on pages that display ads, while Search will remain allowed. Existing customers can opt out of these changes anytime before that date via Security settings.
What are the three AI use cases Cloudflare is distinguishing?
Search (crawling to index content for search results), Agent (automated real-time actions like ChatGPT or Claude browser fetches on behalf of users), and Training (crawling to train or fine-tune models). Website owners can now manage each category separately.
What does the "content use" setting do?
It lets site owners specify three levels of how bots can reuse their content: immediate (interact but store nothing), reference (index, excerpt, and link back), or full (summarize and reproduce). Owners can combine these with bot classifications to express rules like "allow Search bots only up to reference level."

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No discussion yet for this article

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime