Cloudflare Separates AI Training From Search
Subscribe
Cloudflare Separates Search Visibility From AI Training Controls

marketing artificial intelligence

Cloudflare Separates Search Visibility From AI Training Controls

Cloudflare Separates Search Visibility From AI Training Controls

Business Wire

Published on : Sep 16, 2026

Cloudflare is giving website owners a more granular way to manage AI crawler access without forcing them to choose between AI training and traditional search visibility. The company has launched Disallow AI Training, alongside new controls for search, AI training and AI agents.

For years, website owners could largely treat search crawling as a straightforward exchange: crawlers accessed content, and search engines potentially returned visitors to the site. The growth of generative AI has complicated that relationship by introducing new uses for the same web content, including model training, AI-generated summaries and agentic browsing.

Cloudflare is now attempting to separate those uses.

On September 15, the company launched Disallow AI Training, allowing publishers and businesses to prohibit AI training while continuing to permit search crawling. Cloudflare is also introducing an Accountable designation for crawlers that provide controls and transparency around how website content is used.

The change addresses a technical problem created by mixed-use crawlers. A single crawler can sometimes collect content for both search and AI-related purposes, making it difficult for a publisher to block training without also affecting search discovery.

Cloudflare says mixed-use crawlers represented 36.6% of verified crawler traffic on its network, making them the largest crawler category in its data. The company also reports that fewer than 1% of website owners block search crawlers, compared with 17% restricting AI training. These are Cloudflare network measurements rather than industry-wide statistics.

The company's new framework establishes four criteria for its Accountable designation. Crawler operators must provide a mechanism for opting out of AI training, offer controls for AI-generated search summaries, provide URL-level visibility into content usage, and publicly confirm that opting out of training does not affect traditional search rankings.

Cloudflare identifies Apple, Google and Microsoft as Accountable under the new framework, based on its stated criteria and commitments. Google's existing documentation independently confirms that publishers can use the Google-Extended robots.txt token to control whether content is used to train future Gemini models, and that this control does not affect inclusion in Google Search or act as a Google Search ranking signal.

Cloudflare is also replacing its previous single Block AI Bots approach with three separate controls for Search, Training and Agent behavior. Its documentation says that, beginning September 15, new Cloudflare domains have updated defaults in which Training and Agent crawlers are blocked on pages displaying ads while Search remains allowed.

That distinction matters for publishers whose business depends on organic discovery. Blocking AI training no longer necessarily has to mean blocking the crawler responsible for search discovery, provided the operator supports the required controls.

Cloudflare is also introducing Bot Preference Sync, designed to let customers establish crawling preferences once and apply them across supported crawlers. The company says it intends to add more granular controls for AI-generated summaries, allowing publishers to determine how much content can be included rather than relying only on an all-or-nothing setting.

The development comes as the web moves toward more formal mechanisms for expressing AI usage preferences. The IETF's AI Preferences Working Group is developing standards for communicating how digital content may be collected and processed by AI systems. Current drafts address both a vocabulary for AI usage preferences and mechanisms for associating those preferences with content through HTTP and the Robots Exclusion Protocol.

For publishers and SEO teams, the emerging model changes the question from simply asking whether a crawler should be allowed to asking what the crawler is allowed to do.

Market Landscape

The web's crawler ecosystem is becoming more complicated as search engines, AI platforms and autonomous agents access the same publisher infrastructure for different purposes.

Traditional SEO has largely focused on crawlability, indexation and search rankings. AI introduces additional considerations: whether content can be used for model training, whether it can appear in generated answers, and whether agents can retrieve or act on information in real time.

Cloudflare's approach separates these behaviors into distinct categories. Google similarly provides Google-Extended as a control for Gemini-related training while maintaining Google Search access.

The emerging IETF work suggests that the industry is also moving toward standardized ways of expressing these preferences rather than relying solely on vendor-specific controls.

Strategic Outlook

The distinction between search optimization and AI visibility is becoming increasingly important for publishers.

A website may want maximum search discovery while limiting training use. It may allow AI systems to summarize content while restricting autonomous agents. Another publisher may choose to permit all three because its commercial model depends on broad distribution.

That creates a more granular version of technical SEO governance. Robots.txt, crawler identification and access controls are becoming part of a wider content-rights and AI-discovery strategy.

For marketing and publishing teams, the practical priority will be understanding which crawlers access their content, what those crawlers are permitted to do and whether those preferences affect organic discovery.

Cloudflare's announcement also indicates that crawler accountability may become an increasingly important layer between publishers and AI platforms as the web develops standards for AI-era content access.

Top Insights

  • Cloudflare's Disallow AI Training setting separates AI model training controls from search visibility, addressing a key concern for publishers dependent on organic discovery.
  • Mixed-use crawlers create a technical challenge because one crawler can support both search and AI purposes, making granular access controls increasingly important.
  • Cloudflare's Accountable designation establishes requirements around AI training opt-outs, AI summaries, URL-level visibility and search-ranking assurances.
  • Google independently confirms that Google-Extended can restrict Gemini training without affecting a website's inclusion in Google Search or its ranking.
  • The IETF's AI Preferences work indicates that standardized mechanisms for communicating AI content-use preferences are developing alongside vendor-specific controls.

Get in touch with our MarTech Experts.

Looking to publish a press release, guest article, interview or podcast? Connect with us.

GET FEATURED