Cloudflare’s new “Disallow AI Training” setting lets website owners refuse model training without shutting out Google’s search crawler. The company announced its availability on all plans on September 15. Cloudflare’s announcement
That changes the scenario in SEW’s September 9 report. Cloudflare’s earlier plan would have brought mixed-use crawlers—bots used for both search and training—under existing training blocks unless customers opted out. The distinction matters for publishers that want to protect their content without cutting off a source of visitors.
The old block will not become a search block
Existing legacy AI blocks and previous Training “Block” or “Block on pages with ads” selections migrate to Disallow AI Training, according to the launch announcement. Separately configured Search and Agent settings stay unchanged.
Explicitly selecting the new Training “Block” stops Googlebot, Bingbot and Applebot. “Block on pages with ads” does so on ad-serving pages. Disallow is the search-preserving option, provided other controls allow crawling.
It combines robots.txt preferences for participating mixed-use crawlers with network blocks against other training crawlers. Those are different protections: one governs permitted use after access; the other refuses access. Cloudflare’s migration and control details
Google’s training opt-out is not an AI Overviews opt-out
The underlying Google control already exists. Google-Extended is a robots.txt token, not a separate crawler. Google says it controls whether crawled content can train future Gemini models and support grounding in Gemini Apps and specified Vertex AI services. Grounding means supplying source material when the model answers a request.
That makes it broader than a training-only preference, but narrower than a universal Google AI opt-out. Google says Google-Extended does not affect Search inclusion or rankings.
Google’s guidance for AI Overviews and AI Mode instead points publishers to Googlebot access rules and preview controls such as nosnippet, data-nosnippet and max-snippet. Those controls also affect how content can appear in Search; they should not be treated as interchangeable with a training preference.
Apple makes a similar separation: Applebot-Extended controls training use without stopping Applebot itself from crawling. Apple says pages opting out can still appear in search results.
Bing’s automatic opt-out is still months away
Cloudflare says Bing’s robots.txt training opt-out is targeted for early 2027. Until then, choosing Disallow AI Training does not automatically communicate that preference to Bing. Cloudflare’s Bing caveat
The Microsoft guidance linked by Cloudflare, published in 2023, describes NOARCHIVE as excluding content from future generative-model training and from Bing Chat answers, including links in those answers. It says ordinary Bing search results remain available.
There is a less obvious catch: Microsoft’s guidance says that when NOCACHE and NOARCHIVE are both present, it follows NOCACHE. That permits limited answer inclusion and training use of URLs, titles and snippets. Adding more restrictive-looking tags together does not necessarily produce a stricter result.
Publishers should also check the rest of their security configuration. Cloudflare’s firewall documentation warns that earlier rules can still block crawlers marked Allow, and those blocks may not appear in AI Crawl Control analytics. A search-preserving selection is not proof that the crawler can actually fetch a page.
Start the conversation by posting the first comment