Cloudflare puts your robots.txt on autopilot

Publishers can separate AI training from search access, but the available controls differ across OpenAI, Google and Bing.

Cloudflare’s Bot Preference Sync updates a website’s robots.txt from the AI crawler settings in its dashboard. It follows broad categories such as search and training, so publishers with exceptions for individual companies still need to manage those separately. Cloudflare says the sync does not read custom firewall rules.

Cloudflare announced the feature on August 21 for every plan, including Free, with sync enabled by default for new customers. Slobodan Manic examined its limitations in a No Hacks article republished by Search Engine Journal on September 18. The current policy includes a further Cloudflare update published on September 15.

What Cloudflare adds to robots.txt

Cloudflare places its generated instructions above the site’s existing file, inside a marked block, and preserves the original contents below. It periodically updates the generated bot list. A publisher that needs a company-specific exception is directed to turn off sync and maintain its own file.

That file tells cooperating crawlers what the publisher permits. It cannot technically stop a crawler that ignores it. Cloudflare’s robots.txt documentation makes that distinction explicit: refusing a request requires an enforcement control such as AI Crawl Control.

Under Cloudflare’s September 15 update, Disallow AI Training publishes a training preference while allowing qualifying mixed-use crawlers to continue fetching pages for search. Other training crawlers are blocked. Selecting the new Training Block setting also stops Googlebot, Bingbot and Applebot; Block on pages with ads does so on affected pages.

Cloudflare says existing Training blocks migrate to Disallow AI Training, revising the plan described in SEW’s earlier report. Other search restrictions still apply.

ChatGPT search does not require GPTBot access

A publisher deciding which AI companies to admit should check what each crawler does. OpenAI uses OAI-SearchBot for ChatGPT search and GPTBot for content that may be used in model training. Its crawler documentation explicitly permits allowing the former while disallowing the latter.

A site can therefore remain eligible for ChatGPT search without granting GPTBot training access. Allowing search crawling does not guarantee inclusion, citations or referral traffic, but training permission is not a prerequisite in these documented controls.

Google and Bing use different opt-outs

Google’s Google-Extended control governs specified Gemini training and grounding uses. Grounding supplies source material when a model generates an answer. Google-Extended is a robots.txt token, with no separate HTTP crawler, and Google says it does not affect Search inclusion or rankings.

That scope matters for publishers concerned about AI answers replacing visits. Google’s AI Search guidance addresses AI Overviews and AI Mode through Googlebot access and controls such as nosnippet, data-nosnippet, max-snippet and noindex. A Google-Extended training restriction should not be read as an instruction to remove a page from those search features.

Bing has an additional gap. Cloudflare says Microsoft’s robots.txt training opt-out is targeted for early 2027. Until it arrives, Disallow AI Training does not automatically communicate that preference to Bing.

The Microsoft guidance Cloudflare points to offers NOARCHIVE instead. It excludes content from future generative-model training and Bing Chat answers, including links in those answers, while retaining ordinary search availability. The trade-off is wider than training alone. Microsoft also says that if both NOCACHE and NOARCHIVE are present, it follows NOCACHE, which still permits URLs, titles and snippets to be used in training.

A company-wide allowlist can grant more than search access

Cloudflare’s categories can express a useful publisher policy: permit search and refuse training. OpenAI’s separate controls show why a company-wide allowlist can be unnecessarily broad. A publisher seeking ChatGPT referrals has a documented route that does not require admitting its training bot.

A specific training agreement creates a different requirement. Allowing one operator to train while excluding others needs an exception within the Training category, which the sync does not copy from custom rules. Before taking on manual file maintenance, establish whether the exception actually requires training access or only access for that operator’s search crawler.

Want SEW higher in your Google results?Add as a preferred source

More in Development

View more

Start the conversation by posting the first comment

Join the conversation

Posting publicly · your email is never shown