Cloudflare has introduced a new feature called Disallow AI Training, which will apply to websites that already prevent AI training. This feature includes a preference in robots.txt to block training while still permitting Googlebot, Applebot, and Bingbot for search crawling.
The current plan is different from the one mentioned in July, as now selecting the Block option will completely prevent Googlebot, Applebot, and Bingbot from accessing a site, including for search purposes.
What was different on September 15?
Cloudflare’s Training control now includes a feature called Disallow AI Training, introduced in August. This feature allows mixed-use crawlers to continue crawling for search if labeled as “Accountable,” while blocking other training crawlers.
Blocking ads on pages now also applies to mixed-use crawlers, as the company aims to avoid impacting search visibility.
Most customers are not required to make any changes as their existing training preferences will automatically be updated to Disallow AI Training. Sites that previously used the Block or Block on pages with ads options under the Block AI Bots toggle will now have different settings.
Block AI Bots and Cloudflare’s Managed Robots.txt feature are planned to be discontinued. To prevent mixed-use crawlers from accessing a site completely, users must now opt for the Block option.
New domains that use advertisements to make money are provided with the option to block AI Training in their training settings.
Which crawlers maintain search accessibility?
Crawlers that have mixed-use must adhere to Disallow AI Training only if Cloudflare identifies them as Accountable, a status established following discussions with crawler operators since July. Operators must fulfill or agree to fulfill four criteria to be eligible.
An option to decline AI training using robots.txt or a comparable standard.
An option to decline AI summaries can be activated now and will be available through Cloudflare next year.
Visibility at the URL level shows which pages were included for training, along with data on how the content was displayed in search results.
Opting out of training will not impact standard search results.
The company states that Apple, Google, and Microsoft satisfy the criteria, with current features and scheduled commitments for future ones.
It also identifies Amazon, Anthropic, Meta, and OpenAI as accountable due to their separate search and training crawlers, with the training crawlers being restricted under Disallow AI Training.
The Scope of Coverage at Google, Apple, and Bing
Google’s Disallow AI Training is implemented using a Disallow rule for Google-Extended, which is the robots.txt token provided by Google for excluding content from Gemini model training. Google’s crawler documentation clarifies that Google-Extended does not impact a website’s presence in Google Search or its ranking.
A distinct setting in Search Console determines if a website is featured in AI Overviews, AI Mode, and Discover’s generative AI functions, according to Google’s support page, this setting has no impact on AI training.
Apple uses a Disallow rule for Applebot-Extended to prevent it from crawling pages, as stated in Apple’s documentation. This helps to exclude content from AI-generated answers in Siri and Search by using the nosnippet meta tag.
Disallowing AI Training on Bing through robots.txt is not currently effective, as Microsoft has yet to implement support for this preference. Cloudflare has indicated that this functionality is still being developed. The previous Training block on Cloudflare did not impact Bingbot.
Bing’s current method for opting out of training is using the NOARCHIVE meta tag. According to Bing’s guidelines, content labeled with NOARCHIVE is excluded from training Microsoft’s AI models and is not referenced in Chat and Copilot.
Why This is Important
Choosing Block instead of Disallow AI Training on Cloudflare not only prevents AI training crawlers but also hinders Google, Apple, and Bing from crawling your website.
AI Training maps cannot be blocked using the training opt-outs offered by Google and Apple. The decision on whether your pages will be shown in AI Overviews and AI Mode is controlled by Google through Search Console. Bing still uses the NOARCHIVE tag for training opt-outs, which also prevents links from Copilot from being included.
Peering into the future
Google is set to launch URL-level transparency tools for Google-Extended soon, according to Cloudflare. Apple is also developing its own URL-level tool for release next year. Cloudflare reports that Microsoft aims to support a robots.txt no-training preference by early 2027.
Cloudflare is shifting its attention to AI summaries, aiming to streamline content control through a centralized setting on its platform by early next year.
Featured Image by Remo_Designer/Shutterstock