Cloudflare Lets You Block AI Training Without Losing Google Search

Cloudflare Disallow AI Training setting blocks model crawlers while Googlebot keeps crawling

Cloudflare launched a setting on September 15 that blocks AI training crawlers while letting Googlebot, Applebot, and Bingbot keep crawling your site for search. It is called Disallow AI Training, and here is the part most site owners will miss. If you were already using Cloudflare's older Block setting, you were probably migrated to it automatically. Here is what the setting does and what to check on your own site.

The Problem This Solves

Until now the choice was uncomfortable. You could let every crawler in, including the ones training AI models on your content, or you could block them and risk losing search visibility along the way.

The reason that choice was hard is that some crawlers do both jobs. Googlebot crawls your site for search results and Google also trains models. Blocking the crawler to stop one stopped the other.

Cloudflare's new setting splits the two apart. It sits inside their Training control, alongside Search and Agent controls, and it gives you a middle option that did not exist before.

How It Actually Works

The mechanism is refreshingly simple. Disallow AI Training adds a no-training preference to your robots.txt file. It is not a proprietary Cloudflare signal. It is the standard opt-out tokens the platforms already publish.

  • For Google. It publishes a Disallow rule for Google-Extended, which is Google's robots.txt token for opting out of Gemini model training.
  • For Apple. It uses a Disallow rule for Applebot-Extended.
  • For Bing. Support is still pending. Microsoft has not added the capability yet, and Cloudflare points to Bing robots.txt support arriving in early 2027.
  • For everybody else. All other AI training crawlers are blocked outright.

Googlebot, Applebot, and Bingbot keep crawling for regular search, but only when Cloudflare labels the operator "Accountable." That label is not automatic. An operator earns it by meeting four requirements: offering robots.txt opt-out mechanisms, offering AI summary opt-outs, giving URL-level visibility into how content is used for training, and providing assurances that opting out of training does not hurt your search rankings.

The Part You Need To Check

Here is what I would actually go look at today. Most existing Cloudflare customers who were using the previous Block setting, or "Block on pages with ads," were automatically migrated to Disallow AI Training.

That means your crawler policy may have changed without you touching anything.

For most sites that migration is an improvement, because the old Block setting now behaves more aggressively than it used to. Block now completely stops mixed-use crawlers, and that includes their search functions. If you are on Block and you meant to keep search access, you are blocking more than you think.

Cloudflare's own guidance is that most customers do not need to change anything. I would still verify rather than assume. Log in, find your Training control, and confirm which setting your site is actually on. It takes two minutes.

What This Setting Does Not Do

This is where people are going to get confused, so it is worth being precise.

Disallow AI Training does not control whether you appear in AI Overviews. Those are managed separately through Google Search Console. Opting out of Gemini model training and opting out of AI Overviews are two different decisions with two different controls.

Keep in mind, those are different things. Training is about your content being used to build a model. AI Overviews are about your content being summarized in a search result today. Cloudflare has said it plans unified AI summary controls by early next year, but that is not here yet.

So if your goal is to stop losing clicks to AI summaries, this setting does not do that. If your goal is to stop your content being used as training data, it does.

Should You Turn It On

This depends on what your site is for, and I do not think there is one right answer.

  1. If you publish original content as your product. News, research, courses, anything where the writing is the asset. Disallow AI Training is a reasonable default.
  2. If your site exists to generate leads or sales. Your content is a means to an end, and being cited by an AI assistant may send you business. Blocking training buys you very little.
  3. If you are a local business. This is close to irrelevant to you. Spend the time on your Google Business Profile instead.
  4. If you are not sure. Leave it where Cloudflare migrated you and revisit when the unified AI summary controls ship.

The thing I would not do is treat this as a visibility strategy. Blocking training crawlers does not make you more visible anywhere. If you want to show up in AI answers rather than hide from them, that is a completely different job, and it starts with knowing where you currently stand. Our guide to tracking your AI search visibility walks through how to measure that.

It is also worth remembering how fast this area moves. Google has been expanding AI Overviews in search results while these controls are still being built, so whatever you decide today is worth revisiting in a few months.

Question to Answer:

Do you know which crawler setting your site is currently on, and did you choose it or did it change on its own?

In Summary

Cloudflare's Disallow AI Training setting went live on September 15 and adds standard robots.txt opt-out rules for Google-Extended and Applebot-Extended while letting Googlebot, Applebot, and Bingbot keep crawling for search. Bing support is not there yet and is pointed at early 2027. Most customers on the old Block setting were migrated automatically.

Go check which setting your site is actually on, because it may have changed without you doing anything. Then decide deliberately based on whether your content is your product or your marketing. And remember this does not touch AI Overviews, which are still controlled separately in Search Console.

Read the source

0 comments

Leave a comment