News

Blocking AI bots on Cloudflare also blocks Googlebot

Author: Reading time: 6 min 2 views
Blocking AI bots on Cloudflare also blocks Googlebot

Map of this article

Jump straight to the part you came for.

News5 sections
  1. 1One checkbox can remove you from Google

    The two paragraphs that set the date and the multi-purpose crawler behaviour.

  2. 2How many Iranian sites are behind Cloudflare at all

    This story grew large in English-language SEO media, and for the Iranian market it is much smaller than it looks.

  3. 3Our own site is behind Cloudflare

    Which is why we read this with interest.

  4. 4Another figure from the same source

    In a separate post dated 6 August 2026, Cloudflare states that fewer than half of all HTML page requests now come from a human.

From 15 September 2026, Cloudflare changes its default AI traffic settings for newly onboarded domains: training bots and agent bots are blocked by default on pages that display ads, while search bots stay allowed. The announcement on Cloudflare's blog is dated 1 July 2026.

The risk in that news is somewhere else. The same announcement says multi-purpose crawlers, the ones that both search and gather data for model training, will be blocked for customers who have chosen to block Training. Three are named explicitly: Googlebot, Applebot and Bingbot.

One checkbox can remove you from Google

The Setting new defaults section of Cloudflare's blog post, dated 15 September 2026, and the multi-purpose crawler paragraph naming Googlebot
The two paragraphs that set the date and the multi-purpose crawler behaviour. Captured on 16 August 2026.

If your site sits behind Cloudflare and somebody enabled that option to keep your content out of AI training, then by Cloudflare's own text Googlebot lands in the same category. The consequence stops being an AI question. The consequence is that your pages stop being crawled.

Check this today rather than in September.

The path is the security settings of that zone in the Cloudflare dashboard, and the same screen is where the new default can be declined. Cloudflare says the choice is available up to 15 September.

How many Iranian sites are behind Cloudflare at all

This story grew large in English-language SEO media, and for the Iranian market it is much smaller than it looks. On 16 August 2026 we read the response headers of 111 Iranian pages ranking in Google. Six were behind Cloudflare.

The most frequent named server in that sample was ArvanCloud with 17 sites, then nginx with 14, LiteSpeed with 13, BitNinja with 9, Cloudflare with 6 and IIS with 5. Another 33 responses carried no server header at all. Of the 111 pages, 75 were WordPress, and exactly 5 were both WordPress and behind Cloudflare.

So if somebody tells you this change threatens your site, first establish whether your site is behind Cloudflare. One command answers it:

  • run curl -sI https://example.com/ and look for two headers
  • if you see server: cloudflare and cf-ray, you are behind it
  • if you do not, this news has no direct bearing on you

Our own site is behind Cloudflare

Which is why we read this with interest. Our decision was straightforward: the search category stays open, and no option that sweeps Googlebot in with it gets enabled. Refusing to let your content train a model is a defensible goal. It is not worth not being crawled.

That decision does not generalise. If your content is the product, a photo archive or original writing you sell, blocking training may be worth exactly this cost. If your site exists to bring in customers, it is not.

Another figure from the same source

In a separate post dated 6 August 2026, Cloudflare states that fewer than half of all HTML page requests now come from a human. Read that carefully, because it reflects Cloudflare's view of web traffic rather than a statistical sample of the internet. The direction, though, matches what we see in server logs.

The practical consequence for organic search is that separating one kind of bot from another is not a minor setting. Search, agent and training behave differently, and one of the three is what your ranking depends on.

What is worth doing today

Check your headers so you know what you are behind. If it is Cloudflare, open that zone's settings and see which categories are blocked. If Training is blocked anywhere, understand that from September its behaviour towards multi-purpose crawlers changes, and decide whether you want that.

What is not worth doing: switching provider because of this story.

Changing the layer in front of a site carries downtime risk and configuration work, and for something a checkbox resolves that is a bad trade. If you want a real reason to review that layer, speed and caching is a better one.

Does this affect existing sites?

The announcement concerns domains newly onboarded to Cloudflare. For existing domains your current setting stands, which is precisely why checking matters: you may have enabled something months ago yourself.

Does blocking training hurt SEO?

Blocking training is not itself a ranking signal. The harm comes from Googlebot falling into the multi-purpose category and being blocked by the same setting. The problem is the setting, not the policy.

What if I am not behind Cloudflare?

There is no action here for you. But the same three categories are beginning to appear at other providers, so asking yours about it is not a wasted question.

How do I confirm Google is still crawling?

The crawl report in Search Console, plus the last crawl date on individual pages. A sudden crawl drop is the first thing visible after a wrong setting in the front layer, and our SEO guide shows where that number sits in the work.

If your site is behind no layer at all and its security and speed are in your own hands, this is more a reminder than a warning: deciding which bots reach your site is part of site security and belongs in a record somebody can review, not in a forgotten checkbox. It is also one of the things that should be settled from week one when a site gets built.

Method: the Cloudflare announcement was read on 16 August 2026 and the quotations come from that page. The server counts come from the response headers of 111 Iranian pages read the same day from a server in Helsinki; a site that strips its server header does not appear in this count, so the figures are a floor rather than a ceiling.

The default change is announced in Cloudflare's 1 July blog post, and the share of non-human requests appears in its 6 August post.

Hossein Parto

IT engineer and SEO specialist with over 12 years of experience, certified by MOZ, Semrush, and Ahrefs Academy. Founder of RGB.ir, where up-to-date web knowledge is published in plain, actionable language.

Want us to put this knowledge to work for your business?

The RGB team professionally handles everything you just read about, for your own site. Start with a free consultation.

Comments & Questions

Have a question about this article? Ask, we'll answer.

No comments yet; be the first.

Write Your Comment

Your email won't be published. Comments are shown after review.