AI SEO Tools Directory
← Blog
Recorded

Cloudflare Blocks Mixed-Use AI Crawlers on September 15: The Pre-Deadline Site Checklist

Cloudflare blocks Training and Agent crawlers by default on ad pages starting September 15, 2026. An eight-step checklist to check your site's settings first.

Bottom line

Starting September 15, 2026, Cloudflare blocks Training and Agent crawlers by default on any page carrying ads, for every domain onboarding after that date. Search crawlers stay allowed. GPTBot and ClaudeBot, the two largest single-purpose AI crawlers, lose default access on those pages unless a site owner opts in before the deadline.

Last updated September 2026.

GPTBot and ClaudeBot are not minor bots anymore. Cloudflare’s own crawler analysis, published in August 2025, put GPTBot at 28.1% of AI-only crawler traffic in July 2025, up from 11.9% a year earlier. ClaudeBot held 23.3% of that same traffic, up from 15%. Both bots run as single-purpose Training crawlers, so both lose default access to ad-monetized pages the moment the new rule takes effect.

What changes on September 15

Cloudflare confirmed the new defaults in its own policy update, Your site, your rules: new AI traffic options for all customers, published July 1, 2026. The deadline itself sits in the table below.

DetailWhat Cloudflare confirmed
Effective dateSeptember 15, 2026
Who is affected firstDomains onboarding to Cloudflare on or after that date
Blocked by defaultTraining and Agent crawlers, on any page that carries ads
Still allowed by defaultSearch crawlers
Mixed-use bots affectedGooglebot, Applebot, and Bing, once a site already blocks Training
Opt-out windowAny time before September 15, 2026, in Security settings

An ad on a page is the trigger, not the page’s topic. Cloudflare treats an ad as a signal that a human, not a bot, was meant to land there. On pages without ads, none of this changes.

Which crawlers count as mixed-use

Cloudflare sorts every bot into three use cases: Search, Agent, and Training. A bot that runs more than one of those at once, most often Search paired with Training, is what the company calls mixed-use. Cloudflare names Googlebot, Applebot, and Bing directly as its clearest examples, and it enforces the most restrictive rule across all of a mixed-use bot’s behavior.

GPTBot and ClaudeBot are not part of that mixed-use group. Both run Training only, so both already sit inside the category Cloudflare blocks by default on ad-monetized pages, mixed-use label or not.

CrawlerOperatorClassificationEffect on September 15
GPTBotOpenAITrainingBlocked by default on pages that carry ads
ClaudeBotAnthropicTrainingBlocked by default on pages that carry ads
ChatGPT-UserOpenAIAgentBlocked by default on pages that carry ads
GooglebotGoogleSearch plus Training (mixed-use)Loses search crawling too, once the site blocks Training
ApplebotAppleSearch plus Training (mixed-use)Loses search crawling too, once the site blocks Training
BingbotMicrosoftSearch plus Training (mixed-use)Loses search crawling too, once the site blocks Training

The pre-deadline checklist

  1. Confirm whether your domain counts as new. The automatic block on Training and Agent crawlers applies to domains onboarding to Cloudflare on or after September 15, 2026, not to every existing zone by default. Verify it: check the “added to Cloudflare” date on your zone overview page.

  2. Open your AI Bots controls. Find the three toggles under Security > Settings: Search, Agent, and Training. Verify it: record the current state of each toggle with today’s date, so you have a snapshot from before the deadline.

  3. Check your existing Block Training status. If you already block Training, through either the new controls or the legacy Block AI Bots service, mixed-use crawlers such as Googlebot lose their search access too. Verify it: read the classification tags Cloudflare lists for each bot in BotBase.

  4. Decide whether to opt out of the mixed-use change. A site that wants Googlebot, Applebot, and Bing to keep crawling for search, even while blocking Training, has to set that preference before September 15. Verify it: look for the explicit opt-out control next to the AI Bots settings.

  5. Map which pages carry ads. The new default block only applies to ad-monetized pages, so content-only sections of a site are not affected the same way. Verify it: cross-check your ad-tag placements against your sitemap.

  6. Read your robots.txt for Content Signals. Cloudflare-managed robots.txt files now carry search, ai-train, and use fields that state a preference for how a crawler may use your content. Verify it: fetch yourdomain.com/robots.txt and confirm the Content-Signal line matches what you intend.

  7. Pull 30 days of crawler visits by user agent. Isolate GPTBot and ClaudeBot specifically, since they are today’s two largest single-purpose Training crawlers and the most likely to disappear from your logs first. Verify it: export bot traffic from your CDN or from an AI SEO platform’s bot analytics module.

  8. Schedule a post-deadline recount. Set a reminder for the week of September 15 to compare crawler visits before and after the cutover. Verify it: rerun the same 30-day pull and confirm the bots you expected to lose access actually stopped requesting pages.

Confirm the crawl stopped

Reading raw server logs bot by bot is slow. A handful of AI SEO platforms fold crawler evidence into the same dashboard as citation tracking, so a before-and-after comparison takes minutes instead of a log-parsing script.

Temso includes AI bot visit tracking on every plan, from 10,000 monthly visits on Starter to 100 million on Professional, next to citation and content data in the same account. Profound ties its Agent Analytics feature to GA4 and its citation maps, so a drop in GPTBot visits lines up against the URLs that used to get cited. BrightEdge keeps AI-search visibility inside the same enterprise workspace its customers already use for technical crawl reporting. Scrunch’s Agent Traffic module reads crawler visits at the CDN layer, broken out by bot, frequency, and page. Full profiles and Index scores for all four sit in the AI SEO tools index.

Check your Cloudflare Security settings against this list before September 15, 2026, then rerun your crawler count the week after to confirm the block worked. If you want that verification next to your citation and content data instead of a separate export, Temso’s bot traffic tracking is included on every plan, starting at $89 a month.

FAQ

What counts as a mixed-use AI crawler on Cloudflare?

A mixed-use crawler runs more than one purpose in a single bot, most often Search combined with Training. Cloudflare names Googlebot, Applebot, and Bing as its clearest examples. Starting September 15, 2026, these bots get judged by every behavior they run, so a site that already blocks Training loses their search crawling too.

Does the September 15 default apply to every website on Cloudflare?

No. The automatic block on Training and Agent crawlers, on pages that carry ads, applies to domains onboarding to Cloudflare on or after September 15, 2026. Existing domains keep their current settings unless they already block Training, in which case the mixed-use crawler change reaches them too.

Will GPTBot and ClaudeBot lose access to my site?

On pages that carry ads, yes, by default, since both run as single-purpose Training crawlers. GPTBot held 28.1% of AI-only crawler traffic in July 2025, according to Cloudflare's own data, and ClaudeBot held 23.3%. A site owner who wants either bot to keep crawling has to allow Training explicitly before the deadline.

Can a site owner opt out of the new defaults?

Yes. Cloudflare lets any site owner set a preference in Security settings, any time before September 15, 2026, to keep the current behavior for mixed-use crawlers. The opt-out has to be set before the deadline; it does not apply after the fact.

How can a site confirm a crawler actually stopped visiting?

Pull crawler visits by user agent for the 30 days before September 15, then run the same pull again the week after. A gap between the two counts for one bot confirms the block took effect. Several AI SEO platforms, including Temso, Profound, and Scrunch, report this data next to citation tracking in one dashboard.

Does blocking a Training crawler remove a brand from AI answers already generated?

No. A blocked Training crawler stops collecting new pages going forward; it does not erase what a model already learned from older training data. A brand's newer content simply stops feeding future model updates until the site owner allows Training again.