Technical SEO

Is Cloudflare blocking Googlebot? Bot Fight Mode, AI-training blocks and how to check

Three Cloudflare settings can stop Googlebot. Here is how to tell which one is doing it, with curl, Security Events and URL Inspection, and what to change.

By SEOFix TeamPublished
2,431 words · 11 min readMarkdown
Is Cloudflare blocking Googlebot? Bot Fight Mode, AI-training blocks and how to check

Yes, Cloudflare can block Googlebot, and in 2026 there are three common ways it happens: Bot Fight Mode on the Free plan, the AI Training control set to Block (which since 15 September 2026 also stops Googlebot, Bingbot and Applebot), and custom WAF or rate limiting rules you wrote yourself. You can find out which one in about five minutes with Google's URL Inspection live test and Cloudflare's Security Events. Both are described below, together with the fix for each cause.

The stakes are higher than "a few pages not crawled". A Cloudflare challenge answers with HTTP 403, and Google's documentation says that for a 4xx response other than 429, "the indexing pipeline removes the URL from the index if it was previously indexed". A challenge on a whole section can drop that section out of Google.

What changed in 2026

If your site was fine last year and isn't now, this timeline is probably why. The dates come from Cloudflare's own posts, plus one news report.

DateChangeSource
1 July 2026The single "Block AI bots" toggle is replaced by three behaviours, Search, Agent and Training, each set to Allow, Block, or Block on pages with ads. Available on all plans, Free included.Cloudflare changelog, blog
1 July 2026Cloudflare announces that multi-purpose crawlers "such as Googlebot, Applebot, and BingBot" will be blocked for customers who block Training.blog
4 August 2026A Reddit user reports that with AI Training set to Block, Googlebot and Bingbot got 403 on their sitemap.Search Engine Journal
15 September 2026New Disallow AI Training setting. Block and Block on pages with ads now apply to mixed-use crawlers, so Block stops Googlebot outright. Existing Block settings were migrated to Disallow AI Training. The legacy "Block AI bots" setting is deprecated.Cloudflare blog, docs, SEJ

Two consequences follow. First, if you (or a teammate) choose Block under Training after 15 September, Googlebot is blocked by design. Second, the migration moved existing Block settings to Disallow AI Training, so if Googlebot is still blocked on an older zone, check whether someone switched it back to Block afterwards.

How to check if Cloudflare is blocking Googlebot

1. Ask Google: URL Inspection live test

This is the only test that uses real Googlebot from Google's real IP addresses, so it's the one that settles the question.

  1. In Search Console, paste an affected URL into the inspection bar at the top.
  2. Choose Test live URL.
  3. Open View tested page to see the HTTP response, the headers and the raw HTML Google received. Google's help page says it shows "the raw HTML returned, the HTTP headers, JavaScript console output, and any page resources loaded".

If Cloudflare challenged Googlebot, the live test reports that the page couldn't be fetched, and the HTTP response under View tested page shows the 403. In the Page indexing report the same problem shows up as Blocked due to access forbidden (403), which Google describes bluntly: "Googlebot never provides credentials, so your server is returning this error incorrectly".

For a site-wide view, open Settings → Crawl stats and look at the By response breakdown. A sudden rise in "Other 4XX" responses, starting on the day someone changed a Cloudflare setting, is the pattern to look for (Crawl Stats report).

2. Ask Cloudflare: Security Events

Cloudflare logs every request a security feature acted on.

  1. In the Cloudflare dashboard, open your zone, then Security → Analytics and the Events tab. (Cloudflare docs describe it as the zone's Analytics page, Events tab.)
  2. Add a filter: User agent contains Googlebot.
  3. Expand a few events in Sampled logs. The Service field tells you which product acted (Bot Fight Mode, Super Bot Fight Mode, Custom rules, Rate limiting rules and so on), and the Action field tells you whether it was a block or a challenge.

Outside Enterprise you can query about the last 30 days of events (Cloudflare), so check soon after a ranking drop.

Not every request with "Googlebot" in the user agent is Google. Scrapers copy it all the time. Before you change anything, check the source IP of a blocked event. Google's documented method is a reverse DNS lookup that should end in googlebot.com, google.com or googleusercontent.com, followed by a forward lookup that returns the same IP (Google):

host 66.249.66.1
# 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
host crawl-66-249-66-1.googlebot.com
# crawl-66-249-66-1.googlebot.com has address 66.249.66.1

Google also publishes its crawler ranges as JSON at https://developers.google.com/static/crawling/ipranges/common-crawlers.json. If the blocked "Googlebot" requests come from a hosting provider's IP, Cloudflare is doing its job and you can stop here.

3. Ask the server: curl

curl can't impersonate real Googlebot (Cloudflare identifies verified bots by IP and other signals, not by the user agent string), but it shows quickly which paths are challenged at all. Fetch a page from each section of the site and look at the status and the cf-mitigated header:

for path in / /blog/ /jobs/ /sitemap.xml; do
  printf '%s ' "$path"
  curl -s -o /dev/null -D - \
    -A 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)' \
    "https://example.com$path" | grep -i -E '^(HTTP|cf-mitigated)' | tr '\n' ' '
  echo
done

Cloudflare documents that a challenge response carries cf-mitigated: challenge, and that challenge is the header's only valid value (Cloudflare).

Here is what that looks like on jobxdubai.com, a 37,000-page job site we use to test SEOFix, where Cloudflare challenges the /jobs section for anything that isn't a real browser (output from 8 October 2026):

$ curl -s -o /dev/null -D - -A 'Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)' https://jobxdubai.com/jobs
HTTP/2 403
content-type: text/html; charset=UTF-8
cf-mitigated: challenge
server: cloudflare

The body of that response is a page titled Just a moment... with <meta name="robots" content="noindex,nofollow">. A spoofed Googlebot from a laptop being challenged is expected. What matters is whether the real one gets through, and only steps 1 and 2 can tell you that. The curl result is still useful: if /sitemap.xml or a whole section answers 403 with cf-mitigated, you know where to look.

Cause 1: Bot Fight Mode on the Free plan

Bot Fight Mode issues "computationally expensive challenges" to traffic that matches known bot patterns. Cloudflare says verified bots "have been excluded in default bot configurations across all plans", so in theory Googlebot is not affected. In practice there are reports where it was: one site owner, writing on theinfinity.dev on 13 August 2026, traced a drop in daily impressions from 1,153 to 16 and an average position from about 10 to 68 to Bot Fight Mode. That is one site's account, not a measured rate, but the failure mode is real enough to check.

The bigger problem is that you can't make exceptions. Cloudflare's docs say Bot Fight Mode "does not run on the Ruleset Engine", so Skip, Bypass and Allow rules have no effect on it. An IP Access rule that matches first is the only thing that stops it from triggering. You can't allowlist Googlebot, your uptime monitor or your SEO audit tool around it.

Where it lives: Security → Settings, filter by Bot traffic, then Bot fight mode (Cloudflare).

Fix:

(starts_with(http.request.uri.path, "/jobs/") and not cf.client.bot)

Action: Managed Challenge. That is close to the fix the theinfinity.dev author describes.

Cause 2: the AI Training "Block" setting

Since 15 September 2026, the Training control has these options (Cloudflare blog):

Training settingGooglebot, Bingbot, ApplebotDedicated training crawlers (GPTBot, ClaudeBot…)
AllowCrawl normallyAllowed
Disallow AI TrainingKeep crawling for search; a no-training preference is published in your robots.txtBlocked at Cloudflare's edge
Block on pages with adsBlocked on pages with adsBlocked on pages with ads
BlockBlocked everywhere, search includedBlocked

Cloudflare's own words: "Block and 'Block on pages with ads' apply to all training crawlers, including mixed-use crawlers". Search Engine Journal's summary is plainer: selecting Block "will completely prevent Googlebot, Applebot, and Bingbot from crawling".

Where it lives: Security → Settings → Configure AI bot policies (Cloudflare docs).

Fix: if you want to stay in Google and opt out of AI training, choose Disallow AI Training, not Block. Training-only crawlers are still blocked under it: Cloudflare says "Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI". For Googlebot, Bingbot and Applebot it relies on the robots.txt tokens the vendors support. Google honours Google-Extended and Apple honours Applebot-Extended. Microsoft doesn't yet honour a domain-level no-training preference and, per Cloudflare, is building it "targeted for early 2027".

Watch the default on new zones too. Since 15 September, new domains with ads get Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agents (Cloudflare). Search stays allowed, but check the settings once after you add a zone.

A related setting, managed robots.txt, doesn't block anyone. It prepends Disallow: / groups for crawlers such as GPTBot, ClaudeBot, Google-Extended and CCBot, plus a Content-signal: search=yes, ai-train=no, use=reference line, to your existing robots.txt (Cloudflare). It won't hurt Googlebot, but open https://yoursite.com/robots.txt afterwards so you know what it says. The companion article on what to block and allow for GPTBot, OAI-SearchBot and ClaudeBot covers those robots.txt choices.

Cause 3: your own rules

If Security Events point to Custom rules or Rate limiting rules, the cause is a rule on your zone. Common culprits:

  • A rule that challenges a path (/jobs/, /search, /api/) for everyone who isn't a browser, with no exception for verified bots.
  • A country block or a rule with Security Level set high, applied to all traffic.
  • A rate limit tuned for human browsing, which a crawler working through a large section can exceed.

Custom rules live under Security → Security rules (older dashboards: Security → WAF → Custom rules) (Cloudflare). Add and not cf.client.bot to rules that should never apply to search crawlers. For rate limits, a 429 is the gentler answer: Google treats 429 as a signal to "temporarily slow down with crawling" rather than as a missing page.

Cloudflare publishes the list of verified bots it recognises on Cloudflare Radar. If a crawler you care about isn't on it, cf.client.bot won't match it.

Letting SEO audit tools in without opening the door to scrapers

Your own audit tool hits the same rules. An audit that suddenly reports thousands of 403 pages is one way to notice the problem before Search Console does.

Don't allowlist a user agent. Anyone can send AhrefsBot or SEOFixBot in a header. A rule that matches a secret request header is safer, because only the tool you gave it to can send it. With SEOFix, the rule is a Skip rule placed first:

any(http.request.headers["x-seofix-verify"][*] eq "<your crawler secret>")

with the action Skip and these ticked: all remaining custom rules, all managed rules, all Super Bot Fight Mode rules (Pro and above), Browser Integrity Check, Security Level, User Agent Blocking, Hotlink Protection and Zone Lockdown. Rate limiting rules stay in place, and the crawler backs off on 429. The full steps are in Let SEOFix through your firewall, and Connect Cloudflare can create the rule with a scoped API token.

The same caveat applies here as for Googlebot: on the Free plan, Bot Fight Mode ignores Skip rules. If it's on, no allowlist will get an audit tool through, and you'll have to switch it off while the audit runs.

Make sure it doesn't happen silently again

A block can run for days before anyone notices. In the theinfinity.dev case, Googlebot was locked out for more than five days. Three habits catch it sooner:

  1. Change log. Write down when anyone changes Bot Fight Mode, AI bot policies or custom rules. Every cause above starts with a settings change.
  2. Crawl stats after changes. A week after any Cloudflare security change, check Search Console's Crawl stats for a rise in 4xx responses.
  3. A scheduled audit that knows what a firewall looks like. A crawler that doesn't recognise challenges reports a challenged page as a broken 403, which is wrong and noisy. SEOFix records it as Blocked by firewall, leaves it out of the health score, and reports it per section (first path segment), with the provider and sample URLs. When more than 20% of an audit is blocked, the report says so at the top. It detects Cloudflare from the cf-mitigated header or a Server: cloudflare challenge page, including the 1015 and 1020 error pages (Blocked pages in your report).

On jobxdubai.com that difference is the one between "30,000 broken job pages" and "the /jobs section is behind a challenge". Once 50 pages in a section have all been challenged, SEOFix stops fetching that section, so the audit doesn't send 30,000 requests at a firewall that has already said no.

FAQ

Does Cloudflare block Googlebot by default?

No. Verified bots have historically been excluded from Cloudflare's default bot configurations (Cloudflare), and Search stays allowed in the 15 September 2026 defaults for new domains. Googlebot gets blocked when someone chooses Block under Training, turns on a feature like Bot Fight Mode that misfires, or writes a rule without a verified-bot exception.

Will "Disallow AI Training" hurt my Google rankings?

According to Google, the Google-Extended token it relies on "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". It controls Gemini training and grounding, not Search.

Can I allowlist Googlebot by user agent?

You can, but it also lets in every scraper that copies the string. On Pro and above use the verified bot exception in Super Bot Fight Mode or a custom rule with cf.client.bot. On Free, avoid Bot Fight Mode and use custom rules with not cf.client.bot.

My sitemap returns 403 to Googlebot. Is that the same problem?

Usually yes. The 4 August 2026 report that started this news cycle was exactly that: sitemap requests answered with 403 after AI Training was set to Block (SEJ). Run the curl loop above on /sitemap.xml and check Security Events for that path.


If you'd like to see which sections of your site a firewall is challenging, run a free audit. SEOFix reports firewall-blocked sections instead of calling them broken pages, and the first 500 pages are free.

ShareXLinkedIn
Is Cloudflare Blocking Googlebot? How to Check (2026) · SEOFix