Let verified Google crawlers through firewalls, CDNs and bot protection
- Why
Google does not know that a firewall or bot-protection rule is blocking it; it only sees failing requests. Timeouts, connection resets and DNS errors are treated like 5xx errors: crawling slows down at once, and indexed URLs that stay unreachable drop out of the index within days.
- How
Allowlist Google's crawlers in WAF, CDN and bot-management rules by the IP ranges Google publishes for its common and special-case crawlers (common-crawlers.json, special-crawlers.json), or by reverse DNS to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com hosts; never by user-agent string alone (it is trivial to fake), and never by a whole googleusercontent.com host, where Google Cloud customers' machines resolve too. Do not rate-limit or challenge verified Googlebot with CAPTCHAs or JavaScript puzzles, which it cannot solve. Make sure DNS resolves reliably for every host that serves pages, scripts, styles or APIs. Review the rules after every CDN or security-vendor change, including rules written for HTTP/3 traffic or AI scrapers. When Search Console reports DNS or network errors, first ask the hosting provider, CDN or DNS provider what changed and check for recently added firewall rules: said at Search Central Live, most network errors happen between the origin and Google's data centres, where neither Google nor the site owner can see them.
- Test
Search Console Crawl stats report: host status shows no robots.txt fetch, DNS resolution or server connectivity failures, and no Search Console alert for DNS errors, which can get a site removed from Google Search very aggressively and, with it, from every feature that depends on Search. URL Inspection live test succeeds for one URL per template. WAF/CDN logs show no 403, 429 or challenge responses to IPs that verify as Googlebot.
Code · Verify that a request really comes from Googlebot
Before a firewall, CDN or bot-protection rule blocks or challenges a "Googlebot" request, verify it: the user-agent string alone can be faked. In WAF rules, prefer matching the IP against the ranges Google publishes as common-crawlers.json and special-crawlers.json on its page on verifying its crawlers; for log analysis, a reverse DNS lookup followed by a forward lookup works too. Never allowlist a whole googleusercontent.com host: Google Cloud customer VMs resolve there too, and only *.gae.googleusercontent.com belongs to Google's user-triggered fetchers, which are not Googlebot.
# 1. Reverse DNS: Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com
# or geo-crawl-*.geo.googlebot.com; special-case crawlers to rate-limited-proxy-*.google.com.
# A host such as 81.59.117.34.bc.googleusercontent.com is a Google Cloud customer, not Googlebot.
host 66.249.66.1
# -> 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
# 2. Forward DNS: that host name must resolve back to the same IP
host crawl-66-249-66-1.googlebot.com
# -> crawl-66-249-66-1.googlebot.com has address 66.249.66.1The published IP ranges are easier to keep in sync with WAF and CDN allowlists than DNS lookups:
# Common crawlers (Googlebot and others) and special-case crawlers; refresh the lists regularly
curl -s https://developers.google.com/static/crawling/ipranges/common-crawlers.json
curl -s https://developers.google.com/static/crawling/ipranges/special-crawlers.jsonEvidence · 11 claims · 4 Google pages
DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
Google at Search Central Live Deep Dive Europe 2026 (Day 1)
Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.
Google
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
Google at Search Central Live Deep Dive Europe 2026 (Day 2)
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
Google at Search Central Live Deep Dive Europe 2026 (Day 2)
Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.
Google
DNS errors get a site removed from Google Search very aggressively, and Search Console alerts site owners to them.
Google at Search Central Live Deep Dive Europe 2026 (Day 1)
Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.
Google at Search Central Live Deep Dive Europe 2026 (Day 1)
When Search Console reports network errors, the first step Google recommends is to ask the hosting provider, or the CDN, whether they changed anything.
Google at Search Central Live Deep Dive Europe 2026 (Day 1)
Network timeouts are usually caused close to the site, very often by a firewall, a CDN, the hosting provider or the DNS provider, and Google has no visibility into them.
Google at Search Central Live Deep Dive Europe 2026 (Day 1)
Network errors and timeouts, like DNS errors, can get a site removed from Google Search and, with it, from every feature that depends on Search.
Google at Search Central Live Deep Dive Europe 2026 (Day 1)
For DNS and network errors caused by a firewall or CDN, Google advises checking whether new firewall rules were set recently and otherwise asking in the CDN's forum, as Google itself cannot see or help with these errors.
Google at Search Central Live Deep Dive Europe 2026 (Day 1)