Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

Crawl errors

Google slows crawling when it sees 5xx or 429 responses, network timeouts, connection resets or DNS errors, and drops indexed URLs that stay unreachable. It cannot tell that a firewall or CDN rule is the cause, so a misconfigured block looks like a failing server. Other 4xx codes do not change the crawl rate. Day 2 added a second failure mode for bot protection: when a CDN or other bot protection shows Googlebot the same challenge page with a 200 status on many URLs, Google struggles to recognise it as an error and can cluster those pages as duplicates; Google's CDN post says recovery can be slow and recommends a 503 status for bot-verification interstitials. Google also said AI agents browsing for users hit the same bot walls and may go to another site (said at the event, not in Google's docs). Day 3 timed the effect: when a site starts serving 500 errors, Google lowers its crawl capacity within about four hours on average, and recovery takes weeks; Google's crawl rate guide says crawling picks up again automatically once the errors drop and warns against using error codes to slow crawling for longer than 1-2 days. The second recording of Day 1 filled in the status codes. Cherry Prommawin walked through the five classes: a 200 makes the returned resource eligible for indexing (although, as Gary Illyes put it, a 200 only means the server believes it did what was asked), a 204 has nothing to index, Google follows permanent and temporary redirects and indexes what the target returns, 404, 410 and 403 leave nothing to index, and 429 and 5xx make Google slow down so as not to break the site. Gary Illyes called DNS and network errors very common: DNS errors get a site removed from Search very aggressively, network errors and timeouts can also remove it from Search and with it from every feature that depends on Search, and most happen between the origin server and Google's data centers, often at a firewall, CDN, host or DNS provider, where neither side can see them, so the first step is to ask the host or CDN what changed; a rise in HTTP errors can also come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path (Google's CDN post lists both codes as CDN blocks). He added that Google now sees more 403 responses (described on stage as 'authentication required') and, more recently still, more 402 Payment Required responses, and treats both like a 404, and that CDN captcha challenges served with a 200 end up classified as soft 404s; Day 2 said such pages can instead be clustered as duplicates, and Google's CDN post describes both outcomes. Google pointed to the Crawl Stats report (Settings, Crawl stats) and server logs as the main tools for debugging crawl errors, and to the reasons in the Page indexing report for spotting patterns, which is also where soft 404s are looked up. In Lightning session B Dave Smart showed a quieter failure: robots.txt is checked for every URL in a redirect chain, so a redirect through a disallowed URL, an external authorisation service or a step served only to Googlebot stops the crawl and the page is reported as blocked (not in Google's docs). Author’s view: do not answer verified Googlebot with 401, 402 or 403 from a login wall or paywall; and while 4xx responses do not slow crawling, a 404 is still a fetch that counts towards crawl budget.

What to do

  • Serve bot challenges and temporary blocks to crawlers with a 503 (or 429), never as a 200 page; a 200 challenge page on many URLs can get them clustered as duplicates.
  • Ask whoever runs the CDN or firewall to confirm Google's crawler IP ranges are not blocked.
  • Check Crawl stats, Host status, in Search Console for DNS, robots.txt and connectivity failures.
  • Check what the bot protection serves to verified Googlebot and to the AI agents you want to allow.
  • When Search Console reports DNS or network errors, first ask the host, DNS provider or CDN whether anything, such as a firewall rule, changed recently.
  • Debug crawl errors with the Crawl Stats report (Settings, Crawl stats) and your server logs.
  • Do not answer verified Googlebot with 401, 402 or 403 from a login wall or paywall; Google treats them like 404 and drops the pages.

Day 1: Crawling 45

Said on stage 33

StageConsistent with docsD1-C069

DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence notes, transcript

Used byrequirements DEV-MON-04, DEV-SRV-01

  • Extended by D1-C359 Day 1: Most network errors happen somewhere between the site's origin server and Google's data centers, where they…
  • Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
StageConfirmed by docsD1-C354

Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.

Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Things

Used byrequirement DEV-SRV-03

  • Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
StageNot in docsD1-C359

Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Used byrequirement DEV-SRV-01

  • Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
StageConsistent with docsD1-C361

Network timeouts are usually caused close to the site, very often by a firewall, a CDN, the hosting provider or the DNS provider, and Google has no visibility into them.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Things

Used byrequirement DEV-SRV-01

StageConsistent with docsD1-C363

For DNS and network errors caused by a firewall or CDN, Google advises checking whether new firewall rules were set recently and otherwise asking in the CDN's forum, as Google itself cannot see or help with these errors.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Things

Used byrequirement DEV-SRV-01

StageConsistent with docsD1-C365

Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Used byrequirement DEV-SRV-09

  • Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
StageConsistent with docsD1-C366

Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.

“402 will just mean 404 to us”

Wording checked against the slide or recording

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Used byrequirement DEV-SRV-09

  • Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
StageConsistent with docsD1-C367

CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Used byrequirement DEV-SRV-02glossary term Soft 404

  • Extended by D1-C508 Day 1: Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing…
  • Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
  • Extended by D2-C375 Day 2: Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot…
StageConfirmed by docsD1-C368

Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

  • Extended by D1-C508 Day 1: Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing…
  • Extended by D2-C717 Day 2: Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed…
StageNot in docsD1-C509

Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Things

Used byrequirement DEV-SRV-09

  • Extends D1-C366 Day 1: Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
  • Extends D1-C365 Day 1: Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication…
StageNot in docsD1-C534

Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.

Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript

Used byrequirement DEV-CAN-11glossary term Redirect chain

StageNot in docsD1-C536

Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.

Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript

Used byrequirement DEV-CAN-11

StageConsistent with docsD1-C385

4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byrequirement DEV-ERR-01

  • Repeats D1-C072 Day 1: 4xx status codes other than 429 have no effect on crawl rate.

What Google's documentation says 7

DocsSourceD1-C126

Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.

“Google treats network timeouts, connection reset, and DNS errors similarly to 5xx server errors.”

Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search

Things

Used byrequirements DEV-MON-04, DEV-SRV-01, DEV-SRV-03fact F-017

  • Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
DocsSourceD1-C074

If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.

Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search

Used byrequirement DEV-SRV-04

  • Extended by D2-C210 Day 2: An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that…
DocsSourceD1-C138

Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.

Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search

Used byrequirements DEV-MON-04, DEV-SRV-01

DocsSourceD1-C370

Google's crawlers follow up to 10 redirect hops by default (some products' crawlers have other limits); content served by the redirecting URL is ignored and the final target's content is processed instead.

“Any content Google receives from the redirecting URL is ignored, and the final target URL's content is processed instead.”

Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search

Things

Used byrequirement DEV-CAN-01glossary term Redirect chain

Analysis by the author 5

AnalysisD1-C078

For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

Author Ibrahim AnjroAnnotates Day 1, 14:35 · How crawling errors affect Search

Used byrequirements DEV-MON-04, DEV-SRV-02, DEV-SRV-03

  • Extended by D2-C339 Day 2: For soft 404 detection, the position of error text decides: an error in a less important part such as the…
  • Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
  • Extended by D2-C376 Day 2: Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200…
AnalysisD1-C419

A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.

Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget

Day 2: Indexing 5

Said on stage 3

StageConsistent with docsD2-C374

A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
StageConsistent with docsD2-C375

Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Things

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…

What Google's documentation says 1

DocsSourceD2-C376

Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

Publisher Search Central blog (24 December 2024)Annotates Day 2, 11:55 · Handling web duplication

Things

Used byrequirements DEV-ERR-03, DEV-SRV-02

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…

Analysis by the author 1

AnalysisD2-C378

Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.

Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirements DEV-AIF-04, DEV-SRV-02

Day 3: Serving: Ranking, Search Console, and Performance 3

Said on stage 1

StageConsistent with docsD3-C618

When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript

Used byrequirement DEV-SRV-03

  • Extends D1-C126 Day 1: Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down…
  • Extends D1-C354 Day 1: Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the…
  • Extends D1-C440 Day 1: If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its…

What Google's documentation says 1

Analysis by the author 1

Across days and sessions 18

  1. Stage D1-C359 Day 1 · How crawling errors affect Search

    Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.

    extends
    Stage D1-C069 Day 1 · How crawling errors affect Search

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

  2. Stage D1-C508 Day 1 · How crawling errors affect Search

    Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  3. Stage D1-C508 Day 1 · How crawling errors affect Search

    Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.

    extends
    Stage D1-C368 Day 1 · How crawling errors affect Search

    Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.

  4. Stage D1-C509 Day 1 · How crawling errors affect Search

    Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

    extends
    Stage D1-C365 Day 1 · How crawling errors affect Search

    Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.

  5. Stage D1-C509 Day 1 · How crawling errors affect Search

    Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

    extends
    Stage D1-C366 Day 1 · How crawling errors affect Search

    Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.

  6. Analysis D2-C210 Day 2 · Lightning session D: Rendering and JavaScript

    An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.

    extends
    Docs D1-C074 Day 1 · How crawling errors affect Search

    If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.

  7. Stage D2-C339 Day 2 · Understanding what's on a page

    For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  8. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C069 Day 1 · How crawling errors affect Search

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

  9. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  10. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  11. Stage D2-C375 Day 2 · Handling web duplication

    Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  12. Docs D2-C376 Day 2 · Handling web duplication

    Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  13. Stage D2-C377 Day 2 · Handling web duplication

    AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

    extends
    Stage D1-C436 Day 1 · Q&A

    A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

  14. Stage D2-C717 Day 2 · Deciding what goes in the index?

    Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

    extends
    Stage D1-C368 Day 1 · How crawling errors affect Search

    Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.

  15. Stage D3-C618 Day 3 · How long does it take to..?

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    extends
    Docs D1-C126 Day 1 · How crawling errors affect Search

    Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.

  16. Stage D3-C618 Day 3 · How long does it take to..?

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    extends
    Stage D1-C354 Day 1 · How crawling errors affect Search

    Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.

  17. Stage D3-C618 Day 3 · How long does it take to..?

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    extends
    Stage D1-C440 Day 1 · Q&A

    If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.

  18. Stage D1-C385 Day 1 · How Google thinks about crawl budget

    4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.

    repeats
    Docs D1-C072 Day 1 · How crawling errors affect Search

    4xx status codes other than 429 have no effect on crawl rate.

Built on these claims 12

Developer requirements 11

Facts 1

Sources 12