DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
Speaker Gary IllyesEvidence notes, transcript
Used byrequirements DEV-MON-04, DEV-SRV-01
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Day 1 · Wednesday 30 September 2026 · 14:35
Cherry Prommawin presented the status codes and Gary Illyes the DNS, network, CDN and soft 404 errors (on-stage hand-overs; the host thanked both). Transcript from a second attendee recording and a complete audio recording.
DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
Speaker Gary IllyesEvidence notes, transcript
Used byrequirements DEV-MON-04, DEV-SRV-01
Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.
Speaker Gary IllyesEvidence notes, transcript
HTTP status codes fall into five classes, 1xx informational, 2xx success, 3xx redirection, 4xx client error and 5xx server error, and each class affects crawling differently.
Speaker Cherry PrommawinEvidence transcript
Used byglossary term HTTP status code classes
1xx informational status codes, which only say that a request was received and more data is coming, have no meaning of their own for crawling.
Speaker Cherry PrommawinEvidence transcript
When a crawler's request returns 200, the page, file or other resource that comes back is eligible for indexing.
Speaker Cherry PrommawinEvidence transcript
A 204 No Content response is not eligible for indexing, because the server confirms the request but returns no content to index.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-ERR-03
Google follows 3xx redirects, permanent (301, 308) and temporary (302, 307), to the new location, and whether content gets indexed depends on what the redirect target returns.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-CAN-01
A 304 Not Modified response tells Google's crawler that the content has not changed since its last visit.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-SRV-08
404 Not Found and 410 Gone both tell Google there is nothing at the URL, so the URL is not indexable.
Speaker Cherry PrommawinEvidence transcript
Used byrequirements DEV-ERR-01, DEV-SRV-09
A 403 response tells Google's crawler it has no permission to see the requested content, so there is nothing for Google to index.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-SRV-09
429 Too Many Requests is the exception among 4xx codes: instead of meaning there is nothing at the URL, it tells Google's crawler to slow down, and Google slows its crawling.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-SRV-03glossary term HTTP status code classes
Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-SRV-03
A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-ERR-03
Soft 404 pages are still crawled like normal pages; their bigger consequences are for indexing, and they also affect crawling.
Speaker Cherry PrommawinEvidence transcript
Gary Illyes called DNS and network errors a very common issue on the internet nowadays, and DNS problems very pesky to debug.
Speaker Gary IllyesEvidence transcript
DNS errors get a site removed from Google Search very aggressively, and Search Console alerts site owners to them.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-01
Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-01
When Search Console reports network errors, the first step Google recommends is to ask the hosting provider, or the CDN, whether they changed anything.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-01
Network timeouts are usually caused close to the site, very often by a firewall, a CDN, the hosting provider or the DNS provider, and Google has no visibility into them.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-01
Network errors and timeouts, like DNS errors, can get a site removed from Google Search and, with it, from every feature that depends on Search.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-01
For DNS and network errors caused by a firewall or CDN, Google advises checking whether new firewall rules were set recently and otherwise asking in the CDN's forum, as Google itself cannot see or help with these errors.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-01
A rise in HTTP errors in crawl reports can come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path between Google's crawler and the site.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-02
Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-09
Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
“402 will just mean 404 to us”
Wording checked against the slide or recording
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-09
CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-02glossary term Soft 404
Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
Speaker Gary IllyesEvidence transcript
Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-MON-03
Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-09
Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.
“Google treats network timeouts, connection reset, and DNS errors similarly to 5xx server errors.”
Publisher Google
Used byrequirements DEV-MON-04, DEV-SRV-01, DEV-SRV-03fact F-017
5xx and 429 responses prompt Google's crawlers to slow down temporarily. Already indexed URLs are preserved in the index for a while but eventually dropped.
“already indexed URLs are preserved in the index, but eventually dropped”
Publisher Google
Used byrequirement DEV-SRV-03
4xx status codes other than 429 have no effect on crawl rate.
“The 4xx status codes, except 429, have no effect on crawl rate.”
Publisher Google
Used byrequirement DEV-ERR-01
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
Publisher Google, Google Search Central
Used byrequirement DEV-ERR-01glossary term Soft 404
If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.
Publisher Google
Used byrequirement DEV-SRV-04
Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.
Publisher Google
Used byrequirements DEV-MON-04, DEV-SRV-01
Google's status code page says that with a 204 No Content response Google receives no content and cannot process it, and that a 2xx page whose content is empty or an error message shows as a soft 404 in Search Console.
Publisher Google
Used byrequirement DEV-ERR-03
Google's crawlers follow up to 10 redirect hops by default (some products' crawlers have other limits); content served by the redirecting URL is ignored and the final target's content is processed instead.
“Any content Google receives from the redirecting URL is ignored, and the final target URL's content is processed instead.”
Publisher Google
Used byrequirement DEV-CAN-01glossary term Redirect chain
For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
Author Ibrahim Anjro
Used byrequirements DEV-MON-04, DEV-SRV-02, DEV-SRV-03
Day 1 said a CDN captcha page served with HTTP 200 becomes a soft 404, while Day 2 (D2-C374, D2-C375) said such challenge pages are hard to recognise as errors and can be clustered as duplicates; Google's CDN post describes both outcomes, and in both the real pages drop out of Search.
Author Ibrahim Anjro
Do not answer verified Googlebot with 401, 402 or 403, for example from a login wall, a paywall or a pay-per-crawl setup: Google treats them like 404 and drops the pages, and Google's status code page says not to use 401 or 403 to limit crawling; to slow crawling temporarily, return 429 or 503.
Author Ibrahim Anjro
In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.
If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.
A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.
For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.
4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.
4xx status codes other than 429 have no effect on crawl rate.
Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but returns a 200 HTTP status code.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.