Day 1: Crawling 45
Said on stage 33
A community speaker described the retrieval stage as whether an AI requests a site's pages when grounding its answer; if it does not, the cause may be a crawling or an indexing issue.
Speaker not identifiedIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
DNS works as the internet's address book: it is consulted for each request and tells the client at which IP address a host name can be reached.
Speaker Cherry PrommawinIn Day 1, 14:05 · How crawling worksEvidence transcript
An HTTP 200 OK status only means that the server believes it managed to do what the client asked for.
“just means that the server believes that it managed to accomplish whatever the user was asking for”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byrequirement DEV-ERR-03
To debug crawl issues, Google points site owners to the Crawl Stats report in Search Console, found in the Settings page under Crawl stats.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byrequirement DEV-MON-04glossary term Crawl Stats report
Search Console's Crawl Stats report shows how much Google crawls from a site, the errors Google received, and the content types fetched by the different crawlers.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byrequirement DEV-MON-04glossary term Crawl Stats report
Gary Illyes called the Crawl Stats report the most valuable tool for debugging crawl errors, together with server logs, which he called the real thing for those who know how to read them.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence notes, transcript
Used byrequirements DEV-MON-04, DEV-SRV-01
- Extended by D1-C359 Day 1: Most network errors happen somewhere between the site's origin server and Google's data centers, where they…
- Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
HTTP status codes fall into five classes, 1xx informational, 2xx success, 3xx redirection, 4xx client error and 5xx server error, and each class affects crawling differently.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byglossary term HTTP status code classes
1xx informational status codes, which only say that a request was received and more data is coming, have no meaning of their own for crawling.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
When a crawler's request returns 200, the page, file or other resource that comes back is eligible for indexing.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
A 204 No Content response is not eligible for indexing, because the server confirms the request but returns no content to index.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-ERR-03
Google follows 3xx redirects, permanent (301, 308) and temporary (302, 307), to the new location, and whether content gets indexed depends on what the redirect target returns.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-CAN-01
404 Not Found and 410 Gone both tell Google there is nothing at the URL, so the URL is not indexable.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirements DEV-ERR-01, DEV-SRV-09
A 403 response tells Google's crawler it has no permission to see the requested content, so there is nothing for Google to index.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-09
429 Too Many Requests is the exception among 4xx codes: instead of meaning there is nothing at the URL, it tells Google's crawler to slow down, and Google slows its crawling.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-03glossary term HTTP status code classes
Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-03
- Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
Gary Illyes called DNS and network errors a very common issue on the internet nowadays, and DNS problems very pesky to debug.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
DNS errors get a site removed from Google Search very aggressively, and Search Console alerts site owners to them.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-01
Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-01
- Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
When Search Console reports network errors, the first step Google recommends is to ask the hosting provider, or the CDN, whether they changed anything.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-01
Network timeouts are usually caused close to the site, very often by a firewall, a CDN, the hosting provider or the DNS provider, and Google has no visibility into them.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-01
Network errors and timeouts, like DNS errors, can get a site removed from Google Search and, with it, from every feature that depends on Search.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-01
For DNS and network errors caused by a firewall or CDN, Google advises checking whether new firewall rules were set recently and otherwise asking in the CDN's forum, as Google itself cannot see or help with these errors.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-01
A rise in HTTP errors in crawl reports can come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path between Google's crawler and the site.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-02
Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-09
- Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
“402 will just mean 404 to us”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-09
- Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-02glossary term Soft 404
- Extended by D1-C508 Day 1: Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing…
- Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
- Extended by D2-C375 Day 2: Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot…
Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
- Extended by D1-C508 Day 1: Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing…
- Extended by D2-C717 Day 2: Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed…
Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-09
- Extends D1-C366 Day 1: Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
- Extends D1-C365 Day 1: Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication…
Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11glossary term Redirect chain
Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11
HTTP status codes do not all affect crawl budget in the same way, so knowing what each status code class does is one way to manage crawl budget.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirement DEV-ERR-01
- Repeats D1-C072 Day 1: 4xx status codes other than 429 have no effect on crawl rate.
What Google's documentation says 7
Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.
“Google treats network timeouts, connection reset, and DNS errors similarly to 5xx server errors.”
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirements DEV-MON-04, DEV-SRV-01, DEV-SRV-03fact F-017
- Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
5xx and 429 responses prompt Google's crawlers to slow down temporarily. Already indexed URLs are preserved in the index for a while but eventually dropped.
“already indexed URLs are preserved in the index, but eventually dropped”
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirement DEV-SRV-03
4xx status codes other than 429 have no effect on crawl rate.
“The 4xx status codes, except 429, have no effect on crawl rate.”
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirement DEV-ERR-01
- Repeated by D1-C385 Day 1: 4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come…
If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirement DEV-SRV-04
- Extended by D2-C210 Day 2: An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that…
Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirements DEV-MON-04, DEV-SRV-01
Google's status code page says that with a 204 No Content response Google receives no content and cannot process it, and that a 2xx page whose content is empty or an error message shows as a soft 404 in Search Console.
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirement DEV-ERR-03
Google's crawlers follow up to 10 redirect hops by default (some products' crawlers have other limits); content served by the redirecting URL is ignored and the final target's content is processed instead.
“Any content Google receives from the redirecting URL is ignored, and the final target URL's content is processed instead.”
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirement DEV-CAN-01glossary term Redirect chain
Analysis by the author 5
For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
Author Ibrahim AnjroAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirements DEV-MON-04, DEV-SRV-02, DEV-SRV-03
- Extended by D2-C339 Day 2: For soft 404 detection, the position of error text decides: an error in a less important part such as the…
- Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
- Extended by D2-C376 Day 2: Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200…
Day 1 said a CDN captcha page served with HTTP 200 becomes a soft 404, while Day 2 (D2-C374, D2-C375) said such challenge pages are hard to recognise as errors and can be clustered as duplicates; Google's CDN post describes both outcomes, and in both the real pages drop out of Search.
Author Ibrahim AnjroAnnotates Day 1, 14:35 · How crawling errors affect Search
Do not answer verified Googlebot with 401, 402 or 403, for example from a login wall, a paywall or a pay-per-crawl setup: Google treats them like 404 and drops the pages, and Google's status code page says not to use 401 or 403 to limit crawling; to slow crawling temporarily, return 429 or 503.
Author Ibrahim AnjroAnnotates Day 1, 14:35 · How crawling errors affect Search
The stage point that 4xx responses do not affect crawl budget matches Google's documentation that 4xx codes have no effect on crawl rate (D1-C072), with one exception: 429 Too Many Requests counts as a server error and slows crawling like a 5xx (D1-C071, D1-C092).
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget