Day 1: Crawling 22
Shown on screen 3
Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-PRF-01
- Extends D1-C326 Day 1: Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure…
- Extended by D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…
Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
- Extended by D1-C374 Day 1: Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can…
- Extended by D1-C376 Day 1: Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app…
Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirements DEV-MON-04, DEV-PRF-01glossary term Crawl rate limit (hostload)
- Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
Said on stage 12
429 Too Many Requests is the exception among 4xx codes: instead of meaning there is nothing at the URL, it tells Google's crawler to slow down, and Google slows its crawling.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-03glossary term HTTP status code classes
Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-03
- Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
A rise in HTTP errors in crawl reports can come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path between Google's crawler and the site.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-02
Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirement DEV-PRF-01
- Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
- Extended by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
- Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
- Answers D1-C438 Day 1: An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as…
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
- Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
- Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-PRF-01
- Answers D1-C438 Day 1: An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as…
- Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
An audience member asked, for very large sites of about 100 million pages, what signs show that a site is limited by crawl budget, and how to tell a crawl capacity limit problem from a crawl demand problem.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C493 Day 1: There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower…
- Answered by D1-C494 Day 1: To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot…
- Answered by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
- Answers D1-C545 Day 1: An audience member asked, for very large sites of about 100 million pages, what signs show that a site is…
- Extended by D3-C617 Day 3: Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2…
To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot opens to the site: if it has dropped, the capacity limit changed.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
- Answers D1-C545 Day 1: An audience member asked, for very large sites of about 100 million pages, what signs show that a site is…
A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byglossary term Crawl demand
- Extends D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
- Extends D1-C334 Day 1: Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site…
- Repeated by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
- Answers D1-C545 Day 1: An audience member asked, for very large sites of about 100 million pages, what signs show that a site is…
- Extends D1-C374 Day 1: Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can…
- Extended by D1-C547 Day 1: Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the…
What Google's documentation says 3
Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.
“Google treats network timeouts, connection reset, and DNS errors similarly to 5xx server errors.”
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirements DEV-MON-04, DEV-SRV-01, DEV-SRV-03fact F-017
- Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
5xx and 429 responses prompt Google's crawlers to slow down temporarily. Already indexed URLs are preserved in the index for a while but eventually dropped.
“already indexed URLs are preserved in the index, but eventually dropped”
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirement DEV-SRV-03
Google's crawl budget guide says each crawler has its own crawl demand, but the crawl capacity limit (hostload) is shared across all crawlers, so high demand from one crawler can reduce the capacity left for others.
“high demand from one crawler can reduce the capacity available for others”
Publisher GoogleAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
Analysis by the author 4
Because every Google product shares the same host capacity, Ads or Shopping fetches that hit a slow server can reduce how much is crawled for Search.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Do not answer verified Googlebot with 401, 402 or 403, for example from a login wall, a paywall or a pay-per-crawl setup: Google treats them like 404 and drops the pages, and Google's status code page says not to use 401 or 403 to limit crawling; to slow crawling temporarily, return 429 or 503.
Author Ibrahim AnjroAnnotates Day 1, 14:35 · How crawling errors affect Search
The stage point that 4xx responses do not affect crawl budget matches Google's documentation that 4xx codes have no effect on crawl rate (D1-C072), with one exception: 429 Too Many Requests counts as a server error and slows crawling like a 5xx (D1-C071, D1-C092).
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the total time a server spends holding connections open for Google, counting both the number of parallel connections and their duration. That fits the panel's advice to watch how many connections Googlebot opens (D1-C494): fewer connections is the documented form of a lower capacity limit.
Author Ibrahim AnjroAnnotates Day 1, 16:35 · Q&A
Used byrequirement DEV-PRF-01
- Extends D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
Day 3: Serving: Ranking, Search Console, and Performance 6
Shown on screen 1
Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2 weeks, and 1-3 weeks in recovery.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript
- Extends D1-C493 Day 1: There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower…
Said on stage 3
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
Used byrequirement DEV-SRV-03
- Extends D1-C126 Day 1: Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down…
- Extends D1-C354 Day 1: Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the…
- Extends D1-C440 Day 1: If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its…
Increases in crawl capacity take longer than decreases, within one to three weeks, because Google first needs to know that the higher demand will last.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
Used byrequirement DEV-SRV-03
A continuously running process recalculates each site's crawl capacity within a month, so a capacity change takes up to about a month at most.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
What Google's documentation says 1
Google's crawl rate guide says that when a significant number of URLs return 500, 503 or 429, Google reduces the site's crawl rate, which starts increasing again automatically once the errors drop; it warns against doing this for longer than 1-2 days.
Publisher GoogleAnnotates Day 3, 15:45 · How long does it take to..?
Used byrequirement DEV-SRV-03
Analysis by the author 1
Spoken and slide figures for capacity increases differ (one to three weeks, up to a month, versus 1-2 weeks typical and 1-3 weeks in recovery), but the lesson is the same: a burst of 5xx errors cuts crawling within hours and recovery takes weeks, so keep servers stable before launches and migrations.
Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?
Used byrequirement DEV-SRV-03