Google's crawl budget guide says each crawler has its own crawl demand, but the crawl capacity limit (hostload) is shared across all crawlers, so high demand from one crawler can reduce the capacity left for others.
Thing · Metric
Crawl capacity limit
How much Google can crawl a site without overloading its servers (once called crawl rate limit, internally hostload).
- Claims
- 18
- In Google’s docs
- 1
- Said at the event
- 15
- Not in docs
- 2
- Kit items
- 5
Glossary · Crawl rate limit (hostload)
How much crawling a host can take, shared by all Google crawlers. It goes down when connect time, time to first byte, 429 or 5xx responses go up. It applies per host, so a site's CDN, subdomains and main www host may be counted separately. A drop is most of the time an abrupt step down, and Googlebot then opens fewer connections, a Google panelist said.
Google’s documentation 1
Documented in
Said at the event 15
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Not in docs 2
Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2 weeks, and 1-3 weeks in recovery.
A continuously running process recalculates each site's crawl capacity within a month, so a capacity change takes up to about a month at most.
Gary Illyes · Day 3 · How long does it take to..?
Consistent with docs 10
Crawl budget is the combination of crawl rate limit and crawl demand.
Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.
Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.
There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.
To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot opens to the site: if it has dropped, the capacity limit changed.
A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.
A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
Gary Illyes · Day 3 · How long does it take to..?
Increases in crawl capacity take longer than decreases, within one to three weeks, because Google first needs to know that the higher demand will last.
Gary Illyes · Day 3 · How long does it take to..?
Confirmed by docs 2
Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.
Nothing to verify 1
An audience member asked, for very large sites of about 100 million pages, what signs show that a site is limited by crawl budget, and how to tell a crawl capacity limit problem from a crawl demand problem.
From the audience · Day 1 · Q&A
Press and analysis 2
Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.
Ibrahim Anjro · Day 1 · Q&A
Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the total time a server spends holding connections open for Google, counting both the number of parallel connections and their duration. That fits the panel's advice to watch how many connections Googlebot opens (D1-C494): fewer connections is the documented form of a lower capacity limit.
Ibrahim Anjro · Day 1 · Q&A
Built on these claims 5
Kit items about Crawl capacity limit: their own words name it, or several of the claims they rest on do.
Developer requirements 2
Keep connect time and time to first byte low and stable under crawler load
Rests on 13 claims, 8 of them naming Crawl capacity limit; its own words name Crawl capacity limit
Answer planned maintenance and short outages with 503, for a day or two at most
Rests on 11 claims, 2 of them naming Crawl capacity limit; its own words name Crawl capacity limit
Also inglossary terms Crawl rate limit (hostload), Crawl budget, Crawl demand
Connected things 10
Relations
- Part of Crawl budget structure, no claim needed
- Affected by 429
1 claim, 1 documented
- Consistent with docsD1-C092
Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
- Affected by 5xx
1 claim, 1 documented
- Consistent with docsD1-C092
Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
Most often named with it
Things named in the same claim, with the number of claims they share.