The second attendee recording starts mid-talk, so the opening is from slides only.
Shown on screen 13
Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-PRF-01
- Extends D1-C326 Day 1: Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure…
- Extended by D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…
Crawl budget is the combination of crawl rate limit and crawl demand.
“Crawl Budget = Crawl Rate Limit & Crawl Demand”
Wording checked against the slide or recording
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byglossary term Crawl budget
Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
- Extended by D1-C374 Day 1: Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can…
- Extended by D1-C376 Day 1: Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app…
Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirements DEV-MON-04, DEV-PRF-01glossary term Crawl rate limit (hostload)
- Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byglossary term Crawl demand
- Extended by D1-C377 Day 1: The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual…
- Extended by D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
- Extended by D2-C695 Day 2: Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most…
- Extended by D2-C709 Day 2: Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and…
- Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
“If quality or popularity is unknown, use parent root's aggregate quality or popularity is used, then, that path's parent's, and so on.”
Wording checked against the slide or recording
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-10
- Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
- Extended by D2-C685 Day 2: Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies…
Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-08
- Extended by D1-C378 Day 1: Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same…
Six faceted-URL problems were shown: the same facets in a different order, irrelevant or conflicting facet combinations, excessive facet selection, facets on paginated series, facets combined with search queries, and optional facets with default values.
Speaker Cherry PrommawinEvidence slide photo
Used byrequirement DEV-URL-08glossary term Faceted navigation
Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byrequirements DEV-SRV-08, DEV-URL-08
- Extended by D1-C379 Day 1: The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages…
- Extended by D2-C886 Day 2: A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster…
URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-IDX-02fact F-025
- Extended by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-IDX-02fact F-025
- Extended by D2-C022 Day 2: Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a…
- Extended by D2-C069 Day 2: Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a…
- Extended by D2-C697 Day 2: Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing…
The nofollow rule can still consume crawl budget: Google does not crawl through the nofollow link itself, but it still crawls the linked page when it finds that page through other links.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-IDX-02
The non-standard crawl-delay rule is not processed by Googlebot, so it does not save any crawl budget.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirements DEV-IDX-02, DEV-SRV-05
Said on stage 23
Crawl budget was described as the attention span Google gives a website, and better performance increases it.
Speaker Cherry PrommawinEvidence notes, transcript
Used byrequirement DEV-PRF-01glossary term Crawl budget
- Extended by D1-C373 Day 1: Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it…
Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it determines how many of the site's pages are discovered and how often they are revisited.
Speaker Cherry PrommawinEvidence transcript
Used byglossary term Crawl budget
- Extends D1-C096 Day 1: Crawl budget was described as the attention span Google gives a website, and better performance increases it.
Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-PRF-01
- Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
- Extended by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.
Speaker Cherry PrommawinEvidence transcript
- Extends D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
- Extended by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
- Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.
Speaker Cherry PrommawinEvidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
- Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.
Speaker Cherry PrommawinEvidence transcript
Used byrequirements DEV-URL-08, DEV-URL-11
- Extends D1-C099 Day 1: Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and…
The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.
Speaker Cherry PrommawinEvidence transcript
- Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
QQuestion and answer
An audience member asked, in a question submitted before the event, whether crawl budget is still an SEO priority in 2026 or only relevant for very large sites.
From the audienceEvidence transcript
Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.
Speaker Cherry PrommawinEvidence transcript
Used byglossary term Crawl budget
- Repeated by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
HTTP status codes do not all affect crawl budget in the same way, so knowing what each status code class does is one way to manage crawl budget.
Speaker Cherry PrommawinEvidence transcript
Every fetch that returns a 2xx success response consumes crawl budget.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-SRV-08
4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-ERR-01
- Repeats D1-C072 Day 1: 4xx status codes other than 429 have no effect on crawl rate.
QQuestion and answers
An audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.
From the audienceEvidence transcript
AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.
Speaker Cherry PrommawinEvidence transcript
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
Google expects sites to see more crawling overall, because many other services, including AI services, now crawl the web besides Google.
Speaker Cherry PrommawinEvidence transcript
QQuestion and answers
An audience member asked, in a question submitted before the event, whether crawl frequency affects the ranking position of a URL.
From the audienceEvidence transcript
A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.
“if your crawl rate is increased, that doesn't mean that you would rank better.”
Speaker Cherry PrommawinEvidence transcript
Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal that people are interested in the site and that its content is of high quality.
Speaker Cherry PrommawinEvidence transcript
To check whether a site has a crawl budget problem, use Search Console's crawl report (Crawl Stats), which breaks crawl requests down by response and by file type and shows crawl problems Google finds.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-MON-04glossary term Crawl Stats report
The file-type breakdown of crawl requests in Search Console is useful for spotting anomalies, such as most crawling going to images on a site that has no images worth crawling.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-MON-04
Google advised checking crawl data in Search Console now and then without obsessing over crawl budget.
Speaker Cherry PrommawinEvidence transcript
Search Console's crawl data separates discovery (fetches of URLs Google has not seen before) from refresh (fetches of known URLs), and lets you drill down to the problematic URLs.
Speaker Cherry PrommawinEvidence transcript
Used byrequirements DEV-MON-04, DEV-MON-11glossary term Crawl Stats report
What Google's documentation says 8
Google's crawl budget guide says each crawler has its own crawl demand, but the crawl capacity limit (hostload) is shared across all crawlers, so high demand from one crawler can reduce the capacity left for others.
“high demand from one crawler can reduce the capacity available for others”
Publisher Google
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.
Publisher Google
- Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.
Publisher Google
Used byrequirement DEV-URL-08glossary term Faceted navigation
- Extended by D2-C188 Day 2: Hash-fragment links are a problem only where Google should follow them: product, category and language links…
- Repeated by D2-C284 Day 2: URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot…
If faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return a 404 when a filter combination has no results.
Publisher Google
Used byrequirement DEV-URL-09
- Extended by D1-C404 Day 1: For query parameters, a community speaker advised checking that every parameter in a request is actually used…
HTTP caching for crawlers means supporting conditional requests (ETag with If-None-Match, or Last-Modified with If-Modified-Since) and answering 304 Not Modified when nothing changed. Google's crawling team has said it prefers ETag.
Publisher Google, Search Central blog (9 December 2024)
Used byrequirement DEV-SRV-08
Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.
Publisher Google Search Central, Search Central blog (10 September 2019)
Used byrequirements DEV-IDX-02, DEV-IDX-06glossary term nofollow
- Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
Google advises against using noindex to save crawl budget and against using robots.txt to temporarily reallocate budget. Use robots.txt only for pages you never want crawled, and 404 or 410 for removed pages.
Publisher Google
Used byrequirements DEV-ERR-01, DEV-IDX-02
Google's guide to A/B testing says never to show one set of URLs to Googlebot and a different set to humans: that is cloaking, which is against Google's spam policies whether or not a test is running.
Publisher Google Search Central
Used byrequirement DEV-PRF-01
Analysis by the author 8
Because every Google product shares the same host capacity, Ads or Shopping fetches that hit a slow server can reduce how much is crawled for Search.
Author Ibrahim Anjro
New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.
Author Ibrahim Anjro
Used byrequirement DEV-URL-10
- Extended by D2-C686 Day 2: Launch new pages under sections that Google already indexes well, and improve or remove weak sections…
On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Author Ibrahim Anjro
- Extended by D2-C708 Day 2: The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on…
- Extended by D2-C714 Day 2: 'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages…
- Extended by D2-C716 Day 2: Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the…
A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-01
- Extended by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
- Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
- Extended by D2-C849 Day 2: Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can…
- Extended by D2-C058 Day 2: John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use…
- Repeated by D2-C850 Day 2: John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
- Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
The stage rule of thumb (under a few thousand URLs crawl budget is unlikely to be a problem) and Google's crawl budget guide (D1-C097: sites with over a million pages that change about weekly, or over 10,000 pages that change daily) leave a middle range where a site should check Search Console's Crawl Stats report before blaming crawl budget for slow indexing.
Author Ibrahim Anjro
The stage point that 4xx responses do not affect crawl budget matches Google's documentation that 4xx codes have no effect on crawl rate (D1-C072), with one exception: 429 Too Many Requests counts as a server error and slows crawling like a 5xx (D1-C071, D1-C092).
Author Ibrahim Anjro
A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.
Author Ibrahim Anjro
The stage remark that crawl demand follows the quality of the site as a whole sits beside the same talk's slide (D1-C094), which falls back to the parent path's aggregate only when a URL's own quality is unknown, and Google's crawl budget guide lists page quality among the demand factors; read it as site quality setting the baseline while known URL-level signals still count.
Author Ibrahim Anjro