Day 1: Crawling 41
Shown on screen 3
Crawl budget is the combination of crawl rate limit and crawl demand.
“Crawl Budget = Crawl Rate Limit & Crawl Demand”
Wording checked against the slide or recording
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence 2 slide photos, transcript
Used byglossary term Crawl budget
Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-URL-08
- Extended by D1-C378 Day 1: Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same…
Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence 2 slide photos, transcript
Used byrequirements DEV-SRV-08, DEV-URL-08
- Extended by D1-C379 Day 1: The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages…
- Extended by D2-C886 Day 2: A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster…
Said on stage 33
Soft 404 pages are still crawled like normal pages; their bigger consequences are for indexing, and they also affect crawling.
Speaker Cherry PrommawinIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Crawl budget was described as the attention span Google gives a website, and better performance increases it.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence notes, transcript
Used byrequirement DEV-PRF-01glossary term Crawl budget
- Extended by D1-C373 Day 1: Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it…
Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it determines how many of the site's pages are discovered and how often they are revisited.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byglossary term Crawl budget
- Extends D1-C096 Day 1: Crawl budget was described as the attention span Google gives a website, and better performance increases it.
Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirements DEV-URL-08, DEV-URL-11
- Extends D1-C099 Day 1: Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and…
The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
An audience member asked, in a question submitted before the event, whether crawl budget is still an SEO priority in 2026 or only relevant for very large sites.
From the audienceIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Answered by D1-C381 Day 1: Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few…
Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byglossary term Crawl budget
- Answers D1-C380 Day 1: An audience member asked, in a question submitted before the event, whether crawl budget is still an SEO…
- Repeated by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
HTTP status codes do not all affect crawl budget in the same way, so knowing what each status code class does is one way to manage crawl budget.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Every fetch that returns a 2xx success response consumes crawl budget.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirement DEV-SRV-08
4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirement DEV-ERR-01
- Repeats D1-C072 Day 1: 4xx status codes other than 429 have no effect on crawl rate.
An audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.
From the audienceIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Answered by D1-C388 Day 1: AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl…
- Answered by D1-C389 Day 1: Google expects sites to see more crawling overall, because many other services, including AI services, now…
An audience member asked, in a question submitted before the event, whether crawl frequency affects the ranking position of a URL.
From the audienceIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Answered by D1-C391 Day 1: A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.
- Answered by D1-C392 Day 1: Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal…
A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.
“if your crawl rate is increased, that doesn't mean that you would rank better.”
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Answers D1-C390 Day 1: An audience member asked, in a question submitted before the event, whether crawl frequency affects the…
To check whether a site has a crawl budget problem, use Search Console's crawl report (Crawl Stats), which breaks crawl requests down by response and by file type and shows crawl problems Google finds.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirement DEV-MON-04glossary term Crawl Stats report
The file-type breakdown of crawl requests in Search Console is useful for spotting anomalies, such as most crawling going to images on a site that has no images worth crawling.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirement DEV-MON-04
Google advised checking crawl data in Search Console now and then without obsessing over crawl budget.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Google follows robots.txt partly in its own interest: crawling an infinite URL space such as a calendar that robots.txt blocks would waste Google's crawling time as well as the site's resources.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl budget inefficiently on a site.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C442 Day 1: Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the…
- Answered by D1-C443 Day 1: To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued…
- Answered by D1-C444 Day 1: Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily…
- Answered by D1-C445 Day 1: One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it…
Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the site lies; many similar URLs being crawled may not matter if the site is nowhere near that limit.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-04
- Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued over time, and whether the average response time shows the server is fast enough or is limiting Googlebot.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-04
- Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
- Extends D1-C397 Day 1: A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary…
Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
An audience member asked how a large news site can tell whether crawl budget is limiting how fast new articles are discovered (within minutes), and which statistics in Search Console and the logs show this.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
- Answered by D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
- Answered by D1-C468 Day 1: When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google…
- Answered by D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…
Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-11
- Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
- Repeats D1-C381 Day 1: Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few…
- Extended by D3-C606 Day 3: For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading…
To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-11
- Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
- Extended by D1-C543 Day 1: Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises…
When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google crawls and recrawls a lot and judging whether they are useful.
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-11
- Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
- Extended by D1-C470 Day 1: Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless…
John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a rough view of whether discovery crawling is in a reasonable range, though measuring publish-to-first-crawl time is more accurate.
Speaker John MuellerIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-11
- Extended by D1-C544 Day 1: John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering…
An audience member asked, for very large sites of about 100 million pages, what signs show that a site is limited by crawl budget, and how to tell a crawl capacity limit problem from a crawl demand problem.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C493 Day 1: There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower…
- Answered by D1-C494 Day 1: To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot…
- Answered by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Extends D1-C447 Day 1: Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its…
- Extends D1-C446 Day 1: Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a…
Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-11
- Extends D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering new pages, that is a lot of discovery crawling, which suggests that crawling of new pages is not the problem.
Speaker John MuellerIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-11
- Extends D1-C492 Day 1: John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a…
What Google's documentation says 1
Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.
Publisher GoogleAnnotates Day 1, 16:00 · How Google thinks about crawl budget
- Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
Analysis by the author 4
On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
- Extended by D2-C708 Day 2: The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on…
- Extended by D2-C714 Day 2: 'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages…
- Extended by D2-C716 Day 2: Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the…
The stage rule of thumb (under a few thousand URLs crawl budget is unlikely to be a problem) and Google's crawl budget guide (D1-C097: sites with over a million pages that change about weekly, or over 10,000 pages that change daily) leave a middle range where a site should check Search Console's Crawl Stats report before blaming crawl budget for slow indexing.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.
Author Ibrahim AnjroAnnotates Day 1, 16:35 · Q&A
- Extends D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…