Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

Crawl budget

Crawl budget is crawl rate limit combined with crawl demand. Google's crawl budget guide is written for sites with over 1 million unique pages changing weekly, over 10,000 pages changing daily, or many URLs reported as 'Discovered – currently not indexed'. On Day 2 Google called that status a crawl-scheduling state in which it knows a URL but does not want to crawl it yet (said at the event), while the Page indexing help explains it by expected server overload. A Google slide also questioned whether a new URL that fits a known duplicate URL pattern even needs to be crawled (not in Google's docs). Author’s view: rule out slow responses and server errors first, then treat slow crawling of a smaller site as a demand problem, which means quality. Day 3 gave the scale: Google said it knows hundreds of trillions of URLs (as of October 2026) and does not crawl them all, because a much smaller crawl space is enough, and its crawl chart said discovering a new URL can take weeks or never happen. The second recording of Day 1 added the crawl budget talk and the Q&A. Crawl budget is the finite resources Google allocates to crawling one site, which decide how many of its pages are discovered and how often they are revisited; Google crawls every URL that differs even when it leads to the same content, so infinite spaces such as calendars burn it, and the content worth improving or removing includes low-quality, spam, duplicate and soft error pages. Not every site needs to worry, and never did: a site under a few thousand URLs likely has no crawl budget problem, news sites are crawled aggressively anyway, and a higher crawl rate does not make a URL rank better. Every 2xx fetch consumes crawl budget, while 4xx responses were said not to affect it. To check for a problem Google pointed to the Crawl Stats report: whether total crawling has plateaued, whether the average response time limits Googlebot, the file-type breakdown for anomalies and the discovery versus refresh split. Log files alone cannot show inefficiency, because they do not reveal where the crawl limit lies, but they help on sites with filter parameters; many complaints turn out to come from a plugin, such as a WordPress calendar plugin that adds calendar parameters to the URLs of every page, and Google learns that such URLs are useless only from large samples; a panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with an Apache rule, which saves crawl budget for the URLs that matter (the exact mechanism is not clear in the recording). For news sites Gary Illyes advised measuring the time from publishing a URL to Googlebot's first crawl and investigating only a rising trend, adding that for fresh stories something like two hours is probably not great, while John Mueller said that if Crawl Stats shows about half of crawling going to discovering new pages, crawling of new pages is not the problem; Gary Illyes said disallowing a section such as /ads shifts crawl budget to the rest of the site. Author’s view: Google's guide says freed crawl budget shifts only on sites that already hit their capacity limit, a 404 is still a fetch that counts towards crawl budget, and a site between a few thousand URLs and Google's million-page threshold should check Crawl Stats before blaming crawl budget; for a consolidation of about 2,000 URLs the crawl-budget gain is likely minor.

What to do

  • Remove or improve useless pages, fix server errors and close infinite URL spaces.
  • For 'Discovered – currently not indexed', check response times and server errors first, then raise the quality of what is already indexed instead of resubmitting URLs.
  • Avoid many similar-looking URLs that lead to the same content, so new URLs are not written off as duplicates before they are crawled.
  • Check Crawl Stats now and then without obsessing over crawl budget: look for a plateau in total crawling, rising response times and odd file-type shares.
  • On news sites, measure the time from publishing a URL to Googlebot's first crawl and investigate only if it rises; for fresh stories, around two hours is probably too long.
  • Check plugins that generate URLs, such as calendars, before blaming crawl budget, and handle useless parameter URLs at the web server.

Day 1: Crawling 41

Shown on screen 3

SlideConsistent with docsD1-C099

Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C378 Day 1: Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same…
SlideConsistent with docsD1-C103

Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence 2 slide photos, transcript

Used byrequirements DEV-SRV-08, DEV-URL-08

  • Extended by D1-C379 Day 1: The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages…
  • Extended by D2-C886 Day 2: A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster…

Said on stage 33

StageConsistent with docsD1-C096

Crawl budget was described as the attention span Google gives a website, and better performance increases it.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence notes, transcript

Used byrequirement DEV-PRF-01glossary term Crawl budget

  • Extended by D1-C373 Day 1: Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it…
StageConsistent with docsD1-C373

Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it determines how many of the site's pages are discovered and how often they are revisited.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byglossary term Crawl budget

  • Extends D1-C096 Day 1: Crawl budget was described as the attention span Google gives a website, and better performance increases it.
StageConsistent with docsD1-C378

Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byrequirements DEV-URL-08, DEV-URL-11

  • Extends D1-C099 Day 1: Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and…
StageConfirmed by docsD1-C379

The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

  • Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
StageConfirmed by docsD1-C381

Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byglossary term Crawl budget

  • Answers D1-C380 Day 1: An audience member asked, in a question submitted before the event, whether crawl budget is still an SEO…
  • Repeated by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
StageConsistent with docsD1-C385

4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byrequirement DEV-ERR-01

  • Repeats D1-C072 Day 1: 4xx status codes other than 429 have no effect on crawl rate.
StageD1-C387

An audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.

From the audienceIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

  • Answered by D1-C388 Day 1: AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl…
  • Answered by D1-C389 Day 1: Google expects sites to see more crawling overall, because many other services, including AI services, now…
StageD1-C390

An audience member asked, in a question submitted before the event, whether crawl frequency affects the ranking position of a URL.

From the audienceIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

  • Answered by D1-C391 Day 1: A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.
  • Answered by D1-C392 Day 1: Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal…
StageConfirmed by docsD1-C391

A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.

“if your crawl rate is increased, that doesn't mean that you would rank better.”

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

  • Answers D1-C390 Day 1: An audience member asked, in a question submitted before the event, whether crawl frequency affects the…
StageConfirmed by docsD1-C393

To check whether a site has a crawl budget problem, use Search Console's crawl report (Crawl Stats), which breaks crawl requests down by response and by file type and shows crawl problems Google finds.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byrequirement DEV-MON-04glossary term Crawl Stats report

StageD1-C441

An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl budget inefficiently on a site.

From the audienceIn Day 1, 16:35 · Q&AEvidence transcript

  • Answered by D1-C442 Day 1: Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the…
  • Answered by D1-C443 Day 1: To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued…
  • Answered by D1-C444 Day 1: Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily…
  • Answered by D1-C445 Day 1: One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it…
StageNot in docsD1-C442

Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the site lies; many similar URLs being crawled may not matter if the site is nowhere near that limit.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-MON-04

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
StageConsistent with docsD1-C443

To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued over time, and whether the average response time shows the server is fast enough or is limiting Googlebot.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-MON-04

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
StageConsistent with docsD1-C444

Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
  • Extends D1-C397 Day 1: A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary…
StageNot in docsD1-C446

Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
StageConsistent with docsD1-C447

Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
StageD1-C465

An audience member asked how a large news site can tell whether crawl budget is limiting how fast new articles are discovered (within minutes), and which statistics in Search Console and the logs show this.

From the audienceIn Day 1, 16:35 · Q&AEvidence transcript

  • Answered by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
  • Answered by D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
  • Answered by D1-C468 Day 1: When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google…
  • Answered by D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…
StageConsistent with docsD1-C466

Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-MON-11

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Repeats D1-C381 Day 1: Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few…
  • Extended by D3-C606 Day 3: For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading…
StageNot in docsD1-C467

To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.

Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Extended by D1-C543 Day 1: Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises…
StageConsistent with docsD1-C468

When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google crawls and recrawls a lot and judging whether they are useful.

Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-MON-11

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
StageConsistent with docsD1-C469

Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Extended by D1-C470 Day 1: Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless…
StageConsistent with docsD1-C492

John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a rough view of whether discovery crawling is in a reasonable range, though measuring publish-to-first-crawl time is more accurate.

Speaker John MuellerIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-MON-11

  • Extended by D1-C544 Day 1: John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering…
StageD1-C545

An audience member asked, for very large sites of about 100 million pages, what signs show that a site is limited by crawl budget, and how to tell a crawl capacity limit problem from a crawl demand problem.

From the audienceIn Day 1, 16:35 · Q&AEvidence transcript

  • Answered by D1-C493 Day 1: There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower…
  • Answered by D1-C494 Day 1: To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot…
  • Answered by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
StageConsistent with docsD1-C539

A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Extends D1-C447 Day 1: Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its…
  • Extends D1-C446 Day 1: Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a…
StageNot in docsD1-C543

Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Extends D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
StageConsistent with docsD1-C544

John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering new pages, that is a lot of discovery crawling, which suggests that crawling of new pages is not the problem.

Speaker John MuellerIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-MON-11

  • Extends D1-C492 Day 1: John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a…

What Google's documentation says 1

DocsSourceD1-C097

Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.

Publisher GoogleAnnotates Day 1, 16:00 · How Google thinks about crawl budget

  • Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…

Analysis by the author 4

AnalysisD1-C098

On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget

  • Extended by D2-C708 Day 2: The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on…
  • Extended by D2-C714 Day 2: 'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages…
  • Extended by D2-C716 Day 2: Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the…
AnalysisD1-C382

The stage rule of thumb (under a few thousand URLs crawl budget is unlikely to be a problem) and Google's crawl budget guide (D1-C097: sites with over a million pages that change about weekly, or over 10,000 pages that change daily) leave a middle range where a site should check Search Console's Crawl Stats report before blaming crawl budget for slow indexing.

Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget

AnalysisD1-C419

A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.

Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget

AnalysisD1-C470

Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.

Author Ibrahim AnjroAnnotates Day 1, 16:35 · Q&A

  • Extends D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…

Day 2: Indexing 7

Shown on screen 1

SlideNot in docsD2-C369

Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

“Do we even need to crawl /buy/seo-service ?”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo

Used byrequirements DEV-CAN-08, DEV-URL-09

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…

Said on stage 3

StageConsistent with docsD2-C848

Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

  • Repeats D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
StageConsistent with docsD2-C886

A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.

Speaker Martyna AğanoğluIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

  • Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
StageNot in docsD2-C706

'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

“The first one is kind of nastier.”

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed

  • Extends D1-C097 Day 1: Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over…
  • Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…

What Google's documentation says 1

DocsSourceD2-C707

Google's Page indexing report help says a 'Discovered – currently not indexed' page was found but not crawled yet, typically because Google wanted to crawl it but expected the crawl to overload the site, so it rescheduled the crawl.

Publisher Google Search Console HelpAnnotates Day 2, 15:40 · Deciding what goes in the index?

Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed

Analysis by the author 2

AnalysisD2-C914

Google's crawl budget guide is written for sites with over a million pages changing weekly or over 10,000 pages changing daily, so for a consolidation of about 2,000 URLs the crawl-budget gain is likely minor; the benefit of pruning such a site more plausibly comes from one strong URL per intent and consolidated signals.

Author Ibrahim AnjroAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

AnalysisD2-C708

The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

Day 3: Serving: Ranking, Search Console, and Performance 3

Shown on screen 1

Said on stage 2

Across days and sessions 22

  1. Stage D1-C373 Day 1 · How Google thinks about crawl budget

    Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it determines how many of the site's pages are discovered and how often they are revisited.

    extends
    Stage D1-C096 Day 1 · How Google thinks about crawl budget

    Crawl budget was described as the attention span Google gives a website, and better performance increases it.

  2. Stage D1-C378 Day 1 · How Google thinks about crawl budget

    Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

    extends
    Slide D1-C099 Day 1 · How Google thinks about crawl budget

    Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.

  3. Stage D1-C379 Day 1 · How Google thinks about crawl budget

    The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  4. Stage D1-C444 Day 1 · Q&A

    Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

    extends
    Stage D1-C397 Day 1 · Lightning session C: Crawling

    A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.

  5. Analysis D1-C470 Day 1 · Q&A

    Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.

    extends
    Stage D1-C469 Day 1 · Q&A

    Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

  6. Stage D1-C539 Day 1 · Q&A

    A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

    extends
    Stage D1-C446 Day 1 · Q&A

    Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

  7. Stage D1-C539 Day 1 · Q&A

    A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

    extends
    Stage D1-C447 Day 1 · Q&A

    Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.

  8. Stage D1-C543 Day 1 · Q&A

    Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

    extends
    Stage D1-C467 Day 1 · Q&A

    To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.

  9. Stage D1-C544 Day 1 · Q&A

    John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering new pages, that is a lot of discovery crawling, which suggests that crawling of new pages is not the problem.

    extends
    Stage D1-C492 Day 1 · Q&A

    John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a rough view of whether discovery crawling is in a reasonable range, though measuring publish-to-first-crawl time is more accurate.

  10. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  11. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  12. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  13. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Docs D1-C097 Day 1 · How Google thinks about crawl budget

    Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.

  14. Analysis D2-C708 Day 2 · Deciding what goes in the index?

    The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  15. Stage D2-C714 Day 2 · Deciding what goes in the index?

    'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  16. Analysis D2-C716 Day 2 · Deciding what goes in the index?

    Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  17. Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  18. Stage D3-C602 Day 3 · How long does it take to..?

    Google said it knows hundreds of trillions of URLs (as of October 2026).

    extends
    Stage D1-C201 Day 1 · How Search works and where's AI?

    Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.

  19. Stage D3-C606 Day 3 · How long does it take to..?

    For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).

    extends
    Stage D1-C466 Day 1 · Q&A

    Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

  20. Stage D1-C385 Day 1 · How Google thinks about crawl budget

    4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.

    repeats
    Docs D1-C072 Day 1 · How crawling errors affect Search

    4xx status codes other than 429 have no effect on crawl rate.

  21. Stage D1-C466 Day 1 · Q&A

    Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

    repeats
    Stage D1-C381 Day 1 · How Google thinks about crawl budget

    Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.

  22. Stage D2-C848 Day 2 · Welcome to indexing day!

    Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.

    repeats
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

Built on these claims 10

Developer requirements 10

Sources 10