Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 1 · Wednesday 30 September 2026 · 16:00

How Google thinks about crawl budget

Speaker Cherry Prommawin, Search Relations

TalkCoverageSlidesTranscript

The second attendee recording starts mid-talk, so the opening is from slides only.

Shown on screen 13

SlideConsistent with docsD1-C066

Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-PRF-01

  • Extends D1-C326 Day 1: Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure…
  • Extended by D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…
SlideConfirmed by docsD1-C091

Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Extended by D1-C374 Day 1: Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can…
  • Extended by D1-C376 Day 1: Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app…
SlideConsistent with docsD1-C092

Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirements DEV-MON-04, DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
SlideConsistent with docsD1-C093

Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byglossary term Crawl demand

  • Extended by D1-C377 Day 1: The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual…
  • Extended by D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
  • Extended by D2-C695 Day 2: Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most…
  • Extended by D2-C709 Day 2: Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and…
  • Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
SlideNot in docsD1-C094

If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

“If quality or popularity is unknown, use parent root's aggregate quality or popularity is used, then, that path's parent's, and so on.”

Wording checked against the slide or recording

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-10

  • Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
  • Extended by D2-C685 Day 2: Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies…
SlideConsistent with docsD1-C100

Six faceted-URL problems were shown: the same facets in a different order, irrelevant or conflicting facet combinations, excessive facet selection, facets on paginated series, facets combined with search queries, and optional facets with default values.

Speaker Cherry PrommawinEvidence slide photo

Used byrequirement DEV-URL-08glossary term Faceted navigation

SlideConsistent with docsD1-C103

Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirements DEV-SRV-08, DEV-URL-08

  • Extended by D1-C379 Day 1: The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages…
  • Extended by D2-C886 Day 2: A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster…
SlideConfirmed by docsD1-C106

The noindex rule consumes crawl budget, because Google must fetch the page to see it.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-IDX-02fact F-025

  • Extended by D2-C022 Day 2: Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a…
  • Extended by D2-C069 Day 2: Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a…
  • Extended by D2-C697 Day 2: Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing…
SlideConfirmed by docsD1-C107

The nofollow rule can still consume crawl budget: Google does not crawl through the nofollow link itself, but it still crawls the linked page when it finds that page through other links.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-IDX-02

Said on stage 23

StageConsistent with docsD1-C096

Crawl budget was described as the attention span Google gives a website, and better performance increases it.

Speaker Cherry PrommawinEvidence notes, transcript

Used byrequirement DEV-PRF-01glossary term Crawl budget

  • Extended by D1-C373 Day 1: Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it…
StageConsistent with docsD1-C373

Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it determines how many of the site's pages are discovered and how often they are revisited.

Speaker Cherry PrommawinEvidence transcript

Used byglossary term Crawl budget

  • Extends D1-C096 Day 1: Crawl budget was described as the attention span Google gives a website, and better performance increases it.
StageConsistent with docsD1-C374

Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-PRF-01

  • Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
  • Extended by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
StageConfirmed by docsD1-C375

Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

Speaker Cherry PrommawinEvidence transcript

Things
  • Extends D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
  • Extended by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
StageConfirmed by docsD1-C376

Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
StageNot in docsD1-C377

The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.

Speaker Cherry PrommawinEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
  • Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
StageConsistent with docsD1-C378

Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

Speaker Cherry PrommawinEvidence transcript

Used byrequirements DEV-URL-08, DEV-URL-11

  • Extends D1-C099 Day 1: Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and…
StageConfirmed by docsD1-C379

The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.

Speaker Cherry PrommawinEvidence transcript

  • Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…

QQuestion and answer

StageD1-C380

An audience member asked, in a question submitted before the event, whether crawl budget is still an SEO priority in 2026 or only relevant for very large sites.

From the audienceEvidence transcript

StageConfirmed by docsD1-C381

Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.

Speaker Cherry PrommawinEvidence transcript

Used byglossary term Crawl budget

  • Repeated by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…

QQuestion and answers

StageD1-C389

Google expects sites to see more crawling overall, because many other services, including AI services, now crawl the web besides Google.

Speaker Cherry PrommawinEvidence transcript

QQuestion and answers

StageConsistent with docsD1-C392

Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal that people are interested in the site and that its content is of high quality.

Speaker Cherry PrommawinEvidence transcript

What Google's documentation says 8

DocsSourceD1-C132

Google's crawl budget guide says each crawler has its own crawl demand, but the crawl capacity limit (hostload) is shared across all crawlers, so high demand from one crawler can reduce the capacity left for others.

“high demand from one crawler can reduce the capacity available for others”

Publisher Google

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

DocsSourceD1-C097

Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.

Publisher Google

  • Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
DocsSourceD1-C101

Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

Publisher Google

Used byrequirement DEV-URL-08glossary term Faceted navigation

  • Extended by D2-C188 Day 2: Hash-fragment links are a problem only where Google should follow them: product, category and language links…
  • Repeated by D2-C284 Day 2: URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot…
DocsSourceD1-C104

HTTP caching for crawlers means supporting conditional requests (ETag with If-None-Match, or Last-Modified with If-Modified-Since) and answering 304 Not Modified when nothing changed. Google's crawling team has said it prefers ETag.

Publisher Google, Search Central blog (9 December 2024)

Used byrequirement DEV-SRV-08

DocsSourceD1-C127

Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

Publisher Google Search Central, Search Central blog (10 September 2019)

Used byrequirements DEV-IDX-02, DEV-IDX-06glossary term nofollow

  • Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…

Analysis by the author 8

AnalysisD1-C095

New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.

Author Ibrahim Anjro

Used byrequirement DEV-URL-10

  • Extended by D2-C686 Day 2: Launch new pages under sections that Google already indexes well, and improve or remove weak sections…
AnalysisD1-C098

On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

Author Ibrahim Anjro

  • Extended by D2-C708 Day 2: The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on…
  • Extended by D2-C714 Day 2: 'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages…
  • Extended by D2-C716 Day 2: Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the…
AnalysisD1-C110

A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

Author Ibrahim Anjro

Used byrequirement DEV-IDX-01

  • Extended by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
  • Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
  • Extended by D2-C849 Day 2: Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can…
  • Extended by D2-C058 Day 2: John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use…
  • Repeated by D2-C850 Day 2: John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
  • Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
AnalysisD1-C382

The stage rule of thumb (under a few thousand URLs crawl budget is unlikely to be a problem) and Google's crawl budget guide (D1-C097: sites with over a million pages that change about weekly, or over 10,000 pages that change daily) leave a middle range where a site should check Search Console's Crawl Stats report before blaming crawl budget for slow indexing.

Author Ibrahim Anjro

AnalysisD1-C386

The stage point that 4xx responses do not affect crawl budget matches Google's documentation that 4xx codes have no effect on crawl rate (D1-C072), with one exception: 429 Too Many Requests counts as a server error and slows crawling like a 5xx (D1-C071, D1-C092).

Author Ibrahim Anjro

AnalysisD1-C419

A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.

Author Ibrahim Anjro

AnalysisD1-C420

The stage remark that crawl demand follows the quality of the site as a whole sits beside the same talk's slide (D1-C094), which falls back to the parent path's aggregate only when a URL's own quality is unknown, and Google's crawl budget guide lists page quality among the demand factors; read it as site quality setting the baseline while known URL-level signals still count.

Author Ibrahim Anjro

  1. Slide D1-C066 Day 1 · How Google thinks about crawl budget

    Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

    extends
    Stage D1-C326 Day 1 · How crawling works

    Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.

  2. Stage D1-C388 Day 1 · How Google thinks about crawl budget

    AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.

    extends
    Slide D1-C038 Day 1 · How Search works and where's AI?

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

  3. Stage D1-C404 Day 1 · Lightning session C: Crawling

    For query parameters, a community speaker advised checking that every parameter in a request is actually used and in the expected order, and otherwise redirecting to the expected URL with only the used parameters in the correct order.

    extends
    Docs D1-C102 Day 1 · How Google thinks about crawl budget

    If faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return a 404 when a filter combination has no results.

  4. Stage D1-C439 Day 1 · Q&A

    Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  5. Stage D1-C546 Day 1 · Q&A

    A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).

    extends
    Stage D1-C374 Day 1 · How Google thinks about crawl budget

    Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.

  6. Slide D2-C020 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  7. Slide D2-C021 Day 2 · Welcome to indexing day!

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

    extends
    Slide D1-C105 Day 1 · How Google thinks about crawl budget

    URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.

  8. Docs D2-C022 Day 2 · Welcome to indexing day!

    Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  9. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C092 Day 1 · How Google thinks about crawl budget

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

  10. Stage D2-C058 Day 2 · Controlling indexing

    John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  11. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Docs D1-C127 Day 1 · How Google thinks about crawl budget

    Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

  12. Docs D2-C069 Day 2 · Controlling indexing

    Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  13. Analysis D2-C188 Day 2 · Lightning session D: Rendering and JavaScript

    Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.

    extends
    Docs D1-C101 Day 1 · How Google thinks about crawl budget

    Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

  14. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  15. Stage D2-C685 Day 2 · Deciding what goes in the index?

    Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  16. Analysis D2-C686 Day 2 · Deciding what goes in the index?

    Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

    extends
    Analysis D1-C095 Day 1 · How Google thinks about crawl budget

    New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.

  17. Stage D2-C695 Day 2 · Deciding what goes in the index?

    Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  18. Stage D2-C697 Day 2 · Deciding what goes in the index?

    Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  19. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Docs D1-C097 Day 1 · How Google thinks about crawl budget

    Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.

  20. Analysis D2-C708 Day 2 · Deciding what goes in the index?

    The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  21. Stage D2-C709 Day 2 · Deciding what goes in the index?

    Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  22. Stage D2-C714 Day 2 · Deciding what goes in the index?

    'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  23. Analysis D2-C716 Day 2 · Deciding what goes in the index?

    Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  24. Stage D2-C846 Day 2 · Welcome to indexing day!

    Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  25. Analysis D2-C849 Day 2 · Welcome to indexing day!

    Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  26. Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  27. Stage D3-C623 Day 3 · How long does it take to..?

    When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  28. Stage D3-C623 Day 3 · How long does it take to..?

    When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.

    extends
    Stage D1-C377 Day 1 · How Google thinks about crawl budget

    The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.

  29. Stage D3-C624 Day 3 · How long does it take to..?

    Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.

    extends
    Stage D1-C375 Day 1 · How Google thinks about crawl budget

    Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

  30. Stage D1-C385 Day 1 · How Google thinks about crawl budget

    4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.

    repeats
    Docs D1-C072 Day 1 · How crawling errors affect Search

    4xx status codes other than 429 have no effect on crawl rate.

  31. Stage D1-C466 Day 1 · Q&A

    Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

    repeats
    Stage D1-C381 Day 1 · How Google thinks about crawl budget

    Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.

  32. Slide D2-C284 Day 2 · What is Google friendly JavaScript

    URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.

    repeats
    Docs D1-C101 Day 1 · How Google thinks about crawl budget

    Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

  33. Stage D2-C850 Day 2 · Controlling indexing

    John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.

    repeats
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  34. Stage D2-C851 Day 2 · Controlling indexing

    When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

    repeats
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.