Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

Faceted navigation

Filter and sort URLs multiply into near-infinite combinations and are a classic crawl budget leak. Google prefers blocking them in robots.txt and keeping only item pages and one unfiltered listing crawlable. Day 2 added that non-crawlable link markup, such as onclick handlers and hash pseudo-links, is common in single-page apps, so sites with faceted navigation should check their links for it. Author’s view: hash fragments can deliberately keep filter states out of the crawl, as Google's faceted navigation guide allows, as long as product, category and language links use real URLs. Day 1's second recording added that Google crawls every URL that differs even when it leads to the same content, so infinite URL spaces burn crawl budget; a community speaker advised checking that every parameter in a request is used and in the expected order and redirecting otherwise, and in the Q&A Google suggested working out the normalised URL and redirecting parameter variants to it, which costs crawl budget at first but leaves a clean slate, said log files are worth checking on sites with filter parameters, and noted that many crawling complaints trace back to plugins that generate URLs, such as calendars, whose useless parameter URLs can be handled at the web server.

Based on D1-C099, D1-C100, D1-C101, D1-C102, D1-C103, D2-C186, D2-C188, D1-C378, D1-C404, D1-C445, D1-C444, D1-C446, D1-C539

13 claims · raised in 4 sessions · said or shown on Day 1 and Day 2

Open in Reef mapOpen in Graph

What to do

  • Build facet URLs in one fixed order and drop default values.
  • Return 404 for filter combinations with no results.
  • Disallow internal search and sort parameters.
  • Give product, category and language links real <a href> URLs; use fragments only for filter states you do not want crawled.
  • Check server logs for crawling of filter and parameter URLs, and redirect unused or reordered parameters to the normalised URL.

Day 1: Crawling 11

Shown on screen 3

SlideConsistent with docsD1-C099

Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C378 Day 1: Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same…
SlideConsistent with docsD1-C100

Six faceted-URL problems were shown: the same facets in a different order, irrelevant or conflicting facet combinations, excessive facet selection, facets on paginated series, facets combined with search queries, and optional facets with default values.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo

Used byrequirement DEV-URL-08glossary term Faceted navigation

SlideConsistent with docsD1-C103

Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence 2 slide photos, transcript

Used byrequirements DEV-SRV-08, DEV-URL-08

  • Extended by D1-C379 Day 1: The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages…
  • Extended by D2-C886 Day 2: A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster…

Said on stage 6

StageConsistent with docsD1-C378

Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byrequirements DEV-URL-08, DEV-URL-11

  • Extends D1-C099 Day 1: Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and…
StageConsistent with docsD1-C404

For query parameters, a community speaker advised checking that every parameter in a request is actually used and in the expected order, and otherwise redirecting to the expected URL with only the used parameters in the correct order.

Speaker Tobias SchwarzIn Day 1, 16:20 · Lightning session C: CrawlingEvidence transcript

Things

Used byrequirement DEV-URL-11

  • Extends D1-C102 Day 1: If faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return…
StageConsistent with docsD1-C444

Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
  • Extends D1-C397 Day 1: A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary…
StageConsistent with docsD1-C445

One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-11

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
  • Repeats D1-C403 Day 1: A community speaker advised that a web application compute the expected URL for every request, for example…
StageNot in docsD1-C446

Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
StageConsistent with docsD1-C539

A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Extends D1-C447 Day 1: Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its…
  • Extends D1-C446 Day 1: Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a…

What Google's documentation says 2

DocsSourceD1-C101

Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

Publisher GoogleAnnotates Day 1, 16:00 · How Google thinks about crawl budget

Used byrequirement DEV-URL-08glossary term Faceted navigation

  • Extended by D2-C188 Day 2: Hash-fragment links are a problem only where Google should follow them: product, category and language links…
  • Repeated by D2-C284 Day 2: URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot…
DocsSourceD1-C102

If faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return a 404 when a filter combination has no results.

Publisher GoogleAnnotates Day 1, 16:00 · How Google thinks about crawl budget

Used byrequirement DEV-URL-09

  • Extended by D1-C404 Day 1: For query parameters, a community speaker advised checking that every parameter in a request is actually used…

Day 2: Indexing 2

Said on stage 1

StageConsistent with docsD2-C186

Non-crawlable link markup, such as onclick links and hash pseudo-links, is common in single-page web apps, and sites that use faceted navigation should check their links for it.

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Used byrequirement DEV-URL-02glossary term Single-page app (SPA)

Analysis by the author 1

AnalysisD2-C188

Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.

Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-URL-08

  • Extends D1-C101 Day 1: Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable…

Across days and sessions 10

  1. Stage D1-C378 Day 1 · How Google thinks about crawl budget

    Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

    extends
    Slide D1-C099 Day 1 · How Google thinks about crawl budget

    Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.

  2. Stage D1-C379 Day 1 · How Google thinks about crawl budget

    The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  3. Stage D1-C404 Day 1 · Lightning session C: Crawling

    For query parameters, a community speaker advised checking that every parameter in a request is actually used and in the expected order, and otherwise redirecting to the expected URL with only the used parameters in the correct order.

    extends
    Docs D1-C102 Day 1 · How Google thinks about crawl budget

    If faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return a 404 when a filter combination has no results.

  4. Stage D1-C444 Day 1 · Q&A

    Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

    extends
    Stage D1-C397 Day 1 · Lightning session C: Crawling

    A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.

  5. Stage D1-C539 Day 1 · Q&A

    A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

    extends
    Stage D1-C446 Day 1 · Q&A

    Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

  6. Stage D1-C539 Day 1 · Q&A

    A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

    extends
    Stage D1-C447 Day 1 · Q&A

    Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.

  7. Analysis D2-C188 Day 2 · Lightning session D: Rendering and JavaScript

    Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.

    extends
    Docs D1-C101 Day 1 · How Google thinks about crawl budget

    Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

  8. Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  9. Stage D1-C445 Day 1 · Q&A

    One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

    repeats
    Stage D1-C403 Day 1 · Lightning session C: Crawling

    A community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.

  10. Slide D2-C284 Day 2 · What is Google friendly JavaScript

    URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.

    repeats
    Docs D1-C101 Day 1 · How Google thinks about crawl budget

    Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

Built on these claims 5

Developer requirements 5

Sources 7