Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 2 · Thursday 1 October 2026 · 15:40

Deciding what goes in the index?

Speaker Google

TalkCoverageTranscript

Same speaker and recording as Calculating (some) signals.

Said on stage 29

StageConsistent with docsD2-C680

Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

“our index is immense. Like, immense. But it is a finite resource.”

Speaker GoogleEvidence transcript

  • Repeats D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeats D2-C344 Day 2: Google deduplicates pages because many sites have very many pages and Google's index does not have room for…
  • Repeated by D3-C256 Day 3: Google does not index every URL on the web; because it cannot index everything, it has to rank results better…
StageNot in docsD2-C681

Google gives two reasons for not indexing every URL it knows: most of them would not be useful to users, and including URLs in the index that users would never see would be an immense investment.

Speaker GoogleEvidence transcript

StageConsistent with docsD2-C682

Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.

Speaker GoogleEvidence transcript

Used byglossary term Index selection

  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
StageNot in docsD2-C685

Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

“the index selection system is going to be more forgiving when it sees a new URL from your site”

Speaker GoogleEvidence transcript

Used byrequirement DEV-URL-10

  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
StageNot in docsD2-C688

Index selection is the last step before documents enter Google's index.

Speaker GoogleEvidence transcript

Used byglossary term Index selection

  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
StageNot in docsD2-C694

News sites were given as the example of importance at work in index selection: they are generally very important on the web and their pages usually get indexed very fast ('indexed' is a likely but not certain reading of the recording).

Speaker GoogleEvidence transcript

StageConsistent with docsD2-C695

Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.

“focusing on the quality is the most reliable way to get stuff in the index”

Speaker GoogleEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
StageConsistent with docsD2-C696

Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.

Speaker GoogleEvidence transcript

StageConsistent with docsD2-C697

Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).

Speaker GoogleEvidence transcript

  • Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
StageConsistent with docsD2-C698

Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.

Speaker GoogleEvidence transcript

Used byrequirement DEV-IDX-08glossary term unavailable_after

  • Extends D2-C100 Day 2: The unavailable_after rule lets a page drop out of search results after a set date and time, which suits…
StageConsistent with docsD2-C700

Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

Speaker GoogleEvidence transcript

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
StageConsistent with docsD2-C701

When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

Speaker GoogleEvidence transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
StageNot in docsD2-C706

'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

“The first one is kind of nastier.”

Speaker GoogleEvidence transcript

Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed

  • Extends D1-C097 Day 1: Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over…
  • Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
StageConsistent with docsD2-C709

Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.

Speaker GoogleEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
StageNot in docsD2-C714

'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

“most of the time it is actually a quality issue”

Speaker GoogleEvidence transcript

Things

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
StageConsistent with docsD2-C717

Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

Speaker GoogleEvidence transcript

Used byrequirement DEV-MON-03

  • Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…

What Google's documentation says 3

DocsSourceD2-C702

Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.

Publisher Google Search Central

DocsSourceD2-C707

Google's Page indexing report help says a 'Discovered – currently not indexed' page was found but not crawled yet, typically because Google wanted to crawl it but expected the crawl to overload the site, so it rescheduled the crawl.

Publisher Google Search Console Help

Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed

Analysis by the author 7

AnalysisD2-C686

Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

Author Ibrahim Anjro

Used byrequirement DEV-URL-10

  • Extends D1-C095 Day 1: New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under…
AnalysisD2-C692

The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.

Author Ibrahim Anjro

Used byrequirement DEV-INT-10

AnalysisD2-C699

Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.

Author Ibrahim Anjro

Used byrequirement DEV-IDX-08

AnalysisD2-C703

'Only canonicals end up in search results' as said on stage is a simplification: non-canonical duplicates are dropped from the index, but Google's documentation says an alternate from the same cluster can still be shown in some contexts, such as a mobile version to a mobile user.

Author Ibrahim Anjro

AnalysisD2-C708

The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

Author Ibrahim Anjro

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
AnalysisD2-C710

Repeatedly submitting 'Discovered – currently not indexed' URLs does not change why they wait, because the status reflects a crawl-scheduling decision; raise the site's demonstrated quality instead, for example by improving or removing weak pages that are already indexed.

Author Ibrahim Anjro

AnalysisD2-C716

Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

Author Ibrahim Anjro

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
  1. Stage D2-C682 Day 2 · Deciding what goes in the index?

    Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.

    extends
    Stage D1-C211 Day 1 · How Search works and where's AI?

    Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

  2. Stage D2-C684 Day 2 · Deciding what goes in the index?

    Index selection is a predictive AI system that relies heavily on machine learning.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  3. Stage D2-C685 Day 2 · Deciding what goes in the index?

    Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  4. Analysis D2-C686 Day 2 · Deciding what goes in the index?

    Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

    extends
    Analysis D1-C095 Day 1 · How Google thinks about crawl budget

    New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.

  5. Stage D2-C687 Day 2 · Deciding what goes in the index?

    Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.

    extends
    Stage D1-C205 Day 1 · How Search works and where's AI?

    Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.

  6. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Stage D1-C211 Day 1 · How Search works and where's AI?

    Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

  7. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  8. Stage D2-C695 Day 2 · Deciding what goes in the index?

    Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  9. Stage D2-C697 Day 2 · Deciding what goes in the index?

    Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  10. Stage D2-C698 Day 2 · Deciding what goes in the index?

    Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.

    extends
    Stage D2-C100 Day 2 · Controlling indexing

    The unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.

  11. Stage D2-C700 Day 2 · Deciding what goes in the index?

    Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  12. Stage D2-C700 Day 2 · Deciding what goes in the index?

    Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

    extends
    Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

  13. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  14. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

  15. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  16. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Docs D1-C097 Day 1 · How Google thinks about crawl budget

    Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.

  17. Analysis D2-C708 Day 2 · Deciding what goes in the index?

    The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  18. Stage D2-C709 Day 2 · Deciding what goes in the index?

    Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  19. Stage D2-C714 Day 2 · Deciding what goes in the index?

    'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  20. Analysis D2-C716 Day 2 · Deciding what goes in the index?

    Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  21. Stage D2-C717 Day 2 · Deciding what goes in the index?

    Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

    extends
    Stage D1-C368 Day 1 · How crawling errors affect Search

    Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.

  22. Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

    repeats
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  23. Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

    repeats
    Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

  24. Stage D3-C256 Day 3 · What are quality updates

    Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.

    repeats
    Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.