Same speaker and recording as Calculating (some) signals.
Said on stage 29
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
“our index is immense. Like, immense. But it is a finite resource.”
Speaker GoogleEvidence transcript
- Repeats D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
- Repeats D2-C344 Day 2: Google deduplicates pages because many sites have very many pages and Google's index does not have room for…
- Repeated by D3-C256 Day 3: Google does not index every URL on the web; because it cannot index everything, it has to rank results better…
Google gives two reasons for not indexing every URL it knows: most of them would not be useful to users, and including URLs in the index that users would never see would be an immense investment.
Speaker GoogleEvidence transcript
Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.
Speaker GoogleEvidence transcript
Used byglossary term Index selection
- Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
Index selection aims to index only documents that are useful to users now or potentially in the future.
Speaker GoogleEvidence transcript
Index selection is a predictive AI system that relies heavily on machine learning.
Speaker GoogleEvidence transcript
- Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
“the index selection system is going to be more forgiving when it sees a new URL from your site”
Speaker GoogleEvidence transcript
Used byrequirement DEV-URL-10
- Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
Speaker GoogleEvidence transcript
Used byglossary term Index selection
- Extends D1-C205 Day 1: Signals calculated for a page during indexing are stored in the index and used both to decide whether the…
Index selection is the last step before documents enter Google's index.
Speaker GoogleEvidence transcript
Used byglossary term Index selection
- Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
- Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
When Google's coverage of a country is limited, index selection becomes more likely to select documents relevant to that country, even if other signals would suggest otherwise.
Speaker GoogleEvidence transcript
When Google lacks content in a language, such as Basque, index selection becomes more likely to select lower-quality documents in that language.
Speaker GoogleEvidence transcript
Under-served languages were presented as an opportunity: where Google's index holds a lot of spam in a language such as Basque, a site that starts publishing in that language can very likely replace that spam with its own content and rank for those keywords (one clause of the reasoning was inaudible).
Speaker GoogleEvidence transcript
The overall importance of a document or site plays a significant role in index selection.
Speaker GoogleEvidence transcript
News sites were given as the example of importance at work in index selection: they are generally very important on the web and their pages usually get indexed very fast ('indexed' is a likely but not certain reading of the recording).
Speaker GoogleEvidence transcript
Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.
“focusing on the quality is the most reliable way to get stuff in the index”
Speaker GoogleEvidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.
Speaker GoogleEvidence transcript
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
Speaker GoogleEvidence transcript
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.
Speaker GoogleEvidence transcript
Used byrequirement DEV-IDX-08glossary term unavailable_after
- Extends D2-C100 Day 2: The unavailable_after rule lets a page drop out of search results after a set date and time, which suits…
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
Speaker GoogleEvidence transcript
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
- Extends D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Speaker GoogleEvidence transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
Index selection loads most of Google's spam signals and acts as a gatekeeper so that no spam enters search results.
“index selection acts as a gatekeeper, ensuring that no spam enters our search results”
Speaker GoogleEvidence transcript
Index selection also applies other policies, covering egregious violations and, in a less certain reading of the transcript, content Google is legally not allowed to index.
Speaker GoogleEvidence transcript
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
“The first one is kind of nastier.”
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed
- Extends D1-C097 Day 1: Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over…
- Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.
Speaker GoogleEvidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
'Crawled – currently not indexed' in Search Console is an index selection decision: Google crawled and processed the page but decided not to keep it in the index.
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03glossary term Crawled – currently not indexed
If content quality is even across a site, a page reported as 'Crawled – currently not indexed' may be using a different template that keeps Google from understanding where its content is.
Speaker GoogleEvidence transcript
Used byrequirement DEV-HTM-01
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
“most of the time it is actually a quality issue”
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
The first thing to check for pages reported as 'Crawled – currently not indexed' is whether their quality is on par with the parts of the site that Google does index.
Speaker GoogleEvidence transcript
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03
- Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…
Not-indexed reasons in the Page indexing report include pages excluded by a noindex rule and 'Alternate page with proper canonical tag'.
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03
Analysis by the author 7
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
Author Ibrahim Anjro
Used byrequirement DEV-URL-10
- Extends D1-C095 Day 1: New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under…
The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.
Author Ibrahim Anjro
Used byrequirement DEV-INT-10
Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-08
'Only canonicals end up in search results' as said on stage is a simplification: non-canonical duplicates are dropped from the index, but Google's documentation says an alternate from the same cluster can still be shown in some contexts, such as a mobile version to a mobile user.
Author Ibrahim Anjro
The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.
Author Ibrahim Anjro
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Repeatedly submitting 'Discovered – currently not indexed' URLs does not change why they wait, because the status reflects a crawl-scheduling decision; raise the site's demonstrated quality instead, for example by improving or removing weak pages that are already indexed.
Author Ibrahim Anjro
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
Author Ibrahim Anjro
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.