Google detects the language of each document and weights language in index selection so that the index is not dominated by one or two languages.
Thing · Concept
Index selection
Google's decision whether a crawled and processed page is worth keeping in the index.
- Claims
- 30
- In Google’s docs
- 0
- Said at the event
- 25
- Not in docs
- 13
- Kit items
- 6
Glossary · Index selection
The last step before Google's index: it calculates thresholds and decides which documents are kept, using the signals calculated earlier, such as quality, language, country and spam.
Said at the event 25
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Not in docs 13
Without language or country input to index selection, Google's index would become an English index, because so many documents on the web are in English.
Google also takes country into account in index selection, so that countries producing less content in a shared language, such as the UK for English or Switzerland for German, are not disadvantaged against the US or Germany, which produce most of it.
Freshness plays no big role in index selection: whether a page was published 30 years ago or today makes no difference to its chance of being indexed.
Index selection aims to index only documents that are useful to users now or potentially in the future.
Index selection is a predictive AI system that relies heavily on machine learning.
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
Index selection is the last step before documents enter Google's index.
When Google's coverage of a country is limited, index selection becomes more likely to select documents relevant to that country, even if other signals would suggest otherwise.
When Google lacks content in a language, such as Basque, index selection becomes more likely to select lower-quality documents in that language.
The overall importance of a document or site plays a significant role in index selection.
News sites were given as the example of importance at work in index selection: they are generally very important on the web and their pages usually get indexed very fast ('indexed' is a likely but not certain reading of the recording).
Consistent with docs 12
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
Cherry Prommawin · Day 1 · How Search works and where's AI?
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.
Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Index selection loads most of Google's spam signals and acts as a gatekeeper so that no spam enters search results.
Index selection also applies other policies, covering egregious violations and, in a less certain reading of the transcript, content Google is legally not allowed to index.
'Crawled – currently not indexed' in Search Console is an index selection decision: Google crawled and processed the page but decided not to keep it in the index.
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
Press and analysis 5
A site serving a smaller country in a big language, such as British English or Swiss German, should make that country unmistakable (country-code domain or region-specific hreflang, local address, prices and currency) so its pages can benefit from index selection's country balancing instead of competing with the much larger US or German content pool.
Ibrahim Anjro · Day 2 · Calculating (some) signals
Changing a page's date or making cosmetic edits will not help an old page get indexed, because freshness plays no big role in index selection; freshness pays off only when ranking for queries that deserve fresh results.
Ibrahim Anjro · Day 2 · Calculating (some) signals
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
Ibrahim Anjro · Day 2 · Deciding what goes in the index?
The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.
Ibrahim Anjro · Day 2 · Deciding what goes in the index?
Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.
Ibrahim Anjro · Day 2 · Deciding what goes in the index?
Built on these claims 6
Kit items about Index selection: their own words name it, or several of the claims they rest on do.
Developer requirements 3
Use unavailable_after on pages with a known end date
Rests on 4 claims, 2 of them naming Index selection; its own words name Index selection
Watch the Page indexing report and treat its not-indexed reasons as a work list
Rests on 10 claims, 2 of them naming Index selection; its own words name Index selection
Rests on 4 claims, 2 of them naming Index selection; its own words name Index selection
Also inglossary terms Index selection, Crawled – currently not indexed, unavailable_after
Connected things 10
Relations
- Affected by noindex
2 claims, 2 documented
Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
- Affected by Soft 404
2 claims, 2 documented
Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
Most often named with it
Things named in the same claim, with the number of claims they share.
- noindex 2
- Page indexing report 2
- Search Console 2
- Soft 404 2
- Duplicate content 1
- hreflang 1
- Rendering 1
- Scaled content abuse 1
Topics that feature it
- Index selection: deciding what gets indexed 29
- Signals calculated during indexing 4
- Country targeting 3
- Page language and language versions 3
- Robots meta tags and noindex 3
- Crawl demand inherited by path 2
- Indexing statuses in Search Console 2
- Processing: from fetched page to index 2
- Spam detection and SpamBrain 2
- AI systems already inside Search 1
- Canonical selection 1
- Content for people 1
- Duplicate content 1
- Localisation beyond translation 1
- Soft 404s 1
- The pipeline: crawling, indexing, serving 1