Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Topic · The index and its signals
Google's index is immense but finite, so index selection, the last step before it, sets thresholds and keeps only documents likely to be useful to users now or later; Google said the system is predictive and relies heavily on machine learning (not in its docs). Quality is ultimately what decides; Google added that a site's or document's overall importance plays a significant role, as with news sites, and that a site or section already satisfying users gets more forgiving treatment for its new URLs (not in its docs). Language and country are weighted so the index is not dominated by English or by the largest countries, and where Google lacks content in a language such as Basque, lower-quality documents are more likely to be selected, which Google presented as an opening for new publishers (said at the event, not in Google's docs). Index selection also applies blocking signals: a noindex rule (a likely reading of the recording), an expired unavailable_after date, soft 404s, non-canonical duplicates, spam signals and other policies, while freshness and explicit content play no big role in the decision. In Search Console, 'Crawled – currently not indexed' is an index selection decision and, Google said, most of the time a quality issue rather than a technical one: the pages are low quality or useless for the index. Day 3 repeated that Google does not index every URL, which is why it has to rank better and better; Google's indexing chart said end-to-end indexing can take months or never happen, for reasons of quality, and Google said URLs not recrawled for a very long time are dropped from the index (not in its docs). Author’s view: set against the hundreds of trillions of URLs Google said it knows and the hundreds of billions of pages its ranking guide puts in the index, only around one known URL in a thousand is indexed, if both figures hold. A second recording of Day 1 added Cherry Prommawin's version: signals calculated during indexing are stored and used both to decide whether a page is indexed and later for ranking, index selection runs after signals are collected and duplicates dropped, and a selected page takes its whole duplicate cluster into the index with it. Day 2's opening Q&A added that very few robots.txt-disallowed URLs are in the index, and that an important disallowed URL might still be indexed without its content. Author’s view: Google's robots.txt guide names links from elsewhere as the reason, so keep important pages crawlable with noindex if they must stay out of Search.
Things in this topic 10
Counts are claims that name the thing. All things
What to do
Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
When a page is selected for Google's index, its whole duplicate cluster goes into the index with it.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no figure given).
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-IDX-01
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-IDX-01
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence 2 slide photos, transcript
Used byrequirements DEV-ERR-01, DEV-ERR-03glossary term Soft 404
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Storage is a second reason for deduplication: Google's storage has many competing uses and storage prices have risen sharply, so the space for any one use is limited and Google has to draw a line somewhere.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Smaller languages have fewer documents on the web and so worse coverage, from an indexing perspective, than large ones, for example Basque compared with Spanish or Portuguese, Hungarian compared with German, and Sicilian compared with Italian.
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Google detects the language of each document and weights language in index selection so that the index is not dominated by one or two languages.
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Without language or country input to index selection, Google's index would become an English index, because so many documents on the web are in English.
“if you just let the index decide what to rank without language or country input, then you would have an English index”
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Google also takes country into account in index selection, so that countries producing less content in a shared language, such as the UK for English or Switzerland for German, are not disadvantaged against the US or Germany, which produce most of it.
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Freshness plays no big role in index selection: whether a page was published 30 years ago or today makes no difference to its chance of being indexed.
“It doesn't play a big role in index selection.”
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Whether a page's content is explicit does not matter for indexing: judged on content type alone, an explicit page has the same chance of being indexed as a major news site's homepage.
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Spam stops a document at indexing: when Google detects that a document is spam, it does not let it into the index.
“If we detect that something is spam, then we're not letting it into the index.”
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
“our index is immense. Like, immense. But it is a finite resource.”
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Google gives two reasons for not indexing every URL it knows: most of them would not be useful to users, and including URLs in the index that users would never see would be an immense investment.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byglossary term Index selection
Index selection aims to index only documents that are useful to users now or potentially in the future.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Index selection is a predictive AI system that relies heavily on machine learning.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
“the index selection system is going to be more forgiving when it sees a new URL from your site”
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-URL-10
Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byglossary term Index selection
Index selection is the last step before documents enter Google's index.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byglossary term Index selection
When Google's coverage of a country is limited, index selection becomes more likely to select documents relevant to that country, even if other signals would suggest otherwise.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
When Google lacks content in a language, such as Basque, index selection becomes more likely to select lower-quality documents in that language.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Under-served languages were presented as an opportunity: where Google's index holds a lot of spam in a language such as Basque, a site that starts publishing in that language can very likely replace that spam with its own content and rank for those keywords (one clause of the reasoning was inaudible).
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
The overall importance of a document or site plays a significant role in index selection.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
News sites were given as the example of importance at work in index selection: they are generally very important on the web and their pages usually get indexed very fast ('indexed' is a likely but not certain reading of the recording).
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.
“focusing on the quality is the most reliable way to get stuff in the index”
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-IDX-08glossary term unavailable_after
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Index selection loads most of Google's spam signals and acts as a gatekeeper so that no spam enters search results.
“index selection acts as a gatekeeper, ensuring that no spam enters our search results”
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Index selection also applies other policies, covering egregious violations and, in a less certain reading of the transcript, content Google is legally not allowed to index.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
'Crawled – currently not indexed' in Search Console is an index selection decision: Google crawled and processed the page but decided not to keep it in the index.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03glossary term Crawled – currently not indexed
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
“most of the time it is actually a quality issue”
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
A site serving a smaller country in a big language, such as British English or Swiss German, should make that country unmistakable (country-code domain or region-specific hreflang, local address, prices and currency) so its pages can benefit from index selection's country balancing instead of competing with the much larger US or German content pool.
Author Ibrahim AnjroAnnotates Day 2, 15:30 · Calculating (some) signals
Changing a page's date or making cosmetic edits will not help an old page get indexed, because freshness plays no big role in index selection; freshness pays off only when ranking for queries that deserve fresh results.
Author Ibrahim AnjroAnnotates Day 2, 15:30 · Calculating (some) signals
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-URL-10
The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-INT-10
Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-IDX-08
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-MON-03
Google's indexing chart says end-to-end indexing can take months or never happen, for reasons of quality.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo
Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.
Speaker GoogleIn Day 3, 12:00 · What are quality updatesEvidence transcript
Google's index drops URLs that have not been recrawled for a very long time, which is the 'never' end of the refresh estimate.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
Google's ranking systems guide speaks of hundreds of billions of pages in its Search index; by the author's arithmetic, against the hundreds of trillions of URLs Google said it knows, only around one known URL in a thousand is indexed, if both figures are taken at face value.
Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?
Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Google detects the language of each document and weights language in index selection so that the index is not dominated by one or two languages.
Google determines a page's language for indexing from the page content, not from a language code in the URL.
Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
Index selection is a predictive AI system that relies heavily on machine learning.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.
Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.
Index selection is the last step before documents enter Google's index.
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
Index selection is the last step before documents enter Google's index.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.
Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.
The unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
Watch the Page indexing report and treat its not-indexed reasons as a work list
Rests on 10 claims, 4 of them in this topic
Use noindex to keep a page out of Search, and leave that URL crawlable
Rests on 12 claims, 2 of them in this topic
Use unavailable_after on pages with a known end date
Rests on 4 claims, 2 of them in this topic
Rests on 4 claims, 2 of them in this topic
Return 404 or 410 for removed and non-existent URLs
Rests on 7 claims, 1 of them in this topic
Never serve error states, empty results or failed data loads with a 200 status
Rests on 10 claims, 1 of them in this topic
Publish machine-translated language versions only after human review and localisation
Rests on 5 claims, 1 of them in this topic