Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · The index and its signals

Index selection: deciding what gets indexed

Google's index is immense but finite, so index selection, the last step before it, sets thresholds and keeps only documents likely to be useful to users now or later; Google said the system is predictive and relies heavily on machine learning (not in its docs). Quality is ultimately what decides; Google added that a site's or document's overall importance plays a significant role, as with news sites, and that a site or section already satisfying users gets more forgiving treatment for its new URLs (not in its docs). Language and country are weighted so the index is not dominated by English or by the largest countries, and where Google lacks content in a language such as Basque, lower-quality documents are more likely to be selected, which Google presented as an opening for new publishers (said at the event, not in Google's docs). Index selection also applies blocking signals: a noindex rule (a likely reading of the recording), an expired unavailable_after date, soft 404s, non-canonical duplicates, spam signals and other policies, while freshness and explicit content play no big role in the decision. In Search Console, 'Crawled – currently not indexed' is an index selection decision and, Google said, most of the time a quality issue rather than a technical one: the pages are low quality or useless for the index. Day 3 repeated that Google does not index every URL, which is why it has to rank better and better; Google's indexing chart said end-to-end indexing can take months or never happen, for reasons of quality, and Google said URLs not recrawled for a very long time are dropped from the index (not in its docs). Author’s view: set against the hundreds of trillions of URLs Google said it knows and the hundreds of billions of pages its ranking guide puts in the index, only around one known URL in a thousand is indexed, if both figures hold. A second recording of Day 1 added Cherry Prommawin's version: signals calculated during indexing are stored and used both to decide whether a page is indexed and later for ranking, index selection runs after signals are collected and duplicates dropped, and a selected page takes its whole duplicate cluster into the index with it. Day 2's opening Q&A added that very few robots.txt-disallowed URLs are in the index, and that an important disallowed URL might still be indexed without its content. Author’s view: Google's robots.txt guide names links from elsewhere as the reason, so keep important pages crawlable with noindex if they must stay out of Search.

What to do

  • Treat 'Crawled – currently not indexed' as a quality audit: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and template differences.
  • Improve or remove weak sections, and launch new content where Google already indexes well.
  • Use unavailable_after on pages with a known end date, such as events or job ads, so they drop out automatically.
  • In under-served languages, publish content written for that market, not bulk machine translation, which can count as scaled content abuse.
  • For a smaller country in a big language, make the country unmistakable with a country-code domain or regional hreflang, local address, prices and currency.

Day 1: Crawling 3

Said on stage 3

StageConsistent with docsD1-C205

Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.

Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Extended by D2-C687 Day 2: Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
StageConsistent with docsD1-C211

Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Extended by D2-C682 Day 2: Google's index selection system calculates thresholds and decides which documents are kept and which are…
  • Extended by D2-C688 Day 2: Index selection is the last step before documents enter Google's index.

Day 2: Indexing 43

Said on stage 36

StageConsistent with docsD2-C846

Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
StageConfirmed by docsD2-C334

A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence 2 slide photos, transcript

Used byrequirements DEV-ERR-01, DEV-ERR-03glossary term Soft 404

  • Repeats D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D1-C070 Day 1: Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the…
  • Repeats D1-C355 Day 1: A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found'…
  • Extended by D2-C903 Day 2: A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that…
  • Extended by D2-C700 Day 2: Index selection drops soft 404 pages that were not dropped earlier, for example when a document is…
StageConsistent with docsD2-C344

Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
StageNot in docsD2-C651

Google detects the language of each document and weights language in index selection so that the index is not dominated by one or two languages.

Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript

  • Extends D2-C585 Day 2: Google determines a page's language for indexing from the page content, not from a language code in the URL.
StageConsistent with docsD2-C680

Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

“our index is immense. Like, immense. But it is a finite resource.”

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

  • Repeats D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeats D2-C344 Day 2: Google deduplicates pages because many sites have very many pages and Google's index does not have room for…
  • Repeated by D3-C256 Day 3: Google does not index every URL on the web; because it cannot index everything, it has to rank results better…
StageConsistent with docsD2-C682

Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byglossary term Index selection

  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
StageNot in docsD2-C685

Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

“the index selection system is going to be more forgiving when it sees a new URL from your site”

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byrequirement DEV-URL-10

  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
StageNot in docsD2-C688

Index selection is the last step before documents enter Google's index.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byglossary term Index selection

  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
StageNot in docsD2-C691

Under-served languages were presented as an opportunity: where Google's index holds a lot of spam in a language such as Basque, a site that starts publishing in that language can very likely replace that spam with its own content and rank for those keywords (one clause of the reasoning was inaudible).

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

StageConsistent with docsD2-C695

Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.

“focusing on the quality is the most reliable way to get stuff in the index”

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
StageConsistent with docsD2-C696

Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

StageConsistent with docsD2-C697

Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

  • Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
StageConsistent with docsD2-C698

Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byrequirement DEV-IDX-08glossary term unavailable_after

  • Extends D2-C100 Day 2: The unavailable_after rule lets a page drop out of search results after a set date and time, which suits…
StageConsistent with docsD2-C700

Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
StageConsistent with docsD2-C701

When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
StageConsistent with docsD2-C704

Index selection loads most of Google's spam signals and acts as a gatekeeper so that no spam enters search results.

“index selection acts as a gatekeeper, ensuring that no spam enters our search results”

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

StageConsistent with docsD2-C711

'Crawled – currently not indexed' in Search Console is an index selection decision: Google crawled and processed the page but decided not to keep it in the index.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byrequirement DEV-MON-03glossary term Crawled – currently not indexed

StageNot in docsD2-C714

'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

“most of the time it is actually a quality issue”

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Things

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
StageConsistent with docsD2-C717

Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byrequirement DEV-MON-03

  • Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…

Analysis by the author 7

AnalysisD2-C849

Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!

  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
AnalysisD2-C655

A site serving a smaller country in a big language, such as British English or Swiss German, should make that country unmistakable (country-code domain or region-specific hreflang, local address, prices and currency) so its pages can benefit from index selection's country balancing instead of competing with the much larger US or German content pool.

Author Ibrahim AnjroAnnotates Day 2, 15:30 · Calculating (some) signals

AnalysisD2-C686

Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?

Used byrequirement DEV-URL-10

  • Extends D1-C095 Day 1: New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under…
AnalysisD2-C692

The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.

Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?

Used byrequirement DEV-INT-10

AnalysisD2-C699

Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.

Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?

Used byrequirement DEV-IDX-08

AnalysisD2-C716

Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

Day 3: Serving: Ranking, Search Console, and Performance 4

Shown on screen 1

Said on stage 2

StageConsistent with docsD3-C256

Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.

Speaker GoogleIn Day 3, 12:00 · What are quality updatesEvidence transcript

  • Repeats D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…

Analysis by the author 1

AnalysisD3-C677

Google's ranking systems guide speaks of hundreds of billions of pages in its Search index; by the author's arithmetic, against the hundreds of trillions of URLs Google said it knows, only around one known URL in a thousand is indexed, if both figures are taken at face value.

Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?

Across days and sessions 31

  1. Stage D1-C205 Day 1 · How Search works and where's AI?

    Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  2. Stage D1-C211 Day 1 · How Search works and where's AI?

    Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  3. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    extends
    Stage D1-C070 Day 1 · How crawling errors affect Search

    Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.

  4. Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  5. Stage D2-C651 Day 2 · Calculating (some) signals

    Google detects the language of each document and weights language in index selection so that the index is not dominated by one or two languages.

    extends
    Stage D2-C585 Day 2 · Focusing on Internationalisation and Localisation

    Google determines a page's language for indexing from the page content, not from a language code in the URL.

  6. Stage D2-C682 Day 2 · Deciding what goes in the index?

    Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.

    extends
    Stage D1-C211 Day 1 · How Search works and where's AI?

    Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

  7. Stage D2-C684 Day 2 · Deciding what goes in the index?

    Index selection is a predictive AI system that relies heavily on machine learning.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  8. Stage D2-C685 Day 2 · Deciding what goes in the index?

    Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  9. Analysis D2-C686 Day 2 · Deciding what goes in the index?

    Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

    extends
    Analysis D1-C095 Day 1 · How Google thinks about crawl budget

    New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.

  10. Stage D2-C687 Day 2 · Deciding what goes in the index?

    Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.

    extends
    Stage D1-C205 Day 1 · How Search works and where's AI?

    Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.

  11. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Stage D1-C211 Day 1 · How Search works and where's AI?

    Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

  12. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  13. Stage D2-C695 Day 2 · Deciding what goes in the index?

    Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  14. Stage D2-C697 Day 2 · Deciding what goes in the index?

    Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  15. Stage D2-C698 Day 2 · Deciding what goes in the index?

    Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.

    extends
    Stage D2-C100 Day 2 · Controlling indexing

    The unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.

  16. Stage D2-C700 Day 2 · Deciding what goes in the index?

    Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  17. Stage D2-C700 Day 2 · Deciding what goes in the index?

    Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

    extends
    Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

  18. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  19. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

  20. Stage D2-C714 Day 2 · Deciding what goes in the index?

    'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  21. Analysis D2-C716 Day 2 · Deciding what goes in the index?

    Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  22. Stage D2-C717 Day 2 · Deciding what goes in the index?

    Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

    extends
    Stage D1-C368 Day 1 · How crawling errors affect Search

    Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.

  23. Stage D2-C846 Day 2 · Welcome to indexing day!

    Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  24. Stage D2-C846 Day 2 · Welcome to indexing day!

    Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

    extends
    Slide D2-C021 Day 2 · Welcome to indexing day!

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

  25. Analysis D2-C849 Day 2 · Welcome to indexing day!

    Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  26. Stage D2-C903 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.

    extends
    Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

  27. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    repeats
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  28. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    repeats
    Stage D1-C355 Day 1 · How crawling errors affect Search

    A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.

  29. Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

    repeats
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  30. Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

    repeats
    Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

  31. Stage D3-C256 Day 3 · What are quality updates

    Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.

    repeats
    Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

Built on these claims 7

Developer requirements 7

Sources 14