Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 2 · Thursday 1 October 2026 · 16:15

How does the index look like?

Speaker Google

TalkCoverageTranscript

The speaker refers back to the earlier talk on tokenization (Understanding what's on a page). No slides were photographed.

Said on stage 18

StageNot in docsD2-C721

Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.

“we don't have the full content of the page in our index”

Speaker GoogleEvidence transcript

  • Repeats D2-C318 Day 2: Google does not store the complete sentences or the full HTML of a page in the Search index, because large…
StageConsistent with docsD2-C722

Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.

Speaker GoogleEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeated by D3-C076 Day 3: For retrieval, Google uses signals attached individually to each document in the index.
  • Extended by D3-C079 Day 3: To order candidates at retrieval, Google uses signals collected during indexing, and the first two are…
  • Extended by D3-C083 Day 3: Google called quality the most important of the signals used to order candidates at retrieval: a URL of high…
StageNot in docsD2-C724

The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.

“the snippet that you see was reconstructed from these tokens”

Speaker GoogleEvidence transcript

Things
  • Extended by D3-C316 Day 3: Google generates the parts of a text result, such as title link and snippet, from its understanding of the…
StageConsistent with docsD2-C726

AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.

Speaker GoogleEvidence transcript

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
  • Extended by D3-C325 Day 3: AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to…
  • Repeated by D3-C702 Day 3: AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have…
StageConsistent with docsD2-C727

Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

Speaker GoogleEvidence transcript

  • Extends D1-C053 Day 1: Query fan-out means running several related searches at once to gather more results; a question about lawn…
  • Extends D1-C051 Day 1: Three reasons were given: generative AI features are built directly on the core ranking systems, query…
  • Extends D2-C073 Day 2: John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page…
  • Extends D1-C172 Day 1: Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further…
  • Extended by D3-C059 Day 3: Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding…
StageConsistent with docsD2-C732

To find relevant pages, Google's serving system relies on posting lists, a long-established information retrieval structure taught in computer science courses, because simply asking for every page that contains a word would not work.

Speaker GoogleEvidence transcript

Used byglossary term Posting list

  • Extended by D2-C824 Day 2: Google said posting lists, which Google's serving system uses to find the pages that contain a query's words…
StageNot in docsD2-C824

Google said posting lists, which Google's serving system uses to find the pages that contain a query's words, are not new: they are at least 60 years old (as of 2026).

Speaker GoogleEvidence transcript

  • Extends D2-C732 Day 2: To find relevant pages, Google's serving system relies on posting lists, a long-established information…
StageNot in docsD2-C736

In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.

Speaker GoogleEvidence transcript

Used byglossary term Posting list

  • Extended by D3-C075 Day 3: At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the…
StageNot in docsD2-C737

A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

Speaker GoogleEvidence transcript

  • Repeats D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageNot in docsD2-C738

At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.

Speaker GoogleEvidence transcript

  • Extended by D3-C075 Day 3: At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the…
StageConsistent with docsD2-C740

Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.

“just like you build a posting list, you can also build a vector space”

Speaker GoogleEvidence transcript

Used byglossary term Vector embeddings

  • Extended by D3-C077 Day 3: The first condition for retrieving a document is that the query's words, or its concepts in the case of…
  • Extended by D3-C311 Day 3: Google handles an image used as a search query much like a text query interpreted as an embedding: the image…
StageConsistent with docsD2-C741

In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the user's query decides which documents are returned.

Speaker GoogleEvidence transcript

Used byglossary term Vector embeddings

  • Extends D1-C129 Day 1: Google's guide says creating separate content for every variation of how people might search, including…
StageNot in docsD2-C744

Google said a vector space also holds embeddings for associations the web makes with a page, such as what is known about its author; most of them sit far from typical queries, and a query that names the association may move closer to them.

Speaker GoogleEvidence transcript

StageD2-C746

Google said, hedging with 'I think', that because retrieval is still based on content, the content mantra Google started about 25 years ago (as of 2026) still stands.

“this mantra that we started 25 years ago or whatever still stands, whether we like it or not”

Speaker GoogleEvidence transcript

What Google's documentation says 4

DocsSourceD2-C725

Google's snippet documentation says snippets are created automatically, primarily from the page content, to preview the part that best relates to the user's specific search, so one page can get different snippets for different searches; sometimes the meta description is used instead.

Publisher Google Search Central

Things

Used byrequirement DEV-HTM-03

DocsSourceD2-C728

Google's robots meta tag specification says the nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI Mode, and max-snippet also limits how much of it may be used that way.

Publisher Google Search Central

Used byrequirement DEV-IDX-05glossary term max-snippet

DocsSourceD2-C729

Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several data sources, which for AI Mode include the Knowledge Graph and shopping data as well as web content (AI Mode launch post, March 2025).

Publisher Google Search Central, Google blog (5 March 2025)

Used byglossary term Query fan-out

  • Extends D1-C053 Day 1: Query fan-out means running several related searches at once to gather more results; a question about lawn…
  • Extended by D3-C061 Day 3: Google tries to make fan-out queries distinct from each other for better coverage, avoiding asking the same…
DocsSourceD2-C734

Google's How Search Works site describes the Search index as like the index at the back of a book, with an entry for every word seen on every webpage Google indexes.

“It’s like the index in the back of a book - with an entry for every word seen on every webpage we index.”

Publisher Google Search (How Search Works)

Analysis by the author 4

AnalysisD2-C730

There is no separate AI index to optimise for: when a page never shows up as a source in AI Overviews or AI Mode, first check that it is indexed, that no nosnippet rule blocks its snippet and that the site is not excluded in Search Console's generative AI setting, the eligibility conditions Google lists.

Author Ibrahim Anjro

Used byrequirement DEV-AIF-01

AnalysisD2-C731

Treat nosnippet, data-nosnippet and max-snippet as AI visibility settings too: Google lists them as the controls for content in AI features, and if AI answers are built from index snippets, as Google said on stage, a blocked or shortened snippet leaves AI Overviews and AI Mode less to use.

Author Ibrahim Anjro

AnalysisD2-C735

The speaker's 'most of the tokens' is more precise than the public explainer's 'an entry for every word': the explainer simplifies, and the speaker's self-correction suggests some tokens get no posting list, though the speaker did not say which.

Author Ibrahim Anjro

AnalysisD2-C742

Write for both retrieval routes without keyword stuffing: name the page's subject in the plain words people search with, because posting lists match the words on the page, and explain the topic fully, because embedding retrieval matches meaning, so every keyword variation is unnecessary.

Author Ibrahim Anjro

  1. Stage D2-C722 Day 2 · How does the index look like?

    Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  2. Stage D2-C726 Day 2 · How does the index look like?

    AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.

    extends
    Slide D1-C038 Day 1 · How Search works and where's AI?

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

  3. Stage D2-C727 Day 2 · How does the index look like?

    Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

    extends
    Slide D1-C051 Day 1 · How Search works and where's AI?

    Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.

  4. Stage D2-C727 Day 2 · How does the index look like?

    Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

    extends
    Docs D1-C053 Day 1 · How Search works and where's AI?

    Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.

  5. Stage D2-C727 Day 2 · How does the index look like?

    Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

    extends
    Stage D1-C172 Day 1 · Welcome and opening keynotes

    Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.

  6. Stage D2-C727 Day 2 · How does the index look like?

    Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

    extends
    Stage D2-C073 Day 2 · Controlling indexing

    John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.

  7. Docs D2-C729 Day 2 · How does the index look like?

    Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several data sources, which for AI Mode include the Knowledge Graph and shopping data as well as web content (AI Mode launch post, March 2025).

    extends
    Docs D1-C053 Day 1 · How Search works and where's AI?

    Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.

  8. Stage D2-C741 Day 2 · How does the index look like?

    In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the user's query decides which documents are returned.

    extends
    Docs D1-C129 Day 1 · How Search works and where's AI?

    Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.

  9. Stage D3-C013 Day 3 · Making sense of users' queries

    Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

    extends
    Stage D2-C737 Day 2 · How does the index look like?

    A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

  10. Stage D3-C059 Day 3 · Making sense of users' queries

    Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding how normal queries work explains fan-out queries too.

    extends
    Stage D2-C727 Day 2 · How does the index look like?

    Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

  11. Stage D3-C061 Day 3 · Making sense of users' queries

    Google tries to make fan-out queries distinct from each other for better coverage, avoiding asking the same question several times, which would return the same answers.

    extends
    Docs D2-C729 Day 2 · How does the index look like?

    Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several data sources, which for AI Mode include the Knowledge Graph and shopping data as well as web content (AI Mode launch post, March 2025).

  12. Stage D3-C075 Day 3 · Making sense of users' queries

    At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.

    extends
    Stage D2-C736 Day 2 · How does the index look like?

    In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.

  13. Stage D3-C075 Day 3 · Making sense of users' queries

    At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.

    extends
    Stage D2-C738 Day 2 · How does the index look like?

    At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.

  14. Stage D3-C077 Day 3 · Making sense of users' queries

    The first condition for retrieving a document is that the query's words, or its concepts in the case of vectors or embeddings, are in the document or related to it.

    extends
    Stage D2-C740 Day 2 · How does the index look like?

    Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.

  15. Stage D3-C079 Day 3 · Making sense of users' queries

    To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.

    extends
    Stage D2-C722 Day 2 · How does the index look like?

    Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.

  16. Stage D3-C083 Day 3 · Making sense of users' queries

    Google called quality the most important of the signals used to order candidates at retrieval: a URL of high quality is more likely to be retrieved from the index for specific queries.

    extends
    Stage D2-C722 Day 2 · How does the index look like?

    Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.

  17. Stage D3-C311 Day 3 · How Search results are born

    Google handles an image used as a search query much like a text query interpreted as an embedding: the image is broken down into vectors (embeddings) that are then searched for in the index.

    extends
    Stage D2-C740 Day 2 · How does the index look like?

    Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.

  18. Stage D3-C316 Day 3 · How Search results are born

    Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.

    extends
    Stage D2-C724 Day 2 · How does the index look like?

    The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.

  19. Stage D3-C325 Day 3 · How Search results are born

    AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to function and work with the normal text results from Google's index.

    extends
    Stage D2-C726 Day 2 · How does the index look like?

    AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.

  20. Stage D2-C719 Day 2 · How does the index look like?

    The speaker recapped Google's pipeline up to the index: Google crawls pages, processes the fetched documents and then stores them in its index.

    repeats
    Slide D1-C036 Day 1 · How Search works and where's AI?

    Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.

  21. Stage D2-C720 Day 2 · How does the index look like?

    Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.

    repeats
    Stage D2-C323 Day 2 · Understanding what's on a page

    When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

  22. Stage D2-C721 Day 2 · How does the index look like?

    Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.

    repeats
    Stage D2-C318 Day 2 · Understanding what's on a page

    Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

  23. Stage D2-C737 Day 2 · How does the index look like?

    A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

    repeats
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  24. Stage D3-C074 Day 3 · Making sense of users' queries

    Google's index uses posting lists: for each word, a list of the URLs associated with that word.

    repeats
    Stage D2-C733 Day 2 · How does the index look like?

    For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.

  25. Stage D3-C076 Day 3 · Making sense of users' queries

    For retrieval, Google uses signals attached individually to each document in the index.

    repeats
    Stage D2-C722 Day 2 · How does the index look like?

    Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.

  26. Stage D3-C702 Day 3 · Wrapping all up: AI, Search, and making sense of everything.

    AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.

    repeats
    Stage D2-C726 Day 2 · How does the index look like?

    AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.