Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Indexing

Tokenization: how text is stored

Google said, in detail not in its documentation, that the Search index holds neither full pages nor sentences but tokens, the smallest searchable units (words, for languages written with spaces), each stored with its position and metadata, with posting lists of the URLs that contain most tokens. The metadata records where a word appeared (header, main content or 'centerpiece', bold, a heading; a further item, heard as title in one recording and italics in another, is left open) and spam signals such as white-on-white text, and snippets are rebuilt from the stored token positions. Thai, Chinese and other languages written without spaces are segmented with statistical models built from web content in that language, and queries go through the same segmenter so they match the index; retrieval then looks up only the query's important words. Day 2 partly contradicts Day 1's slide that Gemini shares tokenization with Search: Gary Illyes said AI-model tokenization is different ('or mostly'), and his slides showed Search keeping 'robots.txt' and 'tl;dr' whole where the AI tokenizer split them into sub-word pieces, as Google's Gemini documentation describes for long words (the slide also showed each token's numeric ID). Author’s view: read it as a shared step with different outputs, and treat leftover hidden keyword blocks as a liability, since spam metadata is stored with the tokens. Day 3 confirmed the mirror on the query side: Google transforms a query into something that can be matched against the index, removing stop words as part of that, while a phrase whose stop words matter is recognised as a whole and its words are indexed together (said at the event, not in Google's docs). Author’s view: a community speaker's picture of Googlebot seeing a page as ones and zeros blends two steps, since crawling fetches the page and tokenization happens later, when it is processed.

Things in this topic 6

Counts are claims that name the thing. All things

What to do

  • Remove leftover hidden text and keyword blocks, because spam metadata such as white-on-white text is stored with the page's tokens.
  • Put key terms in the title, headings and main content rather than the footer, since token metadata records where each word appeared.
  • Do not judge how Search indexes a page from an AI model's sub-word token counts; Search's word tokens are different units.

Day 2: Indexing 20

Shown on screen 2

SlideConsistent with docsD2-C326

In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byglossary term Tokenization

Said on stage 16

StageNot in docsD2-C318

Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byglossary term Tokenization

  • Repeated by D2-C721 Day 2: Google's Search index does not hold the full content of pages; Google said storing full pages and pulling…
  • Extended by D3-C109 Day 3: The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens…
StageNot in docsD2-C320

Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

  • Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
  • Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
StageNot in docsD2-C321

For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

  • Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageNot in docsD2-C323

When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byrequirement DEV-HTM-07

  • Repeated by D2-C720 Day 2: Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the…
StageConsistent with docsD2-C327

Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byglossary term Tokenization

StageNot in docsD2-C721

Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.

“we don't have the full content of the page in our index”

Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript

  • Repeats D2-C318 Day 2: Google does not store the complete sentences or the full HTML of a page in the Search index, because large…
StageNot in docsD2-C724

The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.

“the snippet that you see was reconstructed from these tokens”

Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript

Things
  • Extended by D3-C316 Day 3: Google generates the parts of a text result, such as title link and snippet, from its understanding of the…
StageNot in docsD2-C737

A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript

  • Repeats D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…

Analysis by the author 2

Day 3: Serving: Ranking, Search Console, and Performance 3

Said on stage 2

StageNot in docsD3-C013

Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript

  • Extends D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
  • Extends D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

Analysis by the author 1

AnalysisD3-C109

The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.

Author Ibrahim AnjroAnnotates Day 3, 10:45 · Lightning session K: Facets of quality

  • Extends D2-C318 Day 2: Google does not store the complete sentences or the full HTML of a page in the Search index, because large…

Across days and sessions 12

  1. Stage D2-C325 Day 2 · Understanding what's on a page

    Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

    contradicts
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  2. Stage D3-C011 Day 3 · Making sense of users' queries

    Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.

    extends
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

  3. Stage D3-C013 Day 3 · Making sense of users' queries

    Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

    extends
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  4. Stage D3-C013 Day 3 · Making sense of users' queries

    Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

    extends
    Stage D2-C737 Day 2 · How does the index look like?

    A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

  5. Stage D3-C075 Day 3 · Making sense of users' queries

    At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.

    extends
    Stage D2-C738 Day 2 · How does the index look like?

    At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.

  6. Analysis D3-C109 Day 3 · Lightning session K: Facets of quality

    The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.

    extends
    Stage D2-C318 Day 2 · Understanding what's on a page

    Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

  7. Stage D3-C316 Day 3 · How Search results are born

    Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.

    extends
    Stage D2-C724 Day 2 · How does the index look like?

    The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.

  8. Stage D2-C720 Day 2 · How does the index look like?

    Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.

    repeats
    Stage D2-C323 Day 2 · Understanding what's on a page

    When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

  9. Stage D2-C721 Day 2 · How does the index look like?

    Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.

    repeats
    Stage D2-C318 Day 2 · Understanding what's on a page

    Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

  10. Stage D2-C737 Day 2 · How does the index look like?

    A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

    repeats
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  11. Slide D3-C070 Day 3 · Making sense of users' queries

    Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.

    repeats
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

  12. Stage D3-C074 Day 3 · Making sense of users' queries

    Google's index uses posting lists: for each word, a list of the URLs associated with that word.

    repeats
    Stage D2-C733 Day 2 · How does the index look like?

    For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.

Built on these claims 3

Developer requirements 3

Sources 3