Day 2: Indexing 20
Shown on screen 2
In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byglossary term Tokenization
Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence 2 slide photos
Said on stage 16
When Google processes a page for indexing, it gives words different weights depending on the part of the page where they appear.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
Used byrequirement DEV-HTM-02
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byglossary term Tokenization
- Repeated by D2-C721 Day 2: Google's Search index does not hold the full content of pages; Google said storing full pages and pulling…
- Extended by D3-C109 Day 3: The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens…
For languages written with spaces between words, such as English and German, Search tokenization splits a sentence into its individual words.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
- Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
- Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
- Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
- Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
Gary Illyes said a colleague, John, would cover how Google interprets the words of a query on the morning of Day 3.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byrequirement DEV-HTM-07
- Repeated by D2-C720 Day 2: Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the…
Google also stores spam metadata with the tokens of a page, for example that text was white on a white background, so ranking can use that information.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
Used byrequirement DEV-HTM-06
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
- Contradicts D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byglossary term Tokenization
Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byrequirement DEV-HTM-07
- Repeats D2-C323 Day 2: When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the…
Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.
“we don't have the full content of the page in our index”
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
- Repeats D2-C318 Day 2: Google does not store the complete sentences or the full HTML of a page in the Search index, because large…
The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.
“the snippet that you see was reconstructed from these tokens”
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
- Extended by D3-C316 Day 3: Google generates the parts of a text result, such as title link and snippet, from its understanding of the…
For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byglossary term Posting list
- Repeated by D3-C074 Day 3: Google's index uses posting lists: for each word, a list of the URLs associated with that word.
A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
- Repeats D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
- Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
- Extended by D3-C075 Day 3: At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the…
Analysis by the author 2
Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.
Author Ibrahim AnjroAnnotates Day 2, 11:30 · Understanding what's on a page
Hidden text is recorded at token level: Google stores spam metadata such as white-on-white text with the tokens, so leftover hidden keyword blocks are a liability, not neutral clutter, and should be removed.
Author Ibrahim AnjroAnnotates Day 2, 11:30 · Understanding what's on a page
Used byrequirement DEV-HTM-06