For languages written with spaces between words, such as English and German, Search tokenization splits a sentence into its individual words.
Gary Illyes · Day 2 · Understanding what's on a page
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Thing · Concept
Splitting text into tokens (words or word parts) for the index or for an AI model.
Glossary · Tokenization
Splitting text into small units. Google Search stores words (tokens) with metadata, while AI models split text into sub-word pieces mapped to numeric IDs.
Documented in
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
For languages written with spaces between words, such as English and German, Search tokenization splits a sentence into its individual words.
Gary Illyes · Day 2 · Understanding what's on a page
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Gary Illyes · Day 2 · Understanding what's on a page
Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.
Gary Illyes · Day 2 · Understanding what's on a page
Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.
A community speaker compared Googlebot to an inspector who cannot enter the shop and only looks through the window, seeing the page as ones and zeros, which the speaker linked to tokenization.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.
Gary Illyes · Day 2 · Understanding what's on a page
Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.
Ibrahim Anjro · Day 2 · Understanding what's on a page
The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.
Ibrahim Anjro · Day 3 · Lightning session K: Facets of quality
Kit items about Tokenization: their own words name it, or several of the claims they rest on do.
Use real heading, title and emphasis elements instead of styling alone
Rests on 2 claims, 1 of them naming Tokenization
Also inglossary term Tokenization
Things named in the same claim, with the number of claims they share.