Google said its index, printed on paper, would reach the Moon and back twelve times.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Topic · The index and its signals
Google's Search index, described on Day 1 as big but not limitless, holds the tokens of each page with their metadata and the signals calculated for the document, not the full page content. Retrieval works through posting lists, kept for most tokens, that list the URLs containing each token: Google intersects the lists of the important query words to get an unranked set of candidates, and its public How Search Works explainer likens the index to the index at the back of a book. Google can also retrieve documents through vector embeddings, where closeness to the query's embedding decides what comes back, and both methods work from the page's content. Snippets are rebuilt from the stored tokens and their positions, and AI Overviews and AI Mode use the same index and snippets: fan-out queries go to the Search index and the returned snippets are the material for the AI answer, so a page that forbids snippets cannot be used for them. The token storage, the posting-list details and snippet reconstruction were said at the event and are not in Google's documentation. Day 3 repeated the posting-list model at retrieval: a document can be retrieved only if it contains the query's words or, for embeddings, related concepts. For scale, Google's first index in 1998 held 26 million pages, Google's own figure (25 million was said on stage). Cherry Prommawin said on Day 1 that Google's index, printed on paper, would reach the Moon and back twelve times (said at the event).
Things in this topic 8
Counts are claims that name the thing. All things
What to do
Google said its index, printed on paper, would reach the Moon and back twelve times.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
“we use the snippet as a way of building out the AI Overviews and the AI Mode answers”
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-05
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byglossary term Tokenization
When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byrequirement DEV-HTM-07
The speaker recapped Google's pipeline up to the index: Google crawls pages, processes the fetched documents and then stores them in its index.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byrequirement DEV-HTM-07
Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.
“we don't have the full content of the page in our index”
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.
“the snippet that you see was reconstructed from these tokens”
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
To find relevant pages, Google's serving system relies on posting lists, a long-established information retrieval structure taught in computer science courses, because simply asking for every page that contains a word would not work.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byglossary term Posting list
Google said posting lists, which Google's serving system uses to find the pages that contain a query's words, are not new: they are at least 60 years old (as of 2026).
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byglossary term Posting list
In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byglossary term Posting list
A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.
“just like you build a posting list, you can also build a vector space”
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byglossary term Vector embeddings
In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the user's query decides which documents are returned.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Used byglossary term Vector embeddings
Google said a vector space also holds embeddings for associations the web makes with a page, such as what is known about its author; most of them sit far from typical queries, and a query that names the association may move closer to them.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Both retrieval methods Google described, the long-established posting lists and the newer vector embeddings, work from the content of the page.
Speaker GoogleIn Day 2, 16:15 · How does the index look like?Evidence transcript
Google's snippet documentation says snippets are created automatically, primarily from the page content, to preview the part that best relates to the user's specific search, so one page can get different snippets for different searches; sometimes the meta description is used instead.
Publisher Google Search CentralAnnotates Day 2, 16:15 · How does the index look like?
Used byrequirement DEV-HTM-03
Google's How Search Works site describes the Search index as like the index at the back of a book, with an entry for every word seen on every webpage Google indexes.
“It’s like the index in the back of a book - with an entry for every word seen on every webpage we index.”
Publisher Google Search (How Search Works)Annotates Day 2, 16:15 · How does the index look like?
There is no separate AI index to optimise for: when a page never shows up as a source in AI Overviews or AI Mode, first check that it is indexed, that no nosnippet rule blocks its snippet and that the site is not excluded in Search Console's generative AI setting, the eligibility conditions Google lists.
Author Ibrahim AnjroAnnotates Day 2, 16:15 · How does the index look like?
Used byrequirement DEV-AIF-01
The speaker's 'most of the tokens' is more precise than the public explainer's 'an entry for every word': the explainer simplifies, and the speaker's self-correction suggests some tokens get no posting list, though the speaker did not say which.
Author Ibrahim AnjroAnnotates Day 2, 16:15 · How does the index look like?
Write for both retrieval routes without keyword stuffing: name the page's subject in the plain words people search with, because posting lists match the words on the page, and explain the topic fully, because embedding retrieval matches meaning, so every keyword variation is unnecessary.
Author Ibrahim AnjroAnnotates Day 2, 16:15 · How does the index look like?
Google's index uses posting lists: for each word, a list of the URLs associated with that word.
Speaker Gary IllyesIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
The first condition for retrieving a document is that the query's words, or its concepts in the case of vectors or embeddings, are in the document or related to it.
Speaker Gary IllyesIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
When Google launched in 1997-98 its index held 25 million pages, at that time the biggest index of any search engine.
Speaker GoogleIn Day 3, 12:00 · What are quality updatesEvidence transcript
A 2008 Official Google Blog post says the first Google index, in 1998, already had 26 million pages, and that the index reached one billion pages by 2000.
Publisher Official Google Blog (25 July 2008)Annotates Day 3, 12:00 · What are quality updates
Quote the size of Google's first index as 26 million pages (1998), the figure in Google's own 2008 blog post; the 25 million said on stage is a rounded version of it.
Author Ibrahim AnjroAnnotates Day 3, 12:00 · What are quality updates
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the user's query decides which documents are returned.
Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.
Google said posting lists, which Google's serving system uses to find the pages that contain a query's words, are not new: they are at least 60 years old (as of 2026).
To find relevant pages, Google's serving system relies on posting lists, a long-established information retrieval structure taught in computer science courses, because simply asking for every page that contains a word would not work.
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding how normal queries work explains fan-out queries too.
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.
In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.
At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.
At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.
The first condition for retrieving a document is that the query's words, or its concepts in the case of vectors or embeddings, are in the document or related to it.
Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.
To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
Google called quality the most important of the signals used to order candidates at retrieval: a URL of high quality is more likely to be retrieved from the index for specific queries.
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
Google handles an image used as a search query much like a text query interpreted as an embedding: the image is broken down into vectors (embeddings) that are then searched for in the index.
Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.
Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.
The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.
AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to function and work with the normal text results from Google's index.
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
The speaker recapped Google's pipeline up to the index: Google crawls pages, processes the fetched documents and then stores them in its index.
Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.
When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.
Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
Google's index uses posting lists: for each word, a list of the URLs associated with that word.
For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.
For retrieval, Google uses signals attached individually to each document in the index.
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
Use real heading, title and emphasis elements instead of styling alone
Rests on 2 claims, 2 of them in this topic
Keep pages indexable and snippet-eligible, the whole technical requirement for AI features
Rests on 5 claims, 1 of them in this topic
Give every indexable page a unique, descriptive title element in the server HTML
Rests on 7 claims, 1 of them in this topic
Do not add nosnippet or a short max-snippet unless you accept losing snippets and AI feature use
Rests on 6 claims, 1 of them in this topic