Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Topic · How Search works
The pipeline: crawling, indexing, serving
Search works in three stages, crawling, indexing and serving, and AI Mode and AI Overviews share the crawling and indexing stages with classic Search, adding grounding on the index and query fan-out only at serving; Day 2 repeated that they are a different experience of the same content. Day 2 filled in the middle: a processing stage between the crawler and the index runs HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection, and sends the links it finds back to the crawl queue. Google's how-Search-works guide describes rendering as part of the crawl, while the event slides drew it inside processing. Gemini is not part of Search but uses Google's crawlers and grounds on the Search index; Day 2 qualified the shared tokenization mentioned on Day 1, showing that Gemini's tokenizer splits text differently from the one Search uses. Day 3 drew the serving stage: query understanding and retrieval lead into the index, then ranking and search features lead back to the user, and AI features send their fan-out queries back through the same search engine. Google said its processes depend on each other, crawling, then indexing, then serving, so a change such as a new canonical cannot be processed until the page is recrawled, and that AI Overviews and AI Mode are built on Search infrastructure 25 to 30 years old with very few processes of their own. Day 1's second recording added that crawling and indexing happen before anyone searches while serving and ranking happen in real time, and that AI Overviews and AI Mode may have extra processes of their own like any search feature but share the bulk of their processing with normal Search.
Treat crawlability, indexability and clear canonical signals as the basis for AI Overviews and AI Mode as well; there is no separate AI pipeline to optimise for.
Debug a missing page stage by stage: was it crawled, rendered, folded into a duplicate cluster, or dropped at index selection?
Make important links real <a href> links in the server HTML, because links found during processing are what feed the crawl queue.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
AI Overviews and AI Mode may have extra processes of their own, like any other search feature, but the bulk of their processing is the same as for normal Search.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
Google's guide to how Search works says rendering happens during the crawl, and describes indexing as analysing a page's text, key tags and attributes such as title elements and alt attributes, images and videos, and deciding whether the page is a duplicate or the canonical.
Google's opening slide for serving day showed serving as the third stage after crawling and indexing, and drew Google's serving infrastructure as query understanding and retrieval leading into the index, then ranking and search features leading back to the user, for the example query 'Where to eat jamon'.
Google's slide on how LLM features with grounding generally work showed a query going to both the search engine and an LLM, the search engine's results going to the LLM, the LLM generating fan-out queries that go back to the search engine, and the LLM returning answers with links.
Google's serving diagram shows the query passing through query understanding and retrieval to the index, then back through ranking and search features to the user, with the Search Features step highlighted (example query 'Where to eat orange').
Google's processes depend on each other: crawling happens first, then indexing, then serving, so a change such as a new canonical cannot be processed until the page has been recrawled.
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
AI Overviews and AI Mode may have extra processes of their own, like any other search feature, but the bulk of their processing is the same as for normal Search.
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements.
To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached system, kept separate because rendering is time-consuming and computationally expensive.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.
Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results, so what works for web results also works for the AI features.
Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.
Among the many signals Google calculates during indexing, the ones singled out as having large effects on search results were country, language, freshness, SafeSearch and spam.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
Google's slide on how LLM features with grounding generally work showed a query going to both the search engine and an LLM, the search engine's results going to the LLM, the LLM generating fan-out queries that go back to the search engine, and the LLM returning answers with links.
Google's closing slide said AI on Google is just SEO: AI features on Google Search use exactly the same processes as traditional results, so no new acronym is needed, as none was for mobile-first indexing or structured data.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
Google's serving diagram shows the query passing through query understanding and retrieval to the index, then back through ranking and search features to the user, with the Search Features step highlighted (example query 'Where to eat orange').
Google's opening slide for serving day showed serving as the third stage after crawling and indexing, and drew Google's serving infrastructure as query understanding and retrieval leading into the index, then ranking and search features leading back to the user, for the example query 'Where to eat jamon'.
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.