Area
How Search works
The pipeline every page goes through, from the crawl queue through processing and the index to serving.
Topics in this area 3
21 claims · 12 sessions
The pipeline: crawling, indexing, serving
Search works in three stages, crawling, indexing and serving, and AI Mode and AI Overviews share the crawling and indexing stages with classic Search, adding grounding on the index and query fan-out only at serving; Day 2 repeated that they are a different experience of the same content. Day 2 filled in the middle: a processing stage between the crawler and the index runs HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection, and sends the links it finds back to the crawl queue. Google's how-Search-works guide describes rendering as part of the crawl, while the event slides drew it inside processing. Gemini is not part of Search but uses Google's crawlers and grounds on the Search index; Day 2 qualified the shared tokenization mentioned on Day 1, showing that Gemini's tokenizer splits text differently from the one Search uses. Day 3 drew the serving stage: query understanding and retrieval lead into the index, then ranking and search features lead back to the user, and AI features send their fan-out queries back through the same search engine. Google said its processes depend on each other, crawling, then indexing, then serving, so a change such as a new canonical cannot be processed until the page is recrawled, and that AI Overviews and AI Mode are built on Search infrastructure 25 to 30 years old with very few processes of their own. Day 1's second recording added that crawling and indexing happen before anyone searches while serving and ranking happen in real time, and that AI Overviews and AI Mode may have extra processes of their own like any search feature but share the bulk of their processing with normal Search.
Day 1Day 2Day 3
6 claims · 5 sessions
Ranking signals by result type
A Day 1 slide said signals differ by result type: web pages rely on text, links and passages; images on resolution, colour and associated text; news on freshness, originality and diversity; local on location, type, rating, reviews and hours; video on language and text from speech. Google's documentation does not list signals this way. Day 2 expanded two of them: the text around an image is critical context for understanding and ranking it, so alt text alone is not enough, which is consistent with Google's image SEO guide; and freshness counts for queries that deserve fresh results, such as a breaking local event, as Google's ranking systems guide confirms. Google added that some freshness signals are calculated during indexing and used in ranking. Day 3's quality talk repeated the split by result type: text, links and passages for web pages and, probably, freshness, diversity and originality for news. A community speaker said Google loves fresh content (not in Google's docs). Author’s view: Google's 'query deserves freshness' systems are narrower than a general preference for fresh content, so refresh pages whose queries expect current information and judge other refreshes by quality.
Day 1Day 2Day 3
38 claims · 2 sessions
How long Google's processes take
Google shared estimates from internal analysis of how long its processes take, as information rather than targets: discovering a new URL takes about 20 hours on average and recrawling a known URL about 30 days, while indexing a document takes about 1.5 hours on average (probably lowered, Google said, by the many news sites). A robots.txt change is picked up in about 24 hours, as Google's docs say, and a sitemap is processed in about 24 hours. After a 404 or noindex a page usually leaves the index within one to three weeks, a canonical change also takes one to three weeks, a site move one to three months and a Search Console removal about 2 hours; snippet and title changes take 1-2 days, removing a manual action 1-2 weeks, and a core update reaches a site in 2 to 4 weeks. Most of these figures are not in Google's documentation, and each step waits for the one before it: a change cannot be processed until the page is recrawled. The robots.txt figure matches Day 1's documented 24-hour cache, and Google's help gives similar ranges elsewhere: a temporary removal usually takes up to a day and lasts about six months, and title link changes a few days to a few weeks. Other estimates said at the event: images are indexed in hours to two days, structured data changes are picked up within hours to two weeks, and a spam update reaches sites within 1-2 days of its rollout. On Day 1 Gary Illyes advised news sites to track their own figure, the time from publishing a URL to Googlebot's first crawl, and to investigate only a rising trend, adding that for fresh stories something like two hours is probably not great (said at the event).
Day 1Day 3
Across days 33
- Stage D2-C325 Day 2 · Understanding what's on a page
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
contradictsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- D2-C026 Day 2 · How is HTML interpreted
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
extendsD1-C036 Day 1 · How Search works and where's AI?Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- D2-C026 Day 2 · How is HTML interpreted
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C073 Day 2 · Controlling indexing
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Docs D2-C074 Day 2 · Controlling indexing
Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C118 Day 2 · Lightning session D: Rendering and JavaScript
To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C120 Day 2 · Lightning session D: Rendering and JavaScript
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C344 Day 2 · Handling web duplication
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C476 Day 2 · What is Structured Data and why we need it on the internet.
The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C603 Day 2 · Focusing on Internationalisation and Localisation
Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C648 Day 2 · Calculating (some) signals
Among the many signals Google calculates during indexing, the ones singled out as having large effects on search results were country, language, freshness, SafeSearch and spam.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C656 Day 2 · Calculating (some) signals
Freshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking event hits a city, such as possible closure of Barcelona's airport, users want really fresh results, not results from two weeks ago.
extendsD1-C045 Day 1 · How Search works and where's AI?Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).
- Stage D2-C669 Day 2 · Calculating (some) signals
SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C684 Day 2 · Deciding what goes in the index?
Index selection is a predictive AI system that relies heavily on machine learning.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C722 Day 2 · How does the index look like?
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C726 Day 2 · How does the index look like?
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- D3-C057 Day 3 · Making sense of users' queries
Google's slide on how LLM features with grounding generally work showed a query going to both the search engine and an LLM, the search engine's results going to the LLM, the LLM generating fan-out queries that go back to the search engine, and the LLM returning answers with links.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Analysis D3-C487 Day 3 · Lightning session L: Understanding SERPs and your users
Google's ranking systems guide describes 'query deserves freshness' systems that show fresher content where it would be expected, which is narrower than a general preference for fresh content; refresh pages whose queries expect current information, and judge other refreshes by quality.
extendsStage D2-C656 Day 2 · Calculating (some) signalsFreshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking event hits a city, such as possible closure of Barcelona's airport, users want really fresh results, not results from two weeks ago.
- Stage D3-C606 Day 3 · How long does it take to..?
For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).
extends - D3-C610 Day 3 · How long does it take to..?
Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.
extendsStage D1-C337 Day 1 · How crawling worksAn XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.
- D3-C611 Day 3 · How long does it take to..?
If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.
extendsStage D1-C337 Day 1 · How crawling worksAn XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.
- D3-C613 Day 3 · How long does it take to..?
Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.
extendsDocs D1-C085 Day 1 · How Google interprets robots.txtGoogle generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
- D3-C617 Day 3 · How long does it take to..?
Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2 weeks, and 1-3 weeks in recovery.
extends - D3-C701 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's closing slide said AI on Google is just SEO: AI features on Google Search use exactly the same processes as traditional results, so no new acronym is needed, as none was for mobile-first indexing or structured data.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C116 Day 2 · Lightning session D: Rendering and JavaScript
AI Overviews and AI Mode are built on top of Search results: they are a different experience of the same content Google already has.
repeatsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C537 Day 2 · Using images to your advantage and Engaging Search users with videos
The text around an image is critical: Google uses it as context to understand the image and to rank it, so an alt attribute alone is not enough.
repeatsD1-C045 Day 1 · How Search works and where's AI?Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).
- Stage D2-C680 Day 2 · Deciding what goes in the index?
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
repeatsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C719 Day 2 · How does the index look like?
The speaker recapped Google's pipeline up to the index: Google crawls pages, processes the fetched documents and then stores them in its index.
repeatsD1-C036 Day 1 · How Search works and where's AI?Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- Stage D2-C847 Day 2 · Welcome to indexing day!
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
repeatsStage D1-C317 Day 1 · How crawling worksGoogle's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
- Stage D3-C142 Day 3 · How Google thinks about Quality
Google's quality talk said ranking signals differ by result type: for web pages they include the text on the page, links and passages, while for news, probably, freshness, diversity and originality become more important.
repeatsD1-C045 Day 1 · How Search works and where's AI?Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).
- Stage D3-C702 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.
repeatsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D3-C702 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.
repeatsStage D2-C726 Day 2 · How does the index look like?AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.