Indexing starts with parsing the fetched HTML, so that elements of the page such as the title can be extracted and accessed easily.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Topic · Indexing
Day 2 opened the stage that Day 1's crawl pipeline hands its fetches to: the crawler passes on a fetch record, shown on a slide as an example (fetch result, connect time, time to first byte, the robots policies that apply, such as a Google-Extended opt-out, and the raw HTTP response), and processing runs HTML parsing, rendering, deduplication, feature extraction, signal extraction and finally index selection. HTML parsing builds a DOM from which Google extracts every element to separate header, navigation and main content, plus the robots meta tag, rel=canonical (used for deduplication and canonical selection), hreflang and links, which go back to the crawl queue. Links are extracted again from the rendered HTML; rendering itself sits among the processing steps on the slides, is a detached system because it is so expensive, Erin Sparling said, and happens during the crawl according to Google's guide to how Search works. Gary Illyes said deduplication runs before feature extraction, so the costly extraction of structured data, images and videos is spent on fewer documents, and images and videos go on to dedicated indexing services. The fetch record, the DOM step and that order were shown or said at the event and are not in Google's documentation, which describes indexing more broadly. Day 3 timed the stage: Google estimated that indexing a document end to end, from entering indexing until it reaches the serving index tokenized and ready to serve, takes about 1.5 hours on average, and that meta annotations such as robots meta tags are typically processed in 45 to 90 minutes, a critical step without which indexing cannot continue (said at the event, not in Google's docs). Rendering happens either right after crawling or later through a queue. Day 1's second recording added that indexing starts with parsing the fetched HTML so that elements such as the title can be extracted, and a second recording of Day 2 that the media indexer attaches each image and video to the URL of the page that hosts it (said at the event).
Things in this topic 14
Counts are claims that name the thing. All things
What to do
Indexing starts with parsing the fetched HTML, so that elements of the page such as the title can be extracted and accessed easily.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
A community speaker described the retrieval stage as whether an AI requests a site's pages when grounding its answer; if it does not, the cause may be a crawling or an indexing issue.
Speaker not identifiedIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo
Used byrequirement DEV-PRF-01
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo
The HTML parsing step turns a fetched page's HTML into a Document Object Model (DOM) tree of elements, attributes and text nodes.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo
Used byrequirement DEV-HTM-05
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript
Used byrequirement DEV-CAN-03
Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript
Used byglossary term hreflang
Links and anchors are among the things Google extracts from a page's HTML, and the slide card for them simply read 'We like links.'
“We like links.”
Wording checked against the slide or recording
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript
Used byrequirement DEV-URL-01
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript
A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.
“Whatever is in the DOM at that moment is what likely gets indexed.”
Wording checked against the slide or recording
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence video, transcript
Used byrequirement DEV-HTM-01
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence slide photo, transcript
Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence slide photo
Besides the visible content, Google extracts elements that site owners add to the HTML, because they are useful for indexing and, for some of them, for ranking.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript
The robots meta element is also extracted when Google processes a page's HTML, and the speaker called it probably one of the most important extracted elements, or one the audience is probably interested in.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript
Google's indexing includes a dedicated system, whose internal name Gary Illyes would not disclose, that extracts the parts of a page that are traditionally expensive to extract.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript, slide photo
Used byrequirement DEV-SDA-10
The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript
For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript
For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript
Used byrequirement DEV-VID-01
Gary Illyes said feature extraction, which extracts page structures into a form Google's systems can consume internally, is still expensive, though not the most expensive operation.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript
Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-IMG-01
Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Index selection is the last step before documents enter Google's index.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byglossary term Index selection
Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.
Publisher Google Search CentralAnnotates Day 2, 10:25 · How is HTML interpreted
Used byrequirements DEV-CAN-03, DEV-HTM-05
Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.
Publisher Google Search CentralAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirement DEV-REN-01glossary term Rendering
Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.
Publisher Google Search CentralAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirement DEV-URL-01
Google's guide to how Search works says rendering happens during the crawl, and describes indexing as analysing a page's text, key tags and attributes such as title elements and alt attributes, images and videos, and deciding whether the page is a duplicate or the canonical.
Publisher Google Search CentralAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!
Used byrequirement DEV-HTM-03
Put rel=canonical and hreflang link elements in the head of the HTML the server sends, not only in JavaScript-rendered HTML, so Google can read them during HTML parsing without depending on rendering.
Author Ibrahim AnjroAnnotates Day 2, 10:25 · How is HTML interpreted
Used byrequirement DEV-CAN-03
Because Google extracts structured data, images and videos only after deduplication, put markup and media on the URL you want as canonical and keep them identical on its duplicates; markup that exists only on a duplicate that loses canonical selection may never be extracted.
Author Ibrahim AnjroAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!
Used byrequirement DEV-SDA-10
Language, country, SafeSearch and spam signals, and some freshness signals, are calculated when a page is indexed, so a fix such as correcting a page's language or removing content that triggers SafeSearch only counts once Google recrawls and reprocesses the page; request recrawling of the most important URLs after the fix.
Author Ibrahim AnjroAnnotates Day 2, 15:30 · Calculating (some) signals
Used byrequirements DEV-IDX-12, DEV-MON-02
Google estimated that indexing a document end to end takes about 1.5 hours on average, with a minimum of seconds.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript
Google defined indexing end to end as the time from a document entering indexing until its critical processes finish and it reaches the serving index, tokenized and ready to be served as a result.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript
Google renders pages in two ways: immediately after crawling, or later through a queue-based process that runs elsewhere.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
Meta annotation processing is a critical step of indexing: until a document passes it, Google cannot go on indexing the document.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.
The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached system, kept separate because rendering is time-consuming and computationally expensive.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.
The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
URL discovery works through links: a homepage links to section pages, which link to further pages.
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.
URL discovery works through links: a homepage links to section pages, which link to further pages.
rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.
Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
Index selection is the last step before documents enter Google's index.
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
Index selection is the last step before documents enter Google's index.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.
A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
Gary Illyes said the main content is what Google considers when ranking a page.
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.
For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.
Most of Google's search features need nothing extra from the site owner; Google generates them from what it extracted from the page during indexing.
The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.
Google renders pages in two ways: immediately after crawling, or later through a queue-based process that runs elsewhere.
After processing, an indexable page is placed in Google's render queue to wait for rendering.
The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.
Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.
Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.
Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.
A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.
For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.
Put exactly one absolute rel=canonical in the server-rendered head of every indexable page
Rests on 8 claims, 3 of them in this topic
Wrap each page's primary content in one clearly delimited main area
Rests on 14 claims, 2 of them in this topic
Write valid HTML and close every element
Rests on 4 claims, 2 of them in this topic
Put the same structured data on duplicate URLs as on the canonical
Rests on 3 claims, 2 of them in this topic
Make every navigational link an <a> element whose href holds a real URL
Rests on 6 claims, 2 of them in this topic
Give every indexable page a unique, descriptive title element in the server HTML
Rests on 7 claims, 1 of them in this topic
Rests on 4 claims, 1 of them in this topic
Put every image that should be found in an <img> element with a src attribute
Rests on 6 claims, 1 of them in this topic
Rests on 7 claims, 1 of them in this topic
Keep connect time and time to first byte low and stable under crawler load
Rests on 13 claims, 1 of them in this topic
Server-render the main content and everything indexing depends on (SSR, static generation or hybrid)
Rests on 11 claims, 1 of them in this topic
Embed each video in a video, iframe, embed or object element that loads without user action
Rests on 5 claims, 1 of them in this topic