Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Day 2 · Thursday 1 October 2026 · 13:30
Finding the gold nuggets: structured data, media, and more!
Speaker Gary Illyes, Search Relations
TalkCoverageTranscriptOne slide
Short introduction to feature extraction (structured data, images, video). Introduced by name by the host and named again in the hand-over at the end of the next talk.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
Speaker Gary IllyesEvidence slide photo, transcript
Google's indexing includes a dedicated system, whose internal name Gary Illyes would not disclose, that extracts the parts of a page that are traditionally expensive to extract.
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
Speaker Gary IllyesEvidence transcript, slide photo
For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.
For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.
Gary Illyes said feature extraction, which extracts page structures into a form Google's systems can consume internally, is still expensive, though not the most expensive operation.
Google's guide to how Search works says rendering happens during the crawl, and describes indexing as analysing a page's text, key tags and attributes such as title elements and alt attributes, images and videos, and deciding whether the page is a duplicate or the canonical.
Google's general structured data guidelines recommend placing the same structured data on all duplicate pages of the same content, not just on the canonical page.
“we recommend placing the same structured data on all page duplicates, not just on the canonical page”
Because Google extracts structured data, images and videos only after deduplication, put markup and media on the URL you want as canonical and keep them identical on its duplicates; markup that exists only on a duplicate that loses canonical selection may never be extracted.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.
For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.