Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 2 · Thursday 1 October 2026 · 13:30

Finding the gold nuggets: structured data, media, and more!

Speaker Gary Illyes, Search Relations

TalkCoverageTranscriptOne slide

Short introduction to feature extraction (structured data, images, video). Introduced by name by the host and named again in the hand-over at the end of the next talk.

Shown on screen 2

SlideNot in docsD2-C441

Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
SlideNot in docsD2-C442

Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

Speaker Gary IllyesEvidence slide photo

  • Repeats D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…

Said on stage 6

StageNot in docsD2-C443

Google's indexing includes a dedicated system, whose internal name Gary Illyes would not disclose, that extracts the parts of a page that are traditionally expensive to extract.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD2-C446

The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C323 Day 3: Most of Google's search features need nothing extra from the site owner; Google generates them from what it…
StageConsistent with docsD2-C447

For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.

Speaker Gary IllyesEvidence transcript

  • Repeated by D2-C524 Day 2: Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly…
StageConsistent with docsD2-C448

For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-01

  • Extended by D2-C930 Day 2: Google's media indexer processes the images and videos that feature extraction passes to it and attaches them…
StageNot in docsD2-C449

Gary Illyes said feature extraction, which extracts page structures into a form Google's systems can consume internally, is still expensive, though not the most expensive operation.

Speaker Gary IllyesEvidence transcript

What Google's documentation says 2

DocsSourceD2-C445

Google's guide to how Search works says rendering happens during the crawl, and describes indexing as analysing a page's text, key tags and attributes such as title elements and alt attributes, images and videos, and deciding whether the page is a duplicate or the canonical.

Publisher Google Search Central

Things

Used byrequirement DEV-HTM-03

DocsSourceD2-C451

Google's general structured data guidelines recommend placing the same structured data on all duplicate pages of the same content, not just on the canonical page.

“we recommend placing the same structured data on all page duplicates, not just on the canonical page”

Publisher Google Search Central

Used byrequirement DEV-SDA-10

Analysis by the author 1

  1. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  2. Stage D2-C444 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  3. Stage D2-C930 Day 2 · Using images to your advantage and Engaging Search users with videos

    Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.

    extends
    Stage D2-C448 Day 2 · Finding the gold nuggets: structured data, media, and more!

    For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.

  4. Stage D3-C323 Day 3 · How Search results are born

    Most of Google's search features need nothing extra from the site owner; Google generates them from what it extracted from the page during indexing.

    extends
    Stage D2-C446 Day 2 · Finding the gold nuggets: structured data, media, and more!

    The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.

  5. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    repeats
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  6. Slide D2-C442 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

    repeats
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

  7. Stage D2-C524 Day 2 · Using images to your advantage and Engaging Search users with videos

    Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.

    repeats
    Stage D2-C447 Day 2 · Finding the gold nuggets: structured data, media, and more!

    For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.