Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Indexing

Processing: from fetched page to index

Day 2 opened the stage that Day 1's crawl pipeline hands its fetches to: the crawler passes on a fetch record, shown on a slide as an example (fetch result, connect time, time to first byte, the robots policies that apply, such as a Google-Extended opt-out, and the raw HTTP response), and processing runs HTML parsing, rendering, deduplication, feature extraction, signal extraction and finally index selection. HTML parsing builds a DOM from which Google extracts every element to separate header, navigation and main content, plus the robots meta tag, rel=canonical (used for deduplication and canonical selection), hreflang and links, which go back to the crawl queue. Links are extracted again from the rendered HTML; rendering itself sits among the processing steps on the slides, is a detached system because it is so expensive, Erin Sparling said, and happens during the crawl according to Google's guide to how Search works. Gary Illyes said deduplication runs before feature extraction, so the costly extraction of structured data, images and videos is spent on fewer documents, and images and videos go on to dedicated indexing services. The fetch record, the DOM step and that order were shown or said at the event and are not in Google's documentation, which describes indexing more broadly. Day 3 timed the stage: Google estimated that indexing a document end to end, from entering indexing until it reaches the serving index tokenized and ready to serve, takes about 1.5 hours on average, and that meta annotations such as robots meta tags are typically processed in 45 to 90 minutes, a critical step without which indexing cannot continue (said at the event, not in Google's docs). Rendering happens either right after crawling or later through a queue. Day 1's second recording added that indexing starts with parsing the fetched HTML so that elements such as the title can be extracted, and a second recording of Day 2 that the media indexer attaches each image and video to the URL of the page that hosts it (said at the event).

What to do

  • Put the robots meta tag, rel=canonical, hreflang and the main links in the HTML the server sends, so Google reads them at HTML parsing without waiting for rendering.
  • Put structured data, images and videos on the URL you want as canonical and keep them identical on its duplicates, because feature extraction runs after deduplication.
  • Mark up images you want indexed as img elements with a src, which is what Google's HTML parser looks for.

Day 1: Crawling 2

Said on stage 2

Day 2: Indexing 34

Shown on screen 16

SlideNot in docsD2-C024

A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo

Used byrequirement DEV-PRF-01

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
SlideConsistent with docsD2-C025

In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
SlideConsistent with docsD2-C026

A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo

  • Extends D1-C036 Day 1: Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extended by D2-C128 Day 2: Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached…
  • Repeated by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
  • Extended by D2-C444 Day 2: Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a…
  • Extended by D2-C688 Day 2: Index selection is the last step before documents enter Google's index.
SlideConsistent with docsD2-C028

Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • Repeated by D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
  • Extended by D2-C868 Day 2: Gary Illyes said the main content is what Google considers when ranking a page.
  • Extended by D2-C474 Day 2: Google's systems sometimes fail to determine a page's main content correctly, and structured data helps…
SlideConfirmed by docsD2-C031

Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript

Used byrequirement DEV-CAN-03

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
SlideConfirmed by docsD2-C033

Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript

Things

Used byglossary term hreflang

  • Extended by D2-C382 Day 2: When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of…
SlideConfirmed by docsD2-C036

Links and anchors are among the things Google extracts from a page's HTML, and the slide card for them simply read 'We like links.'

“We like links.”

Wording checked against the slide or recording

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript

Used byrequirement DEV-URL-01

SlideConfirmed by docsD2-C048

Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
  • Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
  • Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
SlideConsistent with docsD2-C127

A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
SlideConfirmed by docsD2-C167

The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
SlideConfirmed by docsD2-C168

In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

Things
  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extends D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
SlideConsistent with docsD2-C174

Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.

“Whatever is in the DOM at that moment is what likely gets indexed.”

Wording checked against the slide or recording

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

Things
  • Repeated by D2-C262 Day 2: Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.
SlideConsistent with docsD2-C309

A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence video, transcript

Used byrequirement DEV-HTM-01

  • Repeats D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
  • Extended by D2-C861 Day 2: Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not…
SlideNot in docsD2-C441

Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…

Said on stage 11

StageConsistent with docsD2-C049

The robots meta element is also extracted when Google processes a page's HTML, and the speaker called it probably one of the most important extracted elements, or one the audience is probably interested in.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript

StageNot in docsD2-C444

Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.

Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript, slide photo

Used byrequirement DEV-SDA-10

  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
StageConsistent with docsD2-C446

The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.

Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript

  • Extended by D3-C323 Day 3: Most of Google's search features need nothing extra from the site owner; Google generates them from what it…
StageConsistent with docsD2-C447

For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.

Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript

  • Repeated by D2-C524 Day 2: Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly…
StageConsistent with docsD2-C448

For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.

Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript

Used byrequirement DEV-VID-01

  • Extended by D2-C930 Day 2: Google's media indexer processes the images and videos that feature extraction passes to it and attaches them…
StageConsistent with docsD2-C524

Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.

Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript

Used byrequirement DEV-IMG-01

  • Repeats D2-C447 Day 2: For images, Google's feature extraction takes the img element with its src and other attributes, including…
StageNot in docsD2-C688

Index selection is the last step before documents enter Google's index.

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

Used byglossary term Index selection

  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…

What Google's documentation says 4

DocsSourceD2-C831

Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.

Publisher Google Search CentralAnnotates Day 2, 10:25 · How is HTML interpreted

Used byrequirements DEV-CAN-03, DEV-HTM-05

DocsSourceD2-C131

Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.

Publisher Google Search CentralAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-REN-01glossary term Rendering

  • Extended by D3-C676 Day 3: The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page…
DocsSourceD2-C169

Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.

Publisher Google Search CentralAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-URL-01

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
DocsSourceD2-C445

Google's guide to how Search works says rendering happens during the crawl, and describes indexing as analysing a page's text, key tags and attributes such as title elements and alt attributes, images and videos, and deciding whether the page is a duplicate or the canonical.

Publisher Google Search CentralAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!

Things

Used byrequirement DEV-HTM-03

Analysis by the author 3

AnalysisD2-C450

Because Google extracts structured data, images and videos only after deduplication, put markup and media on the URL you want as canonical and keep them identical on its duplicates; markup that exists only on a duplicate that loses canonical selection may never be extracted.

Author Ibrahim AnjroAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!

Used byrequirement DEV-SDA-10

AnalysisD2-C677

Language, country, SafeSearch and spam signals, and some freshness signals, are calculated when a page is indexed, so a fix such as correcting a page's language or removing content that triggers SafeSearch only counts once Google recrawls and reprocesses the page; request recrawling of the most important URLs after the fix.

Author Ibrahim AnjroAnnotates Day 2, 15:30 · Calculating (some) signals

Used byrequirements DEV-IDX-12, DEV-MON-02

Day 3: Serving: Ranking, Search Console, and Performance 4

Shown on screen 2

Said on stage 2

Across days and sessions 35

  1. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  2. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  3. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C092 Day 1 · How Google thinks about crawl budget

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

  4. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  5. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  6. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

  7. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C036 Day 1 · How Search works and where's AI?

    Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.

  8. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  9. Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  10. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  11. Slide D2-C127 Day 2 · Lightning session D: Rendering and JavaScript

    A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  12. Stage D2-C128 Day 2 · Lightning session D: Rendering and JavaScript

    Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached system, kept separate because rendering is time-consuming and computationally expensive.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  13. Slide D2-C167 Day 2 · Lightning session D: Rendering and JavaScript

    The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  14. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  15. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

  16. Docs D2-C169 Day 2 · Lightning session D: Rendering and JavaScript

    Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  17. Stage D2-C379 Day 2 · Handling web duplication

    rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

    extends
    Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

  18. Slide D2-C382 Day 2 · Handling web duplication

    When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

    extends
    Slide D2-C033 Day 2 · How is HTML interpreted

    Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.

  19. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  20. Stage D2-C444 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  21. Stage D2-C474 Day 2 · What is Structured Data and why we need it on the internet.

    Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.

    extends
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  22. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Stage D1-C211 Day 1 · How Search works and where's AI?

    Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

  23. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  24. Stage D2-C861 Day 2 · Understanding what's on a page

    Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.

    extends
    Slide D2-C309 Day 2 · Understanding what's on a page

    A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

  25. Stage D2-C868 Day 2 · Understanding what's on a page

    Gary Illyes said the main content is what Google considers when ranking a page.

    extends
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  26. Stage D2-C930 Day 2 · Using images to your advantage and Engaging Search users with videos

    Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.

    extends
    Stage D2-C448 Day 2 · Finding the gold nuggets: structured data, media, and more!

    For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.

  27. Stage D3-C323 Day 3 · How Search results are born

    Most of Google's search features need nothing extra from the site owner; Google generates them from what it extracted from the page during indexing.

    extends
    Stage D2-C446 Day 2 · Finding the gold nuggets: structured data, media, and more!

    The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.

  28. Stage D3-C626 Day 3 · How long does it take to..?

    Google renders pages in two ways: immediately after crawling, or later through a queue-based process that runs elsewhere.

    extends
    Slide D2-C170 Day 2 · Lightning session D: Rendering and JavaScript

    After processing, an indexable page is placed in Google's render queue to wait for rendering.

  29. Analysis D3-C676 Day 3 · How long does it take to..?

    The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.

    extends
    Docs D2-C131 Day 2 · Lightning session D: Rendering and JavaScript

    Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.

  30. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    repeats
    Stage D1-C328 Day 1 · How crawling works

    During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

  31. Slide D2-C262 Day 2 · What is Google friendly JavaScript

    Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.

    repeats
    Slide D2-C174 Day 2 · Lightning session D: Rendering and JavaScript

    Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.

  32. Slide D2-C309 Day 2 · Understanding what's on a page

    A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

    repeats
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  33. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    repeats
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  34. Slide D2-C442 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

    repeats
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

  35. Stage D2-C524 Day 2 · Using images to your advantage and Engaging Search users with videos

    Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.

    repeats
    Stage D2-C447 Day 2 · Finding the gold nuggets: structured data, media, and more!

    For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.

Built on these claims 12

Developer requirements 12

Sources 15