Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 2 · Thursday 1 October 2026 · 10:25

How is HTML interpreted

Speaker Cherry Prommawin, Search Relations

TalkCoverageTranscriptSlides

Speaker from a back-reference in a later talk ('Cherry already talked about the page ... main content'), heard in a second attendee recording. The author's recording starts mid-talk, so the opening is from slides only.

Shown on screen 14

SlideNot in docsD2-C024

A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

Speaker Cherry PrommawinEvidence slide photo

Used byrequirement DEV-PRF-01

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
SlideConsistent with docsD2-C025

In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

Speaker Cherry PrommawinEvidence slide photo

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
SlideConsistent with docsD2-C026

A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

Speaker Cherry PrommawinEvidence slide photo

  • Extends D1-C036 Day 1: Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extended by D2-C128 Day 2: Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached…
  • Repeated by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
  • Extended by D2-C444 Day 2: Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a…
  • Extended by D2-C688 Day 2: Index selection is the last step before documents enter Google's index.
SlideConsistent with docsD2-C028

Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • Repeated by D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
  • Extended by D2-C868 Day 2: Gary Illyes said the main content is what Google considers when ranking a page.
  • Extended by D2-C474 Day 2: Google's systems sometimes fail to determine a page's main content correctly, and structured data helps…
SlideConfirmed by docsD2-C031

Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirement DEV-CAN-03

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
SlideConfirmed by docsD2-C033

Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.

Speaker Cherry PrommawinEvidence slide photo, transcript

Things

Used byglossary term hreflang

  • Extended by D2-C382 Day 2: When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of…
SlideConfirmed by docsD2-C036

Links and anchors are among the things Google extracts from a page's HTML, and the slide card for them simply read 'We like links.'

“We like links.”

Wording checked against the slide or recording

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirement DEV-URL-01

SlideConfirmed by docsD2-C039

Google can extract links written as an a element with an href attribute that holds an absolute or a relative URL, which the speaker called the good old normal way.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-01

SlideConsistent with docsD2-C040

Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-02

  • Repeats D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
  • Extended by D2-C184 Day 2: Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link…
  • Extended by D2-C287 Day 2: A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google…
SlideConsistent with docsD2-C041

Google cannot extract a link from an href attribute placed on an element other than a, such as a span, because that is not a standard way to make a link.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-02

SlideConfirmed by docsD2-C048

Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

Speaker Cherry PrommawinEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
  • Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
  • Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

Said on stage 6

StageD2-C037

The speaker said links are still an extremely important part of the internet and of most major search and AI systems.

Speaker Cherry PrommawinEvidence transcript

StageConsistent with docsD2-C038

Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-URL-01

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
StageNot in docsD2-C046

The speaker said Google sometimes also extracts URLs that are typed out as plain text on a page without being hyperlinked; the remarks around this point were unclear in the recording.

Speaker Cherry PrommawinEvidence transcript

StageConsistent with docsD2-C049

The robots meta element is also extracted when Google processes a page's HTML, and the speaker called it probably one of the most important extracted elements, or one the audience is probably interested in.

Speaker Cherry PrommawinEvidence transcript

What Google's documentation says 2

DocsSourceD2-C043

Google's link best practices say Google can generally crawl a link only if it is an a element with an href attribute, and list routerLink without href, href on a span, onclick-only a elements and javascript: URLs as not recommended, while noting that Google may still attempt to parse them.

Publisher Google Search Central

Used byrequirements DEV-URL-01, DEV-URL-02

DocsSourceD2-C831

Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.

Publisher Google Search Central

Used byrequirements DEV-CAN-03, DEV-HTM-05

Analysis by the author 5

AnalysisD2-C029

Make the main content of every template easy to separate from the header, navigation and footer, for example as one clearly delimited main area, because Google identifies the main content and treats it as the most important part of the page.

Author Ibrahim Anjro

Used byrequirement DEV-HTM-01

AnalysisD2-C044

Google's link-extraction slide put routerLink, href on a span, onclick-only links and javascript: URLs under 'can not extract', which is stricter than Google's link documentation saying Google may still try to parse them; either way they are not dependable links for discovery.

Author Ibrahim Anjro

AnalysisD2-C045

Audit every template and JavaScript component that outputs links, such as navigation, pagination, filters and product tiles: each needs a real a element with an href, because onclick handlers, routerLink without href, href on a span and javascript: URLs leave the target pages without a link Google can reliably extract.

Author Ibrahim Anjro

AnalysisD2-C047

Do not rely on plain-text URLs for discovery: even if Google sometimes picks them up, a proper a href link is what was described as feeding discovery, site structure and ranking.

Author Ibrahim Anjro

  1. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  2. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  3. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C092 Day 1 · How Google thinks about crawl budget

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

  4. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  5. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  6. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

  7. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C036 Day 1 · How Search works and where's AI?

    Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.

  8. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  9. Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  10. Stage D2-C038 Day 2 · How is HTML interpreted

    Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  11. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  12. Stage D2-C128 Day 2 · Lightning session D: Rendering and JavaScript

    Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached system, kept separate because rendering is time-consuming and computationally expensive.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  13. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

  14. Slide D2-C184 Day 2 · Lightning session D: Rendering and JavaScript

    Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.

    extends
    Slide D2-C040 Day 2 · How is HTML interpreted

    Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.

  15. Stage D2-C287 Day 2 · What is Google friendly JavaScript

    A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.

    extends
    Slide D2-C040 Day 2 · How is HTML interpreted

    Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.

  16. Stage D2-C379 Day 2 · Handling web duplication

    rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

    extends
    Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

  17. Slide D2-C382 Day 2 · Handling web duplication

    When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

    extends
    Slide D2-C033 Day 2 · How is HTML interpreted

    Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.

  18. Stage D2-C444 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  19. Stage D2-C474 Day 2 · What is Structured Data and why we need it on the internet.

    Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.

    extends
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  20. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  21. Stage D2-C868 Day 2 · Understanding what's on a page

    Gary Illyes said the main content is what Google considers when ranking a page.

    extends
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  22. Slide D2-C040 Day 2 · How is HTML interpreted

    Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.

    repeats
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  23. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    repeats
    Stage D1-C328 Day 1 · How crawling works

    During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

  24. Slide D2-C309 Day 2 · Understanding what's on a page

    A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

    repeats
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  25. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    repeats
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  26. Slide D2-C442 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

    repeats
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.