Speaker from a back-reference in a later talk ('Cherry already talked about the page ... main content'), heard in a second attendee recording. The author's recording starts mid-talk, so the opening is from slides only.
Shown on screen 14
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
Speaker Cherry PrommawinEvidence slide photo
Used byrequirement DEV-PRF-01
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
Speaker Cherry PrommawinEvidence slide photo
- Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Extends D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Speaker Cherry PrommawinEvidence slide photo
- Extends D1-C036 Day 1: Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Extended by D2-C128 Day 2: Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached…
- Repeated by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
- Extended by D2-C444 Day 2: Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a…
- Extended by D2-C688 Day 2: Index selection is the last step before documents enter Google's index.
The HTML parsing step turns a fetched page's HTML into a Document Object Model (DOM) tree of elements, attributes and text nodes.
Speaker Cherry PrommawinEvidence slide photo
Used byrequirement DEV-HTM-05
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
- Repeated by D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
- Extended by D2-C868 Day 2: Gary Illyes said the main content is what Google considers when ranking a page.
- Extended by D2-C474 Day 2: Google's systems sometimes fail to determine a page's main content correctly, and structured data helps…
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byrequirement DEV-CAN-03
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
The rel=canonical link element is placed in the head section of the HTML and tells Google that one page is the representative of another.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-CAN-03glossary term rel=canonical
Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byglossary term hreflang
- Extended by D2-C382 Day 2: When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of…
Links and anchors are among the things Google extracts from a page's HTML, and the slide card for them simply read 'We like links.'
“We like links.”
Wording checked against the slide or recording
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byrequirement DEV-URL-01
Google can extract links written as an a element with an href attribute that holds an absolute or a relative URL, which the speaker called the good old normal way.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-01
Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-02
- Repeats D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
- Extended by D2-C184 Day 2: Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link…
- Extended by D2-C287 Day 2: A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google…
Google cannot extract a link from an href attribute placed on an element other than a, such as a span, because that is not a standard way to make a link.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-02
A Google slide also listed as not extractable an a element with a routerLink attribute instead of an href, and javascript: URLs such as javascript:goTo('products') or javascript:window.location.href='/products'.
Speaker Cherry PrommawinEvidence slide photo
Used byrequirement DEV-URL-02
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
Speaker Cherry PrommawinEvidence slide photo, transcript
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
- Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
- Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
Said on stage 6
Besides the visible content, Google extracts elements that site owners add to the HTML, because they are useful for indexing and, for some of them, for ranking.
Speaker Cherry PrommawinEvidence transcript
The speaker said that knowing a page's other-language versions from hreflang is useful in ranking.
Speaker Cherry PrommawinEvidence transcript
The speaker said links are still an extremely important part of the internet and of most major search and AI systems.
Speaker Cherry PrommawinEvidence transcript
Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-URL-01
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
The speaker said Google sometimes also extracts URLs that are typed out as plain text on a page without being hyperlinked; the remarks around this point were unclear in the recording.
Speaker Cherry PrommawinEvidence transcript
The robots meta element is also extracted when Google processes a page's HTML, and the speaker called it probably one of the most important extracted elements, or one the audience is probably interested in.
Speaker Cherry PrommawinEvidence transcript
What Google's documentation says 2
Google's link best practices say Google can generally crawl a link only if it is an a element with an href attribute, and list routerLink without href, href on a span, onclick-only a elements and javascript: URLs as not recommended, while noting that Google may still attempt to parse them.
Publisher Google Search Central
Used byrequirements DEV-URL-01, DEV-URL-02
Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.
Publisher Google Search Central
Used byrequirements DEV-CAN-03, DEV-HTM-05
Analysis by the author 5
Make the main content of every template easy to separate from the header, navigation and footer, for example as one clearly delimited main area, because Google identifies the main content and treats it as the most important part of the page.
Author Ibrahim Anjro
Used byrequirement DEV-HTM-01
Put rel=canonical and hreflang link elements in the head of the HTML the server sends, not only in JavaScript-rendered HTML, so Google can read them during HTML parsing without depending on rendering.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-03
Google's link-extraction slide put routerLink, href on a span, onclick-only links and javascript: URLs under 'can not extract', which is stricter than Google's link documentation saying Google may still try to parse them; either way they are not dependable links for discovery.
Author Ibrahim Anjro
Audit every template and JavaScript component that outputs links, such as navigation, pagination, filters and product tiles: each needs a real a element with an href, because onclick handlers, routerLink without href, href on a span and javascript: URLs leave the target pages without a link Google can reliably extract.
Author Ibrahim Anjro
Do not rely on plain-text URLs for discovery: even if Google sometimes picks them up, a proper a href link is what was described as feeding discovery, site structure and ranking.
Author Ibrahim Anjro