An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
From the audienceEvidence slide photo
- Answered by D2-C002 Day 2: Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
- Answered by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
- Answered by D2-C003 Day 2: Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have…
- Answered by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
“It's hard to make good HTML sitemaps for large sites.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
“Better rely on hubs like category pages that link out to your important pages.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-04
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
- Extends D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
- Extended by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have the extra benefit that they may also be useful to users.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo
Used byrequirement DEV-URL-04
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
Google's answer did not say whether HTML sitemaps pass link equity. On a large site, check that every important page is linked from a category or hub page in the normal navigation instead of relying on an HTML sitemap page to reach it.
Author Ibrahim Anjro
Google's ecommerce documentation recommends linking from menus to category pages, from category pages to sub-category pages and from sub-category pages to all product pages; where not every product can be linked, it recommends a sitemap or a Merchant Center feed.
Publisher Google Search Central
Used byrequirement DEV-URL-04
Google's ecommerce documentation says Google can infer a page's relative importance within a site from its internal links, such as how many links point to the page and how many links Google must follow to reach it.
Publisher Google Search Central
Used byrequirement DEV-URL-04
A 2005 Search Central blog post, now marked as possibly outdated, said Google encouraged HTML sitemaps because they help users navigate a site and a clear hierarchy of text links helps Google index it.
Publisher Search Central blog (26 September 2005)
Google's view of HTML sitemaps has moved: a 2005 blog post encouraged them, the current sitemap and ecommerce documentation does not mention them, and the 2026 Q&A slide pointed large sites to category hub pages instead.
Author Ibrahim Anjro
An audience member asked whether a product detail page that is out of stock for two to three months should keep returning 200 with links to similar products, or be 302-redirected to a similar product or to its parent product listing page.
From the audienceEvidence slide photo, transcript
- Answered by D2-C010 Day 2: Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it…
- Answered by D2-C011 Day 2: Google's Q&A slide on out-of-stock product pages said users might wait months for some products, or even…
- Answered by D2-C841 Day 2: Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to…
Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it depends on how important that product page is to users.
“It depends on the importance of the PDPs to the users.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-ERR-04
- Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
- Extended by D2-C841 Day 2: Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to…
Google's Q&A slide on out-of-stock product pages said users might wait months for some products, or even pre-order them if the site offers pre-ordering.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-ERR-04
- Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
Google's guide to temporarily pausing an online business recommends that a shop expecting to sell again within weeks or months stays online with limited functionality, such as a disabled cart, and updates its Product structured data to show current availability.
Publisher Google Search Central
Used byrequirements DEV-ERR-04, DEV-SDA-08, DEV-SRV-03
Google's redirect documentation says that with a temporary redirect, such as a 302, Google Search shows the source page in search results and does not use the redirect as a signal that the target should be canonical.
Publisher Google Search Central
Used byrequirements DEV-CAN-01, DEV-ERR-04
Keep a product page that is out of stock for a few months live with a 200 status, show its availability on the page and in Product structured data, and offer pre-ordering where possible when users would wait for the product. Redirect it only when users would rather switch to a similar product than wait.
Author Ibrahim Anjro
Used byrequirement DEV-ERR-04
An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.
From the audienceEvidence slide photo, transcript
- Answered by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
- Answered by D2-C018 Day 2: Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in…
- Answered by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
- Answered by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
“If it's included in the robots.txt file, any crawler can pick your sitemaps up”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-05
- Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
- Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
- Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
- Extended by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in Search Console.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-05
- Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
From the audienceEvidence slide photo, transcript
- Answered by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
- Answered by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Answered by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
- Answered by D2-C845 Day 2: Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no…
- Answered by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
“Simply put, it's because of sites that are extremely important and like to disallow their most important pages, either accidentally or out of ignorance.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
- Extended by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D1-C105 Day 1: URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
- Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
- Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
- Extended by D2-C933 Day 2: Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt…
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
Publisher Google Search Central
Used byrequirement DEV-IDX-01glossary term noindex
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-01
Google's sitemap guide says Google uses lastmod only when it is consistently and verifiably accurate, and counts a change to the main content, the structured data or the links of a page as significant, but not a changed copyright date.
Publisher Google Search Central
Used byrequirement DEV-URL-05
Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.
“if you want to make one, knock yourself out, but I will focus on some better things like hub pages, category pages”
Speaker not identifiedEvidence transcript
Used byrequirement DEV-URL-04glossary term Hub pages
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
- Extends D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
An audience member asked how several sloppy migrations on the same domain can affect Googlebot's crawling.
From the audienceEvidence transcript
- Answered by D2-C840 Day 2: Google answered that several sloppy migrations on one domain can cause many effects in the short term, and…
Google answered that several sloppy migrations on one domain can cause many effects in the short term, and noted that migrations concern indexing as well as crawling, a subject a later Day 2 talk would cover.
Speaker not identifiedEvidence transcript
- Answers D2-C839 Day 2: An audience member asked how several sloppy migrations on the same domain can affect Googlebot's crawling.
Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to them, with easily replaced products, such as a particular cheese, for which the buyer simply picks a similar product instead of waiting.
Speaker not identifiedEvidence transcript
- Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
- Extends D2-C010 Day 2: Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it…
Google said listing the sitemap in robots.txt is fine, as many websites do.
“you can include it in robots.txt. No problem whatsoever.”
Speaker not identifiedEvidence transcript
Used byrequirement DEV-URL-05
- Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
- Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.
Speaker not identifiedEvidence transcript
Used byrequirement DEV-URL-05
- Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
- Extends D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.
Speaker not identifiedEvidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no figure given).
Speaker not identifiedEvidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”
Speaker not identifiedEvidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
Speaker not identifiedEvidence transcript
- Repeats D1-C317 Day 1: Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses…
Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.
Speaker not identifiedEvidence transcript
- Repeats D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
Author Ibrahim Anjro
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
Speaker Cherry PrommawinEvidence slide photo
Used byrequirement DEV-PRF-01
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
Speaker Cherry PrommawinEvidence slide photo
- Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Extends D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Speaker Cherry PrommawinEvidence slide photo
- Extends D1-C036 Day 1: Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Extended by D2-C128 Day 2: Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached…
- Repeated by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
- Extended by D2-C444 Day 2: Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a…
- Extended by D2-C688 Day 2: Index selection is the last step before documents enter Google's index.
The HTML parsing step turns a fetched page's HTML into a Document Object Model (DOM) tree of elements, attributes and text nodes.
Speaker Cherry PrommawinEvidence slide photo
Used byrequirement DEV-HTM-05
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
- Repeated by D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
- Extended by D2-C868 Day 2: Gary Illyes said the main content is what Google considers when ranking a page.
- Extended by D2-C474 Day 2: Google's systems sometimes fail to determine a page's main content correctly, and structured data helps…
Make the main content of every template easy to separate from the header, navigation and footer, for example as one clearly delimited main area, because Google identifies the main content and treats it as the most important part of the page.
Author Ibrahim Anjro
Used byrequirement DEV-HTM-01
Besides the visible content, Google extracts elements that site owners add to the HTML, because they are useful for indexing and, for some of them, for ranking.
Speaker Cherry PrommawinEvidence transcript
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byrequirement DEV-CAN-03
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
The rel=canonical link element is placed in the head section of the HTML and tells Google that one page is the representative of another.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-CAN-03glossary term rel=canonical
Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byglossary term hreflang
- Extended by D2-C382 Day 2: When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of…
The speaker said that knowing a page's other-language versions from hreflang is useful in ranking.
Speaker Cherry PrommawinEvidence transcript
Put rel=canonical and hreflang link elements in the head of the HTML the server sends, not only in JavaScript-rendered HTML, so Google can read them during HTML parsing without depending on rendering.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-03
Links and anchors are among the things Google extracts from a page's HTML, and the slide card for them simply read 'We like links.'
“We like links.”
Wording checked against the slide or recording
Speaker Cherry PrommawinEvidence 2 slide photos, transcript
Used byrequirement DEV-URL-01
The speaker said links are still an extremely important part of the internet and of most major search and AI systems.
Speaker Cherry PrommawinEvidence transcript
Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.
Speaker Cherry PrommawinEvidence transcript
Used byrequirement DEV-URL-01
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
Google can extract links written as an a element with an href attribute that holds an absolute or a relative URL, which the speaker called the good old normal way.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-01
Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-02
- Repeats D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
- Extended by D2-C184 Day 2: Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link…
- Extended by D2-C287 Day 2: A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google…
Google cannot extract a link from an href attribute placed on an element other than a, such as a span, because that is not a standard way to make a link.
Speaker Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-02
A Google slide also listed as not extractable an a element with a routerLink attribute instead of an href, and javascript: URLs such as javascript:goTo('products') or javascript:window.location.href='/products'.
Speaker Cherry PrommawinEvidence slide photo
Used byrequirement DEV-URL-02
Google's link best practices say Google can generally crawl a link only if it is an a element with an href attribute, and list routerLink without href, href on a span, onclick-only a elements and javascript: URLs as not recommended, while noting that Google may still attempt to parse them.
Publisher Google Search Central
Used byrequirements DEV-URL-01, DEV-URL-02
Google's link-extraction slide put routerLink, href on a span, onclick-only links and javascript: URLs under 'can not extract', which is stricter than Google's link documentation saying Google may still try to parse them; either way they are not dependable links for discovery.
Author Ibrahim Anjro
Audit every template and JavaScript component that outputs links, such as navigation, pagination, filters and product tiles: each needs a real a element with an href, because onclick handlers, routerLink without href, href on a span and javascript: URLs leave the target pages without a link Google can reliably extract.
Author Ibrahim Anjro
The speaker said Google sometimes also extracts URLs that are typed out as plain text on a page without being hyperlinked; the remarks around this point were unclear in the recording.
Speaker Cherry PrommawinEvidence transcript
Do not rely on plain-text URLs for discovery: even if Google sometimes picks them up, a proper a href link is what was described as feeding discovery, site structure and ranking.
Author Ibrahim Anjro
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
Speaker Cherry PrommawinEvidence slide photo, transcript
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
- Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
- Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
The robots meta element is also extracted when Google processes a page's HTML, and the speaker called it probably one of the most important extracted elements, or one the audience is probably interested in.
Speaker Cherry PrommawinEvidence transcript
Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.
Publisher Google Search Central
Used byrequirements DEV-CAN-03, DEV-HTM-05
John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.
Speaker John MuellerEvidence transcript
- Repeats D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
John Mueller said there is currently no RFC that standardises robots meta tags, unlike robots.txt.
Speaker John MuellerEvidence transcript
John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.
Speaker John MuellerEvidence transcript
- Repeated by D2-C217 Day 2: Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard…
John Mueller opened with a true-or-false quiz slide asking whether robots meta tags can make a page more visible in Search results than having none, and later answered that the statement is true.
“You can use robots meta tags to be more visible in Search results than without robots meta tags.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-04
The robots meta tag is a piece of HTML code in the head section of a page that gives search engine crawlers page-specific instructions.
Speaker John MuellerEvidence slide photo, transcript
Used byglossary term Robots meta tag
Google's robots meta tag specification tells site owners to place the tag in the head section but notes that Google Search does not enforce that placement and also respects robots meta tags in the body of an HTML document.
Publisher Google Search Central
The primary function of the robots meta tag is to control how a page is indexed and shown in search results.
Speaker John MuellerEvidence slide photo, transcript
Used byglossary term Robots meta tag
The robots meta tag is written as <meta name="robots" content="rule1,rule2">, or with the name googlebot instead of robots, with several rules separated by commas.
Speaker John MuellerEvidence slide photo
Used byglossary term Robots meta tag
John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.
Speaker John MuellerEvidence transcript
Used byglossary term robots.txt
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
The robots meta rule all is the default value and does nothing: it places no restrictions on indexing the page or following its links, the same as having no robots meta tag, and it is not an instruction that search engines must index the page.
“This is the default value - it does nothing”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-06
The noindex robots rule tells Google not to show the page in search results.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-01glossary term noindex
A slide named internal admin pages, temporary landing pages and thin content as typical use cases for the noindex rule.
Speaker John MuellerEvidence slide photo
The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page, which John Mueller called a weird and very broad rule that makes the page stand on its own.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-06glossary term nofollow
- Extended by D2-C821 Day 2: John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would…
John Mueller recommends rel=nofollow on individual links instead of the page-level nofollow robots rule, so a site can choose which links are useful and which are not.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-06
John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would probably not be part of them.
Speaker John MuellerEvidence transcript
- Extends D2-C063 Day 2: The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page…
Audit robots meta tags set by templates and plug-ins: an explicit all rule does nothing and can go, while a page-level nofollow strips link signals from every link on the page and is better replaced by qualifying only the specific links that need it.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-06
John Mueller said AI crawlers do not really know what to do with nofollow links, because they look at the content rather than building a link graph; he did not say whether he meant Google's AI systems, other AI crawlers or both.
“AI crawlers don't really know what to do with a nofollow link, because they're looking at the content”
Speaker John MuellerEvidence transcript
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
Author Ibrahim Anjro
- Extends D1-C127 Day 1: Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed…
- Extends D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
The robots rule none is equivalent to noindex plus nofollow, so a page that carries it will not show up in Search, provided Google can see the tag.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-06
Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.
Publisher Google Search Central
Used byrequirements DEV-IDX-01, DEV-IDX-03glossary term X-Robots-Tag
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
The nosnippet rule stops Google from showing a text snippet or video preview for a page in search results, while the page's title is still shown.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-05glossary term nosnippet and data-nosnippet
Google's robots meta tag specification says that with nosnippet a static image thumbnail, if available, may still be shown when that results in a better user experience.
Publisher Google Search Central
The nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI Mode in Search.
“Also prevents the content from being used as a direct input for AI Overviews and AI Mode in Search results.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-05glossary term nosnippet and data-nosnippet
- Extends D1-C034 Day 1: Lead-generation, local-service and e-commerce sites should normally stay included, because AI answers cite…
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
“we use the snippet as a way of building out the AI Overviews and the AI Mode answers”
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-05
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
- Extended by D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements.
Publisher Google Search Central
Used byrequirements DEV-AIF-01, DEV-IDX-05glossary term AI Overviews
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
Before adding nosnippet or a low max-snippet value, weigh the cost in AI features: both limit how much of the page AI Overviews and AI Mode can use, and John Mueller described those answers as built from snippets.
Author Ibrahim Anjro
John Mueller said data-nosnippet is rarely needed but lets a site keep a specific piece of text, such as a business phone number, out of the snippet, so the page can still be found for it while searchers have to visit the page to see it.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-07glossary term nosnippet and data-nosnippet
The data-nosnippet attribute can be used on only a handful of HTML elements.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-07
Broken HTML can widen data-nosnippet: if a div marked data-nosnippet is not closed properly, the rest of the page can be blocked from the snippet as well.
Speaker John MuellerEvidence transcript
Used byrequirements DEV-HTM-05, DEV-IDX-07
Google's robots meta tag specification says data-nosnippet may be extracted both before and after rendering, so the attribute should not be added to or removed from existing elements with JavaScript.
Publisher Google Search Central
Used byrequirements DEV-IDX-07, DEV-REN-05
John Mueller said valid HTML is technically not a ranking factor but does matter for controls such as data-nosnippet.
“valid HTML is technically not an SEO ranking factor, but it does play a role”
Speaker John MuellerEvidence transcript
Used byrequirements DEV-HTM-05, DEV-IDX-07
To hide one detail, such as a phone number, a price or a direct answer, instead of the whole snippet, wrap it in data-nosnippet on a span, div or section in the server HTML and validate the HTML so that an unclosed element cannot hide the rest of the page from snippets.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-07
The max-snippet:[number] rule limits a page's search result snippet to that number of characters (100 characters, for example), and max-snippet:0 shows no snippet, the same as nosnippet.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-04glossary term max-snippet
John Mueller said, hedging with 'I think', that the max-snippet limit also applies to AI Overviews and AI Mode.
Speaker John MuellerEvidence transcript
A max-snippet value too short for a useful snippet may lead Google to show no snippet at all, John Mueller said.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-05
max-snippet:-1 removes the length limit and can produce a longer snippet than having no rule, because by default Google keeps snippets to a length it considers reasonable instead of quoting a page at length.
“max-snippet:-1 = No limit. (can be more than without)”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Google's robots meta tag specification says Google chooses the snippet length when no max-snippet rule is set, and with max-snippet:-1 chooses the length it believes most effective; it does not say that -1 produces longer snippets.
Publisher Google Search Central
Used byrequirement DEV-IDX-04
The max-image-preview rule sets the maximum size of a page's image previews: none shows no preview, standard a default-sized one and large the largest possible, for example in Discover.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-04glossary term max-image-preview
John Mueller said Google Search does not always show an image thumbnail for a result, and when it does, the thumbnail is usually limited to the size of the result entry.
Speaker John MuellerEvidence transcript
max-image-preview:large matters mainly in Discover, where it allows a large image that draws people's attention, so the rule can make a page more visible than leaving it out, John Mueller said.
Speaker John MuellerEvidence transcript
- Repeated by D2-C936 Day 2: Setting the max-image-preview robots meta tag to large can make content perform surprisingly well in…
- Extended by D3-C224 Day 3: Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels…
The max-video-preview rule sets the maximum number of seconds of a video preview in Search; 0 allows only a static image and -1 means no time limit.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-VID-05
John Mueller said he thinks setting max-video-preview to -1 will not make Google show an hour-long video preview in search results.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-VID-05
John Mueller speculated that the use of max-video-preview may change in future, for example in Discover, where letting people view a full video could make sense.
Speaker John MuellerEvidence transcript
The notranslate robots rule tells Google not to offer a translation of the page in search results.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-09
John Mueller said notranslate can also be set in a meta tag named google, and that in that form it also turns off Chrome's automatic translation of the page.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-09
Google's Search documentation gives notranslate as a rule for a robots or googlebot meta tag or an X-Robots-Tag header, and says the rule opts a page out of all translation features in Google Search.
Publisher Google Search Central
Used byrequirements DEV-IDX-03, DEV-IDX-09glossary term X-Robots-Tag
A 2008 Search Central blog post says a meta tag named google with the value notranslate stops Google Translate from translating any of a page's content.
Publisher Search Central blog (14 October 2008)
Used byrequirement DEV-IDX-09
John Mueller said he dislikes notranslate because automatic translation is a form of accessibility for visitors who cannot read the page's language.
Speaker John MuellerEvidence transcript
International sites should remove notranslate from their templates unless translation must be prevented: the robots or googlebot form switches off Google Search's translation features, and John Mueller said the form named google also blocks Chrome's translation for visitors who do not read the language.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-09
The noimageindex rule tells Google not to index any of the images on the page, and John Mueller said he could not see why a site would want that.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-09
- Repeated by D2-C934 Day 2: The noimageindex robots meta tag tells Google not to index the images on the page.
The unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-08glossary term unavailable_after
- Extended by D2-C698 Day 2: Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index…
Google's robots meta tag specification says Googlebot considerably decreases the crawl rate of a URL after the date and time set in its unavailable_after rule.
Publisher Google Search Central
Used byrequirement DEV-IDX-08
John Mueller's slide said Search Console has very few indexing-like settings and listed the Search generative AI control as the one control in his talk that is not a meta tag.
“Very few "indexing-like" settings are in Search Console”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Setting the Search generative AI control in Search Console to exclude keeps a site's links and content out of Search generative AI features such as AI Overviews and AI Mode, so the site gets no traffic or impressions from them; include is the default.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-IDX-10
- Repeats D1-C030 Day 1: The Search generative AI control covers AI Overviews, AI Mode and generative AI features in Discover.…
- Extends D1-C029 Day 1: Search Console has a property setting called Search generative AI that gives direct control over AI Overviews…
To be as visible as possible in Google, John Mueller's closing slide recommended the robots rules max-image-preview:large and max-snippet:-1.
“To be as visible as possible in Google, use: max-image-preview:large, max-snippet:-1”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, video, transcript
Used byrequirement DEV-IDX-04
Set max-snippet:-1 and max-image-preview:large on every indexable template unless licensing requires otherwise, and check that no CMS, plug-in or CDN setting adds lower snippet or image preview limits by default.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-04
A slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-REN-05
- Extended by D2-C853 Day 2: John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding…
Google's meta tags documentation recommends avoiding JavaScript to inject or change meta tags whenever possible and testing the implementation thoroughly when it must be used.
Publisher Google Search Central
Used byrequirement DEV-REN-05
Removing a robots restriction such as noindex with JavaScript does not work, a slide said.
“But... it takes more time, and removing restrictions (like "noindex") doesn't work.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-REN-05
- Extended by D2-C852 Day 2: When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript…
Ship robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may skip rendering a page that arrives with noindex, so the page can stay out of the index even if a script removes the tag later.
Author Ibrahim Anjro
Used byrequirement DEV-REN-05
- Extended by D2-C855 Day 2: On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex)…
John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-01
- Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-IDX-01
- Repeats D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript, so a script cannot switch the page back to indexable, John Mueller said.
“we will see the noindex and say, oh, we will get rid of this page; we won't even process the JavaScript”
Speaker John MuellerEvidence transcript
Used byrequirement DEV-REN-05
- Extends D2-C108 Day 2: Removing a robots restriction such as noindex with JavaScript does not work, a slide said.
John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding them with JavaScript only where that is not possible, as in a JavaScript web app.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-REN-05
- Extends D2-C106 Day 2: A slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.
Google's JavaScript SEO basics guide says that when Google encounters a noindex rule it may skip rendering and JavaScript execution, so using JavaScript to change or remove a noindex robots meta tag may not work as expected.
“it may skip rendering and JavaScript execution”
Publisher Google Search Central
Used byrequirement DEV-REN-05
On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex), while Google's guide says only that it may skip rendering; either way, a noindex in the served HTML must never be one that JavaScript is expected to lift.
Author Ibrahim Anjro
- Extends D2-C109 Day 2: Ship robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may…
Google framed rendering historically: when Google started, most of the web was plain, semantic HTML it could extract directly, but the modern web is built with JavaScript and much content is generated by JavaScript, so Google had to adapt.
Speaker Erin SparlingEvidence transcript
Without its ability to render JavaScript, Google says it would not be able to see most of what is on the web today.
Speaker Erin SparlingEvidence transcript
Google renders pages in order to index what users see: its commitment to reflecting the user experience means understanding what people see when a page loads in a browser and surfacing that in Search.
“Rendering Pages to Index What Users See”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Used byglossary term Rendering
Compared with other search engines and AI crawlers, Google said its effort to mimic what the user sees is essential to keeping its knowledge of the web up to date and comprehensive; it did not say what the others do.
Speaker Erin SparlingEvidence transcript
Being able to render and read JavaScript matters whether the client fetching a page is a search crawler or an AI system, Google said.
Speaker Erin SparlingEvidence transcript
Google described the mission of its rendering as simply executing JavaScript, while the implementation is complex, expensive and difficult.
“our mission is very simple these days: execute JavaScript”
Speaker Erin SparlingEvidence transcript
AI Overviews and AI Mode are built on top of Search results: they are a different experience of the same content Google already has.
Speaker Erin SparlingEvidence transcript
Used bystory angle A-001
- Repeats D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
- Extended by D2-C601 Day 2: Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results…
AI Overviews and AI Mode typically do not ground their answers by reading pages live, unlike Gemini when a user asks about a specific page, Google said.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-PRF-02
To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.
“if it works for Search, it works for Gemini for training”
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-IDX-11
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
If Gemini training reuses the rendering done for Search, a site needs no separate rendering work for Gemini; whether its rendered content is used for training is decided with the Google-Extended token in robots.txt, not by rendering choices.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-11
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-PRF-02
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.
Publisher Google
Used byrequirement DEV-IDX-11glossary term Grounding
- Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
Publisher Google
Used byrequirement DEV-IDX-11
- Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
- Extends D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…
Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.
Author Ibrahim Anjro
Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.
Speaker Erin SparlingEvidence transcript
Used byrequirements DEV-PRF-02, DEV-REN-01
- Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
For pages people are likely to ask an AI assistant about (product, pricing, documentation and policy pages), server-render the main content so a live, user-triggered read does not depend on client-side JavaScript finishing quickly.
Author Ibrahim Anjro
Used byrequirements DEV-PRF-02, DEV-REN-01
Google renders pages with Chromium, the browser technology that also underlies Chrome, Edge and other Chromium-based browsers.
Speaker Erin SparlingEvidence transcript
Used byglossary term Rendering
- Repeats D1-C206 Day 1: Google renders JavaScript-heavy pages from their HTML, CSS and JavaScript as a browser would, using the…
A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.
Speaker Erin SparlingEvidence slide photo, transcript
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached system, kept separate because rendering is time-consuming and computationally expensive.
“it shows rendering as part of indexing, but it's actually a detached system”
Speaker Erin SparlingEvidence transcript
- Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
After the crawler has fetched a page, Google's rendering system executes it, checks that it loads properly and works out what it looks like at different sizes.
Speaker Erin SparlingEvidence transcript
Google said its pipeline needs to ensure that content indexable without JavaScript can pass through without rendering, while content that appears only through JavaScript and CSS takes a longer rendering pass.
Speaker Erin SparlingEvidence transcript
Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.
Publisher Google Search Central
Used byrequirement DEV-REN-01glossary term Rendering
- Extended by D3-C676 Day 3: The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page…
The stage remark that content indexable without JavaScript can pass through without rendering does not mean such pages skip rendering, because Google's guide queues every 200 page for rendering; read it as: content already in the raw HTML does not depend on the slower rendering pass.
Author Ibrahim Anjro
Put everything indexing depends on (title, meta description, canonical, robots meta tag, main text and links) in the raw HTML so it is available without rendering, let JavaScript add enhancements only, and check the rendered HTML in URL Inspection for the rest.
Author Ibrahim Anjro
Used byrequirement DEV-REN-01
Google urged sites to make sure their JavaScript content can be crawled, rendered and indexed, calling this important today and also tomorrow, as AI systems increasingly ground answers to user requests.
Speaker Erin SparlingEvidence transcript
- Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
After Google's opener, Lightning session D had three community talks: rendering and JavaScript execution blind spots (Sören Bendig), debugging JavaScript rendering for Search (Rebecca Yu) and virtualizing the browser (Natalia Venditto).
Speaker Erin SparlingEvidence transcript
Many SEOs still treat everything beyond the raw HTML as the developers' business, but developers often do not handle rendering problems, a community speaker warned.
Speaker Sören BendigEvidence transcript
- Extends D1-C121 Day 1: SEO is not only content; it has many parts.
The first rendering blind spot a community speaker showed was content changes: text, images and recommendations that differ between the raw HTML and the rendered page, illustrated by a shop page whose rendered version had a whole extra section.
Speaker Sören BendigEvidence slide photo, transcript
Used byrequirement DEV-REN-07
Switching JavaScript off and on in the browser gives a quick first impression of which content on a page depends on rendering.
Speaker Sören BendigEvidence transcript
A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.
“everything you want cited, include it in the raw HTML, server-side rendered”
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-REN-01
- Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
In a navigation example, the main navigation was in the raw HTML and would survive a rendering failure, but the whole sub-navigation was built by JavaScript, so all of its links would be inaccessible to bots if rendering broke.
Speaker Sören BendigEvidence transcript
When checking which links bots can see, keep in mind that the rendered HTML can change with user actions, a community speaker cautioned.
Speaker Sören BendigEvidence transcript
One site that builds its whole content by rendering also rendered its meta description with HTML tags inside it, which makes no sense and points to a flawed process.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-HTM-03
- Extended by D2-C858 Day 2: Search engines such as Google and Bing would ignore HTML tags written inside a meta description, a community…
Sites that create their content by rendering should have at least basic monitoring to check that the rendered pages come out right, a community speaker advised.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-MON-05
Browser console messages reveal further problems on a page, Content Security Policy violations among them; few teams analyse console messages at scale, but they should, a community speaker said.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-REN-09
On a travel site, the number of offers differed between the raw HTML and the rendered page, possibly because the two versions drew on different internal databases or applied extra filters.
Speaker Sören BendigEvidence transcript
On a retail brand's page, the rendered version promised a bigger discount for a newsletter sign-up than the non-rendered version that was served, a mismatch that can hurt customer satisfaction; the two recordings disagree on the figure the rendered page promised.
Speaker Sören BendigEvidence transcript
Offer counts, discounts and other data should always be the same in the raw HTML and the rendered page, a community speaker advised.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-REN-07
Diff the raw and the rendered HTML of each key template for prices, offer counts, links and meta tags, and treat any difference in commercial data as a bug, because bots and users may get different versions.
Author Ibrahim Anjro
Used byrequirements DEV-MON-05, DEV-REN-07
A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal systems or APIs can be missing and unresolved placeholders can end up in the final snapshot, a community speaker said.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-REN-06
- Extended by D2-C259 Day 2: Erin Sparling named server-side or hybrid rendering, fallbacks and not leaving placeholders in the DOM, among…
- Extended by D2-C263 Day 2: Content is missing from the rendered HTML either because the server does not serve it or because the content…
In one shop's example, brand names and product descriptions were rendered from a database; the brand name was missing from the meta description, and the shop sometimes appeared with unresolved placeholders and sometimes as intended.
Speaker Sören BendigEvidence transcript
Unresolved placeholders caused by rendering timeouts are hard to catch because the problem moves around: it is not always the same page that is broken.
Speaker Sören BendigEvidence transcript
Used byrequirements DEV-MON-05, DEV-REN-06
Broken titles and snippets caused by unresolved placeholders may hurt business, especially for e-commerce shops and affiliate-heavy sites, a community speaker warned.
Speaker Sören BendigEvidence transcript
To catch intermittent rendering problems, archive the full HTML and the resources of each page during a site audit and analyse them yourself, a community speaker advised.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-MON-05
Add a check for unresolved template placeholders (such as {{brand}}, undefined or null) in rendered titles, meta descriptions and visible text to recurring crawls, and render the same URLs more than once, because such failures are intermittent.
Author Ibrahim Anjro
Used byrequirements DEV-MON-05, DEV-REN-06
Content Security Policy lowers the risk of cross-site scripting and clickjacking by defining trusted hosts; a resource from a host that is not trusted is not used.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-REN-09
Content Security Policy is often misunderstood: on a page of a large German banking group, a video that should have been shown was blocked by a CSP violation, which also appeared in the browser console.
Speaker Sören BendigEvidence transcript
- Extended by D2-C859 Day 2: A rendering failure such as a Content Security Policy blocking a video on a landing page that explains how to…
Never block a resource that a page needs for rendering with robots.txt, a community speaker said, calling this common sense.
Speaker Sören BendigEvidence transcript
When something is missing from a rendered page, check both robots.txt and the Content Security Policy: the URL Inspection live test shows the page resources, the JavaScript console output and a screenshot of the rendered page.
Author Ibrahim Anjro
Used byrequirements DEV-MON-02, DEV-REN-09
A page can render completely blank while its raw HTML looks fine, so a team checking only the raw HTML thinks all is well; causes include server misconfiguration, overload and bad scripts.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-MON-05
A community speaker summarised five rendering blind spots to check: content and markup changes, invisible links, rendering timeouts, Content Security Policy and blocked resources.
Speaker Sören BendigEvidence slide photo
Because even established brands have these rendering problems and every modern website uses a lot of JavaScript, a community speaker argued that most sites probably have them too and that ignoring rendering is reckless.
Speaker Sören BendigEvidence transcript
A community speaker strongly advised SEOs to use Chrome DevTools more often to understand what happens when their pages render.
Speaker Sören BendigEvidence transcript
- Extended by D2-C860 Day 2: Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a…
Third-party systems billed per session that also run when bots render a page can cost a high-traffic site thousands of euros per month, a community speaker said, adding that he sees this regularly.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-PRF-03
Google says Googlebot and its Web Rendering Service identify resources that do not contribute to essential page content, such as reporting and error requests, and may not fetch them, so client-side analytics may not give a full or accurate picture of their activity.
Publisher Google Search Central
Used byrequirement DEV-MON-04
Compare the session counts that per-session-billed tools (chat, personalisation, testing, session replay) charge for with real user sessions, and load such tools only after a user interaction where they add no content that needs to be indexed.
Author Ibrahim Anjro
Used byrequirement DEV-PRF-03
The second community lightning talk of the rendering block covered debugging JavaScript rendering for Search: how Google renders pages, a step-by-step check, four common mistakes and a worked product page example.
Speaker Rebecca YuEvidence transcript
The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.
Speaker Rebecca YuEvidence slide photo, transcript
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
Speaker Rebecca YuEvidence slide photo, transcript
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
- Extends D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.
Publisher Google Search Central
Used byrequirement DEV-URL-01
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
After processing, an indexable page is placed in Google's render queue to wait for rendering.
Speaker Rebecca YuEvidence slide photo, transcript
Used byglossary term Render queue
- Extended by D3-C626 Day 3: Google renders pages in two ways: immediately after crawling, or later through a queue-based process that…
A community speaker's slide put the usual wait in Google's render queue at seconds to a couple of minutes per page.
“Usually from seconds to a couple minutes.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo
Google's JavaScript SEO basics guide says a page may wait in the render queue for a few seconds but that it can take longer, and it gives no upper limit.
“The page may stay on this queue for a few seconds, but it can take longer than that.”
Publisher Google Search Central
Used byrequirements DEV-PRF-02, DEV-REN-01glossary term Render queue
- Extended by D3-C627 Day 3: Content that JavaScript adds to a page is typically seen by Google's indexing system within a few hours, and…
Google's renderer is a headless Chromium that runs the page's JavaScript.
Speaker Rebecca YuEvidence transcript, slide photo
- Extends D1-C206 Day 1: Google renders JavaScript-heavy pages from their HTML, CSS and JavaScript as a browser would, using the…
Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.
“Whatever is in the DOM at that moment is what likely gets indexed.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo, transcript
- Repeated by D2-C262 Day 2: Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.
If a page keeps loading content indefinitely, not all of that content will be indexed, because Google's rendering does not go on forever (this passage of the recording is partly unclear).
Speaker Rebecca YuEvidence transcript
Used byrequirement DEV-REN-06
To check whether JavaScript content is affected, open the URL Inspection tool in Search Console or the Rich Results Test and look for that content in the rendered HTML tab.
Speaker Rebecca YuEvidence transcript
Used byrequirement DEV-MON-02glossary terms Raw HTML and rendered HTML, URL Inspection tool
A community speaker treats the rendered HTML shown in Google's testing tools as the final source of truth, because what Google sees there is what gets indexed.
“I always treat this page as the final source of truth”
Speaker Rebecca YuEvidence transcript
Google's URL Inspection help says the tool's live test examines a URL in real time, so its results can differ from Google's indexed version of the page.
Publisher Google Search Console Help
Used byrequirement DEV-MON-02glossary term URL Inspection tool
Check rendering per template rather than per URL: inspect one URL of each page type (product, category, article) in URL Inspection and confirm that the main content, prices and internal links appear in the rendered HTML.
Author Ibrahim Anjro
Used byrequirement DEV-MON-02
If content is missing from the rendered HTML, find the script responsible in Chrome DevTools: open the Network tab, filter by Fetch/XHR and reload the page.
Speaker Rebecca YuEvidence transcript
For each suspect script or API request in Chrome DevTools, check whether it runs, whether something blocks it and whether robots.txt disallows it.
Speaker Rebecca YuEvidence transcript
If the cause of missing JavaScript content is still unclear after checking the network requests, search the source code for a string related to the missing content.
Speaker Rebecca YuEvidence transcript
A community speaker listed four common but often overlooked JavaScript rendering mistakes: non-crawlable link markup, error pages that return HTTP 200 in single-page apps, content behind user interaction, and JavaScript or API resources blocked in robots.txt.
Speaker Rebecca YuEvidence transcript, 3 slide photos
Used byglossary term Single-page app (SPA)
Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.
Speaker Rebecca YuEvidence slide photo, transcript
Used byrequirement DEV-URL-02
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
- Extends D2-C040 Day 2: Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not…
The crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>, combined with History API routing (window.history.pushState) instead of hash routes.
Speaker Rebecca YuEvidence slide photo
Used byrequirement DEV-URL-03
- Repeated by D2-C285 Day 2: Google recommends the History API to give single-page apps clean URLs instead of fragment-based routes.
- Extended by D2-C288 Day 2: With the History API, a single-page app can use real links and attach event listeners that intercept the…
Non-crawlable link markup, such as onclick links and hash pseudo-links, is common in single-page web apps, and sites that use faceted navigation should check their links for it.
Speaker Rebecca YuEvidence transcript
Used byrequirement DEV-URL-02glossary term Single-page app (SPA)
A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.
“The nav works perfectly for users, and the entire alternate-language cluster is orphaned.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo
Used byrequirement DEV-INT-06
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.
Author Ibrahim Anjro
Used byrequirement DEV-URL-08
- Extends D1-C101 Day 1: Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable…
Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.
Author Ibrahim Anjro
Used byrequirement DEV-INT-06
- Extends D1-C067 Day 1: A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts…
In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.
Speaker Rebecca YuEvidence transcript
Used byrequirement DEV-ERR-02glossary term Single-page app (SPA)
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
- Repeated by D2-C291 Day 2: In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200…
The fix for a client-side soft 404 is to make a missing page end with a real HTTP 404 status code instead of 200.
Speaker Rebecca YuEvidence transcript
Used byrequirement DEV-ERR-02
- Extended by D2-C292 Day 2: The fix for a client-side soft 404 is to serve a real 404 error page where appropriate; how to detect URLs…
Google's JavaScript guides say that when client-side routing makes a real 404 status impractical, a single-page app can avoid soft 404s by redirecting with JavaScript to a URL whose server returns 404, or by adding a robots noindex meta tag with JavaScript.
Publisher Google Search Central
Used byrequirement DEV-ERR-02
In a single-page app, let the server return 404 for unknown routes where possible; otherwise use one of the two client-side fixes Google documents for error views: a JavaScript redirect to a URL that returns 404, or a robots noindex added with JavaScript.
Author Ibrahim Anjro
A community speaker called not hiding content behind user interaction the most important of the four JavaScript rendering mistakes.
Speaker Rebecca YuEvidence transcript
Content that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.
Speaker Rebecca YuEvidence slide photo, transcript
Used byrequirement DEV-REN-02
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
- Repeated by D2-C268 Day 2: Content that loads only when a user clicks an element is not supported in the way Google renders pages for…
For rendering, what matters is whether content is present in the DOM, not whether it is visible on screen: hidden content can be indexed, absent content cannot.
“Not visible versus hidden. Present versus absent in the DOM.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo
Used byrequirement DEV-REN-02
Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.
“Googlebot doesn't scroll. Never runs.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo, transcript
Used byrequirement DEV-REN-03glossary term Lazy loading
- Extends D1-C114 Day 1: A Google panelist called pagination one of the trickiest things in web development and said switching to…
- Repeated by D2-C271 Day 2: Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on…
Lazy loading should be triggered by an Intersection Observer that watches a sentinel element instead of by a scroll event listener.
Speaker Rebecca YuEvidence slide photo, transcript
Used byrequirement DEV-REN-03glossary term Lazy loading
An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders with is tall, so content lazy-loaded this way can load during rendering.
“Fires on viewport entry. Viewport is tall.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo
Used byrequirement DEV-REN-03
- Extended by D2-C272 Day 2: Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.
- Repeated by D2-C274 Day 2: Content loaded as elements enter the viewport, for example with an Intersection Observer, does load when…
In Google's March 2023 SEO office hours, John Mueller said Google handles infinite scroll with viewport expansion, rendering a page like a very long phone, which is not very efficient and can miss content, so pagination links are strongly recommended.
“this is done through a technique called "viewport expansion", where we render a page like a very long phone.”
Publisher Google Search Central (SEO office hours transcript, March 2023)
Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () => fetch('/api/specs'), does not exist for Google until someone clicks.
“Doesn't exist until someone clicks”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo
Used byrequirement DEV-REN-02
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden attribute; Google indexes such hidden content.
“CSS-hidden is fine. Google indexes it.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo, transcript
Used byrequirement DEV-REN-02
- Extended by D2-C866 Day 2: Content inside tabs, for example separate tabs for a product description and a manufacturer description…
Audit tabs, accordions and 'load more' lists for content fetched on click or scroll: put tab content in the initial DOM and hide it with CSS, trigger lazy loading with an Intersection Observer, and give long lists paginated URLs linked with <a href>.
Author Ibrahim Anjro
Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.
“If the renderer can't fetch it, the renderer can't run it.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo, transcript
Used byrequirement DEV-REN-04
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Repeated by D2-C265 Day 2: Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the…
- Repeated by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
Disallowing a script folder such as /static/js/ or a generic /api/ folder in robots.txt can block the endpoints that supply a page's content, leaving blank modules and missing content after rendering.
Speaker Rebecca YuEvidence slide photo, transcript
Google says its Web Rendering Service fetches the resources a page references through Googlebot, including JavaScript, CSS and XHR requests to APIs, but not images or videos.
Publisher Search Central blog (3 December 2024), Search Central blog (31 March 2026)
Used byrequirement DEV-REN-04
If an API folder must stay blocked in robots.txt, allow the endpoints that rendering needs, for example Disallow: /api/ together with Allow: /api/products/.
“Carve out only what rendering needs.”
Wording checked against the slide or recording
Speaker Rebecca YuEvidence slide photo, transcript
Used byrequirement DEV-REN-04
A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.
Author Ibrahim Anjro
- Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
robots.txt rules apply per host, so the rules in example.com/robots.txt do not apply at all to an API served from api.example.com.
Speaker Rebecca YuEvidence slide photo
Used byrequirement DEV-SRV-04glossary term robots.txt
An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.
Author Ibrahim Anjro
Used byrequirement DEV-SRV-04
- Extends D1-C074 Day 1: If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the…
On a product page whose data is fetched client-side, the rendering mistakes described can leave Google seeing no content at all while users still see the price and images.
Speaker Rebecca YuEvidence transcript
In the speaker's product page example, three of the mistakes combine: product links behind onclick handlers keep the product detail pages hidden, the product API is blocked in robots.txt, and some content waits for a user interaction.
Speaker Rebecca YuEvidence transcript
A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.
Speaker Rebecca YuEvidence transcript
Used byrequirements DEV-ERR-03, DEV-REN-04
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
- Extended by D2-C336 Day 2: Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty…
A community speaker summed up JavaScript rendering issues as cases where search engines cannot execute, access or trigger the code needed to display a page's core content.
“JavaScript rendering issues happen when search engines cannot execute, access or trigger the code required to display your core content.”
Speaker Rebecca YuEvidence transcript
To find rendering issues, compare a page's raw source code with its rendered versions, including the rendered HTML that Google's testing tools show; the differences point to the problems.
Speaker Rebecca YuEvidence transcript
Used byrequirement DEV-MON-02glossary term Raw HTML and rendered HTML
Natalia Venditto, a Principal Software Engineer at Adobe, gave a seven-minute community lightning talk on JavaScript and web standards, built around the Web Fragments library.
Speaker Natalia VendittoEvidence transcript
Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard, and that this is what the Web Fragments work is trying to do.
Speaker Natalia VendittoEvidence transcript
- Repeats D2-C052 Day 2: John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet…
Natalia Venditto said teams are increasingly asked to integrate AI-generated content and applications into existing web applications and want to do so without breaking the host application, a situation she called a typical micro-frontend scenario.
Speaker Natalia VendittoEvidence transcript
Natalia Venditto named three typical failure modes by which an embedded app can break the application hosting it: collisions in the global JavaScript scope and module registry, CSS bleed, and fate sharing.
Speaker Natalia VendittoEvidence transcript
CSS bleed, as Natalia Venditto described it, means styles leak between an embedded app and its host, so that suddenly everything looks like the embedded app or the embedded app takes on the look of the host.
Speaker Natalia VendittoEvidence transcript
Fate sharing means a failure in an embedded app spreads to its host, for example an unhandled exception in the embedded app that ends up breaking the host application.
Speaker Natalia VendittoEvidence transcript
Pages that embed third-party or AI-generated apps can share their fate, since an unhandled exception in the embedded code can break the host page; check the rendered HTML and JavaScript console output of such pages in Google's testing tools to make sure the host's own content still renders.
Author Ibrahim Anjro
Used byrequirement DEV-REN-09
A slide titled Orchestration approaches compared four ways to combine micro-frontends in one page (iframe, Shadow DOM, rewriting everything, and reframing with Web Fragments) on three criteria (JavaScript isolated, styles isolated, part of the page) and named the catch of each.
Speaker Natalia VendittoEvidence slide photo, transcript
An iframe, the long-established way to embed an app, isolates both JavaScript and styles but is not part of the page: it is walled off from the host's DOM, navigation and layout, which Natalia Venditto said brings many problems with accessibility, layout and navigation.
“Walled off from DOM, navigation, layout”
Wording checked against the slide or recording
Speaker Natalia VendittoEvidence slide photo, transcript
Before embedding an app or widget in an iframe, test it with a keyboard and a screen reader as well as for layout and navigation, because the isolation an iframe gives comes at those costs.
Author Ibrahim Anjro
Google says its systems generally try to index the content of a page embedded with an iframe as part of the page that embeds it, but this is not guaranteed because both pages are also normal HTML pages on their own.
Publisher Google Search Central (SEO office hours transcript, December 2023), Search Central blog (21 January 2022)
Used byrequirement DEV-REN-10
Shadow DOM isolates styles and keeps the embedded content part of the page, but it does not isolate JavaScript: the embedded code still shares the host's JavaScript globals.
“Still shares JS globals”
Wording checked against the slide or recording
Speaker Natalia VendittoEvidence slide photo, transcript
Used byglossary term Shadow DOM
Rewriting everything into one application keeps the result part of the page, but JavaScript and style isolation then have to be done by hand.
Speaker Natalia VendittoEvidence slide photo, transcript
The catch the comparison slide gave for rewriting everything was that you cannot rewrite code you did not write; Natalia Venditto explained that AI cannot rewrite code already in place because it does not know its requirements, design system or other constraints.
“Can't rewrite code you didn't write”
Wording checked against the slide or recording
Speaker Natalia VendittoEvidence slide photo, transcript
Reframing with Web Fragments was the only approach on the comparison slide ticked for all three criteria (JavaScript isolated, styles isolated, part of the page); its stated catch is that it relies on browser patches today.
“Needs patches today”
Wording checked against the slide or recording
Speaker Natalia VendittoEvidence slide photo, transcript
Web Fragments is a lightweight library that puts a hidden iframe in the host page and uses it as an isolated sandbox to run the embedded app's JavaScript, not to render the app.
Speaker Natalia VendittoEvidence transcript
Web Fragments fetches the embedded app's assets from a remote endpoint through middleware that acts as a gateway.
Speaker Natalia VendittoEvidence transcript
Web Fragments reframes the embedded app's DOM by placing it inside a shadow DOM in the host page, while the app's JavaScript runs in the hidden iframe.
Speaker Natalia VendittoEvidence transcript
Google's documentation says Google supports web components and flattens shadow DOM and light DOM content when it renders a page, and that content not visible in the rendered HTML cannot be indexed.
Publisher Google Search Central
Used byrequirement DEV-REN-10glossary term Shadow DOM
With Web Fragments the page ends up as a single document in which the host application does not know the embedded app runs inside it and the embedded app does not know it lives in another document.
Speaker Natalia VendittoEvidence transcript
Web Fragments monkey-patches browser APIs to containerize the browser, recreating inside it an architecture similar to Docker containers on the back end.
“we monkey patch the browser to containerize it”
Speaker Natalia VendittoEvidence transcript
The browser APIs Web Fragments patches include document, history and location; Natalia Venditto said they are virtualized rather than hijacked, so controls work exactly as in a normal application.
Speaker Natalia VendittoEvidence transcript
Web Fragments retargets dispatchEvent calls to the fragment's shadow root and resolves DOM calls such as appendChild and element lookups against the main document.
Speaker Natalia VendittoEvidence transcript
Because Web Fragments monkey-patches core browser APIs such as document, history and location, test pages that use it in Google's rendering (the rendered HTML in Search Console's URL Inspection tool or the Rich Results Test), not only in a normal desktop browser.
Author Ibrahim Anjro
Used byrequirement DEV-REN-10
Natalia Venditto said Web Fragments suits AI-generated apps because a web fragment used as a custom element sandboxes the app's rendering, avoiding collisions and CSS bleed between the app and its host in either direction.
Speaker Natalia VendittoEvidence transcript
A slide set out Web Fragments in five steps: an LLM writes the app and adds a custom element such as <web-fragment fragment-id="some-id"> to the HTML, a standalone HTTP endpoint is set up, the library is imported and initialized with initializeWebFragments(), the fragment is registered in the gateway, and step 5 is a celebration emoji.
Speaker Natalia VendittoEvidence slide photo, transcript
Natalia Venditto said the five Web Fragments steps are all that is needed to run a containerized application fully on the client side.
Speaker Natalia VendittoEvidence transcript
A web fragment is served from its own standalone HTTP endpoint, so the embedded app can be deployed anywhere, separately from the host page.
Speaker Natalia VendittoEvidence slide photo, transcript
Make sure robots.txt does not block the URLs from which the browser loads a fragment's assets (the gateway paths on the host page's origin, or the endpoint's own host if assets load from there), because Google does not render JavaScript from blocked files and robots.txt rules apply per host.
Author Ibrahim Anjro
Natalia Venditto said Web Fragments is framework-agnostic and vendor-agnostic and works the same whatever JavaScript framework the embedded app uses.
Speaker Natalia VendittoEvidence transcript
Natalia Venditto recommended that web fragments be server-rendered.
Speaker Natalia VendittoEvidence transcript
Used byrequirement DEV-REN-01
Natalia Venditto reported an enormous performance increase (no figure given) when an application is reframed inside a fully client-side host, because the reframed app is rendered and interactive immediately.
Speaker Natalia VendittoEvidence transcript
Natalia Venditto said a server-rendered app reframed into the host page is fully indexable.
Speaker Natalia VendittoEvidence transcript
Used byrequirement DEV-REN-01
Because Google flattens shadow DOM when it renders, content that Web Fragments places in a shadow root can in principle be indexed with the host page, but only if it appears in the rendered HTML; check that in URL Inspection rather than relying on a library's promise of indexability.
Author Ibrahim Anjro
Used byrequirement DEV-REN-10
Natalia Venditto said the goal of the Web Fragments team is not to maintain monkey patches but to standardize the approach.
Speaker Natalia VendittoEvidence transcript, slide photo
Natalia Venditto said the case for standardizing Web Fragments is that it already relies on many existing browser APIs, built that way so as not to rewrite much or reinvent the wheel.
Speaker Natalia VendittoEvidence transcript
Natalia Venditto said ShadowRealm, a JavaScript standard proposal, was at stage 2.7 at the time of the talk, which she glossed as meaning that only implementation is missing.
“the only thing missing is implementation”
Speaker Natalia VendittoEvidence transcript
Natalia Venditto said the ShadowRealm proposal would do much of what Web Fragments does today with patches: containerizing and isolating execution.
Speaker Natalia VendittoEvidence transcript
- Extended by D2-C281 Day 2: Polyfills work in two directions: they patch backwards where a browser lacks support, or patch forward to…
Natalia Venditto said the Web Fragments team supports the ShadowRealm standard proposal.
Speaker Natalia VendittoEvidence transcript
Websites that rely heavily on JavaScript and rendering have common blind spots, because these features often break in subtle ways, a community speaker said.
Speaker Sören BendigEvidence transcript
None of the rendering problems a community speaker showed from established brands' sites had been fixed promptly: some were fixed by the time of the talk and some were still live.
Speaker Sören BendigEvidence transcript
Search engines such as Google and Bing would ignore HTML tags written inside a meta description, a community speaker said.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-HTM-03
- Extends D2-C142 Day 2: One site that builds its whole content by rendering also rendered its meta description with HTML tags inside…
A rendering failure such as a Content Security Policy blocking a video on a landing page that explains how to open a business account could have a drastic impact for a purely online business, a community speaker warned.
Speaker Sören BendigEvidence transcript
- Extends D2-C156 Day 2: Content Security Policy is often misunderstood: on a page of a large German banking group, a video that…
Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a modern crawler.
Speaker Sören BendigEvidence transcript
Used byrequirement DEV-MON-05
- Extends D2-C162 Day 2: A community speaker strongly advised SEOs to use Chrome DevTools more often to understand what happens when…
Google renders nearly all of the web by replicating what a browser does, using a real browser's rendering engine.
“Google renders nearly all of the internet by replicating browser actions”
Speaker Erin SparlingEvidence transcript
- Extended by D3-C628 Day 3: Gary Illyes said Google keeps saying it renders every URL on the internet and that he would say this is not…
Google's JavaScript SEO basics says every page with a 200 status code is queued for rendering, whether or not it uses JavaScript, unless a robots meta tag or header says not to index it; for non-200 pages such as 404 error pages, rendering might be skipped.
Publisher Google Search Central
Used byrequirement DEV-REN-05
If JavaScript rendering fails, what Google indexes for a page differs from what the page shows once it is rendered.
Speaker Erin SparlingEvidence transcript
Erin Sparling named server-side or hybrid rendering, fallbacks and not leaving placeholders in the DOM, among other measures, as ways to guard against failed JavaScript rendering.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-REN-06
- Extends D2-C149 Day 2: A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal…
Google listed four common JavaScript problems for indexing: content not present in the rendered HTML, using # in URLs for content changes, soft 404s and blocked resources.
Speaker Erin SparlingEvidence slide photo, transcript
Google's slide called content missing from the rendered HTML the most common JavaScript issue for indexing.
“The most common issue.”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirement DEV-REN-02
Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.
“If it's not in the final DOM, Google can't see it.”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirement DEV-REN-02
- Repeats D2-C174 Day 2: Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.
Content is missing from the rendered HTML either because the server does not serve it or because the content has still not appeared after some period of time.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-REN-06
- Extends D2-C149 Day 2: A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal…
Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but returns a 200 HTTP status code.
“Your application serves a "Not Found" message but returns a 200 HTTP status code.”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirement DEV-ERR-02
- Repeats D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the crawling of critical JavaScript files or API endpoints.
“robots.txt disallowing crawling of critical .js or API endpoints.”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirement DEV-REN-04
- Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
Google named three causes of content missing from the rendered HTML: JavaScript inaccessible to Googlebot, JavaScript DOM event triggers and disabled browser APIs.
Speaker Erin SparlingEvidence slide photo, video, transcript
When the JavaScript is not accessible to Googlebot, Google cannot render the client-side DOM that the page builds asynchronously, so parts of the page may be present while the main content is absent.
Speaker Erin SparlingEvidence transcript
Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-REN-02glossary term Lazy loading
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
- Repeats D2-C195 Day 2: Content that loads only after a user action such as a click or a scroll is not in the DOM while Google…
Browser APIs that need user permission, such as payments and location, are not supported when Google renders a page.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-REN-08
Google said there is no single fix for content missing from the rendered HTML: the solution varies by case, and the advice is to follow best practices.
“Solution: varies by case, but follow best practices.”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirements DEV-REN-03, DEV-URL-07
- Extends D1-C114 Day 1: A Google panelist called pagination one of the trickiest things in web development and said switching to…
- Repeats D2-C197 Day 2: Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll'…
- Extends D1-C491 Day 1: If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them…
Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.
“it renders it in a very tall viewport. Specifically, around 10,000 pixels is what the viewport gets rendered as.”
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirements DEV-REN-03, DEV-URL-07glossary term Viewport expansion
- Extends D2-C199 Day 2: An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders…
Google's March 2023 SEO office hours say Google sees infinite-scroll content through viewport expansion, rendering a page like a very long phone, which is not particularly efficient and can miss infinite content.
“a technique called "viewport expansion", where we render a page like a very long phone”
Publisher Google Search Central (SEO office hours transcript, March 2023)
Used byrequirement DEV-URL-07glossary term Viewport expansion
Content loaded as elements enter the viewport, for example with an Intersection Observer, does load when Google renders a page, because Google's rendering viewport is very tall.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-REN-03
- Repeats D2-C199 Day 2: An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders…
Treat the 10,000-pixel viewport as a safety margin, not a target: load lazy content when it enters the viewport (Intersection Observer or native lazy loading, never scroll listeners), and give infinite-scroll content paginated URLs linked with <a href> so anything below the first rendered screen stays reachable.
Author Ibrahim Anjro
To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each chunk its own persistent, unique URL, link sequentially to those URLs, and update the displayed URL with the History API when a new chunk becomes the main visible element.
Publisher Google Search Central
Used byrequirement DEV-URL-07
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
WebGL does not work well in Google's rendering: a WebGL shader that makes a page look like it is underwater will not be applied when Google renders the page.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-REN-08
For effects that need WebGL, which Googlebot does not support, Google's guide to fixing JavaScript problems suggests skipping the effect or prerendering it with server-side rendering so that the content is accessible to Googlebot.
Publisher Google Search Central
Used byrequirement DEV-REN-08
Erin Sparling recommended differential serving and polyfills to keep a site compatible with Google's rendering engine where it lacks a browser feature.
Speaker Erin SparlingEvidence transcript
Google's JavaScript SEO basics recommends differential serving and polyfills when feature detection finds a missing browser API, and warns that some browser features cannot be polyfilled.
“We recommend using differential serving and polyfills if you feature-detect a missing browser API that you need.”
Publisher Google Search Central
Used byrequirement DEV-REN-08
Polyfills work in two directions: they patch backwards where a browser lacks support, or patch forward to build now for a feature expected in the future, as Web Fragments does.
Speaker Erin SparlingEvidence transcript
- Extends D2-C254 Day 2: Natalia Venditto said the ShadowRealm proposal would do much of what Web Fragments does today with patches…
Wrap permission-based features (location, payments, camera) and WebGL effects in feature detection with a fallback, so the main content renders when the API is missing or the permission is declined; never make indexable content wait for a permission prompt.
Author Ibrahim Anjro
Used byrequirement DEV-REN-08
Erin Sparling called URL fragments (the part of a URL after #) the next most common JavaScript issue seen, after rendering problems and blocked rendering.
Speaker Erin SparlingEvidence transcript
URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.
“Fragment identifiers (#) are often ignored by crawlers.”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirement DEV-URL-03
- Repeats D1-C101 Day 1: Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable…
Google recommends the History API to give single-page apps clean URLs instead of fragment-based routes.
“Use the History API for clean URLs in SPAs.”
Wording checked against the slide or recording
Speaker Erin SparlingEvidence slide photo, transcript
Used byrequirement DEV-URL-03glossary term History API
- Repeats D2-C185 Day 2: The crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>…
Moving a single-page app from fragment-based routes to real paths with the History API keeps the same client-side behaviour and was described as not free but relatively straightforward.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-URL-03
A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-URL-01
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
- Extends D2-C040 Day 2: Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not…
With the History API, a single-page app can use real links and attach event listeners that intercept the click, rewrite the URL with pushState and load the new content, so Google can follow the links while users avoid full page reloads.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-URL-03glossary terms History API, Single-page app (SPA)
- Extends D2-C185 Day 2: The crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>…
Real URLs also work as deep links: Android, iOS and desktop operating systems accept full URLs as keys to specific content in an app, so clean URLs simplify the cross-platform user experience, not only indexing.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-URL-03
Audit a single-page app by searching both the templates and the rendered HTML for href="#..." routes, <a> elements without href and onclick navigation; every view that should rank needs an <a href> link to a real path.
Author Ibrahim Anjro
In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-ERR-02
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
- Repeats D2-C190 Day 2: In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because…
The fix for a client-side soft 404 is to serve a real 404 error page where appropriate; how to detect URLs that do not exist depends on how the app works and where it is hosted.
Speaker Erin SparlingEvidence transcript
- Extends D2-C191 Day 2: The fix for a client-side soft 404 is to make a missing page end with a real HTTP 404 status code instead of…
In an app routed with the History API, when a user reaches a page that does not exist, Erin Sparling suggested asking how they got there and redirecting to a real 404 page.
Speaker Erin SparlingEvidence transcript
For client-side rendered single-page apps, where meaningful status codes can be impossible or impractical, Google's documentation gives two ways to avoid soft 404s: a JavaScript redirect to a URL that returns a 404 status, or a noindex robots meta tag added with JavaScript.
Publisher Google Search Central
Used byrequirement DEV-ERR-02
After moving to real paths, set the server or hosting rewrite rules so a direct request to every valid path returns 200 with the content (ideally server-rendered) and an unknown path returns a 404 status, which removes client-side soft 404s at the source.
Author Ibrahim Anjro
Used byrequirement DEV-ERR-02
If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-REN-04
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
Google's introduction to robots.txt says unimportant image, script or style files may be blocked, but resources whose absence makes a page harder for Google to understand should not be blocked.
Publisher Google Search Central
Used byrequirement DEV-REN-04
Check robots.txt for Disallow rules covering JavaScript bundles, build folders or API endpoints the page calls while rendering, including on separate API or CDN hostnames with their own robots.txt, then confirm in URL Inspection's live test that the rendered HTML contains the main content.
Author Ibrahim Anjro
A headless content management system serves its content as an API; even WordPress, often seen as one monolithic application, can be used headless, with only its editor or only its renderer.
Speaker Erin SparlingEvidence transcript
Erin Sparling's practice with a headless CMS is to give an AI agent the content's JSON Schema and have it build a throwaway prototype front end, to see the content before deciding how to render it; the prototype is not meant for launch.
Speaker Erin SparlingEvidence transcript
For headless CMS content, Erin Sparling presented rendering on the server, not only in the browser, as a way to avoid indexing issues.
Speaker Erin SparlingEvidence transcript
When one content schema feeds several front ends (the example was a news UI, then a recipe UI), Erin Sparling had an AI agent add automated browser UI tests and agent controls so the interfaces could be managed and tested.
Speaker Erin SparlingEvidence transcript
Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.
Speaker Erin SparlingEvidence transcript
- Extends D1-C281 Day 1: A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions…
When Search Console reports that content is not rendering, Erin Sparling suggested putting AI agents on the problem, but only after doing your own due diligence.
Speaker Erin SparlingEvidence transcript
With a headless CMS, choose the rendering per template before launch (server-rendered or pre-rendered for pages that must rank), and make any agent-built test suite assert that the main content and <a href> links are present in the rendered HTML.
Author Ibrahim Anjro
To use Search Console to debug what Google can see on a site and why, Erin Sparling said the first step is to get access to the site by verifying ownership (the speaker's words were authorized domains).
Speaker Erin SparlingEvidence transcript
Used byrequirement DEV-MON-01
Erin Sparling showed a Search Console result saying a URL would be indexed only under certain conditions, one of which had not been met, and contrasted it with the result for an available URL whose content loads.
Speaker Erin SparlingEvidence transcript
Google's URL Inspection help says a valid live test only confirms that Google can access a page for indexing; the page must still meet other conditions to be indexed, such as having no manual action, not being a duplicate and being of high enough quality.
Publisher Google Search Console Help
Used byrequirement DEV-MON-02
Google's guide to fixing Search-related JavaScript problems says its Web Rendering Service keeps no state across page loads: local storage, session storage and HTTP cookies are cleared.
Publisher Google Search Central
Used byrequirement DEV-REN-06
Google's guide to fixing Search-related JavaScript problems says the Web Rendering Service may ignore caching headers and so use outdated JavaScript or CSS, and recommends content fingerprinting, which puts a hash of the content in the file name, as in main.2bb85551.js.
Publisher Google Search Central
Used byrequirement DEV-REN-06
A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
Speaker Gary IllyesEvidence video, transcript
Used byrequirement DEV-HTM-01
- Repeats D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
- Extended by D2-C861 Day 2: Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not…
Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.
“When Google indexes a page, it determines the primary content (or centerpiece) of each page.”
Publisher Google Search Central
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.
Speaker Gary IllyesEvidence transcript
- Extended by D3-C170 Day 3: Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content…
When Google processes a page for indexing, it gives words different weights depending on the part of the page where they appear.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-02
Words in the footer of a page get a lower weight, so text placed in the footer is unlikely to contribute much to ranking the page.
“if you put something in a footer, it's more likely that it's not going to contribute much to ranking”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-02
On Google's example blog page, the post title and opening sentence counted as important because they sit in the main content, in front of the user, while the site tagline, the 'Categories' sidebar and category links such as 'Hugo (7)' counted as less important supplementary text.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-HTM-02
To make a word count for ranking a page, Gary Illyes said the simplest step is to move it into the main content, because where text sits on a page already contributes quite a bit to ranking.
“where you position text on a page will already contribute quite a bit to ranking that page”
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-HTM-02
Gary Illyes said not everything on a page can or should be important: if everything were placed in the main content, nothing would stand out as main content, which he called working as intended.
Speaker Gary IllyesEvidence transcript
Put the terms a page should rank for in its main content (title, headings, opening paragraphs) rather than only in sidebars, tag lists or footers; footer keyword blocks are unlikely to add ranking value.
Author Ibrahim Anjro
Used byrequirement DEV-HTM-02
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
Speaker Gary IllyesEvidence slide photo, transcript
Used byglossary term Tokenization
- Repeated by D2-C721 Day 2: Google's Search index does not hold the full content of pages; Google said storing full pages and pulling…
- Extended by D3-C109 Day 3: The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens…
For languages written with spaces between words, such as English and German, Search tokenization splits a sentence into its individual words.
Speaker Gary IllyesEvidence transcript
Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
Speaker Gary IllyesEvidence transcript
- Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
- Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
Speaker Gary IllyesEvidence transcript
- Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
- Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
Gary Illyes said a colleague, John, would cover how Google interprets the words of a query on the morning of Day 3.
Speaker Gary IllyesEvidence transcript
When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-HTM-07
- Repeated by D2-C720 Day 2: Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the…
Google also stores spam metadata with the tokens of a page, for example that text was white on a white background, so ranking can use that information.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-06
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Speaker Gary IllyesEvidence slide photo, transcript
- Contradicts D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.
Speaker Gary IllyesEvidence slide photo, transcript
Used byglossary term Tokenization
Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.
Speaker Gary IllyesEvidence slide photo, transcript
Used byglossary term Tokenization
Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.
Speaker Gary IllyesEvidence 2 slide photos
Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.
Author Ibrahim Anjro
Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real, but it matters at the level of an AI model's context window.
Speaker Gary IllyesEvidence transcript
Used bymyth M-002
- Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
Speaker Gary IllyesEvidence transcript
Used bymyth M-002
- Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
- Extended by D2-C869 Day 2: Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps…
Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-AIF-03myth M-002
- Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
- Extended by D2-C825 Day 2: Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window…
Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-AIF-03
- Extends D2-C332 Day 2: Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller…
Do not rewrite pages into short, self-contained chunks for AI systems; Google says Gemini reads context windows of millions of tokens, so structure content for readers, with clear headings and complete explanations.
Author Ibrahim Anjro
Used byrequirement DEV-AIF-03
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Speaker Gary IllyesEvidence 2 slide photos, transcript
Used byrequirements DEV-ERR-01, DEV-ERR-03glossary term Soft 404
- Repeats D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
- Extends D1-C070 Day 1: Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the…
- Repeats D1-C355 Day 1: A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found'…
- Extended by D2-C903 Day 2: A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that…
- Extended by D2-C700 Day 2: Index selection drops soft 404 pages that were not dropped earlier, for example when a document is…
Because error pages are worded in endless variations, of which 'page not found' is only the classic one, Google cannot detect soft 404s with simple error, word or keyword matching.
Speaker Gary IllyesEvidence transcript
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-ERR-03
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
- Extends D2-C213 Day 2: A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as…
Mistakes in Google's own systems are a further cause of soft 404s, and Gary Illyes asked site owners to report such mistakes in Google's forums.
“BONUS: Mistakes in Google's systems (that you should notify us about)”
Wording checked against the slide or recording
Speaker Gary IllyesEvidence slide photo, transcript
Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.
“This is basically an LLM thing, something like BERT, that is specifically trained to understand page structure”
Speaker Gary IllyesEvidence transcript
- Extends D1-C042 Day 1: BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at…
For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-ERR-03
- Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
Google's soft 404 detection tells the page chrome, such as navigation and footer, apart from the main content by analysing the page's visual hierarchy alongside its text.
Speaker Gary IllyesEvidence transcript
Gary Illyes said soft 404 detection keeps searchers from clicking into dead ends and avoids wasting site owners' resources on visitors sent to error pages.
Speaker Gary IllyesEvidence transcript
Return real error status codes for error states, 404 or 410 for missing content and 503 for outages such as a failed database connection, also in single-page apps; a 200 page whose main content is only an error message is treated as a soft 404 even when header and navigation look normal.
Author Ibrahim Anjro
Used byrequirements DEV-ERR-03, DEV-SRV-03
Hidden text is recorded at token level: Google stores spam metadata such as white-on-white text with the tokens, so leftover hidden keyword blocks are a liability, not neutral clutter, and should be removed.
Author Ibrahim Anjro
Used byrequirement DEV-HTM-06
Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
- Extends D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
Gary Illyes defined a page's main content as any part of the page that directly helps the page achieve its purpose, what it was built for.
“Main content is any part of the page that directly helps the page achieve its purpose”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
Main content is not only text: images, videos, a tool or anything else that helps a page achieve its purpose can be main content.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
Content created by other users can be main content: on a user-generated content site, the user-generated content can be the page's main content.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
A comment section below a blog post can still be part of the page's main content and can contribute to Google's understanding of the page.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-REN-02
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
- Extends D2-C202 Day 2: Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden…
A page's main content includes all of its headings and its visible title.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
Gary Illyes said the main content is what Google considers when ranking a page.
“It's the main content that we consider for ranking.”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
- Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.
“the context window is perhaps 900,000 or even closer to a million big”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-AIF-03
- Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
- Extends D2-C331 Day 2: Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
Google's Search Quality Rater Guidelines define main content as any part of the page that directly helps it achieve its purpose, including text, images, videos, page features such as calculators and content created by users, and they count the title at the top of the page as part of it.
“Main Content is any part of the page that directly helps the page achieve its purpose.”
Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
Google's Search Quality Rater Guidelines say navigation links are a common type of supplementary content, and that content behind tabs and user reviews or comments may count as main content on some pages and as supplementary content on others, depending on the page's purpose.
Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)
Used byrequirement DEV-HTM-01
- General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
The talk gave Gemini's context window both as millions of tokens and as roughly 900,000 to a million; Google's long-context docs say Gemini models have context windows of 1 million or more tokens (about eight average novels per million), so plan with about one million tokens as the documented floor rather than several million.
Author Ibrahim Anjro
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
Speaker John MuellerEvidence transcript
- Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
- Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
Speaker John MuellerEvidence slide photo, transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
- Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.
Speaker John MuellerEvidence transcript
Used byglossary term Duplicate cluster
- Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
Google's speaker said what SEOs call the canonical is, for Google, the representative of a cluster of duplicate pages: the URL Google would ideally show.
Speaker John MuellerEvidence transcript
Used byglossary term Canonical
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
Speaker John MuellerEvidence slide photo, transcript
Used byglossary term Canonical
- Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
- Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
The first reason Google deduplicates is that users do not want to see the same page repeated in the search results, even if site owners would like it to rank ten times on page one.
Speaker John MuellerEvidence transcript
Storage is a second reason for deduplication: Google's storage has many competing uses and storage prices have risen sharply, so the space for any one use is limited and Google has to draw a line somewhere.
Speaker John MuellerEvidence transcript
Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-09
- Extended by D3-C642 Day 3: Google treats a site move as a complex canonicalization process in which every signal of the old site is…
Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.
Speaker John MuellerEvidence slide photo, transcript
Used byglossary term Alternate names
Alternate names also serve localization: if Google knows that country versions such as you.de and you.at are equivalent, it can pick the right version to show using hreflang.
Speaker John MuellerEvidence slide photo
Alternate names are why a site: query for an old domain still shows the old domain's URLs after a site migration, which site owners often misread as a migration that is not working.
“FYI "alternate names" is why you see old domains in site:-queries”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-09
When someone searches for an old domain after a migration, Google shows the old URL as an alternate version because that is what was searched for, and relies on the redirect to take the user to the new site.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-09
After a rebrand that changes the domain (the slide's example was johns-bikes to slow-bikes), Google can still show the old domain to people who search for the old brand by name, in navigational and branded queries.
Speaker John MuellerEvidence slide photo, transcript
Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.
“This is normal and as users get used to the new domain name, the alternate names will fade away without you doing anything.”
Publisher Google Search Central
Used byrequirements DEV-CAN-01, DEV-CAN-09glossary term Alternate names
After a migration, judge success by the new domain's indexing and traffic in Search Console rather than by a site: query on the old domain, and keep the old domain's redirects in place long term so searches for the old brand still reach the new site.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-09
Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.
Speaker John MuellerEvidence transcript
Google builds duplicate clusters from four kinds of input: redirects, content, rel=canonical, and a 'magic bucket' of other things.
Speaker John MuellerEvidence transcript
Google trusts redirects very much for clustering, because a redirect is a clear sign that there is one version of the content; Google keeps track of both URLs but stores only one copy of the content.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-01
Whether a redirect is permanent or temporary matters only for choosing the canonical, not for clustering the URLs together.
“the permanent redirect really only matters for canonicalization, not for clustering.”
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-01
Use permanent redirects (301 or 308) for moves you want reflected in Search: a temporary redirect still groups the URLs, but it changes which URL Google is likely to pick as the canonical.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-01
Google clusters duplicate pages by content in four ways: exact matches, near matches, structurally similar content, and soft 404s.
Speaker John MuellerEvidence transcript
Exact-match duplicates, such as the www and non-www versions of the same page, are clustered and Google keeps only one of them.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-02
Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.
“When main content is the same, pages may be clustered.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-07
Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-ERR-01
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
“Do we even need to crawl /buy/seo-service ?”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo
Used byrequirements DEV-CAN-08, DEV-URL-09
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-08
To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.
“Misleading site structure (use clear signals!)”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
Make location and variant pages differ in their main content (local stock, staff, prices, addresses) and return 404 for empty or invalid combinations; otherwise every URL that fits the pattern can be folded into one canonical.
Author Ibrahim Anjro
Used byrequirements DEV-CAN-08, DEV-URL-09
Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.
Publisher Google Search Central
Used byrequirement DEV-CAN-08
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
Speaker John MuellerEvidence transcript
Used byrequirements DEV-SRV-01, DEV-SRV-02
- Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
- Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
- Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
Speaker John MuellerEvidence transcript
Used byrequirements DEV-SRV-01, DEV-SRV-02
- Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
Publisher Search Central blog (24 December 2024)
Used byrequirements DEV-ERR-03, DEV-SRV-02
- Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-AIF-04
- Extends D1-C436 Day 1: A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to…
Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.
Author Ibrahim Anjro
Used byrequirements DEV-AIF-04, DEV-SRV-02
rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-03
- Extends D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.
Speaker John MuellerEvidence transcript
Used byglossary term rel=canonical
- Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
- Repeated by D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-08
When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.
“We try to use hreflang alternates.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-08
- Extends D2-C033 Day 2: Google extracts hreflang annotations, through which site owners specify the language variants of their…
Google's speaker advised against 'clever' geo-redirecting, because it very often goes wrong.
“clever" geo-redirecting (is often not so clever)”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-02
For German-language sites serving Germany, Austria and Switzerland, add hreflang with region codes (de-DE, de-AT, de-CH) and make the country pages differ in more than boilerplate (prices, shipping, legal details), or expect Google to cluster them and show one.
Author Ibrahim Anjro
Used byrequirement DEV-INT-08
Google's canonical guide says that for canonicalization Google prefers URLs that are part of hreflang clusters: if German pages for Germany and Switzerland point to each other with hreflang but not to the Austrian page, the German and Swiss pages are preferred as canonicals.
Publisher Google Search Central
Used byrequirement DEV-INT-08
Google picks the canonical from a variety of criteria and uses some kind of machine learning to decide how much weight each criterion gets; the weighting changes from time to time.
“we use some kind of machine learning to understand how strong these criteria should be. And this changes from time to time.”
Speaker John MuellerEvidence transcript
Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-01, DEV-CAN-02, DEV-CAN-07
Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-07
Whether a page can load is really important for canonical selection: a broken certificate, failing JavaScript or a page that cannot be loaded counts against a URL, and the slide also listed meta refresh and security.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-02
Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-05, DEV-URL-05
Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.
Publisher Google Search Central
Used byrequirements DEV-CAN-05, DEV-URL-05glossary term rel=canonical
Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.
Author Ibrahim Anjro
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-URL-06
- Contradicts D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.
Author Ibrahim Anjro
Used byrequirement DEV-URL-06
Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.
Publisher Search Central blog (8 April 2013)
Used byrequirement DEV-URL-06
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-02
- Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…
Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.
“Don't block agents.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-AIF-04
- Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-MON-06
- Extended by D2-C408 Day 2: A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying…
When all canonical signals point to the same URL, Google follows what the site owner says; when they point in different directions, Google cannot tell what the owner wants.
“if there are multiple things in different directions, we don't know what to do.”
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-05
Before a migration or template change, align every canonical signal for the preferred URL (redirects, rel=canonical, internal links, sitemap entries, hreflang and working HTTPS); Google says it follows the site owner only when the signals agree.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-05
Google's pagination guide says the pages of a paginated sequence may share the same title and description, and suggests linking every page of the sequence back to the first page.
Publisher Google Search Central
Used byrequirement DEV-URL-06
Broken canonical tags can make the wrong pages of a site show up in search results.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-MON-06
- Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.
Speaker Tobias SchwarzEvidence transcript
- Repeats D2-C380 Day 2: Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google…
- Extended by D2-C873 Day 2: A community speaker said that, under the canonical link specification, an improperly declared canonical tag…
Google's guide to specifying canonical URLs says Google supports explicit rel=canonical link annotations as described in RFC 6596, the canonical link relation specification.
Publisher Google Search Central
A broken canonical setup does not fail visibly: pages still load normally while the choice of URL passes to the search engine's own heuristics, so canonical tags need a regular audit with a crawler and the URL Inspection tool.
Author Ibrahim Anjro
Used byrequirement DEV-MON-06
Checking canonicals pairwise (page A has a canonical to page B, page B a self-referencing canonical) is common but, in a community speaker's experience, misses the bigger picture, such as other links pointing to A or B.
Speaker Tobias SchwarzEvidence transcript
In a community speaker's terms, a canonical group is the set of all URLs connected through canonical links, and its canonical leader is the page that all the group's canonical links ultimately point to.
Speaker Tobias SchwarzEvidence transcript
A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-MON-06
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D2-C398 Day 2: Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure…
A community speaker said it helps to view small canonical groups as a graph that shows any redirects between members, together with the indexability of each member page.
Speaker Tobias SchwarzEvidence transcript
A crawler report shown on screen drew one canonical group of five URLs as a link graph with the leader marked by a crown, and listed each member's HTTP status, indexability, document language, leader flag, incoming and outgoing links, and hints.
Speaker Tobias SchwarzEvidence slide photo, transcript
The crawler report shown on screen ran group-level checks on each canonical group: an overall group status, whether the leader has self links, whether languages are consistent across members, and whether the group contains a chain.
Speaker Tobias SchwarzEvidence slide photo
In the real-world canonical group shown on screen (a German-language travel magazine section, site not named), a feed URL's canonical pointed to a URL on another subdomain that 301-redirected to the leader, a canonical chain through a server-side redirect.
Speaker Tobias SchwarzEvidence slide photo
The crawler report shown on screen rated a canonical group member that issues a redirect instead of displaying content as an error, and flagged canonical links whose target returns a non-200 status or is itself a redirect.
Speaker Tobias SchwarzEvidence slide photo
Google's 2009 Search Central blog post introducing rel=canonical says a canonical URL may be a URL that redirects: Google then processes the redirect as usual.
Publisher Search Central blog (12 February 2009)
The crawler report shown on screen flagged a URL discovered only through a canonical tag, with no internal a href links leading to it, as a phantom document outside the visible site structure.
“a phantom document not part of the visible site structure”
Wording checked against the slide or recording
Speaker Tobias SchwarzEvidence slide photo
To audit canonicals at scale, resolve every crawled URL to the final URL its canonical links and redirects lead to, group URLs by that leader, and flag groups with more than one hop, a loop, several canonicals on one page, or a leader that is noindexed, blocked, non-200 or redirecting.
Author Ibrahim Anjro
Used byrequirements DEV-CAN-04, DEV-MON-06
Multiple canonical declarations on one page conflict and are invalid, so the search engine applies its own heuristic and picks the canonical for the site owner; a page should declare only one canonical.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-03
Google's 2013 Search Central blog post on rel=canonical mistakes says to specify no more than one rel=canonical per page: when a page has more than one, Google will likely ignore all of them, and any benefit of a legitimate canonical is lost.
Publisher Search Central blog (8 April 2013)
Used byrequirement DEV-CAN-03
Canonical chains, where a page's canonical target leads on to yet another URL, are named as improper use in the canonical link specification, according to a community speaker.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-04glossary term Canonical chain and canonical loop
Google's 2009 Search Central blog post introducing rel=canonical says Google's algorithm is lenient and can follow canonical chains, but strongly recommends updating links to point to a single canonical page for optimal canonicalization.
Publisher Search Central blog (12 February 2009)
Used byrequirement DEV-CAN-04glossary term Canonical chain and canonical loop
Canonical chains can run through canonical links alone, through a server-side redirect between members, or through a client-side redirect between members.
Speaker Tobias SchwarzEvidence transcript
The crawler report shown on screen flagged a canonical chain as a problem: canonical links that form a chain instead of all pointing directly to the group's canonical leader.
“Canonical links form a chain rather than all pointing directly to the leader.”
Wording checked against the slide or recording
Speaker Tobias SchwarzEvidence slide photo
After a migration or a subdomain change, update rel=canonical tags to the new final URLs: canonicals still pointing at old, redirecting URLs create chains through server-side redirects like the one in the real-world example shown.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-04
Canonical loops, such as two HTML pages whose canonical tags point to each other, are a structural conflict: the group has no canonical leader, so its canonical information cannot be used.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-04glossary term Canonical chain and canonical loop
Canonical loops can also run through server-side or client-side redirects, not only through canonical tags.
Speaker Tobias SchwarzEvidence transcript
A community speaker said that at best a search engine's heuristic would treat a canonical loop through redirects as self-referencing canonicals, and doubted that this is often done when the loop runs through a client-side redirect.
“But regarding the client-side redirect, I highly doubt that this is often done.”
Speaker Tobias SchwarzEvidence transcript
When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.
“you're telling the search engine to index something that a human can't reach”
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-04
- Extends D1-C067 Day 1: A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts…
For a canonical leader that users cannot reach through links, a community speaker suggested revisiting the canonical graph and probably making the internally linked page the leader instead.
Speaker Tobias SchwarzEvidence transcript
A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.
Speaker Tobias SchwarzEvidence transcript
- Contradicts D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…
Google's guide to specifying canonical URLs says a canonical helps consolidate signals for duplicate pages: links to a duplicate URL are consolidated with links to the preferred URL once the preferred URL becomes canonical.
Publisher Google Search Central
Used byrequirement DEV-CAN-05
A leader reachable only through a canonical and a leader with weak link equity share one fix: point internal links, sitemap entries and redirects at the URL chosen as canonical, so the leader is both reachable for users and the strongest page in its group.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-05
The canonical leader of a group should be indexable: it should carry no noindex robots directive, return no error status code and not be blocked in robots.txt.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-04
Canonical links between different language versions of a page are a misuse: if a search engine accepted such a canonical group, only one language version would rank.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-06
Language versions of a page should be connected with hreflang annotations instead of canonical links.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-06
Google's guide to specifying canonical URLs says that on pages using hreflang, the canonical should be a page in the same language, or the best possible substitute language if no canonical page exists in the same language.
Publisher Google Search Central
Used byrequirement DEV-CAN-06
Check that every canonical group contains a single document language: a group spanning languages means canonicals are joining versions that hreflang should connect, and each language version should keep its own self-referencing canonical.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-06
In a community speaker's experience, real-world canonical graphs are far more complex than the isolated examples in the talk and commonly mix several canonical problems at once.
Speaker Tobias SchwarzEvidence transcript
A community speaker advised fixing canonical chains by cleaning up the canonical graph so that every canonical link points directly to the canonical leader.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-04
The crawler report shown at the talk rated a redirecting canonical target as an error, while Google's 2009 guidance accepts a canonical URL that redirects and says Google can follow canonical chains; the gap is one of severity, since both recommend pointing every canonical link straight at the final URL.
Author Ibrahim Anjro
Google says canonicalization consolidates link signals from duplicates into the chosen canonical, so a leader with few links of its own is not necessarily weaker once Google accepts it; the practical risk of a weakly linked leader is that Google, treating rel=canonical as a hint, picks a different page as canonical.
Author Ibrahim Anjro
A community speaker said that, under the canonical link specification, an improperly declared canonical tag can also be ignored completely by the application that processes it, not only replaced by that application's own heuristic.
Speaker Tobias SchwarzEvidence transcript
Used byrequirement DEV-CAN-03
- Extends D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
In a community case study, two of the first loan-comparison websites in Poland did exactly the same thing, so they competed with each other for the same users, the same keywords and the same space in search results.
Speaker Martyna AğanoğluEvidence transcript
A community speaker said running two competing sites in one market meant keeping up twice with changes in how people search and in what Google values, instead of putting all resources behind one site.
Speaker Martyna AğanoğluEvidence transcript
In a community case study, an acquisition by a group of loan-comparison brands (operating in eleven markets) prompted the merger of two competing portals into one brand, which the speaker framed as creating order rather than only a technical migration.
Speaker Martyna AğanoğluEvidence transcript
A community speaker set three goals before any technical work on a two-site consolidation: lower costs (infrastructure, maintenance and content paid for only once), one strong brand focused on one specialization, and a simple structure instead of hundreds of similar URLs competing with each other.
Speaker Martyna AğanoğluEvidence transcript
Before choosing which of two domains to keep in a consolidation, a community speaker's team checked both domains for past penalties and compared their traffic, conversion and revenue.
Speaker Martyna AğanoğluEvidence transcript
In a community case study, every URL of both domains was exported, compared on traffic, conversion, backlinks, revenue and rankings, and given one of three decisions: keep, merge or remove.
Speaker Martyna AğanoğluEvidence transcript
A community speaker's test for keeping a page during a consolidation was whether it had real potential to make money and really fitted what the business does; pages that failed were removed.
Speaker Martyna AğanoğluEvidence transcript
In a community case study, the kept pages that shared the same search intent were merged into one strong article each, taking the merged site from over 2,000 URLs to about 100, a cut of about 95% of URLs.
Speaker Martyna AğanoğluEvidence transcript
In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs with no same-intent match were removed with an error status instead of being redirected (the exact code is unclear in the recording).
Speaker Martyna AğanoğluEvidence transcript
Used byrequirement DEV-CAN-10
- Extends D2-C396 Day 2: Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't…
A community speaker's team never redirected removed URLs to the homepage, even those that still brought traffic, because the aim was a clear signal about what the site is rather than keeping every visitor.
Speaker Martyna AğanoğluEvidence transcript
Used byrequirement DEV-CAN-10
The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.
Speaker Martyna AğanoğluEvidence transcript
Used byrequirement DEV-CAN-09glossary term Change of Address tool
- Extends D1-C540 Day 1: Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is…
During a domain consolidation, a community speaker's team merged the analytics of both sites so that no historical data was lost.
Speaker Martyna AğanoğluEvidence transcript
A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.
Speaker Martyna AğanoğluEvidence transcript
- Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
As part of a consolidation, a community speaker's team introduced the real people behind the site, financial experts visibly responsible for what it publishes, which the speaker counted as important as the technical work.
Speaker Martyna AğanoğluEvidence transcript
After a domain consolidation went live, a community speaker's team checked Search Console every day, ran full site crawls and compared the numbers with the pre-migration baseline for months.
Speaker Martyna AğanoğluEvidence transcript
Used byrequirement DEV-CAN-09
A community speaker reported that after consolidating two loan-comparison sites into one domain, traffic grew from the first day, with no drop or waiting period.
Speaker Martyna AğanoğluEvidence transcript
A community speaker reported that revenue grew 100% month over month after the consolidation (the speaker's figure, period not stated), although the site had been cut from over 2,000 URLs to about 100.
Speaker Martyna AğanoğluEvidence transcript
A community speaker credited a consolidation's success to doing everything at once: one clear topic, visible expertise instead of an anonymous content factory, and crawling spent only on pages that matter, growing through quality rather than more content.
Speaker Martyna AğanoğluEvidence transcript
A community talk on a still unfinished migration, after a client bought a company whose site had to fit into one product line, presented the decisions behind the migration rather than its results.
Speaker David Carrasco PamiesEvidence transcript
A community speaker argued that migrations are hard because of people rather than technology: two companies, two teams, a dozen stakeholders and a board that set a deadline before anyone opened the CMS.
“I think there are no complex migrations. There are complex people.”
Speaker David Carrasco PamiesEvidence transcript
A community speaker named four kinds of post-acquisition site migration: absorb (one site integrates into the other), keep apart, merge (two or three sites into a new one), and plug in (a full site becomes a section of the other site).
Speaker David Carrasco PamiesEvidence transcript
A community speaker said a migration should be run around one document, the redirect map, which gives every old URL an approved destination.
Speaker David Carrasco PamiesEvidence transcript
Used byrequirement DEV-CAN-10glossary term Redirect map
A community speaker's four tips before mapping a migration: agree what to keep and retire with an owner for every decision, review what to retire with data, prioritise by business value and not only traffic, and set expectations with an owner to escalate to.
Speaker David Carrasco PamiesEvidence transcript
A community speaker advised asking support and sales which pages make buyers buy, and preparing the destination with relevant content and internal links before a migration launches.
Speaker David Carrasco PamiesEvidence transcript
A community speaker said most migration failures look technical but trace back to a decision nobody made; in the migration presented, more than 7,000 URLs still had no approved destination at the time of the talk.
Speaker David Carrasco PamiesEvidence transcript
A community speaker recorded each migration decision per URL: old URL, new URL, what must survive (for example a product answer and a demo request), who approves, and the test that proves it works.
Speaker David Carrasco PamiesEvidence transcript
Used byglossary term Redirect map
A community speaker advised asking other teams during a migration about everything a crawler cannot reveal, such as sales workflows or legal requirements.
Speaker David Carrasco PamiesEvidence transcript
For migrations with many URLs, a community speaker suggested letting a similarity score propose redirect destinations and having people review the critical pages.
Speaker David Carrasco PamiesEvidence transcript
Before a migration launches, a community speaker advised freezing the inventory (CMS export, crawl, sitemap, search data and logs), saving all content and approving the redirect map.
Speaker David Carrasco PamiesEvidence transcript
Used byrequirement DEV-CAN-10
A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.
Speaker David Carrasco PamiesEvidence transcript
Used byrequirement DEV-CAN-10
- Extends D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
A community speaker's three post-migration checks per URL: the route (the old URL reaches its destination), the answer (the approved content is still there) and the lead (a demo request still reaches the CRM); automate them, but have people check critical journeys.
Speaker David Carrasco PamiesEvidence transcript
A community speaker advised giving moved content its own baseline, because site-wide totals can hide lost pages and lost search queries.
Speaker David Carrasco PamiesEvidence transcript
A community speaker cited a recent study (not named) that reported a median recovery time of 304 days after a site migration.
Speaker David Carrasco PamiesEvidence transcript
A community speaker said a long expected recovery after a migration is no reason to wait: broken redirect routes and blocked pages need action immediately.
Speaker David Carrasco PamiesEvidence transcript
A community speaker advised telling users of a migrated brand that they are in the right place, through emails, a press release and a blog post, to keep the trust built with the old brand.
Speaker David Carrasco PamiesEvidence transcript
A community speaker said reputation travels with a migrated product: a good product takes its reviews and links along, and a product nobody likes takes its complaints along.
Speaker David Carrasco PamiesEvidence transcript
A community speaker summed up migration planning in three questions: what survives, who approves, and what test shows the migration actually works.
Speaker David Carrasco PamiesEvidence transcript
Google's site move guide says not to redirect many old URLs to one irrelevant destination such as the new site's home page, which can confuse users and might be treated as a soft 404, and to return a 404 or 410 for deleted or merged content that is not moved to the new site.
Publisher Google Search Central
Used byrequirement DEV-CAN-10
Google's site move guide says to submit a Change of Address in Search Console when moving from one domain or subdomain to another, to submit the new sitemap, and to change internal links on the new site from the old URLs to the new ones.
Publisher Google Search Central
Used byrequirement DEV-CAN-09glossary term Change of Address tool
For URLs removed in a consolidation, return 404 or 410: Google's site move guide names those two codes, and Google's crawlers treat every 4xx code except 429 the same way, as content that does not exist, so the choice between them matters less than not redirecting to an unrelated page.
Author Ibrahim Anjro
Google's crawl budget guide is written for sites with over a million pages changing weekly or over 10,000 pages changing daily, so for a consolidation of about 2,000 URLs the crawl-budget gain is likely minor; the benefit of pruning such a site more plausibly comes from one strong URL per intent and consolidated signals.
Author Ibrahim Anjro
The 304-day median recovery a community speaker cited from an unnamed study and Google's site move guide (a few weeks or more for a medium-sized site until Google shows the new URLs, longer for larger sites) measure different things: indexing can switch within weeks while traffic recovery can take far longer, so plan and report on both.
Author Ibrahim Anjro
Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
Speaker Ryan LeveringEvidence transcript
- Extends D1-C011 Day 1: Ecosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such…
- Extends D1-C158 Day 1: Google said that even the classic ten-blue-links layout was settled only after millions of experiments, as…
AI Overviews and AI Mode launched as fairly text-heavy answers with little image content and few tables, and they have become more structured over time because that is what users want.
Speaker Ryan LeveringEvidence transcript
- Extends D1-C011 Day 1: Ecosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such…
Structured data turns the loosely structured web into structured information that powers visual search features such as review stars and recipe filters (for example by preparation time).
Speaker Ryan LeveringEvidence transcript
Used byglossary term Rich results
- Extended by D3-C337 Day 3: Review structured data lets a site specify how its users rated something; the review snippet shows an average…
Describing a page's content in a structured way lets Google understand and interpret that content more accurately.
Speaker Ryan LeveringEvidence transcript
Used byglossary term Structured data
Structured data can bring more qualified traffic, because the site can be shown in new, more interesting ways to more people, who are then more likely to click.
Speaker Ryan LeveringEvidence transcript
Inside Google, views on structured data split into two camps: one says it is useless because models can generate it or read the page directly, the other says it is the future of machines talking to each other through MCP servers and new standards; the speaker said the truth is in the middle.
Speaker Ryan LeveringEvidence transcript
Google gave four reasons why structured data is still valuable even though models can extract information from pages: precision, extra content, efficiency and focus.
Speaker Ryan LeveringEvidence slide photo, transcript
Structured data gives the high precision that complex schemas such as sale pricing need, with higher accuracy than large-scale extraction by large language models (LLMs).
“Structured data provides the high precision needed for complex schema (sale pricing), achieving higher accuracy than large-scale LLM extraction.”
Wording checked against the slide or recording
Speaker Ryan LeveringEvidence slide photo, transcript
In the speaker's own tests, even the latest LLMs asked to generate schema.org markup for a page often invent properties that do not exist, get deeply nested schemas such as complex pricing models wrong, and duplicate content across several fields.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-11
Automatic LLM extraction of structured information makes a great demo but is not yet good enough for extraction that aims at something like 99.9% accuracy, the speaker said.
Speaker Ryan LeveringEvidence transcript
Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
Publisher Google Search Central
Used byrequirement DEV-SDA-11
- Extends D1-C250 Day 1: In a community demo, a fix loop gave a large LLM (Claude Opus) the current markup, the errors Google's test…
- Extended by D2-C947 Day 2: In a community speaker's case, an LLM asked to write alt text for a product image of a veterinary anxiety…
- Extended by D2-C952 Day 2: A community speaker warned that articles still recommend rewriting alt text with AI, called AI a tool, and…
- Extended by D3-C686 Day 3: Hallucinations can happen with any AI model, and with current training methods there is no way to get rid of…
Structured data often carries non-visible metadata that the page text lacks, such as full ISO dates or stable identifiers for user-generated content.
“It often contains non-visible metadata, such as full ISO dates or stable identifiers for UGC, that is not present in the page text.”
Wording checked against the slide or recording
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirements DEV-SDA-04, DEV-SDA-05
Google's own extraction systems focus heavily on what is visible on a page, which improves precision but means they can miss content or interpret it incorrectly.
Speaker Ryan LeveringEvidence transcript
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-HTM-04
- Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
Markup can state a date as a full ISO 8601 value, which disambiguates a visible date whose local time zone Google might not detect correctly; this matters for event extraction and wherever the date must be exactly right.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-05
Google's structured data documentation says not to mark up content that is not visible to readers of the page, and not to add structured data about information that users cannot see even if it is accurate.
“don't add structured data about information that is not visible to the user, even if the information is accurate”
Publisher Google Search Central
Used byrequirement DEV-SDA-02
The 'non-visible metadata' argument is not a licence to mark up hidden content, since Google's structured data guidelines still require markup to describe what users can see; use markup for the machine-precise form of facts the page shows, such as a full ISO date with time zone for a visible event date or a homepage url for a named organisation.
Author Ibrahim Anjro
Used byrequirements DEV-SDA-02, DEV-SDA-04, DEV-SDA-05
Parsing structured data is significantly cheaper and more efficient than relying on LLMs for every complex extraction task.
Speaker Ryan LeveringEvidence slide photo, transcript
- Repeated by D3-C405 Day 3: Web markup is an efficient and unambiguous way for sites to share product data with Google, Google Shopping…
Even Google cannot afford to run complex AI models on every page in its index, and distilling them into cheaper models makes them less precise.
Speaker Ryan LeveringEvidence transcript
Rule-based parsing of markup is nearly free by comparison with AI models, so Google will always prefer extracting information from structured data over model-based extraction.
“So we're always going to prefer that particular approach.”
Speaker Ryan LeveringEvidence transcript
Do not drop markup on the assumption that AI reads the page anyway: by Google's own account LLM extraction is not precise enough for prices and nested offers and too costly to run on every page, so explicit markup remains the dependable route for those facts.
Author Ibrahim Anjro
Structured data explicitly points to the pertinent data on a page, which reduces noise and stops Google's systems from pulling in extraneous information such as prices of related products.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-07
Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-07
- Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
On product pages with related-product or recently-viewed carousels, make sure the Product markup describes only the main item and its price; Google's own example of what markup prevents was pulling a price from related products.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-07
The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-AIF-02
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
As far as the speaker knows, in most of Google's main AI uses a page's schema.org markup is not turned into text and put directly into the model's context; the data is first sorted out, checked for quality and indexed before it is passed on as grounding context.
Speaker Ryan LeveringEvidence transcript
- Extends D1-C061 Day 1: Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in…
The speaker said there are concerns that feeding raw schema.org markup straight into an AI model's context would be a big abuse vector, one reason Google does not do so.
Speaker Ryan LeveringEvidence transcript
There is no separate markup for AI features: by Google's account the same processed markup feeds classic results and AI answers, and raw schema.org is generally not passed into model context, so invest in the types Google documents for its features rather than in extra markup written for AI.
Author Ibrahim Anjro
Used byrequirement DEV-AIF-02
Google's guide to optimizing for generative AI features lists 'overfocusing on structured data' among the things site owners don't need to do: structured data is not required for generative AI search and no special schema.org markup is needed, though it remains worth using because it helps pages become eligible for rich results.
Publisher Google Search Central
Used byrequirement DEV-AIF-02
- Extends D1-C061 Day 1: Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in…
Google's developer site has a number of case studies that show traffic and session-duration gains from structured data.
Speaker Ryan LeveringEvidence transcript
A Google case study shown on stage (published May 2018) reported that Rakuten Recipes got 2.7 times more traffic from search engines and a 1.5 times increase in session duration after implementing recipe structured data.
Speaker Ryan LeveringEvidence slide photo, transcript
The speaker called the Rakuten study old and said the world has changed a lot, but expects the link between structured data and more interactive, visually appealing results to hold for the foreseeable future.
Speaker Ryan LeveringEvidence transcript
Schema.org is a common vocabulary that several major search engines started together so that site owners can mark up pages and every consumer interprets the markup the same way; the speaker put its start 15 to 20 years ago (Google, Bing and Yahoo! announced it in June 2011).
Speaker Ryan LeveringEvidence transcript
Used byglossary term Schema.org
Schema.org is an open public collaboration, and the speaker invited people who enjoy ontologies and data to join it.
Speaker Ryan LeveringEvidence transcript
Marking something up with only a very generic schema.org type says little about it and is very hard for Google to use.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Google uses schema.org as its markup vocabulary but mostly consumes only the subsets that its developer documentation declares, for the features it builds.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03glossary term Schema.org
Google will never penalise a site just for having more structured data on its pages than Google uses; extra markup does not hurt.
“we will never penalize you for having more structured data on your pages”
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Describing every semantic detail of a page in markup is probably not worth the effort; focus on the structured data that Google or other consumers actually use.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
- Extended by D3-C378 Day 3: Google Shopping called rich data with light structure the AI sweet spot: a little structure that helps the AI…
Google recommends using the Search gallery in its developer documentation to find the structured data features that suit a site; the gallery shows each feature and how Google uses the markup.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-03
- Repeated by D3-C329 Day 3: Google's structured data feature guide lists the kinds of structured data Google supports with a search…
Google's structured data types fall roughly into two groups: broad page-level types such as breadcrumbs and article, and vertical-specific types such as recipes, events and products.
Speaker Ryan LeveringEvidence transcript
Pick only the structured data types that are relevant to a page instead of adding every type Google is interested in.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Most sites now run on content management systems whose plug-ins generate the markup, so check that the plug-ins in your stack produce good markup, for example a calendar plug-in that outputs event markup.
Speaker Ryan LeveringEvidence transcript
Google accepts three structured data syntaxes, JSON-LD, microdata and RDFa, which are all valid and are extracted at the very start into the same pipelines, so they are interpreted identically downstream.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-01glossary term JSON-LD
Google recommends JSON-LD because it is one contiguous block that is easier to author, so people make fewer mistakes with it.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-01glossary term JSON-LD
Microdata has an advantage when payload size matters: embedded in the existing HTML, it avoids duplicating the page content in a separate JSON-LD block.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-01
Default to JSON-LD and switch to microdata only where page weight is critical, since Google interprets all three syntaxes identically and JSON-LD is the one people get wrong least often.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-01
Google recommends testing and previewing structured data in the Rich Results Test, which shows the rich result features it detected and whether the markup is valid.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-11glossary term Rich Results Test
Test structured data manually first and only then put it into the site's templates.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-11
Structured data that is not relevant to the page's content can be treated as abusive: Google's filters make it ineffective, and egregious cases can lead to a manual action.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-02
- Extended by D3-C647 Day 3: Google may never use structured data from a site it does not trust: once it sees markup it does not trust, it…
Google's general structured data guidelines say a structured data manual action removes a page's eligibility for rich results but does not affect how the page ranks in web search.
Publisher Google Search Central
Used byrequirement DEV-SDA-02
Use unique identifiers in structured data, for example the homepage URL in an organisation's url property, so Google can tell which specific entity is meant rather than reading just a name string.
Speaker Ryan LeveringEvidence transcript
Used byrequirements DEV-SDA-04, DEV-SDA-09
Several plug-ins emitting the same markup type is one of the most common structured data problems: the duplicates can make an event details page look like a list of events and change how Google interprets the page.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-06
Audit CMS templates for duplicate markup by running one page per template through the Rich Results Test and checking whether an SEO plug-in and a theme or events plug-in emit the same type twice; 'more markup never hurts' covers relevant, non-duplicated markup only.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-06
Google's shopping structured data launches of the previous year (2025) added support for merchant loyalty programs and shipping policies, letting merchants define a policy at organisation level and specify details at product level.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-09
- Extended by D3-C383 Day 3: In product markup, an offer's shipping and return information can point through a JSON-LD identifier (@id) to…
Google's structured data speaker said that in the months before the event Google added support for validity dates on sale prices in product structured data, so merchants no longer need to rush to remove a sale price when the sale ends for fear it shows wrongly in snippets.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-08
- Extended by D3-C396 Day 3: Google works to make sure that what merchants can express in its shopping feeds can also be expressed in…
Google's documentation updates log calls the July 2026 sale price change a clarification: a new Sale duration section of the merchant listing guide explains validFrom with validThrough or priceValidUntil, aligned with Merchant Center's sale_price_effective_date attribute.
Publisher Google Search Central
Used byrequirement DEV-SDA-08
For time-limited sales, give the sale price its validity dates in the product markup (Google's merchant listing documentation describes validFrom, validThrough and priceValidUntil) instead of editing markup by hand when the sale ends, and keep the dates aligned with the Merchant Center feed.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-08
More shopping structured data news was left for a Day 3 talk by a Google colleague, Alex.
Speaker Ryan LeveringEvidence transcript
- Extended by D3-C365 Day 3: Google added six Merchant Center feed attributes for AI shopping experiences: question and answer, documents…
Schema.org, in which Google is a major participant, began publishing usage statistics for all its types and properties in 2026, showing in buckets how many domains use each one; the data is also in schema.org's GitHub repository.
Speaker Ryan LeveringEvidence transcript
Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a chicken-and-egg problem that the schema.org usage statistics are meant to help break.
Speaker Ryan LeveringEvidence transcript
- Extends D1-C120 Day 1: To get Google to build something, such as an addition to an API, report and request it publicly and in volume.
Schema.org added support for RDF lists and sets, which give a way to express ordered values, because RDF triples are not ordered by nature.
Speaker Ryan LeveringEvidence transcript
Markup for schema.org types that Google does not use yet is a bet on future features; check a type's usage bucket on schema.org before investing, since adoption is what Google says it watches before building a feature.
Author Ibrahim Anjro
Google announced server-side structured data validation as coming soon: it will publish downloadable validation rules in SHACL on each structured data feature guide, with other kinds of checks to follow.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-12glossary term SHACL
Google's planned validation workflow has five steps: download the rules from the feature guide, generate the JSON or embedded microdata or RDFa, run the rules against the generated server-side markup as a first check, deploy and test in the Rich Results Test, and monitor ongoing performance in Search Console.
Speaker Ryan LeveringEvidence slide photo
Used byrequirement DEV-SDA-12
The SHACL rules are meant to run inside a site's content generation, so markup is sanity-checked before it is published and does not silently regress later, a breakage site owners might otherwise discover only through a Search Console report.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-12glossary term SHACL
The SHACL rules will not replace Search Console as the canonical place for structured data reports, because some checks use Google's internal libraries and cannot be expressed in SHACL, but they will catch problems such as a missing required field.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-12
Google's example SHACL shape for Event requires a name, treats a missing description only as a warning, and accepts the image either as an ImageObject or as a URL.
Speaker Ryan LeveringEvidence slide photo, transcript
Google plans a SHACL rule set for each of its structured data feature types and will release the rule sets gradually once they have been checked.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-12
When the SHACL rules appear, wire them into the build or CMS publishing step as an automated test, so a template change that drops a required property fails before deployment instead of surfacing weeks later in Search Console.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-12
Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.
Speaker Ryan LeveringEvidence transcript
Used byglossary terms Rich results, Structured data
- Extends D1-C122 Day 1: Check and focus on rich results for Google.
- Extended by D3-C324 Day 3: Rich results differ from other search features because Google builds them from extra data that site owners…
Google keeps adding structured data features and recommendations but also removes them: in the previous year (2025) it removed several features that brought little benefit and were little used.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Google's merchant listing guide says Product rich results only support pages that focus on a single product, or on several variants of the same product.
Publisher Google Search Central
Used byrequirement DEV-SDA-07
Google's merchant listing guide warns that a listing may not display if its priceValidUntil property indicates a past date.
Publisher Google Search Central
Used byrequirement DEV-SDA-08
Google's Event structured data guide says to give a date without a time, such as 2019-08-15, when the start hour is not known, and to include the UTC or GMT offset whenever a time is given.
Publisher Google Search Central
Used byrequirement DEV-SDA-05
Gary Illyes said images and videos drive a large amount of traffic to publishers.
Speaker Gary IllyesEvidence transcript
- Extends D1-C116 Day 1: Images were stressed as important.
Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-01
- Repeats D2-C447 Day 2: For images, Google's feature extraction takes the img element with its src and other attributes, including…
An image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-03
- Extends D1-C116 Day 1: Images were stressed as important.
- Extends D1-C117 Day 1: With one in six AI Mode searches being multimodal, original images with descriptive file names, alt text and…
- Extended by D3-C319 Day 3: Image results shown among web results come from Google's image index and are roughly the same images that…
Google supports the picture element only because a picture element must contain an img element, and that img element is what Google extracts and passes to its media indexer.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-01
One of the most reliable ways to make sure Google finds an image is to include it in an img element in the HTML; Gary Illyes named image sitemaps as the other method.
Speaker Gary IllyesEvidence transcript
Used byrequirements DEV-IMG-01, DEV-IMG-04
Google does not extract CSS background images.
“we don't support CSS background extraction”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-01
Gary Illyes suggested a div with a CSS background image as a way to keep an image from being picked up by Google.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-06
Google's documented ways to keep a site's images out of search results are a robots.txt disallow rule (for example for Googlebot-Image) or a noindex X-Robots-Tag HTTP header, with the Removals tool for emergencies.
Publisher Google Search Central
Used byrequirements DEV-IDX-03, DEV-IMG-06glossary term X-Robots-Tag
Hiding an image as a CSS background is a fragile way to keep it out of Google: it only stops extraction from that page, so the same image URL used in an img element elsewhere or listed in a sitemap can still be indexed; the documented robots.txt or noindex X-Robots-Tag methods are the reliable route.
Author Ibrahim Anjro
Used byrequirement DEV-IMG-06
Of the attributes the HTML standard defines for the img element, Google uses three and ignores the rest; Gary Illyes named src and alt but not the third.
Speaker Gary IllyesEvidence transcript
The src attribute is the most important img attribute, because without it Google does not know where the image bytes are and cannot index the image.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-01
Gary Illyes ranked the alt attribute below src in importance, while saying alt attributes are still important.
Speaker Gary IllyesEvidence transcript
Google's image SEO guide calls alt text the most important attribute for providing more metadata about an image, and says Google uses it together with computer vision and the page content to understand the image.
“The most important attribute when it comes to providing more metadata for an image is the alt text”
Publisher Google Search Central
Used byrequirement DEV-IMG-02
Alt text should describe the image in words for someone who cannot see it.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-02
Google uses the words in an alt attribute to understand the image, may attach them to the image at serving time, and the image can rank for concepts the alt text describes.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-02
The text around an image is critical: Google uses it as context to understand the image and to rank it, so an alt attribute alone is not enough.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-03
- Repeats D1-C045 Day 1: Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour…
- Extended by D2-C950 Day 2: Because not every generated alt text can be checked by hand, a community speaker's safer script adds the page…
Images that should rank need an img element with a real src; hero or product images set as CSS backgrounds are invisible to Google Images and image features, so keep CSS backgrounds for decorative images you do not need indexed.
Author Ibrahim Anjro
Used byrequirement DEV-IMG-01
Give important images a caption or explanatory sentence next to them rather than relying on alt text alone, since Gary Illyes ranked src above alt and called the surrounding text critical for ranking the image.
Author Ibrahim Anjro
Used byrequirement DEV-IMG-03
Besides img elements, image sitemaps tell Google about images: they are XML sitemaps that list, under a page's loc entry, the locations of the images on that page.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-04
Google supports pretty much all of the most popular image formats on the web.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-05
Gary Illyes said the AVIF image format currently has hiccups and Google may have problems ingesting it, although it should technically be supported.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-05
Google's image SEO guide lists AVIF among the supported image formats (with BMP, GIF, JPEG, PNG, WebP and SVG), and Google announced in August 2024 that AVIF files need nothing special to be indexed.
Publisher Google Search Central, Search Central blog (30 August 2024)
Used byrequirement DEV-IMG-05
Use high-quality modern image formats such as WebP for a good balance of quality and compression, and keep image files small so people can enjoy them.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-05
While the AVIF ingestion hiccup Gary Illyes mentioned lasts, serve AVIF only through source elements inside picture and keep a WebP or JPEG file in the img src, which is the URL Google extracts, even though Google's documentation lists AVIF as supported; then check in Google Images that key images appear.
Author Ibrahim Anjro
Used byrequirement DEV-IMG-05
Gary Illyes cited a figure that more than 40% of Southeast Asian shoppers rely on videos to make purchase decisions.
Speaker Gary IllyesEvidence transcript
Gary Illyes said over 219 million people in Southeast Asia consume content on or through YouTube daily (the daily scope was heard in two independent recordings but is not verified).
Speaker Gary IllyesEvidence transcript
Gary Illyes said there are over 150 streaming apps in Southeast Asia.
Speaker Gary IllyesEvidence transcript
Videos can appear in the main search results, in video-specific result tabs and in Discover (the tab names are unclear in the recording).
Speaker Gary IllyesEvidence transcript
Google has video-specific features, such as key moments and previews, that help users interact with videos more easily.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-03
Google strongly suggests describing a video's metadata with JSON-LD structured data, on top of placing the video in an HTML element on the page.
Speaker Gary IllyesEvidence transcript
Used byrequirements DEV-VID-01, DEV-VID-03
Google's slide listed seven key factors for video SEO success: high-quality video content, a dedicated watch page, compelling titles and descriptions, relevant thumbnails, video markup, fast-loading pages and sitemap inclusion.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-VID-02
Google's video SEO slide said each video must be on its own dedicated HTML watch page.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-VID-02
Google's video SEO guide says a video must be embedded on an indexed watch page to be eligible for video features and recommends a dedicated watch page per video where it makes sense for the business; a non-watch page with the video can still appear as a text result or a Google Images result with a video badge.
Publisher Google Search Central
Used byrequirement DEV-VID-02
Google's video SEO slide warned that spending money on a generative AI subscription does not mean AI-generated videos will do well, because engaging, informative, well-produced content is foundational.
“just because you spent money on a genAI subscription, that does not mean you are going to do well with AI generated videos”
Wording checked against the slide or recording
Speaker Gary IllyesEvidence slide photo
Including videos in a sitemap helps Google discover all of a site's video content.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-VID-04glossary term Video sitemap
Give each important video its own watch page where the video is the main content, add VideoObject markup and list the video in a video sitemap; a video buried in a long article lacks the dedicated watch page that Google listed as a key factor.
Author Ibrahim Anjro
Used byrequirement DEV-VID-02
Gary Illyes said a video without a dedicated watch page is not going to be indexed.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-02glossary term Video watch page
Google knows from experiments that when a video's thumbnail is wrong, the share of viewers who drop out at the start of the video is extremely high.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-03
Gary Illyes called a video container in the page's HTML critical: without one, Google treats the page as having no video.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-01
Gary Illyes called the 'fast-loading pages' factor for video SEO a misnomer: what matters is that the video itself loads fast, because people no longer have the patience to wait for videos to load.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-06
Serving videos from a content delivery network (CDN) that loads them faster than the site's own server is a win for video SEO, Gary Illyes said.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-06
Google supports the most popular video file formats on the web.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-06
Google recommends the MP4 container for videos because of its browser and device compatibility.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-06
Encode videos with standard codecs, because some people will not be able to play videos that use unusual ones.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-06
For a video to be discovered it must be embedded prominently, above the fold; Gary Illyes said a video placed below the fold is not going to be indexed.
“If it's not above the fold, then you basically lost the game.”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-01
Video sitemaps are not critical but good to have, because Google ingests video sitemaps much more often than it can process HTML pages.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-04glossary term Video sitemap
A video sitemap or video feed tells Google which URLs carry videos so it can visit those URLs to double-check; without one, Google has to check every page individually.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-04glossary term Video sitemap
Descriptive text around a video helps Google rank and retrieve the video.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-02
Gary Illyes suggested considering video hosting platforms such as YouTube or Vimeo, because they solved video search years ago.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-06
The more people talk about a site's videos, the more likely Google is to surface them in search results, so videos should be promoted.
Speaker Gary IllyesEvidence transcript
Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.
Speaker Gary IllyesEvidence transcript
- Extends D2-C448 Day 2: For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense…
Googlebot-Image and Googlebot-Video are the Google crawlers that fetch images and videos, and robots.txt rules addressed to their user agent tokens control how Google indexes a site's images and videos.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-06glossary term Googlebot-Image and Googlebot-Video
If robots.txt disallows the location of an image or video file, Google does not index that image or video.
Speaker Gary IllyesEvidence transcript
Used byrequirements DEV-IMG-06, DEV-VID-06glossary term Googlebot-Image and Googlebot-Video
Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-VID-06
- Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
The noimageindex robots meta tag tells Google not to index the images on the page.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IDX-09glossary term noimageindex
- Repeats D2-C099 Day 2: The noimageindex rule tells Google not to index any of the images on the page, and John Mueller said he could…
Gary Illyes said noimageindex also affects videos on the page, because Google has to index a video's thumbnail, which is an image.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IDX-09glossary term noimageindex
Setting the max-image-preview robots meta tag to large can make content perform surprisingly well in Discover, Gary Illyes said.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IDX-04
- Repeats D2-C089 Day 2: max-image-preview:large matters mainly in Discover, where it allows a large image that draws people's…
- Repeated by D3-C224 Day 3: Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels…
To the frequent question about AI-generated images on a site, Google's answer is that it is up to the site owner.
Speaker Gary IllyesEvidence transcript
Used byrequirements DEV-IMG-08, DEV-SPM-05
Google serves AI-generated images and videos in search results when users are specifically looking for them.
“We will serve users AI-generated images and videos in search results if they are looking for them specifically.”
Speaker Gary IllyesEvidence transcript
The diffusion models that generate images were built to generate images, not text, so they are typically poor at rendering text inside an image.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-IMG-08
- Extended by D3-C688 Day 3: Generative models, including image diffusion models, make things up when they lack information or because of…
Gary Illyes showed a timeline image that Google made with an image generation model in 2025 and deliberately left unfixed as a counterexample: its text was garbled and its years jumped from 1994 to 1995 and back to 1994.
Speaker Gary IllyesEvidence transcript
Sites that use AI-generated images or videos should make sure they work for users, check them for hallucinations and regenerate them where needed.
Speaker Gary IllyesEvidence transcript
Used byrequirements DEV-IMG-08, DEV-SPM-05
- Extended by D3-C700 Day 3: Google's closing slide said to use AI responsibly because AI hallucinates, and, especially when creating…
Google's Video indexing report help says only videos on a watch page are eligible for indexing, and flags 'Cannot determine video position and size' when the player is not on the page at load, for example behind a click-to-play image, asking for the player to load at its real size and position without user interaction.
Publisher Google Search Console Help
Used byrequirement DEV-VID-01glossary term Video watch page
Google's image metadata guide says Google Images supports the IPTC Digital Source Type values for algorithmically created images, such as trainedAlgorithmicMedia, and can show C2PA details in 'About this image', such as whether an image was created or edited with AI tools.
Publisher Google Search Central
Used byrequirement DEV-IMG-08
Gary Illyes said a video below the fold is not indexed, but Google's video documentation only requires the player to be present at load, at a position and size Google can determine and not hidden behind other elements; to be safe, put the main video of a watch page in the first viewport and load the player without a click-to-play placeholder.
Author Ibrahim Anjro
Do not set noimageindex on video watch pages: Gary Illyes said it also stops the video, through its thumbnail, and although the robots meta tag specification mentions only images, the Video indexing report treats a missing or blocked thumbnail as a reason a video is not indexed.
Author Ibrahim Anjro
Google may find a site's different language versions on its own, without the site owner doing anything, but owners can take steps to make this easier for Google.
Speaker GoogleEvidence transcript
Each language version should have its own URL; switching languages with cookies or by updating the content in place makes the versions very difficult for Google to handle.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-01
Google's documentation says Googlebot mostly crawls from US IP addresses and sends no Accept-Language header, so pages that change content or redirect by the visitor's perceived country or language may not have every version crawled, indexed or ranked; it recommends separate URLs annotated with hreflang.
Publisher Google Search Central
Used byrequirements DEV-INT-01, DEV-INT-02
Language versions can also be language-plus-region variants: Google's example site had a generic English page (en), a UK English page (en-gb) and a Spanish page (es), each on its own URL.
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-01
- Extended by D2-C627 Day 2: A community speaker said language versions need adjusting for regional variants of one language: 90 is…
Google's hreflang guide says language-specific subdomains such as en, en-gb or de in a URL are not used to determine a page's target audience; site owners must map the audience explicitly, for example with hreflang.
Publisher Google Search Central
Google's hreflang example lists every language version in the page's HTML head as a link element with rel=alternate, the version's full URL and its hreflang code (en, en-gb, es), including the page's own URL.
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-03glossary term hreflang
Missing return links are a common hreflang mistake: if page X names page Y as a language version, page Y must link back to page X.
“If page X links to page Y, page Y must link back to page X.”
Wording checked against the slide or recording
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-03glossary term hreflang return link
Each page in an hreflang set must also list its own URL, and a missing self-reference is a common mistake Google sees.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-03
When the HTML head is not suitable for a page, hreflang can be given by other methods instead, such as listing the language versions in a sitemap.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-05
Any one hreflang method is fine, but a site should use only one: Google often sees sites using two or three methods at the same time, with annotations that conflict.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-05
Google's hreflang guide says its three methods (link elements in the HTML head, an HTTP Link header, which suits non-HTML files such as PDFs, and an XML sitemap) are equivalent; using several at once is allowed but brings no benefit in Search and is harder to manage.
Publisher Google Search Central
Used byrequirement DEV-INT-05
Google ignores hreflang annotations that point only one way, where page A names page B as a language version but page B does not name page A back.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-03glossary term hreflang return link
Google's hreflang guide says a site that finds it hard to keep every language pair bidirectional may omit some languages on some pages, because Google still processes the pairs that point to each other; new language versions should at least link both ways with the original or dominant language.
Publisher Google Search Central
Used byrequirement DEV-INT-03
Google requires hreflang return links because an alternate URL is a full URL that can be on any domain, so without them any site could declare itself a language version of someone else's website.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-03
Google's hreflang guide says alternate URLs must be fully qualified, including the protocol (https://example.com/foo, not //example.com/foo or /foo), and need not be on the same domain.
Publisher Google Search Central
Used byrequirement DEV-INT-03
Audit hreflang per cluster rather than per page: check that every page lists itself and all its alternates, that every alternate links back, and that only one method (HTML head, HTTP header or sitemap) supplies the annotations.
Author Ibrahim Anjro
Used byrequirement DEV-INT-03
hreflang language codes must be ISO 639-1 codes: Google's list of common mistakes marked se, dk and cz (the country codes of Sweden, Denmark and Czechia) as wrong and sv, da and cs (Swedish, Danish, Czech) as right.
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-04
- Extended by D2-C962 Day 2: Google said incorrect hreflang language or region codes are a mistake it finds very often, with the same…
Swedish content for Sweden is marked sv-SE, not se-SE: the language part must be the Swedish language code sv, even though Sweden's country code is SE.
Speaker GoogleEvidence transcript, slide photo
Used byrequirement DEV-INT-04
Site owners do not need to memorise hreflang language and region codes, but should double-check every code when implementing hreflang.
Speaker GoogleEvidence transcript
Czech is a frequent hreflang error: CZ is the country code and ccTLD of Czechia, but the language code for Czech is cs.
Speaker GoogleEvidence transcript, slide photo
Used byrequirement DEV-INT-04
hreflang region codes must be ISO 3166-1 Alpha-2 country codes; frequent mistakes are UK instead of GB for the United Kingdom, SW instead of CH for Switzerland and GE instead of DE for Germany.
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-04
EU is not a valid hreflang region code: sites use it hoping to target the European Union, but hreflang targeting works only by country.
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-04
Google also flagged LA used as a region code for Los Angeles (a city, not a country) and SA used for South Africa as wrong hreflang region codes.
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-04
A region code cannot be used on its own in hreflang: en and en-GB are valid values, but GB alone is not, because the value must name a language.
Speaker GoogleEvidence slide photo, transcript
Used byrequirement DEV-INT-04
Validate hreflang values against the ISO 639-1 and ISO 3166-1 Alpha-2 lists in a crawler rule or build check rather than by eye: reserved codes such as UK and EU are simply ignored, but some plausible mistakes are valid codes for something else (se is Northern Sami, GE Georgia, LA Laos, SA Saudi Arabia) and silently point the annotation at the wrong audience.
Author Ibrahim Anjro
Used byrequirement DEV-INT-04
Google does not use hreflang to determine a page's language when indexing the page.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-07glossary term hreflang
Google determines a page's language for indexing from the page content, not from a language code in the URL.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-07
- Extended by D2-C651 Day 2: Google detects the language of each document and weights language in index selection so that the index is not…
Google annotates each page with only one language in the index.
Speaker GoogleEvidence transcript
Pages that mix several languages make it very hard for Google to decide which language a page is in, so each page should make its target language obvious and avoid mixing languages.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-07
Scattered boilerplate in a second language is sometimes fine but still not recommended, partly because mixed languages on one page are odd for users.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-07
Google's hreflang guide says pages that translate only the template (navigation, footer) around main content in one language still count as duplicates, because localized versions are duplicates only if the main content stays untranslated; it recommends hreflang for such pages.
Publisher Google Search Central
Used byrequirement DEV-INT-07
Because Google reads a page's language from its content and stores one language per page, translate templates (navigation, footer, buttons, legal notices) together with the main text; large untranslated blocks risk the page being annotated with the wrong language.
Author Ibrahim Anjro
Used byrequirement DEV-INT-07
The ccTLD (country-code top-level domain) is one of the most important and strongest signals Google uses to decide which country a site targets; Google also considers other signals, which matter less.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-09glossary term ccTLD
- Extended by D2-C982 Day 2: At the point where one recording names the ccTLD as the most important country-targeting signal, a second…
Server location is not really a reliable country-targeting signal nowadays, so Google does not use it much; the recording is unclear on the word 'server'.
Speaker GoogleEvidence transcript
At the point where one recording names the ccTLD as the most important country-targeting signal, a second attendee recording heard 'IP address' instead; Google's multi-regional sites documentation names the ccTLD as a strong signal of a site's target country and server location only as a possible, not definitive one, so the ccTLD is the better-supported reading.
Author Ibrahim Anjro
- Extends D2-C591 Day 2: The ccTLD (country-code top-level domain) is one of the most important and strongest signals Google uses to…
Google's documentation calls a ccTLD a strong signal that a site is meant for a certain country and still lists server location as a possible signal of a site's audience, though not a definitive one because sites use CDNs or are hosted abroad; it does not rank the signals.
Publisher Google Search Central
Used byglossary term ccTLD
For country versions of a site, a ccTLD, subdomains or subdirectories are all acceptable choices; the right one depends on the site's needs, goals and resources.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-09
- Extended by D2-C822 Day 2: Of the URL structures for country versions in a table on Google's slide (not photographed), the speaker…
Google's table of URL structures for country targeting gives pros and cons for a country-specific domain (example.de), a subdomain (de.example.com) and a subdirectory (example.com/de/) on a generic domain, and marks URL parameters (site.com?loc=de) as not recommended.
Publisher Google Search Central
Used byrequirements DEV-INT-01, DEV-INT-09
Of the URL structures for country versions in a table on Google's slide (not photographed), the speaker called only one a wrong choice: the last option, which the speaker did not name but said they really do not recommend.
Speaker GoogleEvidence transcript
- Extends D2-C594 Day 2: For country versions of a site, a ccTLD, subdomains or subdirectories are all acceptable choices; the right…
A separate ccTLD for each country version is not necessarily better, even though the ccTLD is one of the strongest country signals, and sites should not move their domain for that reason.
“Don't try to move your domain now.”
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-09
Whether to use ccTLDs depends on whether the domain can be obtained, what it costs, local laws or regulations, and how committed the business is to expanding in that country.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-09
For a business testing new countries, subdirectories on the existing domain are usually the cheapest start; a ccTLD's stronger country signal pays off only where domain availability, cost, local rules and long-term commitment justify it, and an established domain should not be moved just for the signal.
Author Ibrahim Anjro
Used byrequirement DEV-INT-09
Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.
Speaker GoogleEvidence transcript
- Extended by D2-C965 Day 2: Users of non-Latin-script languages do not always search in their own script: the same Persian query may be…
- Extended by D3-C051 Day 3: Users expect content written the way they search: in some languages they search in Latin characters, in…
Google usually understands a query word whether it is written with or without diacritics (accents).
Speaker GoogleEvidence transcript
- Extended by D3-C048 Day 3: Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a…
Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results, so what works for web results also works for the AI features.
Speaker GoogleEvidence transcript
Used byrequirement DEV-AIF-01story angle A-001
- Repeats D1-C050 Day 1: Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
- Repeats D1-C051 Day 1: Three reasons were given: generative AI features are built directly on the core ranking systems, query…
- Extends D2-C116 Day 2: AI Overviews and AI Mode are built on top of Search results: they are a different experience of the same…
In classic results users can usually tell when content in their language is missing, poorly translated or from another country, but AI answers synthesize many sources and hide this; the speaker called it an invisible gap.
“an invisible gap”
Speaker GoogleEvidence transcript
Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.
“AI is still language-dependent”
Speaker GoogleEvidence transcript
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
- Extended by D2-C972 Day 2: In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English…
Google's Search Help says the language a query is typed in is an important factor in choosing the language of results, alongside interface and device language, location and site owners' annotations, and that Google may also show results in other languages when they are helpful or when there is not enough information in the query's language.
Publisher Google Search Help
In the speaker's own experience, the differences between languages are larger in AI answers than in classic search results.
Speaker GoogleEvidence transcript
Site owners may want to research whether they show up in AI answers in each language and market they serve, to find blind spots.
Speaker GoogleEvidence transcript
Check AI Overviews and AI Mode with native-language queries in each target market, not with translated English keywords, and compare with the country breakdown of Search Console's generative AI performance report; topics where competitors are cited and you are not point to missing or weak local content.
Author Ibrahim Anjro
Used byrequirement DEV-MON-07
- Extends D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
Whether machine-translated content is acceptable depends on the case and is the site owner's decision, after weighing three things machine translation can miss: translation quality, local conventions and cultural adaptation.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-10
To show uneven machine-translation quality, the speaker showed Google Translate output of attendee questions from a visit to China with a Google colleague, Gary, some clear and some unintelligible, and a restaurant menu whose English listed an item called 'aggressive element'.
Speaker GoogleEvidence transcript
Localisation should account for local conventions such as date formats, which differ between Europe, the UK, the US and other countries, and calendars (in Thailand the current year is 2569), or a date can point users to the wrong day.
Speaker GoogleEvidence transcript
Used byrequirements DEV-INT-10, DEV-SDA-05
Google's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.
Publisher Google Search Central
Used byrequirement DEV-INT-10glossary term Scaled content abuse
- Extended by D3-C695 Day 3: The cheaper tokens become, the more AI slop is created, and Google counts AI slop as scaled content abuse.
- Extended by D3-C697 Day 3: Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem…
Treat machine translation as a first draft: have a native speaker review it and adapt dates, calendars and units before publishing, because unreviewed bulk translation that adds little value can also fall under Google's scaled content abuse policy.
Author Ibrahim Anjro
Used byrequirement DEV-INT-10
- Extends D1-C048 Day 1: The answer is not a claim that Google detects AI text. It says ranking favours text that reads as natural to…
According to the speaker, shoppers in Europe and the US pay a lot of attention to promotions and discounts when deciding to buy (no source for this was captured), one of the cultural factors localisation should take into account.
Speaker GoogleEvidence transcript
According to consumer data shown on a slide in the talk (source not captured), US consumers judge product quality more by user feedback and reviews, while European shoppers seem to look more at brand reputation.
Speaker GoogleEvidence transcript
- Extended by D3-C341 Day 3: Some countries rely a lot on social proof such as reviews, Google's speaker said, recalling a point from…
According to consumer data shown on a slide in the talk (source not captured), German shoppers also rely heavily on expert recommendations and certifications when judging product quality.
Speaker GoogleEvidence transcript
According to consumer data shown on a slide in the talk (source not captured), French shoppers particularly want to know where a product comes from.
Speaker GoogleEvidence transcript
Localise the trust signals on product pages, not only the words: if the findings shown hold for your category, lead with reviews in the US, brand reputation elsewhere in Europe, expert certifications in Germany and origin information in France, and test the order per market.
Author Ibrahim Anjro
Google's hreflang guide says the ISO 639-1 language code can be followed by a script in ISO 15924 format, such as zh-Hant or zh-Hans, and by an optional ISO 3166-1 Alpha-2 region code, as in zh-Hans-US.
Publisher Google Search Central
Used byrequirement DEV-INT-04
Google said that working out a site's localized pages and which countries it targets is one of the difficult, complex tasks for it in indexing.
Speaker GoogleEvidence transcript
Google said incorrect hreflang language or region codes are a mistake it finds very often, with the same wrong codes coming up again and again, especially in Europe, because site owners assume the codes are easy.
Speaker GoogleEvidence transcript
Used byrequirement DEV-INT-04
- Extends D2-C575 Day 2: hreflang language codes must be ISO 639-1 codes: Google's list of common mistakes marked se, dk and cz (the…
In a community speaker's audit case, a premium multifunctional-furniture maker about to be acquired had multilingual sites with clean translations by native speakers and hreflang tags that were in fairly good shape.
Speaker Alizée BaudezEvidence transcript
In the same furniture-maker audit, the Dutch site had only a fraction of the conversions of the brand's other language sites, although it showed the same catalogue, the same hero products and the same photography.
Speaker Alizée BaudezEvidence transcript
Checking a competitor during the furniture-maker audit, a community speaker found that a large furniture retailer's French and Dutch sites differ a little in the products they show and especially in their photography, with items apparently shown in much smaller rooms on the Dutch site.
Speaker Alizée BaudezEvidence transcript
A community speaker said her first guess, that Dutch homes are smaller than French ones, was wrong: Dutch and French homes have about the same floor area per person, and what differs is the type of house.
Speaker Alizée BaudezEvidence transcript
A community speaker said about 19% of homes in Europe are terraced houses on several floors, against 58% in the Netherlands; the matching Eurostat figures are 2019 shares of the EU and Dutch population living in semi-detached or terraced houses, not shares of homes.
Speaker Alizée BaudezEvidence transcript
In a community speaker's comparison, a typical Dutch home is narrow and spread over several levels, often with very steep, narrow staircases that make furniture harder to install upstairs, while a typical French home has a wide ground floor and usually wide stairs.
Speaker Alizée BaudezEvidence transcript
The lesson of the furniture-maker case was that the translation was correct but the target region had never been checked.
“the translation was right, but the region was never checked”
Speaker Alizée BaudezEvidence transcript
When translations and hreflang are clean but one country version converts far worse, compare that market's product range, photography and delivery or installation constraints with local competitors before changing anything technical.
Author Ibrahim Anjro
A community speaker framed international SEO around three signals to settle before working on hreflang: language (what people can read), region (who the content is for) and authority (why local people should trust the site).
Speaker Alizée BaudezEvidence transcript
A community speaker said language versions need adjusting for regional variants of one language: 90 is 'quatre-vingt-dix' in France but 'nonante' in Belgium, and Swiss German differs from German, for example in whether the letter ß (Eszett) is used.
Speaker Alizée BaudezEvidence transcript
- Extends D2-C562 Day 2: Language versions can also be language-plus-region variants: Google's example site had a generic English page…
Google's documentation recommends hreflang annotations even when a site's versions differ only by small regional variations within one language, for example English-language content targeted to the US, GB and Ireland.
Publisher Google Search Central
Used byrequirement DEV-INT-08
Settling the region means deciding who the content is actually for, which shapes how the catalogue is organised, how prices look and which payment options each market gets.
Speaker Alizée BaudezEvidence transcript
Of language, region and authority, a community speaker called language easy to settle and region solvable with a little research, while authority (why local people should trust a brand) has to be earned and built over time.
Speaker Alizée BaudezEvidence transcript
A community speaker stressed that an hreflang annotation for a country (her example was German for Austria, de-AT) does not by itself make a site established or trusted in that country; local authority has to be worked for.
Speaker Alizée BaudezEvidence transcript
Google's documentation lists the signals it uses to decide which locale a page targets: ccTLDs, hreflang statements, server location, and other signals such as local addresses and phone numbers, local language and currency, links from other local sites and Business Profile signals.
Publisher Google Search Central
Used byrequirement DEV-INT-08
A community speaker cited a Semrush study that found at least one hreflang error on 75% of multilingual websites; the study, reported in February 2017, audited 20,000 sites with several language versions (the speaker recalled it as about five years old).
Speaker Alizée BaudezEvidence transcript
A community speaker said the SEO industry has treated hreflang as the hard part of international SEO; hreflang is complicated, but the real hard part is the homework of understanding consumers in each market.
“the hard part is actually doing the homework to understand your consumers”
Speaker Alizée BaudezEvidence transcript
A community speaker said the answers about what consumers in each market want expire, because regional differences change over time.
Speaker Alizée BaudezEvidence transcript
A community speaker showed a chart of search interest over time for ceiling fans in northern versus southern Europe, saying southern Europeans commonly have them at home while northern Europeans paid them no attention until the previous summer's heat waves (the talk was on 1 October 2026).
Speaker Alizée BaudezEvidence transcript
Re-check regional demand for core product categories at least once a year, for example by comparing regions and seasons in Google Trends, because climate, housing or habits can shift what a market wants.
Author Ibrahim Anjro
The first of three questions to ask before the technical work of international SEO is which languages, and which language variants, the site actually serves, keeping languages and variants apart.
Speaker Alizée BaudezEvidence transcript
The second question before international SEO work is what genuinely differs between the regions of each language, such as photography, content, vocabulary, the way products are presented, pricing, catalogues and payment options.
Speaker Alizée BaudezEvidence transcript
The third question is whether you can name three local sources that already treat the brand as present in a country; if you cannot, a community speaker advised not to expand there but to fix the markets you already serve.
“don't bother expanding to another country; just fix the ones you already have”
Speaker Alizée BaudezEvidence transcript
Treat hreflang as a way to show the right language version, not as a market-entry signal: before launching a country version, make sure local publications, partners or directories already mention the brand, and localise pricing and payment options.
Author Ibrahim Anjro
List language variants such as fr-FR and fr-BE or de-DE and de-CH separately in a site's language inventory, and give each its own vocabulary and spelling review instead of reusing one translation for every country.
Author Ibrahim Anjro
A community talk, presented on its author's behalf by a colleague, was about multilingual SEO for languages written in non-Latin scripts, such as Persian and Arabic.
Speaker a second community speakerEvidence transcript
Arabic and Persian contain letters that look identical to users but are different characters to software, with different Unicode code points; the example given was the letter ye, which has an Arabic and a Persian form.
Speaker a second community speakerEvidence transcript
Because Arabic and Persian lookalike letters have different code points, a searcher may type one variant of a word while a website's text uses the other.
Speaker a second community speakerEvidence transcript
Lookalike Arabic and Persian character variants cause problems for data analysis and keyword research: data for one term can be split across the variants, which makes keyword research less accurate.
Speaker a second community speakerEvidence transcript
For Persian and Arabic keyword research, look up each lookalike-character spelling of a term separately and add up the volumes, and check which variant the site's own content uses.
Author Ibrahim Anjro
Persian and Arabic are written right to left and English left to right, so a title that mixes the two scripts can display in a confusing, unpredictable order even when its content is correct.
Speaker a second community speakerEvidence transcript
Used byrequirement DEV-INT-12glossary term Bidirectional text
The display problem of mixing right-to-left and left-to-right text also affects product titles and URLs, the presenter of the non-Latin-script talk said, showing an example from a large Iranian e-commerce site.
Speaker a second community speakerEvidence transcript
Used byrequirement DEV-INT-12glossary term Bidirectional text
Users of non-Latin-script languages do not always search in their own script: the same Persian query may be typed in Persian script or in Latin letters, with the same intent and the same expected results.
Speaker a second community speakerEvidence transcript
Used byrequirement DEV-INT-11glossary term Transliterated queries
- Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…
Keyword research for non-Latin-script languages should ask how people actually type and spell in local search, not only what the right keyword is, because switching keyboards is inconvenient, especially when typing on mobile.
Speaker a second community speakerEvidence transcript
The presenter of the non-Latin-script talk said that in competitive Persian, Turkish and Arabic searches, bought backlinks and paid editorial content still visibly influence rankings and are widespread (an observation; no data was shown); the presenter stressed this described the situation and was not a recommendation.
Speaker a second community speakerEvidence transcript
The presenter of the non-Latin-script talk said Google announced in October 2023 that it had improved its spam-detection coverage for languages: the spam policy was global, but the coverage improvement was language-specific.
Speaker a second community speakerEvidence transcript
The presenter of the non-Latin-script talk said Persian offers far fewer natural link opportunities than English, with fewer publications, niche blogs and websites.
Speaker a second community speakerEvidence transcript
For SEO in languages such as Persian, the presenter of the non-Latin-script talk advised accepting different competitive dynamics and basing strategy on the actual maturity of that language's market, not only on the global policy.
Speaker a second community speakerEvidence transcript
Google search features often launch in some languages or countries first and expand later (the presenter's example was site names, launched in several languages and then extended to all languages in 2023), so comparisons of performance across languages and countries should not assume a feature is live everywhere at once.
Speaker a second community speakerEvidence transcript
In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English sources; the presenter explained that Persian content on a topic is often thinner, of lower quality or less relevant, so the systems retrieve from languages with better content, which the presenter called cross-lingual retrieval.
Speaker a second community speakerEvidence transcript
- Extends D2-C603 Day 2: Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is…
The non-Latin-script talk concluded that multilingual SEO is not just translation: script, language behaviour and market maturity also matter, and basic SEO tactics, though the same everywhere, must be fitted to each market.
Speaker a second community speakerEvidence transcript
The presenter of the non-Latin-script talk recalled that Google officially treats buying links for ranking as spam (the year of the announcement was not clear in the recording).
Speaker a second community speakerEvidence transcript
Google's October 2023 spam update post (4 October 2023) says the update improved coverage in many languages and spam types, cleaning up spam reported in Turkish, Vietnamese, Indonesian, Hindi, Chinese and other languages, particularly cloaking, hacked, auto-generated and scraped spam.
Publisher Search Central blog (4 October 2023)
The October 2023 spam update cited in the non-Latin-script talk names neither Persian nor Arabic (only 'other languages') and lists cloaking, hacked, auto-generated and scraped spam, not link spam, so it does not show that the bought links the presenter saw working in Persian or Arabic search were addressed.
Author Ibrahim Anjro
Google's site names were introduced on mobile results for select languages in October 2022, expanded to desktop in March 2023, and on 7 September 2023 became available in all languages where Google Search is available, on mobile and desktop.
Publisher Search Central blog (7 September 2023)
Google's post on multilingual searches (8 September 2023) says that, because of typing difficulty on some keyboards, a person in India might search in Hindi using Latin rather than Devanagari characters and want and receive Hindi results written either way.
Publisher Search Central blog (8 September 2023)
Used byrequirement DEV-INT-11glossary term Transliterated queries
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
“our index is immense. Like, immense. But it is a finite resource.”
Speaker GoogleEvidence transcript
- Repeats D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
- Repeats D2-C344 Day 2: Google deduplicates pages because many sites have very many pages and Google's index does not have room for…
- Repeated by D3-C256 Day 3: Google does not index every URL on the web; because it cannot index everything, it has to rank results better…
Google gives two reasons for not indexing every URL it knows: most of them would not be useful to users, and including URLs in the index that users would never see would be an immense investment.
Speaker GoogleEvidence transcript
Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.
Speaker GoogleEvidence transcript
Used byglossary term Index selection
- Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
Index selection aims to index only documents that are useful to users now or potentially in the future.
Speaker GoogleEvidence transcript
Index selection is a predictive AI system that relies heavily on machine learning.
Speaker GoogleEvidence transcript
- Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
“the index selection system is going to be more forgiving when it sees a new URL from your site”
Speaker GoogleEvidence transcript
Used byrequirement DEV-URL-10
- Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
Author Ibrahim Anjro
Used byrequirement DEV-URL-10
- Extends D1-C095 Day 1: New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under…
Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
Speaker GoogleEvidence transcript
Used byglossary term Index selection
- Extends D1-C205 Day 1: Signals calculated for a page during indexing are stored in the index and used both to decide whether the…
Index selection is the last step before documents enter Google's index.
Speaker GoogleEvidence transcript
Used byglossary term Index selection
- Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
- Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
When Google's coverage of a country is limited, index selection becomes more likely to select documents relevant to that country, even if other signals would suggest otherwise.
Speaker GoogleEvidence transcript
When Google lacks content in a language, such as Basque, index selection becomes more likely to select lower-quality documents in that language.
Speaker GoogleEvidence transcript
Under-served languages were presented as an opportunity: where Google's index holds a lot of spam in a language such as Basque, a site that starts publishing in that language can very likely replace that spam with its own content and rank for those keywords (one clause of the reasoning was inaudible).
Speaker GoogleEvidence transcript
The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.
Author Ibrahim Anjro
Used byrequirement DEV-INT-10
The overall importance of a document or site plays a significant role in index selection.
Speaker GoogleEvidence transcript
News sites were given as the example of importance at work in index selection: they are generally very important on the web and their pages usually get indexed very fast ('indexed' is a likely but not certain reading of the recording).
Speaker GoogleEvidence transcript
Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.
“focusing on the quality is the most reliable way to get stuff in the index”
Speaker GoogleEvidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.
Speaker GoogleEvidence transcript
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
Speaker GoogleEvidence transcript
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.
Speaker GoogleEvidence transcript
Used byrequirement DEV-IDX-08glossary term unavailable_after
- Extends D2-C100 Day 2: The unavailable_after rule lets a page drop out of search results after a set date and time, which suits…
Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-08
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
Speaker GoogleEvidence transcript
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
- Extends D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Speaker GoogleEvidence transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.
Publisher Google Search Central
'Only canonicals end up in search results' as said on stage is a simplification: non-canonical duplicates are dropped from the index, but Google's documentation says an alternate from the same cluster can still be shown in some contexts, such as a mobile version to a mobile user.
Author Ibrahim Anjro
Index selection loads most of Google's spam signals and acts as a gatekeeper so that no spam enters search results.
“index selection acts as a gatekeeper, ensuring that no spam enters our search results”
Speaker GoogleEvidence transcript
Index selection also applies other policies, covering egregious violations and, in a less certain reading of the transcript, content Google is legally not allowed to index.
Speaker GoogleEvidence transcript
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
“The first one is kind of nastier.”
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed
- Extends D1-C097 Day 1: Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over…
- Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
Google's Page indexing report help says a 'Discovered – currently not indexed' page was found but not crawled yet, typically because Google wanted to crawl it but expected the crawl to overload the site, so it rescheduled the crawl.
Publisher Google Search Console Help
Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed
The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.
Author Ibrahim Anjro
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.
Speaker GoogleEvidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
Repeatedly submitting 'Discovered – currently not indexed' URLs does not change why they wait, because the status reflects a crawl-scheduling decision; raise the site's demonstrated quality instead, for example by improving or removing weak pages that are already indexed.
Author Ibrahim Anjro
'Crawled – currently not indexed' in Search Console is an index selection decision: Google crawled and processed the page but decided not to keep it in the index.
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03glossary term Crawled – currently not indexed
Google's Page indexing report help says a 'Crawled – currently not indexed' page was crawled but not indexed, may or may not be indexed in the future, and does not need to be resubmitted for crawling.
Publisher Google Search Console Help
Used byrequirement DEV-MON-03glossary term Crawled – currently not indexed
If content quality is even across a site, a page reported as 'Crawled – currently not indexed' may be using a different template that keeps Google from understanding where its content is.
Speaker GoogleEvidence transcript
Used byrequirement DEV-HTM-01
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
“most of the time it is actually a quality issue”
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
The first thing to check for pages reported as 'Crawled – currently not indexed' is whether their quality is on par with the parts of the site that Google does index.
Speaker GoogleEvidence transcript
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
Author Ibrahim Anjro
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03
- Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…
Not-indexed reasons in the Page indexing report include pages excluded by a noindex rule and 'Alternate page with proper canonical tag'.
Speaker GoogleEvidence transcript
Used byrequirement DEV-MON-03
Google Trends is a free Google tool that gives access to a sample of aggregated, anonymized and categorized searches.
Speaker Omri WeismanEvidence transcript
Google Trends data covers searches on Google Search, including web, image, shopping and news search, plus searches on YouTube.
Speaker Omri WeismanEvidence transcript
Google Discover is not included in Google Trends, because Discover is a content feed with no search box, while Trends counts only the queries people type into Google Search or YouTube.
Speaker Omri WeismanEvidence transcript
Google's Trends FAQ says Google Trends filters out searches made by Google products and services, including the internal searches made by AI Mode and AI Overviews.
“This includes internal searches made by AI Mode and AI Overviews.”
Publisher Google Trends Help
Google Trends cannot show demand that reaches people through the Discover feed, so Discover visibility has to be judged from Search Console's Discover performance report rather than from Trends.
Author Ibrahim Anjro
Google handles more than five trillion searches a year, a figure Google announced publicly at the beginning of 2025.
Speaker Omri WeismanEvidence transcript
- Repeats D1-C006 Day 1: Google reports more than 5 trillion searches per year.
By the author's arithmetic, more than five trillion Google searches a year means on average more than 13.7 billion searches a day, or roughly 160,000 a second.
Author Ibrahim Anjro
Google Trends data can be viewed for the whole world, for a single country or for a region within a country, for example Spain or Catalonia.
Speaker Omri WeismanEvidence transcript
Google said Google Trends data is almost real time, reaching up to a few minutes before the moment of viewing.
“we have data that goes all the way back from 2004 and up until three minutes ago”
Speaker Omri WeismanEvidence transcript
Google Trends offers time ranges from the past hour or the past day back to 2004, so search patterns can be compared across two decades.
Speaker Omri WeismanEvidence transcript
Google argued that people search authentically, for health conditions, career troubles or home repairs, so aggregated searches reflect what a country or the world cares about better than the curated picture on social media.
Speaker Omri WeismanEvidence transcript
Google Trends has two flagship experiences: the Explore page and Trending now.
Speaker Omri WeismanEvidence transcript
The Google Trends Explore page, which Google called the heart of Trends, shows search interest in a query or topic and how it changes over time.
Speaker Omri WeismanEvidence transcript
- Extended by D3-C015 Day 3: Google said the difference between words and entities can be seen in Google Trends, where a term can be…
Google launched a brand-new Google Trends Explore page at the beginning of 2026.
Speaker Omri WeismanEvidence transcript
The legacy Google Trends Explore page is still available but will be retired: Google is adding its features to the new Explore page until everyone can move to the new one (no date was given).
“the legacy Explore page, which is still available, but not for a long time”
Speaker Omri WeismanEvidence transcript
Google's Trends Help says Google Trends compares up to 8 groups of terms at once, with up to 50 terms in each group, while Classic explore (the legacy page) supports 5 groups of up to 25 terms.
Publisher Google Trends Help
Trending now, a tab at the top of Google Trends, shows the topics people are searching for on Google right now.
Speaker Omri WeismanEvidence transcript
Used byglossary term Trending now
Google said Trends' Trending now covered 125 countries at the time of the talk (October 2026), plus regions within many of those countries.
Speaker Omri WeismanEvidence transcript
Google's August 2024 blog post said Trending now was available in 125 countries, with regional trends in 40 of them; the current Trends Help page says 100+ countries and regions.
Publisher Google blog (14 August 2024), Google Trends Help
Google's documentation says the Trending now page shows an approximate, bucketed search volume for each trend, with its growth against a predicted baseline.
Publisher Google Trends Help, Google Search Central
Used byglossary term Trending now
A Google Trends Explore chart presented on stage put worldwide search interest in Gemini above ChatGPT for the first time ever in September 2025, during a sharp surge for Gemini.
Speaker Omri WeismanEvidence transcript
Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of Nano Banana (the image editing model in the Gemini app), as people searched for both Gemini and Nano Banana.
Speaker Omri WeismanEvidence transcript
- Extends D1-C162 Day 1: Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the…
Google announced Nano Banana, a new image editing model from Google DeepMind in the Gemini app, on 26 August 2025.
Publisher Google blog (26 August 2025)
On stage the Nano Banana launch was placed in September 2025, but Google announced it on 26 August 2025, so the September 2025 surge in Gemini search interest described in the talk came in the weeks after the launch.
Author Ibrahim Anjro
After the Nano Banana surge faded, worldwide baseline search interest in Gemini stayed much higher than before, which Google read as a sustained gain in brand recognition.
Speaker Omri WeismanEvidence transcript
Real-world events such as product launches show up directly in Google Trends search interest graphs.
Speaker Omri WeismanEvidence transcript
To judge a launch or campaign in Google Trends, compare the brand's baseline search interest before and after the spike, not only the height of the spike, because a lasting baseline lift is the durable gain.
Author Ibrahim Anjro
A Google Trends chart of worldwide search interest in AI from 2004 to mid-2025 showed a huge rise in the two to three years before mid-2025.
Speaker Omri WeismanEvidence transcript
Google said worldwide search interest in AI has climbed far higher since mid-2025, beyond the range of the 2004 to mid-2025 chart shown in the talk.
Speaker Omri WeismanEvidence transcript
An early-2012 worldwide spike in Google Trends search interest for 'ai' was caused by the global hit song 'Ai Se Eu Te Pego', whose title contains the Portuguese word 'ai', not by interest in artificial intelligence.
Speaker Omri WeismanEvidence transcript
A Google Trends engineer found the cause of the 2012 'ai' spike by zooming in on that time range on the Explore page and reading the related queries.
Speaker Omri WeismanEvidence transcript
Google advised being curious and digging deeper when explaining a pattern in Google Trends, because context matters and the obvious explanation is sometimes wrong.
Speaker Omri WeismanEvidence transcript
Google's Trends FAQ says Google Trends adds statistical noise to protect privacy, so one-off spikes on queries with low or no search interest may be noise rather than real search activity.
Publisher Google Trends Help
Before treating a spike in an ambiguous keyword as demand for your topic, zoom in on the period in Google Trends and check the commonly searched queries, since another meaning of the word (as with 'ai' in 2012) may be driving it.
Author Ibrahim Anjro
Worldwide Google Trends search interest for ski since 2004 shows strong seasonality, peaking every winter.
Speaker Omri WeismanEvidence transcript
Worldwide Google Trends search interest for ski also shows a mostly steady decline since 2004, which can be misread as falling search volume.
Speaker Omri WeismanEvidence transcript
Google Trends shows search interest, not search volume: each point on its 0 to 100 scale represents the share of searches for the term out of all searches.
“Search interest is not search volume.”
Speaker Omri WeismanEvidence transcript
Used byglossary term Search interest
Because total Google searches were far fewer in 2004 than today, a falling search interest share in Google Trends can hide a rising number of searches.
Speaker Omri WeismanEvidence transcript
Google said its internal numbers show the actual volume of ski-related searches is going up, even though ski's share of all searches in Google Trends has declined since 2004.
“we can tell you from our own internal numbers that we're seeing on Google Trends: the actual volume is, in fact, going up.”
Speaker Omri WeismanEvidence transcript
Google stressed that the difference between search interest and search volume matters most when reading very long time ranges in Google Trends.
Speaker Omri WeismanEvidence transcript
A Google Trends line that falls over many years is not proof that a market is shrinking; compare terms against each other in the same chart, and take absolute numbers from keyword tools or Search Console impressions.
Author Ibrahim Anjro
Google presented three uses of Google Trends for content creation, marketing and SEO: keyword selection, scheduling and ideation.
Speaker Omri WeismanEvidence transcript
- Extended by D3-C104 Day 3: To choose what to cover in depth, a community speaker's team used Google Trends to find what people in their…
For keyword selection, Google recommended entering candidate keywords on the Google Trends Explore page, setting the relevant country, region and time frame, and seeing which terms have the higher search interest.
Speaker Omri WeismanEvidence transcript
Google recommended using the keywords with higher search interest in Google Trends in page titles and marketing campaigns.
Speaker Omri WeismanEvidence transcript
Google's Search Central guide to Google Trends says to use Trends to inform a content strategy only where a topic fits the business and its users, and to focus on terms where the site has expertise and experience.
“you shouldn't write about something just because it's trending”
Publisher Google Search Central
Compare candidate keywords in Google Trends for the exact country or region where the page should rank, because the higher-interest wording can differ between markets that share a language.
Author Ibrahim Anjro
The 'Commonly searched queries' section of the new Google Trends Explore page, called related queries in the legacy Explore page, lists queries that people typed in the same search sessions as the entered term.
Speaker Omri WeismanEvidence transcript
Google Trends' commonly searched queries come as top queries, with the highest search interest or volume, and rising queries, with the largest increase from the previous time frame to the current one.
Speaker Omri WeismanEvidence transcript
Google recommended checking both the top and the rising commonly searched queries in Google Trends when selecting keywords.
Speaker Omri WeismanEvidence transcript
In Google Trends' commonly searched queries, rising queries point to emerging sub-topics worth a new page or section, while top queries show the established wording to reuse in titles and headings.
Author Ibrahim Anjro
For scheduling, Google recommended checking on the Google Trends Explore page whether the topics relevant to your domain show seasonality; many topics do, though not all.
Speaker Omri WeismanEvidence transcript
Where a topic is seasonal in Google Trends, Google said the pattern predicts the next wave of search interest, which is the time to run a marketing campaign or have content ready.
Speaker Omri WeismanEvidence transcript
A Google Trends chart of US search interest over five years showed umbrella and sunscreen correlated, both peaking in June, against the common assumption that umbrella searches peak in winter.
Speaker Omri WeismanEvidence transcript
Google warned that assumptions about when search interest peaks are sometimes wrong, and advised checking the pattern in Google Trends for your own target market.
Speaker Omri WeismanEvidence transcript
Publish or refresh seasonal content a few weeks before the seasonal rise that Google Trends shows for each target country, so the pages are crawled and indexed before demand peaks.
Author Ibrahim Anjro
For ideation, Google recommended Google Trends' Trending now, where a real-world event shows up within about ten minutes as people hear the news and search for it.
Speaker Omri WeismanEvidence transcript
Google said journalists use Google Trends' Trending now every morning to find content gaps: topics people are searching for that their publication may need to cover.
Speaker Omri WeismanEvidence transcript
The new Google Trends Explore page has a 'Suggest search terms' button that opens an optional Gemini-based side panel.
Speaker Omri WeismanEvidence transcript
Google Trends' Suggest search terms picks the search terms or topics to compare from a plain-language request, which helps in a field you know little about, for example the most popular jazz singers in Japan.
Speaker Omri WeismanEvidence transcript
Google Trends' Suggest search terms combines Gemini's knowledge of the world with Google Trends data.
Speaker Omri WeismanEvidence transcript
Google Trends' Suggest search terms compares up to eight search terms at once.
Speaker Omri WeismanEvidence transcript
Google said the search terms returned by Google Trends' Suggest search terms are ranked by search interest.
Speaker Omri WeismanEvidence transcript
Quick comparisons, a recent Google Trends Explore page feature, show year-over-year, month-over-month and week-over-week comparisons as chips when a single search term or topic is entered.
Speaker Omri WeismanEvidence transcript
Clicking a Google Trends quick comparison chip overlays the search interest of the two time periods being compared.
Speaker Omri WeismanEvidence transcript
Google said colour maps of the geographic breakdown of search interest (regions within a country, or countries worldwide) had launched in the new Google Trends Explore page about a month before the October 2026 talk.
Speaker Omri WeismanEvidence transcript
Search Engine Roundtable reported on 19 August 2026 that Google Trends had added explore maps and a regional breakdown that morning, as announced by Google on LinkedIn.
Reported by Search Engine Roundtable (19 August 2026)
The category filter of the legacy Google Trends Explore page has been reintroduced in the new Explore page.
Speaker Omri WeismanEvidence transcript
Search Engine Roundtable reported on 2 September 2026 that Google had announced category filters for the new Google Trends Explore page, usable with or without a query.
Reported by Search Engine Roundtable (2 September 2026)
The regional search interest maps in Google Trends give a quick first check of demand by country or region before localising content for a new market.
Author Ibrahim Anjro
Trends TV, at trends.google.com/tv, is a little-known Google Trends dashboard of what is trending on Google right now, meant for an office screen or as a screensaver.
Speaker Omri WeismanEvidence transcript
Google has published a YouTube series of about eight episodes on using Google Trends for research, journalism, marketing and SEO.
Speaker Omri WeismanEvidence transcript