An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
From the audienceEvidence slide photo
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Day 2 · Thursday 1 October 2026 · 10:15
Presented by Gary Illyes and Cherry Prommawin (named in the host's hand-over; turns not attributable). Highly rated Slido questions skipped on Day 1 for lack of time: four on slides with written answers, one more (sloppy migrations) only spoken, then a recap of Day 1. Transcript from a second attendee recording.
QQuestion and answers
An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
From the audienceEvidence slide photo
Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
“It's hard to make good HTML sitemaps for large sites.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
“Better rely on hubs like category pages that link out to your important pages.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-04
Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have the extra benefit that they may also be useful to users.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo
Used byrequirement DEV-URL-04
Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.
“if you want to make one, knock yourself out, but I will focus on some better things like hub pages, category pages”
Speaker not identifiedEvidence transcript
Used byrequirement DEV-URL-04glossary term Hub pages
QQuestion and answers
An audience member asked whether a product detail page that is out of stock for two to three months should keep returning 200 with links to similar products, or be 302-redirected to a similar product or to its parent product listing page.
From the audienceEvidence slide photo, transcript
Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it depends on how important that product page is to users.
“It depends on the importance of the PDPs to the users.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-ERR-04
Google's Q&A slide on out-of-stock product pages said users might wait months for some products, or even pre-order them if the site offers pre-ordering.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-ERR-04
Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to them, with easily replaced products, such as a particular cheese, for which the buyer simply picks a similar product instead of waiting.
Speaker not identifiedEvidence transcript
QQuestion and answers
An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.
From the audienceEvidence slide photo, transcript
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
“If it's included in the robots.txt file, any crawler can pick your sitemaps up”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-05
Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in Search Console.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-URL-05
Google said listing the sitemap in robots.txt is fine, as many websites do.
“you can include it in robots.txt. No problem whatsoever.”
Speaker not identifiedEvidence transcript
Used byrequirement DEV-URL-05
Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.
Speaker not identifiedEvidence transcript
Used byrequirement DEV-URL-05
QQuestion and answers
An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
From the audienceEvidence slide photo, transcript
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
“Simply put, it's because of sites that are extremely important and like to disallow their most important pages, either accidentally or out of ignorance.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-IDX-01
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript
Used byrequirement DEV-IDX-01
Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.
Speaker not identifiedEvidence transcript
Used byrequirement DEV-IDX-01
Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no figure given).
Speaker not identifiedEvidence transcript
Used byrequirement DEV-IDX-01
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”
Speaker not identifiedEvidence transcript
Used byrequirement DEV-IDX-01
QQuestion and answer
An audience member asked how several sloppy migrations on the same domain can affect Googlebot's crawling.
From the audienceEvidence transcript
Google answered that several sloppy migrations on one domain can cause many effects in the short term, and noted that migrations concern indexing as well as crawling, a subject a later Day 2 talk would cover.
Speaker not identifiedEvidence transcript
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
Speaker not identifiedEvidence transcript
Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.
Speaker not identifiedEvidence transcript
Google's ecommerce documentation recommends linking from menus to category pages, from category pages to sub-category pages and from sub-category pages to all product pages; where not every product can be linked, it recommends a sitemap or a Merchant Center feed.
Publisher Google Search Central
Used byrequirement DEV-URL-04
Google's ecommerce documentation says Google can infer a page's relative importance within a site from its internal links, such as how many links point to the page and how many links Google must follow to reach it.
Publisher Google Search Central
Used byrequirement DEV-URL-04
A 2005 Search Central blog post, now marked as possibly outdated, said Google encouraged HTML sitemaps because they help users navigate a site and a clear hierarchy of text links helps Google index it.
Publisher Search Central blog (26 September 2005)
Google's guide to temporarily pausing an online business recommends that a shop expecting to sell again within weeks or months stays online with limited functionality, such as a disabled cart, and updates its Product structured data to show current availability.
Publisher Google Search Central
Used byrequirements DEV-ERR-04, DEV-SDA-08, DEV-SRV-03
Google's redirect documentation says that with a temporary redirect, such as a 302, Google Search shows the source page in search results and does not use the redirect as a signal that the target should be canonical.
Publisher Google Search Central
Used byrequirements DEV-CAN-01, DEV-ERR-04
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
Publisher Google Search Central
Used byrequirement DEV-IDX-01glossary term noindex
Google's sitemap guide says Google uses lastmod only when it is consistently and verifiably accurate, and counts a change to the main content, the structured data or the links of a page as significant, but not a changed copyright date.
Publisher Google Search Central
Used byrequirement DEV-URL-05
Google's answer did not say whether HTML sitemaps pass link equity. On a large site, check that every important page is linked from a category or hub page in the normal navigation instead of relying on an HTML sitemap page to reach it.
Author Ibrahim Anjro
Google's view of HTML sitemaps has moved: a 2005 blog post encouraged them, the current sitemap and ecommerce documentation does not mention them, and the 2026 Q&A slide pointed large sites to category hub pages instead.
Author Ibrahim Anjro
Keep a product page that is out of stock for a few months live with a 200 status, show its availability on the page and in Product structured data, and offer pre-ordering where possible when users would wait for the product. Redirect it only when users would rather switch to a similar product than wait.
Author Ibrahim Anjro
Used byrequirement DEV-ERR-04
Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.
Author Ibrahim Anjro
Used byrequirement DEV-IDX-01
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
Author Ibrahim Anjro
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
URL discovery works through links: a homepage links to section pages, which link to further pages.
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.
Google said listing the sitemap in robots.txt is fine, as many websites do.
Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.
Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.
When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.