Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

Sitemaps

A sitemap line in robots.txt lets any crawler find the sitemap; a sitemap not listed there has to be submitted, for example in Search Console. For HTML sitemaps on large sites, a Google Q&A slide said good HTML sitemaps are hard to build at that scale (said at the event) and advised relying instead on category hub pages that link to the important pages; Google's ecommerce guide recommends the same linking, with a sitemap or Merchant Center feed where not every product can be linked. As a canonical signal, listing only the preferred URL in sitemaps was named on stage among the signals that make a big difference, but Google's canonical guide rates sitemap inclusion as weak next to redirects and rel=canonical. Image and video sitemaps help Google find media, and an XML sitemap is one of three equivalent ways to declare hreflang. Author’s view: sitemaps support discovery but do not replace internal links, because pages known only from sitemaps are found slowly. Day 3 timed sitemaps and tied them to quality: Google estimated that processing a sitemap takes about 24 hours on average and that it tries to refetch a useful sitemap of a high-quality site within 14 days at most (neither figure is in Google's docs), while it may never fetch a lower-quality site's sitemap again. The second recordings added Google's framing: Gary Illyes said Google finds what to crawl mainly by extracting URLs from crawled pages and additionally from sitemaps, an XML format from around 2005 that is 'nothing fancy' but still used, and in the Day 1 Q&A Google said there is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that Google can find each of them; for a migration, listing the new URLs in a sitemap cannot hurt but is not the main tool. Day 2's opening Q&A said a site that wants an HTML sitemap can make one but the effort is better spent on hub pages, and that listing a sitemap in robots.txt is fine but makes it public. For migrations, Google's site move guide and a community case study both submit the new sitemap. Gary Illyes said video sitemaps are not critical but good to have, because Google ingests them much more often than it can process HTML pages and they tell it which URLs carry videos (said at the event).

What to do

  • Reference every XML sitemap in robots.txt, unless it must stay private, and submit it in Search Console as well.
  • List only canonical URLs in sitemaps, and back them with redirects and rel=canonical, which Google rates as stronger signals.
  • On large sites, build category hubs instead of an HTML sitemap page.
  • Add image and video entries, or separate media sitemaps, for important images and videos.

Day 1: Crawling 5

Said on stage 5

StageConfirmed by docsD1-C336

Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
StageConsistent with docsD1-C337

An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Extended by D3-C610 Day 3: Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.
  • Extended by D3-C611 Day 3: If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within…
StageD1-C474

An audience member asked where the sweet spot is for a sitemap that misses nothing but does not overflow Search Console.

From the audienceIn Day 1, 16:35 · Q&AEvidence transcript

  • Answered by D1-C475 Day 1: There is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that…
StageConfirmed by docsD1-C475

There is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that Google can find each of them.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Answers D1-C474 Day 1: An audience member asked where the sweet spot is for a sitemap that misses nothing but does not overflow…
StageConsistent with docsD1-C540

Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Things

Used byrequirement DEV-CAN-09

  • Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
  • Extended by D2-C884 Day 2: The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a…

Day 2: Indexing 26

Shown on screen 8

SlideD2-C001

An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.

From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo

Things
  • Answered by D2-C002 Day 2: Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
  • Answered by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
  • Answered by D2-C003 Day 2: Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have…
  • Answered by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
SlideNot in docsD2-C002

Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.

“It's hard to make good HTML sitemaps for large sites.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
SlideConsistent with docsD2-C820

Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

“Better rely on hubs like category pages that link out to your important pages.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

Used byrequirement DEV-URL-04

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extends D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
  • Extended by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
SlideD2-C015

An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.

From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

  • Answered by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
  • Answered by D2-C018 Day 2: Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in…
  • Answered by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
  • Answered by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
SlideConfirmed by docsD2-C017

Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

“If it's included in the robots.txt file, any crawler can pick your sitemaps up”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extended by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
SlideConsistent with docsD2-C018

Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in Search Console.

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
SlideConsistent with docsD2-C390

Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirements DEV-CAN-05, DEV-URL-05

Said on stage 8

StageConsistent with docsD2-C838

Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.

“if you want to make one, knock yourself out, but I will focus on some better things like hub pages, category pages”

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-URL-04glossary term Hub pages

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
  • Extends D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
StageConfirmed by docsD2-C842

Google said listing the sitemap in robots.txt is fine, as many websites do.

“you can include it in robots.txt. No problem whatsoever.”

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
StageConsistent with docsD2-C843

Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
StageConfirmed by docsD2-C884

The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.

Speaker Martyna AğanoğluIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

Used byrequirement DEV-CAN-09glossary term Change of Address tool

  • Extends D1-C540 Day 1: Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is…
StageConsistent with docsD2-C926

A video sitemap or video feed tells Google which URLs carry videos so it can visit those URLs to double-check; without one, Google has to check every page individually.

Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript

Things

Used byrequirement DEV-VID-04glossary term Video sitemap

What Google's documentation says 6

DocsSourceD2-C005

Google's ecommerce documentation recommends linking from menus to category pages, from category pages to sub-category pages and from sub-category pages to all product pages; where not every product can be linked, it recommends a sitemap or a Merchant Center feed.

Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!

Used byrequirement DEV-URL-04

DocsSourceD2-C391

Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.

Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirements DEV-CAN-05, DEV-URL-05glossary term rel=canonical

DocsSourceD2-C912

Google's site move guide says to submit a Change of Address in Search Console when moving from one domain or subdomain to another, to submit the new sitemap, and to change internal links on the new site from the old URLs to the new ones.

Publisher Google Search CentralAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

Used byrequirement DEV-CAN-09glossary term Change of Address tool

DocsSourceD2-C569

Google's hreflang guide says its three methods (link elements in the HTML head, an HTTP Link header, which suits non-HTML files such as PDFs, and an XML sitemap) are equivalent; using several at once is allowed but brings no benefit in Search and is harder to manage.

Publisher Google Search CentralAnnotates Day 2, 14:15 · Focusing on Internationalisation and Localisation

Used byrequirement DEV-INT-05

Analysis by the author 4

AnalysisD2-C004

Google's answer did not say whether HTML sitemaps pass link equity. On a large site, check that every important page is linked from a category or hub page in the normal navigation instead of relying on an HTML sitemap page to reach it.

Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!

AnalysisD2-C392

Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.

Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication

Day 3: Serving: Ranking, Search Console, and Performance 3

Shown on screen 3

SlideConsistent with docsD3-C612

Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.

Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript

Things
  • Extends D1-C331 Day 1: Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically…

Across days and sessions 12

  1. Stage D1-C336 Day 1 · How crawling works

    Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  2. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Slide D1-C079 Day 1 · How Google interprets robots.txt

    The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

  3. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  4. Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  5. Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

    extends
    Stage D1-C202 Day 1 · How Search works and where's AI?

    Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.

  6. Stage D2-C838 Day 2 · Welcome to indexing day!

    Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.

    extends
    Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

  7. Stage D2-C842 Day 2 · Welcome to indexing day!

    Google said listing the sitemap in robots.txt is fine, as many websites do.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  8. Stage D2-C843 Day 2 · Welcome to indexing day!

    Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.

    extends
    Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

  9. Stage D2-C884 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.

    extends
    Stage D1-C540 Day 1 · Q&A

    Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

  10. Slide D3-C610 Day 3 · How long does it take to..?

    Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.

    extends
    Stage D1-C337 Day 1 · How crawling works

    An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

  11. Slide D3-C611 Day 3 · How long does it take to..?

    If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.

    extends
    Stage D1-C337 Day 1 · How crawling works

    An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

  12. Slide D3-C612 Day 3 · How long does it take to..?

    Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.

    extends
    Stage D1-C331 Day 1 · How crawling works

    Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.

Built on these claims 8

Developer requirements 8

Sources 14