Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

URL discovery

Google finds new URLs by following links from pages it already knows. Author’s view: pages without internal links depend on sitemaps alone and are found slowly. Day 2 added that Google uses extracted links for discovery, for understanding a site's structure and for ranking, and extracts them from the HTML before rendering and again from the rendered HTML, sending new URLs back to the crawl queue. For large sites, a Google Q&A slide said good HTML sitemaps are hard to do well at that scale (said at the event) and advised relying instead on category hub pages that link to important pages, the kind of hub page Google's guide to how Search works also describes; Google's ecommerce guide recommends the same menu-to-category-to-product linking, with a sitemap or Merchant Center feed where not every product can be linked. Google said it sometimes picks up URLs written as plain text, but the remark was unclear in the recording and is not documented. Day 3 added scale and timing: Google said it knows hundreds of trillions of URLs (as of October 2026) but does not crawl them all, and estimated that discovering a new URL takes about 20 hours on average, from seconds to weeks or never (not in Google's docs). Author’s view: since discovery and recrawling take far longer than indexing, internal links from often-crawled pages and accurate sitemaps are what speed results up. The second recording of Day 1 added that there are trillions of URLs on the internet or more, that even Google cannot tell how many, and that some may never be discovered; Google finds what to crawl mainly by extracting URLs from crawled pages and passing them back to the scheduler, and additionally from sitemaps, and may visit hub pages such as the homepage and category pages more often because they link to new or updated pages. Search Console's crawl data separates discovery fetches of new URLs from refreshes of known ones. Day 2's opening Q&A repeated that the effort an HTML sitemap takes on a large site is better spent on hub pages such as category pages.

What to do

  • Link every important page from a category or hub page in the normal navigation; on a large site, do not rely on an HTML sitemap page to reach it.
  • Use real <a href> links, ideally in the server HTML; plain-text URLs are not a dependable discovery route.
  • Do not leave a canonical target reachable only through a rel=canonical link; link to it internally.

Day 1: Crawling 7

Shown on screen 1

SlideConfirmed by docsD1-C040

URL discovery works through links: a homepage links to section pages, which link to further pages.

Speaker Cherry Prommawin, Gary IllyesIn Day 1, 11:45 · How Search works and where's AI?Evidence slide photo, transcript

Used byrequirement DEV-URL-04

  • Extended by D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
  • Extended by D1-C336 Day 1: Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from…
  • Extended by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
  • Extended by D2-C038 Day 2: Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure…
  • Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
  • Extended by D2-C169 Day 2: Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before…
  • Extended by D2-C287 Day 2: A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google…

Said on stage 5

StageNot in docsD1-C201

Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.

Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript

  • Extended by D3-C602 Day 3: Google said it knows hundreds of trillions of URLs (as of October 2026).
StageConsistent with docsD1-C202

Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.

Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript

Used byrequirement DEV-URL-04glossary term Hub pages

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extended by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
StageConfirmed by docsD1-C336

Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
StageConfirmed by docsD1-C396

Search Console's crawl data separates discovery (fetches of URLs Google has not seen before) from refresh (fetches of known URLs), and lets you drill down to the problematic URLs.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Used byrequirements DEV-MON-04, DEV-MON-11glossary term Crawl Stats report

Analysis by the author 1

AnalysisD1-C067

A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

Author Ibrahim AnjroAnnotates Day 1, 14:05 · How crawling works

Used byrequirements DEV-URL-04, DEV-URL-05

  • Extended by D2-C189 Day 2: Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script…
  • Extended by D2-C427 Day 2: When internal links point only to page A and the canonical leader is reached only through A's canonical link…

Day 2: Indexing 20

Shown on screen 7

SlideD2-C001

An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.

From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo

Things
  • Answered by D2-C002 Day 2: Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
  • Answered by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
  • Answered by D2-C003 Day 2: Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have…
  • Answered by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
SlideNot in docsD2-C002

Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.

“It's hard to make good HTML sitemaps for large sites.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
SlideConsistent with docsD2-C820

Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

“Better rely on hubs like category pages that link out to your important pages.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

Used byrequirement DEV-URL-04

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extends D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
  • Extended by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
SlideD2-C003

Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have the extra benefit that they may also be useful to users.

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo

Used byrequirement DEV-URL-04

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
SlideConfirmed by docsD2-C048

Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
  • Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
  • Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
SlideConfirmed by docsD2-C168

In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

Things
  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extends D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
SlideConsistent with docsD2-C415

The crawler report shown on screen flagged a URL discovered only through a canonical tag, with no internal a href links leading to it, as a phantom document outside the visible site structure.

“a phantom document not part of the visible site structure”

Wording checked against the slide or recording

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence slide photo

Said on stage 6

StageConsistent with docsD2-C838

Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.

“if you want to make one, knock yourself out, but I will focus on some better things like hub pages, category pages”

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-URL-04glossary term Hub pages

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
  • Extends D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
StageConsistent with docsD2-C038

Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript

Used byrequirement DEV-URL-01

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
StageConsistent with docsD2-C427

When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.

“you're telling the search engine to index something that a human can't reach”

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

Used byrequirement DEV-CAN-04

  • Extends D1-C067 Day 1: A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts…
StageConsistent with docsD2-C926

A video sitemap or video feed tells Google which URLs carry videos so it can visit those URLs to double-check; without one, Google has to check every page individually.

Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript

Things

Used byrequirement DEV-VID-04glossary term Video sitemap

What Google's documentation says 3

DocsSourceD2-C005

Google's ecommerce documentation recommends linking from menus to category pages, from category pages to sub-category pages and from sub-category pages to all product pages; where not every product can be linked, it recommends a sitemap or a Merchant Center feed.

Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!

Used byrequirement DEV-URL-04

Analysis by the author 4

AnalysisD2-C004

Google's answer did not say whether HTML sitemaps pass link equity. On a large site, check that every important page is linked from a category or hub page in the normal navigation instead of relying on an HTML sitemap page to reach it.

Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!

AnalysisD2-C431

A leader reachable only through a canonical and a leader with weak link equity share one fix: point internal links, sitemap entries and redirects at the URL chosen as canonical, so the leader is both reachable for users and the strongest page in its group.

Author Ibrahim AnjroAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

Used byrequirement DEV-CAN-05

Day 3: Serving: Ranking, Search Console, and Performance 6

Shown on screen 2

Said on stage 2

Analysis by the author 2

AnalysisD3-C677

Google's ranking systems guide speaks of hundreds of billions of pages in its Search index; by the author's arithmetic, against the hundreds of trillions of URLs Google said it knows, only around one known URL in a thousand is indexed, if both figures are taken at face value.

Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?

Across days and sessions 16

  1. Stage D1-C202 Day 1 · How Search works and where's AI?

    Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  2. Stage D1-C336 Day 1 · How crawling works

    Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  3. Stage D2-C038 Day 2 · How is HTML interpreted

    Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  4. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  5. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  6. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

  7. Docs D2-C169 Day 2 · Lightning session D: Rendering and JavaScript

    Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  8. Analysis D2-C189 Day 2 · Lightning session D: Rendering and JavaScript

    Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.

    extends
    Analysis D1-C067 Day 1 · How crawling works

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

  9. Stage D2-C287 Day 2 · What is Google friendly JavaScript

    A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  10. Stage D2-C427 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.

    extends
    Analysis D1-C067 Day 1 · How crawling works

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

  11. Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  12. Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

    extends
    Stage D1-C202 Day 1 · How Search works and where's AI?

    Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.

  13. Stage D2-C838 Day 2 · Welcome to indexing day!

    Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.

    extends
    Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

  14. Stage D3-C602 Day 3 · How long does it take to..?

    Google said it knows hundreds of trillions of URLs (as of October 2026).

    extends
    Stage D1-C201 Day 1 · How Search works and where's AI?

    Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.

  15. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    repeats
    Stage D1-C328 Day 1 · How crawling works

    During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

  16. Slide D2-C442 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

    repeats
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

Built on these claims 8

Developer requirements 8

Sources 11