Day 1: Crawling 7
Shown on screen 1
URL discovery works through links: a homepage links to section pages, which link to further pages.
Speaker Cherry Prommawin, Gary IllyesIn Day 1, 11:45 · How Search works and where's AI?Evidence slide photo, transcript
Used byrequirement DEV-URL-04
- Extended by D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
- Extended by D1-C336 Day 1: Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from…
- Extended by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
- Extended by D2-C038 Day 2: Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure…
- Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
- Extended by D2-C169 Day 2: Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before…
- Extended by D2-C287 Day 2: A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google…
Said on stage 5
Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
- Extended by D3-C602 Day 3: Google said it knows hundreds of trillions of URLs (as of October 2026).
Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Used byrequirement DEV-URL-04glossary term Hub pages
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
- Extended by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Repeated by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byrequirement DEV-URL-05
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
Search Console's crawl data separates discovery (fetches of URLs Google has not seen before) from refresh (fetches of known URLs), and lets you drill down to the problematic URLs.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Used byrequirements DEV-MON-04, DEV-MON-11glossary term Crawl Stats report
Analysis by the author 1
A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
Author Ibrahim AnjroAnnotates Day 1, 14:05 · How crawling works
Used byrequirements DEV-URL-04, DEV-URL-05
- Extended by D2-C189 Day 2: Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script…
- Extended by D2-C427 Day 2: When internal links point only to page A and the canonical leader is reached only through A's canonical link…
Day 2: Indexing 20
Shown on screen 7
An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo
- Answered by D2-C002 Day 2: Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
- Answered by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
- Answered by D2-C003 Day 2: Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have…
- Answered by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
“It's hard to make good HTML sitemaps for large sites.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
“Better rely on hubs like category pages that link out to your important pages.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
Used byrequirement DEV-URL-04
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
- Extends D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
- Extended by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have the extra benefit that they may also be useful to users.
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo
Used byrequirement DEV-URL-04
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
- Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
- Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
- Extends D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
The crawler report shown on screen flagged a URL discovered only through a canonical tag, with no internal a href links leading to it, as a phantom document outside the visible site structure.
“a phantom document not part of the visible site structure”
Wording checked against the slide or recording
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence slide photo
Said on stage 6
Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.
“if you want to make one, knock yourself out, but I will focus on some better things like hub pages, category pages”
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-URL-04glossary term Hub pages
- Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
- Extends D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
The speaker said links are still an extremely important part of the internet and of most major search and AI systems.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript
Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript
Used byrequirement DEV-URL-01
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
The speaker said Google sometimes also extracts URLs that are typed out as plain text on a page without being hyperlinked; the remarks around this point were unclear in the recording.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript
When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.
“you're telling the search engine to index something that a human can't reach”
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-CAN-04
- Extends D1-C067 Day 1: A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts…
A video sitemap or video feed tells Google which URLs carry videos so it can visit those URLs to double-check; without one, Google has to check every page individually.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-VID-04glossary term Video sitemap
What Google's documentation says 3
Google's ecommerce documentation recommends linking from menus to category pages, from category pages to sub-category pages and from sub-category pages to all product pages; where not every product can be linked, it recommends a sitemap or a Merchant Center feed.
Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-URL-04
Google's ecommerce documentation says Google can infer a page's relative importance within a site from its internal links, such as how many links point to the page and how many links Google must follow to reach it.
Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-URL-04
A 2005 Search Central blog post, now marked as possibly outdated, said Google encouraged HTML sitemaps because they help users navigate a site and a clear hierarchy of text links helps Google index it.
Publisher Search Central blog (26 September 2005)Annotates Day 2, 10:15 · Welcome to indexing day!
Analysis by the author 4
Google's answer did not say whether HTML sitemaps pass link equity. On a large site, check that every important page is linked from a category or hub page in the normal navigation instead of relying on an HTML sitemap page to reach it.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
Google's view of HTML sitemaps has moved: a 2005 blog post encouraged them, the current sitemap and ecommerce documentation does not mention them, and the 2026 Q&A slide pointed large sites to category hub pages instead.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
Do not rely on plain-text URLs for discovery: even if Google sometimes picks them up, a proper a href link is what was described as feeding discovery, site structure and ranking.
Author Ibrahim AnjroAnnotates Day 2, 10:25 · How is HTML interpreted
A leader reachable only through a canonical and a leader with weak link equity share one fix: point internal links, sitemap entries and redirects at the URL chosen as canonical, so the leader is both reachable for users and the strongest page in its group.
Author Ibrahim AnjroAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves
Used byrequirement DEV-CAN-05
Day 3: Serving: Ranking, Search Console, and Performance 6
Shown on screen 2
Google estimated that discovering a new URL takes about 20 hours on average across all sites it knows of, with a minimum of seconds.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence 2 slide photos, transcript
Google's crawl chart says the discovery of a new URL can take weeks or never happen, because the web is huge.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript
Said on stage 2
Google said it knows hundreds of trillions of URLs (as of October 2026).
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
- Extends D1-C201 Day 1: Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many…
Google does not crawl all the URLs it knows, because crawling everything would generally waste resources and a much smaller crawl space is enough.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
Analysis by the author 2
By Google's own averages, getting found and crawled takes far longer than indexing (about 20 hours to discover a URL and 30 days to refresh one, against 1.5 hours to index), so for faster results work on discovery: internal links from often-crawled pages and accurate sitemaps.
Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?
Google's ranking systems guide speaks of hundreds of billions of pages in its Search index; by the author's arithmetic, against the hundreds of trillions of URLs Google said it knows, only around one known URL in a thousand is indexed, if both figures are taken at face value.
Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?