Day 1: Crawling 22
Shown on screen 4
The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence 3 slide photos, transcript
- Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
- Extended by D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
- Extended by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
- Extended by D2-C127 Day 2: A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing…
- Extended by D2-C167 Day 2: The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler…
- Extended by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
“Ensure we don't break the internet”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence slide photo, transcript
- Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
- Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
- Extended by D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
- Extended by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence slide photo, transcript
- Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
- Extended by D3-C385 Day 3: Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh…
Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-PRF-01
- Extends D1-C326 Day 1: Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure…
- Extended by D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…
Said on stage 13
A crawler is software that downloads pages, extracts their links and repeats the process on the links it extracted; Googlebot is the main crawler of Google Search.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
Speaker Cherry PrommawinIn Day 1, 14:05 · How crawling worksEvidence transcript
- Repeated by D2-C847 Day 2: Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including…
Googlebot is an ordinary HTTP client with nothing special about it: like a browser, it fetches a URL it was given and returns the fetched bytes to Google's servers.
“Googlebot is just a client. It is an HTTP client. There's nothing all that much special about it.”
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Google runs probably hundreds, if not thousands, of crawlers on its crawler infrastructure; some of them are named and some are not.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Googlebot is the crawler Google uses for web search, including Search's AI features.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Extended by D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
Google said it does not usually talk publicly about crawl components such as the scheduler and the crawl queue, because the details get confusing and taken out of context.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byglossary term Crawl scheduler
During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Repeated by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
The scheduler hands the crawler an ordered list of URLs from the crawl queue, and the crawler works through the list from top to bottom.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byglossary term Crawl scheduler
Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
Google described its crawling as a large-scale distributed swarm of simple HTTP clients, roughly what one would get by deploying many wget or curl libraries on cloud compute instances.
“a large-scale distributed swarm of simple HTTP clients”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Extends D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
- Extended by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Answers D1-C387 Day 1: An audience member asked, in questions submitted before the event, how often Googlebot should be expected to…
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
What Google's documentation says 3
Google's Inside Googlebot post (March 2026) says Googlebot currently fetches only the first 2MB of each URL, HTTP headers included (64MB for PDFs); bytes past that cutoff are not fetched, rendered or indexed, and each resource the page loads has its own separate limit.
Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works
Used byrequirement DEV-PRF-04
Google's Inside Googlebot post warns that bloated inline base64 images, large blocks of inline CSS or JavaScript, or megabytes of menus can push a page's text or structured data past Googlebot's 2MB cutoff, and advises moving heavy CSS and JavaScript to external files and placing meta tags, the title, the canonical and essential structured data high in the HTML.
Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works
Used byrequirement DEV-PRF-04
Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.
“Googlebot is just a user of something that resembles a centralized crawling platform”
Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works
Analysis by the author 2
Because every Google product shares the same host capacity, Ads or Shopping fetches that hit a slow server can reduce how much is crawled for Search.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
On stage Google spoke of probably hundreds, if not thousands, of crawlers, while its Inside Googlebot post speaks of dozens of other clients; both agree that only the larger crawlers are documented, so a Google user agent missing from the public lists is not proof of a fake request, and reverse DNS or Google's published IP ranges are the test.
Author Ibrahim AnjroAnnotates Day 1, 14:05 · How crawling works
Day 2: Indexing 5
Shown on screen 4
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo
Used byrequirement DEV-PRF-01
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
- Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
- Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
- Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
- Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence slide photo
- Repeats D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
Said on stage 1
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
- Repeats D1-C317 Day 1: Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses…
Day 3: Serving: Ranking, Search Console, and Performance 6
Said on stage 3
Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.
Speaker Alex JansenIn Day 3, 13:40 · Shopping on Search: Beyond the blue linksEvidence transcript
Used byrequirement DEV-SHP-02glossary term Storebot-Google
- Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
Storebot-Google is used to validate merchant feeds: when a feed gives a price, Google double-checks it on the merchant's site.
Speaker Alex JansenIn Day 3, 13:40 · Shopping on Search: Beyond the blue linksEvidence transcript
Used byrequirement DEV-SHP-02
Storebot-Google also does deep crawls, for example of checkout pages, which do not belong in the search index; this is another reason it is kept separate from Googlebot.
Speaker Alex JansenIn Day 3, 13:40 · Shopping on Search: Beyond the blue linksEvidence transcript
Used byrequirement DEV-SHP-02
What Google's documentation says 2
Google's Merchant Center help says the StoreBot crawler goes through product detail, cart and checkout pages, can fill in checkout forms, and records price, shipping, availability, coupons and payment methods to verify the data merchants share in Merchant Center.
Publisher Google Merchant Center HelpAnnotates Day 3, 13:40 · Shopping on Search: Beyond the blue links
Used byrequirement DEV-SHP-02glossary term Storebot-Google
Google's crawler list says robots.txt rules addressed to the Storebot-Google user agent affect all surfaces of Google Shopping, such as the Shopping tab in Google Search.
Publisher GoogleAnnotates Day 3, 13:40 · Shopping on Search: Beyond the blue links
Used byrequirement DEV-SHP-02glossary term Storebot-Google
Analysis by the author 1
Treat Storebot-Google separately from Googlebot in robots.txt and bot protection: a shop that blocks it on product, cart or checkout paths can undermine Shopping price and availability checks, even when Googlebot is allowed.
Author Ibrahim AnjroAnnotates Day 3, 13:40 · Shopping on Search: Beyond the blue links
Used byrequirement DEV-SHP-02