Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

The crawl pipeline

URLs wait in a crawl queue, a shared scheduler decides what to fetch and when, and the crawler fetches while enforcing robots.txt and protecting servers. One centralised infrastructure serves Search, Ads, Shopping and Images. Day 2 showed an example of what the crawler hands to processing: a fetch record with the fetch result, connect time and time to first byte in milliseconds, the robots policies that apply (such as a Google-Extended opt-out) and the raw HTTP response; this record was shown on a slide and is not in Google's docs. Links extracted during processing go back to the crawl queue, and the renderer fetches JavaScript and CSS through the same crawler. Day 3 named a second crawler on that shared infrastructure: Google Shopping uses Storebot-Google instead of Googlebot because it needs fresh prices, availability and shipping details and crawls more often; it validates merchant feeds by double-checking prices on the site and crawls product, cart and checkout pages, as Google's Merchant Center help confirms, and its robots.txt group governs all Shopping surfaces. The second recording of Day 1 added the crawling talk itself. Cherry Prommawin defined a crawler as software that downloads pages, extracts their links and repeats the process, and opened with how the internet works, from TCP/IP to DNS and HTTP. Gary Illyes said Googlebot is an ordinary HTTP client with nothing special about it, and that Google runs probably hundreds, if not thousands, of crawlers but does not let teams build their own: all share one infrastructure and obey Google's internal crawling policies (Google's Inside Googlebot post speaks of dozens of other clients); he described it as a large distributed swarm of simple HTTP clients, roughly many wget or curl instances. The scheduler hands the crawler an ordered list that it works through from top to bottom, and teams prioritise differently: web search cares a lot about site and content quality, while Ads checks every publisher page as it comes in (said at the event). URLs extracted during indexing go back to the scheduler, and the crawl budget talk added that a request from any Google user agent is routed through the shared platform. Author’s view: a Google user agent missing from the public lists is not proof of a fake request; verify with reverse DNS or Google's IP ranges.

What to do

  • Keep connect time and time to first byte low and stable; Google's example fetch record carries both, and they drive how fast Google crawls.
  • Keep the JavaScript and CSS files a page needs crawlable, because the renderer fetches them through the crawler.
  • Allow Storebot-Google in robots.txt and bot protection on product, cart and checkout paths, even where other bots are blocked.

Day 1: Crawling 22

Shown on screen 4

SlideConsistent with docsD1-C063

The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence 3 slide photos, transcript

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
  • Extended by D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extended by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
  • Extended by D2-C127 Day 2: A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing…
  • Extended by D2-C167 Day 2: The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler…
  • Extended by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
SlideConsistent with docsD1-C064

The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

“Ensure we don't break the internet”

Wording checked against the slide or recording

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence slide photo, transcript

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
  • Extended by D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
  • Extended by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
SlideConsistent with docsD1-C065

The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence slide photo, transcript

  • Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
  • Extended by D3-C385 Day 3: Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh…
SlideConsistent with docsD1-C066

Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript

Used byrequirement DEV-PRF-01

  • Extends D1-C326 Day 1: Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure…
  • Extended by D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…

Said on stage 13

StageD1-C317

Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.

Speaker Cherry PrommawinIn Day 1, 14:05 · How crawling worksEvidence transcript

  • Repeated by D2-C847 Day 2: Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including…
StageConsistent with docsD1-C322

Googlebot is an ordinary HTTP client with nothing special about it: like a browser, it fetches a URL it was given and returns the fetched bytes to Google's servers.

“Googlebot is just a client. It is an HTTP client. There's nothing all that much special about it.”

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things
StageConsistent with docsD1-C325

Googlebot is the crawler Google uses for web search, including Search's AI features.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things
  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageConsistent with docsD1-C326

Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

  • Extended by D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
StageConsistent with docsD1-C334

Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

  • Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
StageNot in docsD1-C338

Google described its crawling as a large-scale distributed swarm of simple HTTP clients, roughly what one would get by deploying many wget or curl libraries on cloud compute instances.

“a large-scale distributed swarm of simple HTTP clients”

Wording checked against the slide or recording

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

StageConfirmed by docsD1-C375

Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Things
  • Extends D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
  • Extended by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
StageConsistent with docsD1-C388

AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Things
  • Answers D1-C387 Day 1: An audience member asked, in questions submitted before the event, how often Googlebot should be expected to…
  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…

What Google's documentation says 3

DocsSourceD1-C136

Google's Inside Googlebot post (March 2026) says Googlebot currently fetches only the first 2MB of each URL, HTTP headers included (64MB for PDFs); bytes past that cutoff are not fetched, rendered or indexed, and each resource the page loads has its own separate limit.

Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works

Things

Used byrequirement DEV-PRF-04

DocsSourceD1-C137

Google's Inside Googlebot post warns that bloated inline base64 images, large blocks of inline CSS or JavaScript, or megabytes of menus can push a page's text or structured data past Googlebot's 2MB cutoff, and advises moving heavy CSS and JavaScript to external files and placing meta tags, the title, the canonical and essential structured data high in the HTML.

Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works

Used byrequirement DEV-PRF-04

DocsSourceD1-C342

Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.

“Googlebot is just a user of something that resembles a centralized crawling platform”

Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works

Analysis by the author 2

Day 2: Indexing 5

Shown on screen 4

SlideNot in docsD2-C024

A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo

Used byrequirement DEV-PRF-01

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
SlideConfirmed by docsD2-C048

Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
  • Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
  • Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
SlideConsistent with docsD2-C127

A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…

Said on stage 1

Day 3: Serving: Ranking, Search Console, and Performance 6

Said on stage 3

StageConsistent with docsD3-C385

Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.

Speaker Alex JansenIn Day 3, 13:40 · Shopping on Search: Beyond the blue linksEvidence transcript

Used byrequirement DEV-SHP-02glossary term Storebot-Google

  • Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
StageConfirmed by docsD3-C387

Storebot-Google also does deep crawls, for example of checkout pages, which do not belong in the search index; this is another reason it is kept separate from Googlebot.

Speaker Alex JansenIn Day 3, 13:40 · Shopping on Search: Beyond the blue linksEvidence transcript

Used byrequirement DEV-SHP-02

What Google's documentation says 2

DocsSourceD3-C388

Google's Merchant Center help says the StoreBot crawler goes through product detail, cart and checkout pages, can fill in checkout forms, and records price, shipping, availability, coupons and payment methods to verify the data merchants share in Merchant Center.

Publisher Google Merchant Center HelpAnnotates Day 3, 13:40 · Shopping on Search: Beyond the blue links

Used byrequirement DEV-SHP-02glossary term Storebot-Google

DocsSourceD3-C389

Google's crawler list says robots.txt rules addressed to the Storebot-Google user agent affect all surfaces of Google Shopping, such as the Shopping tab in Google Search.

Publisher GoogleAnnotates Day 3, 13:40 · Shopping on Search: Beyond the blue links

Used byrequirement DEV-SHP-02glossary term Storebot-Google

Analysis by the author 1

Across days and sessions 23

  1. Slide D1-C066 Day 1 · How Google thinks about crawl budget

    Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

    extends
    Stage D1-C326 Day 1 · How crawling works

    Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.

  2. Stage D1-C325 Day 1 · How crawling works

    Googlebot is the crawler Google uses for web search, including Search's AI features.

    extends
    Slide D1-C038 Day 1 · How Search works and where's AI?

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

  3. Stage D1-C375 Day 1 · How Google thinks about crawl budget

    Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

    extends
    Slide D1-C066 Day 1 · How Google thinks about crawl budget

    Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

  4. Stage D1-C388 Day 1 · How Google thinks about crawl budget

    AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.

    extends
    Slide D1-C038 Day 1 · How Search works and where's AI?

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

  5. Stage D1-C538 Day 1 · Q&A

    A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.

    extends
    Stage D1-C334 Day 1 · How crawling works

    Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.

  6. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  7. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  8. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C092 Day 1 · How Google thinks about crawl budget

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

  9. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  10. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  11. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  12. Slide D2-C127 Day 2 · Lightning session D: Rendering and JavaScript

    A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  13. Slide D2-C167 Day 2 · Lightning session D: Rendering and JavaScript

    The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  14. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

  15. Slide D2-C204 Day 2 · Lightning session D: Rendering and JavaScript

    Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  16. Stage D2-C296 Day 2 · What is Google friendly JavaScript

    If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  17. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  18. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  19. Stage D3-C385 Day 3 · Shopping on Search: Beyond the blue links

    Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  20. Stage D3-C624 Day 3 · How long does it take to..?

    Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.

    extends
    Stage D1-C375 Day 1 · How Google thinks about crawl budget

    Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

  21. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    repeats
    Stage D1-C328 Day 1 · How crawling works

    During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

  22. Slide D2-C442 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

    repeats
    Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

  23. Stage D2-C847 Day 2 · Welcome to indexing day!

    Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.

    repeats
    Stage D1-C317 Day 1 · How crawling works

    Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.

Built on these claims 3

Developer requirements 3

Sources 10