Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 1 · Wednesday 30 September 2026 · 14:05

How crawling works

Speakers Cherry Prommawin, Gary Illyes, Search Relations

TalkCoverageSlidesTranscript

Speakers from on-stage hand-overs between the two and the host's thanks to Gary and Cherry. Transcript from a second attendee recording; two audio recordings cover the talk except for about three minutes before the crawling infrastructure part.

Shown on screen 3

SlideConsistent with docsD1-C063

The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

Speaker Gary IllyesEvidence 3 slide photos, transcript

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
  • Extended by D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extended by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
  • Extended by D2-C127 Day 2: A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing…
  • Extended by D2-C167 Day 2: The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler…
  • Extended by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
SlideConsistent with docsD1-C064

The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

“Ensure we don't break the internet”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
  • Extended by D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
  • Extended by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
SlideConsistent with docsD1-C065

The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
  • Extended by D3-C385 Day 3: Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh…

Said on stage 25

StageD1-C317

Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.

Speaker Cherry PrommawinEvidence transcript

  • Repeated by D2-C847 Day 2: Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including…
StageNot in docsD1-C318

A URL states how a resource is requested (the protocol, HTTP or HTTPS), where (the host, meaning which computer on the network) and what (the path to the exact page or file).

Speaker Cherry PrommawinEvidence transcript

StageConsistent with docsD1-C320

An HTTP 200 OK status only means that the server believes it managed to do what the client asked for.

“just means that the server believes that it managed to accomplish whatever the user was asking for”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-ERR-03

StageConsistent with docsD1-C322

Googlebot is an ordinary HTTP client with nothing special about it: like a browser, it fetches a URL it was given and returns the fetched bytes to Google's servers.

“Googlebot is just a client. It is an HTTP client. There's nothing all that much special about it.”

Speaker Gary IllyesEvidence transcript

Things
StageNot in docsD1-C323

AI agents are, technically, the same thing as crawlers: HTTP clients that accomplish something on behalf of a user or a service.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD1-C326

Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.

Speaker Gary IllyesEvidence transcript

  • Extended by D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
StageD1-C327

Google said it does not usually talk publicly about crawl components such as the scheduler and the crawl queue, because the details get confusing and taken out of context.

Speaker Gary IllyesEvidence transcript

Used byglossary term Crawl scheduler

StageConsistent with docsD1-C328

During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

Speaker Gary IllyesEvidence transcript

  • Repeated by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
StageConsistent with docsD1-C329

Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

Speaker Gary IllyesEvidence transcript

  • Extended by D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
  • Repeated by D2-C848 Day 2: Google said it strongly believes site owners should be able to opt out of crawling and control how their site…
StageConsistent with docsD1-C330

The only Google-owned crawlers that do not obey robots.txt are contractual crawlers, which crawl a site whose owner has agreed that Google may crawl it however it likes.

Speaker Gary IllyesEvidence transcript

Used byglossary term Contractual crawlers (special-case crawlers)

StageConsistent with docsD1-C331

Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C612 Day 3: Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower…
StageConsistent with docsD1-C332

The crawl scheduler logs how often each page changes and crawls frequently changing pages first: a news site's homepage, which changes very often, is prioritised over its terms of service page, which may change once a year.

Speaker Gary IllyesEvidence transcript

Used byglossary term Crawl scheduler

StageNot in docsD1-C333

The scheduler hands the crawler an ordered list of URLs from the crawl queue, and the crawler works through the list from top to bottom.

Speaker Gary IllyesEvidence transcript

Used byglossary term Crawl scheduler

StageConsistent with docsD1-C334

Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.

Speaker Gary IllyesEvidence transcript

  • Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
StageNot in docsD1-C335

Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

“the number of tokens is actually more important”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

Things
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageConfirmed by docsD1-C336

Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
StageConsistent with docsD1-C337

An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Extended by D3-C610 Day 3: Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.
  • Extended by D3-C611 Day 3: If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within…
StageNot in docsD1-C338

Google described its crawling as a large-scale distributed swarm of simple HTTP clients, roughly what one would get by deploying many wget or curl libraries on cloud compute instances.

“a large-scale distributed swarm of simple HTTP clients”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

StageD1-C341

Gary Illyes called the Crawl Stats report the most valuable tool for debugging crawl errors, together with server logs, which he called the real thing for those who know how to read them.

Speaker Gary IllyesEvidence transcript

What Google's documentation says 4

DocsSourceD1-C136

Google's Inside Googlebot post (March 2026) says Googlebot currently fetches only the first 2MB of each URL, HTTP headers included (64MB for PDFs); bytes past that cutoff are not fetched, rendered or indexed, and each resource the page loads has its own separate limit.

Publisher Search Central blog (31 March 2026)

Things

Used byrequirement DEV-PRF-04

DocsSourceD1-C137

Google's Inside Googlebot post warns that bloated inline base64 images, large blocks of inline CSS or JavaScript, or megabytes of menus can push a page's text or structured data past Googlebot's 2MB cutoff, and advises moving heavy CSS and JavaScript to external files and placing meta tags, the title, the canonical and essential structured data high in the HTML.

Publisher Search Central blog (31 March 2026)

Used byrequirement DEV-PRF-04

DocsSourceD1-C342

Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.

“Googlebot is just a user of something that resembles a centralized crawling platform”

Publisher Search Central blog (31 March 2026)

DocsSourceD1-C343

Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.

Publisher Google

Used byglossary term Contractual crawlers (special-case crawlers)

Analysis by the author 2

AnalysisD1-C067

A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

Author Ibrahim Anjro

Used byrequirements DEV-URL-04, DEV-URL-05

  • Extended by D2-C189 Day 2: Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script…
  • Extended by D2-C427 Day 2: When internal links point only to page A and the canonical leader is reached only through A's canonical link…
AnalysisD1-C344

On stage Google spoke of probably hundreds, if not thousands, of crawlers, while its Inside Googlebot post speaks of dozens of other clients; both agree that only the larger crawlers are documented, so a Google user agent missing from the public lists is not proof of a fake request, and reverse DNS or Google's published IP ranges are the test.

Author Ibrahim Anjro

Things
  1. Slide D1-C066 Day 1 · How Google thinks about crawl budget

    Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

    extends
    Stage D1-C326 Day 1 · How crawling works

    Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.

  2. Stage D1-C325 Day 1 · How crawling works

    Googlebot is the crawler Google uses for web search, including Search's AI features.

    extends
    Slide D1-C038 Day 1 · How Search works and where's AI?

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

  3. Stage D1-C335 Day 1 · How crawling works

    Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  4. Stage D1-C336 Day 1 · How crawling works

    Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  5. Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

    extends
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

  6. Stage D1-C538 Day 1 · Q&A

    A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.

    extends
    Stage D1-C334 Day 1 · How crawling works

    Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.

  7. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  8. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  9. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  10. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  11. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  12. Slide D2-C127 Day 2 · Lightning session D: Rendering and JavaScript

    A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  13. Slide D2-C167 Day 2 · Lightning session D: Rendering and JavaScript

    The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  14. Analysis D2-C189 Day 2 · Lightning session D: Rendering and JavaScript

    Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.

    extends
    Analysis D1-C067 Day 1 · How crawling works

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

  15. Slide D2-C204 Day 2 · Lightning session D: Rendering and JavaScript

    Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  16. Stage D2-C296 Day 2 · What is Google friendly JavaScript

    If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  17. Stage D2-C427 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.

    extends
    Analysis D1-C067 Day 1 · How crawling works

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

  18. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  19. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  20. Stage D3-C385 Day 3 · Shopping on Search: Beyond the blue links

    Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  21. Slide D3-C610 Day 3 · How long does it take to..?

    Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.

    extends
    Stage D1-C337 Day 1 · How crawling works

    An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

  22. Slide D3-C611 Day 3 · How long does it take to..?

    If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.

    extends
    Stage D1-C337 Day 1 · How crawling works

    An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

  23. Slide D3-C612 Day 3 · How long does it take to..?

    Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.

    extends
    Stage D1-C331 Day 1 · How crawling works

    Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.

  24. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    repeats
    Stage D1-C328 Day 1 · How crawling works

    During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

  25. Stage D2-C847 Day 2 · Welcome to indexing day!

    Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.

    repeats
    Stage D1-C317 Day 1 · How crawling works

    Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.

  26. Stage D2-C848 Day 2 · Welcome to indexing day!

    Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.

    repeats
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.