Speakers from on-stage hand-overs between the two and the host's thanks to Gary and Cherry. Transcript from a second attendee recording; two audio recordings cover the talk except for about three minutes before the crawling infrastructure part.
Said on stage 25
Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
Speaker Cherry PrommawinEvidence transcript
- Repeated by D2-C847 Day 2: Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including…
A URL states how a resource is requested (the protocol, HTTP or HTTPS), where (the host, meaning which computer on the network) and what (the path to the exact page or file).
Speaker Cherry PrommawinEvidence transcript
DNS works as the internet's address book: it is consulted for each request and tells the client at which IP address a host name can be reached.
Speaker Cherry PrommawinEvidence transcript
An HTTP 200 OK status only means that the server believes it managed to do what the client asked for.
“just means that the server believes that it managed to accomplish whatever the user was asking for”
Wording checked against the slide or recording
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-ERR-03
Gary Illyes pointed to the Last-Modified and ETag headers as the fields of an HTTP response that matter for caching.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-SRV-08
Googlebot is an ordinary HTTP client with nothing special about it: like a browser, it fetches a URL it was given and returns the fetched bytes to Google's servers.
“Googlebot is just a client. It is an HTTP client. There's nothing all that much special about it.”
Speaker Gary IllyesEvidence transcript
AI agents are, technically, the same thing as crawlers: HTTP clients that accomplish something on behalf of a user or a service.
Speaker Gary IllyesEvidence transcript
Google runs probably hundreds, if not thousands, of crawlers on its crawler infrastructure; some of them are named and some are not.
Speaker Gary IllyesEvidence transcript
Googlebot is the crawler Google uses for web search, including Search's AI features.
Speaker Gary IllyesEvidence transcript
- Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.
Speaker Gary IllyesEvidence transcript
- Extended by D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
Google said it does not usually talk publicly about crawl components such as the scheduler and the crawl queue, because the details get confusing and taken out of context.
Speaker Gary IllyesEvidence transcript
Used byglossary term Crawl scheduler
During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.
Speaker Gary IllyesEvidence transcript
- Repeated by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.
Speaker Gary IllyesEvidence transcript
- Extended by D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
- Repeated by D2-C848 Day 2: Google said it strongly believes site owners should be able to opt out of crawling and control how their site…
The only Google-owned crawlers that do not obey robots.txt are contractual crawlers, which crawl a site whose owner has agreed that Google may crawl it however it likes.
Speaker Gary IllyesEvidence transcript
Used byglossary term Contractual crawlers (special-case crawlers)
Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.
Speaker Gary IllyesEvidence transcript
- Extended by D3-C612 Day 3: Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower…
The crawl scheduler logs how often each page changes and crawls frequently changing pages first: a news site's homepage, which changes very often, is prioritised over its terms of service page, which may change once a year.
Speaker Gary IllyesEvidence transcript
Used byglossary term Crawl scheduler
The scheduler hands the crawler an ordered list of URLs from the crawl queue, and the crawler works through the list from top to bottom.
Speaker Gary IllyesEvidence transcript
Used byglossary term Crawl scheduler
Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.
Speaker Gary IllyesEvidence transcript
- Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.
“the number of tokens is actually more important”
Wording checked against the slide or recording
Speaker Gary IllyesEvidence transcript
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-URL-05
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-URL-05
- Extended by D3-C610 Day 3: Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.
- Extended by D3-C611 Day 3: If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within…
Google described its crawling as a large-scale distributed swarm of simple HTTP clients, roughly what one would get by deploying many wget or curl libraries on cloud compute instances.
“a large-scale distributed swarm of simple HTTP clients”
Wording checked against the slide or recording
Speaker Gary IllyesEvidence transcript
To debug crawl issues, Google points site owners to the Crawl Stats report in Search Console, found in the Settings page under Crawl stats.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-MON-04glossary term Crawl Stats report
Search Console's Crawl Stats report shows how much Google crawls from a site, the errors Google received, and the content types fetched by the different crawlers.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-MON-04glossary term Crawl Stats report
Gary Illyes called the Crawl Stats report the most valuable tool for debugging crawl errors, together with server logs, which he called the real thing for those who know how to read them.
Speaker Gary IllyesEvidence transcript
What Google's documentation says 4
Google's Inside Googlebot post (March 2026) says Googlebot currently fetches only the first 2MB of each URL, HTTP headers included (64MB for PDFs); bytes past that cutoff are not fetched, rendered or indexed, and each resource the page loads has its own separate limit.
Publisher Search Central blog (31 March 2026)
Used byrequirement DEV-PRF-04
Google's Inside Googlebot post warns that bloated inline base64 images, large blocks of inline CSS or JavaScript, or megabytes of menus can push a page's text or structured data past Googlebot's 2MB cutoff, and advises moving heavy CSS and JavaScript to external files and placing meta tags, the title, the canonical and essential structured data high in the HTML.
Publisher Search Central blog (31 March 2026)
Used byrequirement DEV-PRF-04
Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.
“Googlebot is just a user of something that resembles a centralized crawling platform”
Publisher Search Central blog (31 March 2026)
Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.
Publisher Google
Used byglossary term Contractual crawlers (special-case crawlers)