Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Thing · Crawler

Googlebot

Google Search's main crawler, which fetches pages for the index as a smartphone or desktop user agent.

Claims
68
In Google’s docs
13
Said at the event
43
Not in docs
5
Kit items
26

Google’s documentation 13

Documented in

DocsSourceD1-C136

Google's Inside Googlebot post (March 2026) says Googlebot currently fetches only the first 2MB of each URL, HTTP headers included (64MB for PDFs); bytes past that cutoff are not fetched, rendered or indexed, and each resource the page loads has its own separate limit.

Search Central blog (31 March 2026) · Day 1 · How crawling works

DocsSourceD1-C137

Google's Inside Googlebot post warns that bloated inline base64 images, large blocks of inline CSS or JavaScript, or megabytes of menus can push a page's text or structured data past Googlebot's 2MB cutoff, and advises moving heavy CSS and JavaScript to external files and placing meta tags, the title, the canonical and essential structured data high in the HTML.

Search Central blog (31 March 2026) · Day 1 · How crawling works

DocsSourceD1-C342

Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.

Search Central blog (31 March 2026) · Day 1 · How crawling works

DocsSourceD1-C138

Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.

Google · Day 1 · How crawling errors affect Search

DocsSourceD1-C141

Google's guide to A/B testing says never to show one set of URLs to Googlebot and a different set to humans: that is cloaking, which is against Google's spam policies whether or not a test is running.

Google Search Central · Day 1 · How Google thinks about crawl budget

DocsSourceD2-C164

Google says Googlebot and its Web Rendering Service identify resources that do not contribute to essential page content, such as reporting and error requests, and may not fetch them, so client-side analytics may not give a full or accurate picture of their activity.

Google Search Central · Day 2 · Lightning session D: Rendering and JavaScript

DocsSourceD2-C561

Google's documentation says Googlebot mostly crawls from US IP addresses and sends no Accept-Language header, so pages that change content or redirect by the visitor's perceived country or language may not have every version crawled, indexed or ranked; it recommends separate URLs annotated with hreflang.

Google Search Central · Day 2 · Focusing on Internationalisation and Localisation

Said at the event 43

Slide and stage claims that name it, the ones Google’s documentation does not cover first.

Not in docs 5

SlideNot in docsD1-C123

Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

Day 1 · How Google interprets robots.txt

StageNot in docsD1-C536

Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.

Dave Smart · Day 1 · Lightning session B: Robots.txt

StageNot in docsD1-C467

To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.

Gary Illyes · Day 1 · Q&A

StageNot in docsD1-C543

Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

Gary Illyes · Day 1 · Q&A

Consistent with docs 22

SlideConsistent with docsD1-C037

For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

Day 1 · How Search works and where's AI?

StageConsistent with docsD1-C322

Googlebot is an ordinary HTTP client with nothing special about it: like a browser, it fetches a URL it was given and returns the fetched bytes to Google's servers.

Gary Illyes · Day 1 · How crawling works

StageConsistent with docsD1-C519

A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C531

In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C443

To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued over time, and whether the average response time shows the server is fast enough or is limiting Googlebot.

Day 1 · Q&A

StageConsistent with docsD1-C469

Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

Gary Illyes · Day 1 · Q&A

StageConsistent with docsD1-C473

Files such as cats.txt have no importance for Google Search: if a site links to one, Googlebot will find and crawl it, but it has no other effect.

Day 1 · Q&A

StageConsistent with docsD1-C494

To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot opens to the site: if it has dropped, the capacity limit changed.

Day 1 · Q&A

StageConsistent with docsD1-C114

A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

Day 1 · Q&A

StageConsistent with docsD1-C546

A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).

Day 1 · Q&A

StageConsistent with docsD2-C374

A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

Day 2 · Handling web duplication

StageConsistent with docsD2-C375

Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

Day 2 · Handling web duplication

Confirmed by docs 8

StageConfirmed by docsD1-C203

A crawler is software that downloads pages, extracts their links and repeats the process on the links it extracted; Googlebot is the main crawler of Google Search.

Cherry Prommawin · Day 1 · How Search works and where's AI?

SlideConfirmed by docsD1-C087

Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.

Day 1 · How Google interprets robots.txt

StageConfirmed by docsD1-C533

Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

Dave Smart · Day 1 · Lightning session B: Robots.txt

StageConfirmed by docsD1-C375

Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

Day 1 · How Google thinks about crawl budget

StageConfirmed by docsD1-C491

If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.

Day 1 · Q&A

Nothing to verify 8

StageD1-C481

An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that site owners can control AI access separately from Googlebot and analyse it in their logs.

From the audience · Day 1 · Q&A

StageD1-C484

A Google panelist called blocking all AI crawlers while allowing search crawlers such as Googlebot, Bingbot and Applebot a personal, philosophical decision that every site owner can make.

Day 1 · Q&A

SlideD2-C001

An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.

From the audience · Day 2 · Welcome to indexing day!

Press and analysis 12

AnalysisD1-C344

On stage Google spoke of probably hundreds, if not thousands, of crawlers, while its Inside Googlebot post speaks of dozens of other clients; both agree that only the larger crawlers are documented, so a Google user agent missing from the public lists is not proof of a fake request, and reverse DNS or Google's published IP ranges are the test.

Ibrahim Anjro · Day 1 · How crawling works

AnalysisD1-C372

Do not answer verified Googlebot with 401, 402 or 403, for example from a login wall, a paywall or a pay-per-crawl setup: Google treats them like 404 and drops the pages, and Google's status code page says not to use 401 or 403 to limit crawling; to slow crawling temporarily, return 429 or 503.

Ibrahim Anjro · Day 1 · How crawling errors affect Search

AnalysisD1-C083

Under the example file, Googlebot may crawl /, /politics/eu-vote and /sports/live/, and is blocked from /?utm_source=x, /index.html, /politics (no trailing slash), /live/ and /sports/live-score. An 'allow: /$' rule does not cover the homepage with tracking parameters.

Ibrahim Anjro · Day 1 · How Google interprets robots.txt

AnalysisD1-C419

A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.

Ibrahim Anjro · Day 1 · How Google thinks about crawl budget

AnalysisD1-C486

The panel spoke of blocking Google's 'AI crawling or training' with Google-Extended, but Google documents Google-Extended as a usage token, not a crawler: disallowing it does not stop Googlebot fetching pages and only controls use for Gemini training and grounding.

Ibrahim Anjro · Day 1 · Q&A

AnalysisD1-C547

Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the total time a server spends holding connections open for Google, counting both the number of parallel connections and their duration. That fits the panel's advice to watch how many connections Googlebot opens (D1-C494): fewer connections is the documented form of a lower capacity limit.

Ibrahim Anjro · Day 1 · Q&A

AnalysisD2-C098

International sites should remove notranslate from their templates unless translation must be prevented: the robots or googlebot form switches off Google Search's translation features, and John Mueller said the form named google also blocks Chrome's translation for visitors who do not read the language.

Ibrahim Anjro · Day 2 · Controlling indexing

AnalysisD2-C378

Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.

Ibrahim Anjro · Day 2 · Handling web duplication

AnalysisD3-C301

For developers, Google's five listed spam types map to concrete checks: serve Googlebot the same content as users (cloaking), avoid near-identical location or keyword pages (doorways), add value to reused content (scraped content), and patch and monitor the CMS against injected pages and links (hacked content).

Ibrahim Anjro · Day 3 · What are quality updates

Built on these claims 26

Kit items about Googlebot: their own words name it, or several of the claims they rest on do.

Developer requirements 22

17 more

Also inglossary terms User-agent group, cats.txt, Crawl rate limit (hostload), Storebot-Google

Connected things 35

Most often named with it

Things named in the same claim, with the number of claims they share.

19 more things

Topics that feature it