Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Not in docs 5
Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.
Day 1 · How Google interprets robots.txt
Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.
Dave Smart · Day 1 · Lightning session B: Robots.txt
To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.
Gary Illyes · Day 1 · Q&A
Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).
Gary Illyes · Day 1 · Q&A
A community speaker compared Googlebot to an inspector who cannot enter the shop and only looks through the window, seeing the page as ones and zeros, which the speaker linked to tokenization.
Day 3 · Lightning session K: Facets of quality
Consistent with docs 22
For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
Day 1 · How Search works and where's AI?
Googlebot is an ordinary HTTP client with nothing special about it: like a browser, it fetches a URL it was given and returns the fetched bytes to Google's servers.
Gary Illyes · Day 1 · How crawling works
Googlebot is the crawler Google uses for web search, including Search's AI features.
Gary Illyes · Day 1 · How crawling works
Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
Gary Illyes · Day 1 · How crawling errors affect Search
CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
Gary Illyes · Day 1 · How crawling errors affect Search
A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.
Day 1 · How Google interprets robots.txt
In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.
Day 1 · How Google interprets robots.txt
To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued over time, and whether the average response time shows the server is fast enough or is limiting Googlebot.
Day 1 · Q&A
Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.
Gary Illyes · Day 1 · Q&A
Files such as cats.txt have no importance for Google Search: if a site links to one, Googlebot will find and crawl it, but it has no other effect.
Day 1 · Q&A
To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot opens to the site: if it has dropped, the capacity limit changed.
Day 1 · Q&A
A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.
Day 1 · Q&A
A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).
Day 1 · Q&A
Googlebot does not support HTTP/3 today.
Day 1 · session not recorded
Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
Day 2 · How is HTML interpreted
Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.
Rebecca Yu · Day 2 · Lightning session D: Rendering and JavaScript
Google named three causes of content missing from the rendered HTML: JavaScript inaccessible to Googlebot, JavaScript DOM event triggers and disabled browser APIs.
Erin Sparling · Day 2 · What is Google friendly JavaScript
When the JavaScript is not accessible to Googlebot, Google cannot render the client-side DOM that the page builds asynchronously, so parts of the page may be present while the main content is absent.
Erin Sparling · Day 2 · What is Google friendly JavaScript
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
Day 2 · Handling web duplication
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
Day 2 · Handling web duplication
Google's long-standing advice to write for people and give them what they want has become true in practice because Googlebot has become more and more human, a community speaker argued.
Day 3 · Lightning session K: Facets of quality
Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.
Alex Jansen · Day 3 · Shopping on Search: Beyond the blue links
Confirmed by docs 8
A crawler is software that downloads pages, extracts their links and repeats the process on the links it extracted; Googlebot is the main crawler of Google Search.
Cherry Prommawin · Day 1 · How Search works and where's AI?
Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.
Day 1 · How Google interprets robots.txt
Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.
Dave Smart · Day 1 · Lightning session B: Robots.txt
The non-standard crawl-delay rule is not processed by Googlebot, so it does not save any crawl budget.
Day 1 · How Google thinks about crawl budget
Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.
Day 1 · How Google thinks about crawl budget
If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.
Day 1 · Q&A
The robots meta tag is written as <meta name="robots" content="rule1,rule2">, or with the name googlebot instead of robots, with several rules separated by commas.
John Mueller · Day 2 · Controlling indexing
Storebot-Google also does deep crawls, for example of checkout pages, which do not belong in the search index; this is another reason it is kept separate from Googlebot.
Alex Jansen · Day 3 · Shopping on Search: Beyond the blue links
Nothing to verify 8
The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
Day 1 · How Google interprets robots.txt
An audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.
From the audience · Day 1 · How Google thinks about crawl budget
An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that site owners can control AI access separately from Googlebot and analyse it in their logs.
From the audience · Day 1 · Q&A
A Google panelist called blocking all AI crawlers while allowing search crawlers such as Googlebot, Bingbot and Applebot a personal, philosophical decision that every site owner can make.
Day 1 · Q&A
An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
From the audience · Day 2 · Welcome to indexing day!
An audience member asked how several sloppy migrations on the same domain can affect Googlebot's crawling.
From the audience · Day 2 · Welcome to indexing day!
In Lightning session K (Facets of quality), a community speaker argued that SEO is for humans: write for people and the problems they have, rather than to please Googlebot.
Day 3 · Lightning session K: Facets of quality
Because Googlebot is not a human, the speaker's team at the time thought more about Googlebot than about human readers, a community speaker said.
Day 3 · Lightning session K: Facets of quality