Area
Crawling
How Google discovers and fetches URLs, and what gets in the way.
Topics in this area 18
33 claims · 8 sessions
The crawl pipeline
URLs wait in a crawl queue, a shared scheduler decides what to fetch and when, and the crawler fetches while enforcing robots.txt and protecting servers. One centralised infrastructure serves Search, Ads, Shopping and Images. Day 2 showed an example of what the crawler hands to processing: a fetch record with the fetch result, connect time and time to first byte in milliseconds, the robots policies that apply (such as a Google-Extended opt-out) and the raw HTTP response; this record was shown on a slide and is not in Google's docs. Links extracted during processing go back to the crawl queue, and the renderer fetches JavaScript and CSS through the same crawler. Day 3 named a second crawler on that shared infrastructure: Google Shopping uses Storebot-Google instead of Googlebot because it needs fresh prices, availability and shipping details and crawls more often; it validates merchant feeds by double-checking prices on the site and crawls product, cart and checkout pages, as Google's Merchant Center help confirms, and its robots.txt group governs all Shopping surfaces. The second recording of Day 1 added the crawling talk itself. Cherry Prommawin defined a crawler as software that downloads pages, extracts their links and repeats the process, and opened with how the internet works, from TCP/IP to DNS and HTTP. Gary Illyes said Googlebot is an ordinary HTTP client with nothing special about it, and that Google runs probably hundreds, if not thousands, of crawlers but does not let teams build their own: all share one infrastructure and obey Google's internal crawling policies (Google's Inside Googlebot post speaks of dozens of other clients); he described it as a large distributed swarm of simple HTTP clients, roughly many wget or curl instances. The scheduler hands the crawler an ordered list that it works through from top to bottom, and teams prioritise differently: web search cares a lot about site and content quality, while Ads checks every publisher page as it comes in (said at the event). URLs extracted during indexing go back to the scheduler, and the crawl budget talk added that a request from any Google user agent is routed through the shared platform. Author’s view: a Google user agent missing from the public lists is not proof of a fake request; verify with reverse DNS or Google's IP ranges.
Day 1Day 2Day 3
33 claims · 9 sessions
URL discovery
Google finds new URLs by following links from pages it already knows. Author’s view: pages without internal links depend on sitemaps alone and are found slowly. Day 2 added that Google uses extracted links for discovery, for understanding a site's structure and for ranking, and extracts them from the HTML before rendering and again from the rendered HTML, sending new URLs back to the crawl queue. For large sites, a Google Q&A slide said good HTML sitemaps are hard to do well at that scale (said at the event) and advised relying instead on category hub pages that link to important pages, the kind of hub page Google's guide to how Search works also describes; Google's ecommerce guide recommends the same menu-to-category-to-product linking, with a sitemap or Merchant Center feed where not every product can be linked. Google said it sometimes picks up URLs written as plain text, but the remark was unclear in the recording and is not documented. Day 3 added scale and timing: Google said it knows hundreds of trillions of URLs (as of October 2026) but does not crawl them all, and estimated that discovering a new URL takes about 20 hours on average, from seconds to weeks or never (not in Google's docs). Author’s view: since discovery and recrawling take far longer than indexing, internal links from often-crawled pages and accurate sitemaps are what speed results up. The second recording of Day 1 added that there are trillions of URLs on the internet or more, that even Google cannot tell how many, and that some may never be discovered; Google finds what to crawl mainly by extracting URLs from crawled pages and passing them back to the scheduler, and additionally from sitemaps, and may visit hub pages such as the homepage and category pages more often because they link to new or updated pages. Search Console's crawl data separates discovery fetches of new URLs from refreshes of known ones. Day 2's opening Q&A repeated that the effort an HTML sitemap takes on a large site is better spent on hub pages such as category pages.
Day 1Day 2Day 3
53 claims · 7 sessions
Crawl errors
Google slows crawling when it sees 5xx or 429 responses, network timeouts, connection resets or DNS errors, and drops indexed URLs that stay unreachable. It cannot tell that a firewall or CDN rule is the cause, so a misconfigured block looks like a failing server. Other 4xx codes do not change the crawl rate. Day 2 added a second failure mode for bot protection: when a CDN or other bot protection shows Googlebot the same challenge page with a 200 status on many URLs, Google struggles to recognise it as an error and can cluster those pages as duplicates; Google's CDN post says recovery can be slow and recommends a 503 status for bot-verification interstitials. Google also said AI agents browsing for users hit the same bot walls and may go to another site (said at the event, not in Google's docs). Day 3 timed the effect: when a site starts serving 500 errors, Google lowers its crawl capacity within about four hours on average, and recovery takes weeks; Google's crawl rate guide says crawling picks up again automatically once the errors drop and warns against using error codes to slow crawling for longer than 1-2 days. The second recording of Day 1 filled in the status codes. Cherry Prommawin walked through the five classes: a 200 makes the returned resource eligible for indexing (although, as Gary Illyes put it, a 200 only means the server believes it did what was asked), a 204 has nothing to index, Google follows permanent and temporary redirects and indexes what the target returns, 404, 410 and 403 leave nothing to index, and 429 and 5xx make Google slow down so as not to break the site. Gary Illyes called DNS and network errors very common: DNS errors get a site removed from Search very aggressively, network errors and timeouts can also remove it from Search and with it from every feature that depends on Search, and most happen between the origin server and Google's data centers, often at a firewall, CDN, host or DNS provider, where neither side can see them, so the first step is to ask the host or CDN what changed; a rise in HTTP errors can also come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path (Google's CDN post lists both codes as CDN blocks). He added that Google now sees more 403 responses (described on stage as 'authentication required') and, more recently still, more 402 Payment Required responses, and treats both like a 404, and that CDN captcha challenges served with a 200 end up classified as soft 404s; Day 2 said such pages can instead be clustered as duplicates, and Google's CDN post describes both outcomes. Google pointed to the Crawl Stats report (Settings, Crawl stats) and server logs as the main tools for debugging crawl errors, and to the reasons in the Page indexing report for spotting patterns, which is also where soft 404s are looked up. In Lightning session B Dave Smart showed a quieter failure: robots.txt is checked for every URL in a redirect chain, so a redirect through a disallowed URL, an external authorisation service or a step served only to Googlebot stops the crawl and the page is reported as blocked (not in Google's docs). Author’s view: do not answer verified Googlebot with 401, 402 or 403 from a login wall or paywall; and while 4xx responses do not slow crawling, a 404 is still a fetch that counts towards crawl budget.
Day 1Day 2Day 3
40 claims · 10 sessions
Soft 404s
A soft 404 returns a success code while its main content looks like an error or an empty page; Google keeps it out of the index but keeps crawling it, which wastes crawl budget. On Day 2 Google explained that it detects soft 404s with a BERT-like language model trained on page structure and layout, so an error message alone in the main content makes a soft 404 while an error in the navigation need not, and simple keyword matching would not work (said at the event, not in Google's docs). Google listed four common causes that its troubleshooting guide also names (pages that look like errors, thin or empty content, server or CMS misconfiguration, JavaScript content that fails to load) and a fifth said only at the event, mistakes in Google's own systems, which it asked site owners to report in its forums. Single-page apps are a frequent source because the server returns 200 for every route, and Google documents a real 404, a JavaScript redirect to a URL that returns 404, or a noindex added with JavaScript as fixes. Soft 404s are also clustered together as duplicates, and index selection drops any that earlier steps missed. On Day 3 an audience member reported that on one large client site, pages clustered as soft 404s were still recrawled, but only every 160 to 190 days (one site's observation, not Google data). The second recording of Day 1 added the full definitions: Cherry Prommawin called a soft 404 a 404 in disguise, a 200 page whose content says 'not found', information the site should have sent as the status, and said soft 404s are still crawled like normal pages, with their bigger consequences in indexing; Gary Illyes called them one of the biggest problems on the internet right now for crawling and showing up in Search, and said a 200 only means the server believes it did what was asked. CDN captcha pages served with a 200 end up classified as soft 404s (Day 2 added that they can also be clustered as duplicates; Google's CDN post describes both outcomes), and Google's status code page says a 204 or an empty 2xx page also shows as a soft 404. Day 2's second recording confirmed that a client-rendered page that renders empty for Google is seen as thin content before it is treated as a soft 404, and Google's site move guide warns that redirecting many old URLs to an irrelevant page such as the home page might be treated as a soft 404. An audio recording of Day 1 added that soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.
Day 1Day 2Day 3
7 claims · 1 session
HTTP versions and Googlebot
Google's crawlers use HTTP/1.1 and HTTP/2, not HTTP/3. Author’s view: serving HTTP/3 to browsers is fine as long as the older versions keep working. Day 1's second recording added the basics behind it: a URL states how a resource is requested (HTTP or HTTPS), where (the host) and what (the path), and DNS, the internet's address book, is consulted for each request to find the host's IP address.
Day 1
96 claims · 13 sessions
robots.txt rules
A crawler obeys only the most specific group that names it; within the group the longest matching path wins, and ties go to the less restrictive rule. An audio recording of Google's Day 1 robots.txt talk added the history and the basics. Robots.txt began in 1994, when Martijn Koster proposed a text file of access rules because bots were crashing servers; Google has supported it since it started crawling in 1996, and it is now an IETF standard, RFC 9309, the Robots Exclusion Protocol, which Google follows because anyone who wants to opt out of crawling should be able to. Google stressed that the protocol only controls which automated clients may access what, not how the content is used, and that it is not a security measure: the file always sits at the root of the host, where anyone can read the paths it disallows, so a secret folder needs authentication. Everything is implicitly allowed, so an allow rule only re-opens a path inside a disallowed one; * matches any characters and $ ends the match, so user-agent: *, disallow: / and allow: /$ let unnamed crawlers fetch only the homepage. Comments start with #, and the talk pointed to the ASCII-art robots.txt of thebestfriedchickenever.com to show why they are useful. Google called the format extremely forgiving: lines a parser cannot read are skipped, a typo in a path blocks only the wrong path and a misspelt rule name makes Google ignore that line (not in Google's docs). Search Console's robots.txt report shows the file as Google last fetched it, with a version history and the errors and successes of each fetch; Google said the report uses its open-source parser and helps catch CDNs that change robots.txt without the owner knowing and hosts that cloak the file, which happens more often than people think (not in Google's docs). In Lightning session B Dave Smart added that robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one: a page that redirected through a disallowed /cart/ URL to set the local currency was reported as blocked, and Search Console names only the first URL of the chain, which is not itself disallowed; the same happens with an external authorisation service blocked by its own robots.txt, content that moved through several URLs and redirects served only to Googlebot (the per-URL check is not in Google's docs). Google supports only user-agent, allow, disallow and sitemap, and Day 2 confirmed that a sitemap line can be picked up by any crawler, while a sitemap not listed there has to be submitted, for example in Search Console. A disallow is not a noindex: Google does not index a disallowed page's content, but its URL can still appear in results. Google said it will not treat a disallow as noindex because some very important sites block their most important pages by accident (said at the event, not in Google's docs). Day 2 also showed that disallowing JavaScript or API endpoints a page needs for rendering leaves its content missing, and that rules apply per host, so an API or CDN on its own hostname has its own robots.txt. Day 3 added timing and Shopping: Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for and its robots.txt guide confirms, though delays happen, and that requesting a recrawl in Search Console's robots.txt report refreshes it sooner. Rules addressed to Storebot-Google affect all Google Shopping surfaces, such as the Shopping tab. The second recordings added Google's own stance. Gary Illyes said all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling, except contractual crawlers that crawl by agreement; on Day 2 Google repeated that site owners should be able to opt out, for legal reasons or for crawl budget, and in the Day 1 Q&A said it follows robots.txt partly in its own interest, since crawling a blocked infinite URL space would waste its time too. Mainstream crawlers from search engines and AI companies try to follow robots.txt, so correct rules are the control; user-initiated fetchers and agents acting for a user generally do not check it, and a panelist noted that much of robots.txt is written for search engines, which should never add items to a cart, while an agent probably should. A panelist also mentioned crawler best practices in progress that would exempt research, malware-scanning and privacy crawlers, apparently because they sometimes need to ignore robots.txt or probe URLs other crawlers would not touch. Day 2's opening Q&A said listing a sitemap in robots.txt is fine but makes the sitemap public, and explained why a disallow is not a noindex with a national tax authority that blocks very important PDFs: Google can show their URLs but cannot index their content; very few disallowed URLs are in the index, and John Mueller added that robots meta rules only work if robots.txt lets Google fetch the page. For media, rules for Googlebot-Image and Googlebot-Video control image and video indexing, and a video blocked by robots.txt is not shown by its bare URL at all (said at the event, not in Google's docs). Author’s view: the best practices mentioned match the public IETF draft 'Crawler best practices', not Google documentation; and Google's docs name links from elsewhere as the reason a disallowed URL can still be indexed, so keep an important page crawlable with noindex if it must stay out of Search. Author’s view: Google's open-source parser in fact accepts common misspellings of disallow and user-agent, though not of allow, and other crawlers may be stricter, so spell rule names correctly.
Day 1Day 2Day 3
9 claims · 2 sessions
robots.txt group parsing differences
A user-agent line with its allow and disallow rules forms a group, and one group can name several crawlers; Google's Day 1 robots.txt talk advised giving a named crawler its own group only when it needs rules the * group should not grant to every crawler. Groups are not additive: a crawler that a named group matches follows only that group, so a Googlebot group that blocks /dogs/ leaves Googlebot free to crawl what the * group blocks, as Dave Smart showed in Lightning session B and Google's specification confirms. The talk's quiz made the same point: to let Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed, the answer was a Googlebot group with both the disallow and the allow, and Google's specification combines all groups that name the same user agent into one. Crawlers disagree on how to treat an unknown line between user-agent lines: Google said that while writing RFC 9309 it asked people whether the first crawler should inherit the rules that follow and opinions split roughly 50-50 (not in Google's docs). Googlebot ignores the line and merges the groups, as Google's robots.txt specification documents, while a Google slide showed Bing ending the group there instead, so the same file can block one crawler and not another. Author’s view: Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files, so put any non-standard line after a group's rules and test the file with each search engine's tools.
Day 1
51 claims · 8 sessions
Crawl budget
Crawl budget is crawl rate limit combined with crawl demand. Google's crawl budget guide is written for sites with over 1 million unique pages changing weekly, over 10,000 pages changing daily, or many URLs reported as 'Discovered – currently not indexed'. On Day 2 Google called that status a crawl-scheduling state in which it knows a URL but does not want to crawl it yet (said at the event), while the Page indexing help explains it by expected server overload. A Google slide also questioned whether a new URL that fits a known duplicate URL pattern even needs to be crawled (not in Google's docs). Author’s view: rule out slow responses and server errors first, then treat slow crawling of a smaller site as a demand problem, which means quality. Day 3 gave the scale: Google said it knows hundreds of trillions of URLs (as of October 2026) and does not crawl them all, because a much smaller crawl space is enough, and its crawl chart said discovering a new URL can take weeks or never happen. The second recording of Day 1 added the crawl budget talk and the Q&A. Crawl budget is the finite resources Google allocates to crawling one site, which decide how many of its pages are discovered and how often they are revisited; Google crawls every URL that differs even when it leads to the same content, so infinite spaces such as calendars burn it, and the content worth improving or removing includes low-quality, spam, duplicate and soft error pages. Not every site needs to worry, and never did: a site under a few thousand URLs likely has no crawl budget problem, news sites are crawled aggressively anyway, and a higher crawl rate does not make a URL rank better. Every 2xx fetch consumes crawl budget, while 4xx responses were said not to affect it. To check for a problem Google pointed to the Crawl Stats report: whether total crawling has plateaued, whether the average response time limits Googlebot, the file-type breakdown for anomalies and the discovery versus refresh split. Log files alone cannot show inefficiency, because they do not reveal where the crawl limit lies, but they help on sites with filter parameters; many complaints turn out to come from a plugin, such as a WordPress calendar plugin that adds calendar parameters to the URLs of every page, and Google learns that such URLs are useless only from large samples; a panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with an Apache rule, which saves crawl budget for the URLs that matter (the exact mechanism is not clear in the recording). For news sites Gary Illyes advised measuring the time from publishing a URL to Googlebot's first crawl and investigating only a rising trend, adding that for fresh stories something like two hours is probably not great, while John Mueller said that if Crawl Stats shows about half of crawling going to discovering new pages, crawling of new pages is not the problem; Gary Illyes said disallowing a section such as /ads shifts crawl budget to the rest of the site. Author’s view: Google's guide says freed crawl budget shifts only on sites that already hit their capacity limit, a 404 is still a fetch that counts towards crawl budget, and a site between a few thousand URLs and Google's million-page threshold should check Crawl Stats before blaming crawl budget; for a consolidation of about 2,000 URLs the crawl-budget gain is likely minor.
Day 1Day 2Day 3
29 claims · 5 sessions
Crawl rate limit (hostload)
Hostload, the crawl capacity limit, is a host-wide limit shared by all Google crawlers, so heavy crawling by one Google product can leave less capacity for the others. It falls when connect time or time to first byte rise, or when the server returns 429, 5xx or network errors. On Day 2 a Google slide showed an example fetch record, as passed on to processing, that includes connect time and time to first byte in milliseconds (not in Google's docs). Day 3 gave timings: when a site starts serving 500 errors, Google lowers its crawl capacity within about four hours on average, while increases take one to three weeks because Google first needs to know that higher demand will last, and a continuously running process recalculates capacity within about a month (the spoken and slide figures differ slightly and are not in Google's docs). Google's crawl rate guide confirms that many 500, 503 or 429 responses reduce the crawl rate, which recovers automatically once the errors drop. The second recording of Day 1 added that crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle, and that it applies per host, not necessarily per site: a site's CDN, app, subdomains and www host may be different hosts. Cherry Prommawin said 429 is the one 4xx code that tells Google to slow down, and that a 5xx makes Google slow down so as not to break the site; a rise in errors can come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path. In the Q&A Google said that when higher demand brings more crawling than a server can cope with, Google reduces it again, and, asked how a site of about 100 million pages can tell a crawl capacity problem from a crawl demand problem, that a capacity drop is most of the time an abrupt step down, illustrated as hostload falling from 10 to 5, which a panelist said means up to five requests per second, and that a fall in the number of connections Googlebot opens shows the capacity limit changed. Google added that demand from Search, Google Ads or another Google product raises crawling only as far as the site's crawl capacity allows. Author’s view: Google's crawl budget guide does not express the capacity limit in requests per second but as the total time a server spends holding connections open for Google, counting parallel connections and their duration, so fewer connections is the documented form of a lower limit. Author’s view: to slow crawling for a short time return 429 or 503, never 401, 402 or 403, which Google treats like 404.
Day 1Day 2Day 3
34 claims · 6 sessions
Crawl demand
Crawl demand is how much Google wants a site's URLs, and each Google crawler has its own. It rises with site quality, how often URLs change and how popular they are on the web. On Day 2 Google said site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing that its content is useful to users. Google's robots meta tag specification adds that Googlebot considerably lowers the crawl rate of a URL after the date in its unavailable_after rule. Day 3 put demand in time: Google estimated that recrawling a known URL takes about 30 days on average, from seconds (for news sites, a likely reading of the recording) to never, since URLs not recrawled for a very long time are dropped from the index, and that a crawl demand update driven by Search takes about 20 hours (estimates not in Google's docs). When the web gets excited about a few URLs, Google raises crawl demand for the whole site, and other Google products such as Shopping create their own demand; increases in crawl capacity take one to three weeks, longer than decreases. An audience member's large client site showed recrawl frequency following site hierarchy, demand and internal linking, with pages clustered as soft 404s recrawled only every 160 to 190 days. The second recording of Day 1 added how the scheduler sets demand: it logs how often each page changes and crawls frequently changing pages first (a news homepage before its terms of service page), very likely deprioritises URLs or sites known to be historically spammy, and visits hub pages such as the homepage and category pages more often because they link to new or updated pages. The crawl budget talk said the quality that drives demand is that of the site as a whole, and that frequent crawling is a consequence of interest and quality, not a cause of better rankings. In the Q&A Google said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, it raises the site's crawl demand and crawls up to it as far as crawl capacity allows, which Day 3 repeated with Shopping as the example. It also said very large, constantly changing sites get no special handling: frequently changing URLs, or URLs useful to users, raise demand and Google tries to raise its capacity, the same logic as for small sites; news sites are crawled aggressively anyway; and, asked how a very large site can tell the two apart, that a drop in crawling is hard to attribute to lower demand or a lower capacity limit, although a capacity drop is most of the time an abrupt step down. Author’s view: read site quality as the baseline, with known URL-level signals still counting, as the same talk's slide and Google's crawl budget guide suggest.
Day 1Day 2Day 3
5 claims · 2 sessions
Crawl demand inherited by path
When Google does not know a URL's quality or popularity, it uses the aggregate of its parent path, then that path's parent. This was shown on a slide and is not in Google's public documentation. Day 2 extended the idea to indexing: Google said index selection is more forgiving with new URLs from a site, or even a section, that already satisfies users well (said at the event, not in Google's docs). Author’s view: new content therefore starts with the reputation of the folder it sits in, for crawling and for indexing. Day 1's second recording added a stage remark that the quality driving crawl demand is that of the site as a whole. Author’s view: read it beside the slide's fallback to the parent path, which applies only when a URL's own quality is unknown: site quality sets the baseline while known URL-level signals still count.
Day 1Day 2
13 claims · 4 sessions
Faceted navigation
Filter and sort URLs multiply into near-infinite combinations and are a classic crawl budget leak. Google prefers blocking them in robots.txt and keeping only item pages and one unfiltered listing crawlable. Day 2 added that non-crawlable link markup, such as onclick handlers and hash pseudo-links, is common in single-page apps, so sites with faceted navigation should check their links for it. Author’s view: hash fragments can deliberately keep filter states out of the crawl, as Google's faceted navigation guide allows, as long as product, category and language links use real URLs. Day 1's second recording added that Google crawls every URL that differs even when it leads to the same content, so infinite URL spaces burn crawl budget; a community speaker advised checking that every parameter in a request is used and in the expected order and redirecting otherwise, and in the Q&A Google suggested working out the normalised URL and redirecting parameter variants to it, which costs crawl budget at first but leaves a clean slate, said log files are worth checking on sites with filter parameters, and noted that many crawling complaints trace back to plugins that generate URLs, such as calendars, whose useless parameter URLs can be handled at the web server.
Day 1Day 2
5 claims · 3 sessions
HTTP caching for crawlers
Supporting conditional requests lets Google skip re-downloading unchanged pages. Answer with 304 Not Modified, and prefer ETag. Day 1's second recording added that Gary Illyes named Last-Modified and ETag as the response headers that matter for caching, and that Cherry Prommawin described a 304 as telling Google's crawler that the content has not changed since its last visit.
Day 1Day 2
32 claims · 5 sessions
Directives and crawl budget
Disallowed URLs are not fetched, so they cost no crawl budget. Pages with noindex must be fetched, so they do; nofollow is a hint, not a block, so the linked URLs can still be crawled, and crawl-delay is ignored by Googlebot. Day 2 stressed that robots.txt does not control indexing: a disallowed URL can still appear in results without its content, and a noindex on a blocked URL is never seen. John Mueller advised against the page-level nofollow robots rule, which stops signals passing through every link on the page, in favour of rel=nofollow on individual links. He also said AI crawlers do not really know what to do with nofollow, without saying whose crawlers he meant (not in Google's docs). A second recording of Day 1 explained the slide on nofollow: Google does not crawl through the nofollow link itself but still crawls the linked page when it finds it through other links, as Google's 2017 crawl budget post also says. A second recording of Day 2 confirmed John Mueller's remark that if robots meta tags were reinvented today the page-level nofollow rule would probably not be part of them. Day 2's opening Q&A illustrated why a disallow is not a noindex with a national tax authority that blocks very important PDFs, whose URLs Google can show without their content, and John Mueller said robots meta rules only work if robots.txt allows the fetch. Disallowing an image or video file keeps it out of Google, and a disallowed video is not shown by its bare URL. In the Day 1 Q&A Gary Illyes said disallowing a section such as /ads shifts crawl budget to the rest of the site. Author’s view: Google's guide says freed crawl budget shifts only on sites already at their capacity limit, so block only sections you never want crawled.
Day 1Day 2
10 claims · 2 sessions
Pagination and "load more"
Google's crawlers do not click buttons, so each page in a series needs its own URL and a normal <a href> link to the next. Google's pagination guide says later pages should not use page 1 as their canonical. On Day 2 Google said on stage that pointing paginated pages' canonical to page 1 can sometimes make sense, for example to make a category page more visible, but that it affects canonicalization and deduplication; this contradicts Day 1 and is not in Google's docs, and a 2013 Google post says the content of the later pages would then not be indexed at all. For infinite scroll, Google has said it renders with a very tall viewport that can miss content, and its lazy-loading guide asks for paginated loading with a unique URL per chunk. Moving to infinite scroll or a load-more button was called risky, depending on what you want to achieve, because Googlebot does not click buttons. In the Day 1 Q&A, captured by a second recording, Google answered a question about replacing thousands of paginated category pages with a JavaScript load-more button: without crawlable links to the further pages, Google will not see them at all.
Day 1Day 2
38 claims · 6 sessions
Crawlable links
Google extracts links from a page's HTML ('We like links.') to discover pages, understand site structure and rank, sending new URLs back to the crawl queue, which extends Day 1's point that discovery runs on links. Only an a element whose href holds a real URL is dependable: Google's slide put onclick-only links (Googlebot does not click, as on Day 1), routerLink without href, href on a span and javascript: URLs under 'can not extract', while its documentation calls them not recommended and says Google may still try to parse them. Links are extracted before and after rendering, so JavaScript-added a href links can count, but a community speaker showed navigation that exists only after rendering being lost when rendering fails. Day 2 repeats Day 1 on fragments: Google cannot request the part of a URL after #, which Google called the next most common JavaScript issue, so single-page apps should use real paths with the History API, a change it called not free but relatively straightforward. A community speaker showed a language selector built as a button that orphaned a whole set of language versions; Google said it sometimes extracts plain-text URLs too, but that passage of the recording is unclear. Day 3 timed link processing: Google estimated that link annotations are processed in minutes to 1-3 weeks, with an end point of about a year (not in Google's docs), and an audience member's large site showed recrawl frequency following site hierarchy and internal linking. In the Day 1 Q&A Google said that if paginated pages are replaced by a load-more button or infinite scroll without crawlable links, Google will not see the further pages at all; on Day 2 migration speakers stressed updating internal links so moved pages do not rely on redirects alone.
Day 1Day 2Day 3
34 claims · 8 sessions
Sitemaps
A sitemap line in robots.txt lets any crawler find the sitemap; a sitemap not listed there has to be submitted, for example in Search Console. For HTML sitemaps on large sites, a Google Q&A slide said good HTML sitemaps are hard to build at that scale (said at the event) and advised relying instead on category hub pages that link to the important pages; Google's ecommerce guide recommends the same linking, with a sitemap or Merchant Center feed where not every product can be linked. As a canonical signal, listing only the preferred URL in sitemaps was named on stage among the signals that make a big difference, but Google's canonical guide rates sitemap inclusion as weak next to redirects and rel=canonical. Image and video sitemaps help Google find media, and an XML sitemap is one of three equivalent ways to declare hreflang. Author’s view: sitemaps support discovery but do not replace internal links, because pages known only from sitemaps are found slowly. Day 3 timed sitemaps and tied them to quality: Google estimated that processing a sitemap takes about 24 hours on average and that it tries to refetch a useful sitemap of a high-quality site within 14 days at most (neither figure is in Google's docs), while it may never fetch a lower-quality site's sitemap again. The second recordings added Google's framing: Gary Illyes said Google finds what to crawl mainly by extracting URLs from crawled pages and additionally from sitemaps, an XML format from around 2005 that is 'nothing fancy' but still used, and in the Day 1 Q&A Google said there is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that Google can find each of them; for a migration, listing the new URLs in a sitemap cannot hurt but is not the main tool. Day 2's opening Q&A said a site that wants an HTML sitemap can make one but the effort is better spent on hub pages, and that listing a sitemap in robots.txt is fine but makes it public. For migrations, Google's site move guide and a community case study both submit the new sitemap. Gary Illyes said video sitemaps are not critical but good to have, because Google ingests them much more often than it can process HTML pages and they tell it which URLs carry videos (said at the event).
Day 1Day 2Day 3
10 claims · 1 session
Similar URLs and URL normalisation
On Day 1 community speaker Tobias Schwarz treated similar URLs, technically different URLs that differ in only a few characters, as an indicator of duplicate content, broken links, unnecessary redirects and other technical issues. He showed five kinds on well-known brand sites: a different protocol or host (HTTP vs HTTPS, www vs non-www), capitalisation, the number of delimiters such as slashes, a space encoded as %20 in one URL and + in another, and parameters in a different order. To find them he normalises a URL list twice, first without changing the URLs' meaning and then into a deliberately uniform form that groups similar URLs together. Because anyone can link to a variant, he advised that the application compute the expected URL for every request and redirect, or return an error, when the requested URL differs, including checks on unused or reordered parameters and on numeric IDs written with leading zeros. In the Day 1 Q&A Google likewise suggested working out the normalised URL and redirecting parameter variants to it, which costs crawl budget at first but leaves a clean slate, and a panelist said Tobias Schwarz's slides had pleased the engineer side of his brain. Author’s view: of the two answers, Google's canonicalization guide favours the redirect, a strong signal that its target should be canonical, while an error page throws away the links pointing at the variant.
Day 1
Across days 107
- Stage D2-C393 Day 2 · Handling web duplication
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
contradictsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C017 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
extendsD1-C079 Day 1 · How Google interprets robots.txtThe robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
- D2-C017 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
extendsDocs D1-C084 Day 1 · How Google interprets robots.txtGoogle supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
- D2-C020 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- D2-C021 Day 2 · Welcome to indexing day!
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
extendsD1-C105 Day 1 · How Google thinks about crawl budgetURLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
- Docs D2-C022 Day 2 · Welcome to indexing day!
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
extendsD1-C106 Day 1 · How Google thinks about crawl budgetThe noindex rule consumes crawl budget, because Google must fetch the page to see it.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C092 Day 1 · How Google thinks about crawl budgetHostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
- D2-C025 Day 2 · How is HTML interpreted
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- D2-C026 Day 2 · How is HTML interpreted
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C038 Day 2 · How is HTML interpreted
Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- D2-C048 Day 2 · How is HTML interpreted
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C058 Day 2 · Controlling indexing
John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Analysis D2-C067 Day 2 · Controlling indexing
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
extendsDocs D1-C127 Day 1 · How Google thinks about crawl budgetGoogle treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.
- Analysis D2-C067 Day 2 · Controlling indexing
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
extends - Docs D2-C069 Day 2 · Controlling indexing
Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.
extendsD1-C106 Day 1 · How Google thinks about crawl budgetThe noindex rule consumes crawl budget, because Google must fetch the page to see it.
- Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
extends - D2-C127 Day 2 · Lightning session D: Rendering and JavaScript
A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C167 Day 2 · Lightning session D: Rendering and JavaScript
The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C168 Day 2 · Lightning session D: Rendering and JavaScript
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Docs D2-C169 Day 2 · Lightning session D: Rendering and JavaScript
Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- D2-C184 Day 2 · Lightning session D: Rendering and JavaScript
Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C187 Day 2 · Lightning session D: Rendering and JavaScript
A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Analysis D2-C188 Day 2 · Lightning session D: Rendering and JavaScript
Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.
extendsDocs D1-C101 Day 1 · How Google thinks about crawl budgetGoogle's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.
- Analysis D2-C189 Day 2 · Lightning session D: Rendering and JavaScript
Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.
extendsAnalysis D1-C067 Day 1 · How crawling worksA page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
- Stage D2-C190 Day 2 · Lightning session D: Rendering and JavaScript
In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C195 Day 2 · Lightning session D: Rendering and JavaScript
Content that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C197 Day 2 · Lightning session D: Rendering and JavaScript
Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.
extends - D2-C201 Day 2 · Lightning session D: Rendering and JavaScript
Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () => fetch('/api/specs'), does not exist for Google until someone clicks.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C204 Day 2 · Lightning session D: Rendering and JavaScript
Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- Analysis D2-C208 Day 2 · Lightning session D: Rendering and JavaScript
A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.
extendsDocs D1-C081 Day 1 · How Google interprets robots.txtWhen matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.
- Analysis D2-C210 Day 2 · Lightning session D: Rendering and JavaScript
An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.
extendsDocs D1-C074 Day 1 · How crawling errors affect SearchIf robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.
- Stage D2-C213 Day 2 · Lightning session D: Rendering and JavaScript
A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C268 Day 2 · What is Google friendly JavaScript
Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Stage D2-C271 Day 2 · What is Google friendly JavaScript
Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.
extends - Stage D2-C271 Day 2 · What is Google friendly JavaScript
Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.
extends - Docs D2-C276 Day 2 · What is Google friendly JavaScript
To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each chunk its own persistent, unique URL, link sequentially to those URLs, and update the displayed URL with the History API when a new chunk becomes the main visible element.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Stage D2-C287 Day 2 · What is Google friendly JavaScript
A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Stage D2-C291 Day 2 · What is Google friendly JavaScript
In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C296 Day 2 · What is Google friendly JavaScript
If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
extendsStage D1-C070 Day 1 · How crawling errors affect SearchSoft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.
- D2-C336 Day 2 · Understanding what's on a page
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C338 Day 2 · Understanding what's on a page
Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.
extendsD1-C042 Day 1 · How Search works and where's AI?BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.
- Stage D2-C339 Day 2 · Understanding what's on a page
For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C367 Day 2 · Handling web duplication
Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C369 Day 2 · Handling web duplication
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
extendsD1-C094 Day 1 · How Google thinks about crawl budgetIf the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
- D2-C369 Day 2 · Handling web duplication
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsStage D1-C069 Day 1 · How crawling errors affect SearchDNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsStage D1-C367 Day 1 · How crawling errors affect SearchCDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
- Stage D2-C375 Day 2 · Handling web duplication
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
extendsStage D1-C367 Day 1 · How crawling errors affect SearchCDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
- Docs D2-C376 Day 2 · Handling web duplication
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C377 Day 2 · Handling web duplication
AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.
extends - Docs D2-C395 Day 2 · Handling web duplication
Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Stage D2-C427 Day 2 · Lightning session E: Managing Duplicates and Site Moves
When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.
extendsAnalysis D1-C067 Day 1 · How crawling worksA page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
- D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C685 Day 2 · Deciding what goes in the index?
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
extendsD1-C094 Day 1 · How Google thinks about crawl budgetIf the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
- Analysis D2-C686 Day 2 · Deciding what goes in the index?
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
extendsAnalysis D1-C095 Day 1 · How Google thinks about crawl budgetNew content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.
- Stage D2-C695 Day 2 · Deciding what goes in the index?
Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D2-C697 Day 2 · Deciding what goes in the index?
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
extendsD1-C106 Day 1 · How Google thinks about crawl budgetThe noindex rule consumes crawl budget, because Google must fetch the page to see it.
- Stage D2-C700 Day 2 · Deciding what goes in the index?
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C706 Day 2 · Deciding what goes in the index?
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
extendsD1-C065 Day 1 · How crawling worksThe scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.
- Stage D2-C706 Day 2 · Deciding what goes in the index?
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
extendsDocs D1-C097 Day 1 · How Google thinks about crawl budgetGoogle's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.
- Analysis D2-C708 Day 2 · Deciding what goes in the index?
The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Stage D2-C709 Day 2 · Deciding what goes in the index?
Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D2-C714 Day 2 · Deciding what goes in the index?
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Analysis D2-C716 Day 2 · Deciding what goes in the index?
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Stage D2-C717 Day 2 · Deciding what goes in the index?
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
extendsStage D1-C368 Day 1 · How crawling errors affect SearchSearch Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
- D2-C820 Day 2 · Welcome to indexing day!
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- D2-C820 Day 2 · Welcome to indexing day!
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
extendsStage D1-C202 Day 1 · How Search works and where's AI?Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.
- Stage D2-C842 Day 2 · Welcome to indexing day!
Google said listing the sitemap in robots.txt is fine, as many websites do.
extendsDocs D1-C084 Day 1 · How Google interprets robots.txtGoogle supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
- Stage D2-C846 Day 2 · Welcome to indexing day!
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Analysis D2-C849 Day 2 · Welcome to indexing day!
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D2-C884 Day 2 · Lightning session E: Managing Duplicates and Site Moves
The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.
extendsStage D1-C540 Day 1 · Q&AAsked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.
- Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.
extendsD1-C103 Day 1 · How Google thinks about crawl budgetFour ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.
- Stage D3-C385 Day 3 · Shopping on Search: Beyond the blue links
Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.
extendsD1-C065 Day 1 · How crawling worksThe scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.
- Stage D3-C602 Day 3 · How long does it take to..?
Google said it knows hundreds of trillions of URLs (as of October 2026).
extendsStage D1-C201 Day 1 · How Search works and where's AI?Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.
- Stage D3-C606 Day 3 · How long does it take to..?
For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).
extends - D3-C610 Day 3 · How long does it take to..?
Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.
extendsStage D1-C337 Day 1 · How crawling worksAn XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.
- D3-C611 Day 3 · How long does it take to..?
If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.
extendsStage D1-C337 Day 1 · How crawling worksAn XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.
- D3-C612 Day 3 · How long does it take to..?
Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.
extendsStage D1-C331 Day 1 · How crawling worksGoogle's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.
- D3-C613 Day 3 · How long does it take to..?
Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.
extendsDocs D1-C085 Day 1 · How Google interprets robots.txtGoogle generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
- Stage D3-C614 Day 3 · How long does it take to..?
Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.
extendsDocs D1-C085 Day 1 · How Google interprets robots.txtGoogle generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
- Stage D3-C615 Day 3 · How long does it take to..?
A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.
extendsDocs D1-C140 Day 1 · How Google interprets robots.txtSearch Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.
- D3-C617 Day 3 · How long does it take to..?
Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2 weeks, and 1-3 weeks in recovery.
extends - Stage D3-C618 Day 3 · How long does it take to..?
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
extendsDocs D1-C126 Day 1 · How crawling errors affect SearchGoogle treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.
- Stage D3-C618 Day 3 · How long does it take to..?
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
extendsStage D1-C354 Day 1 · How crawling errors affect SearchGoogle slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.
- Stage D3-C618 Day 3 · How long does it take to..?
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
extends - Stage D3-C623 Day 3 · How long does it take to..?
When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D3-C623 Day 3 · How long does it take to..?
When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.
extendsStage D1-C377 Day 1 · How Google thinks about crawl budgetThe quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.
- Stage D3-C623 Day 3 · How long does it take to..?
When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.
extends - Stage D3-C624 Day 3 · How long does it take to..?
Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.
extendsStage D1-C375 Day 1 · How Google thinks about crawl budgetGooglebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.
- Docs D3-C667 Day 3 · How long does it take to..?
Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.
extendsStage D1-C528 Day 1 · How Google interprets robots.txtSearch Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.
- D2-C040 Day 2 · How is HTML interpreted
Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
repeatsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C048 Day 2 · How is HTML interpreted
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
repeatsStage D1-C328 Day 1 · How crawling worksDuring indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.
- Stage D2-C050 Day 2 · Controlling indexing
John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.
repeatsStage D1-C512 Day 1 · How Google interprets robots.txtRobots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
- Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
repeatsStage D1-C273 Day 1 · Lightning session A: Automation and AIA community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.
- D2-C264 Day 2 · What is Google friendly JavaScript
Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but returns a 200 HTTP status code.
repeatsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C284 Day 2 · What is Google friendly JavaScript
URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.
repeatsDocs D1-C101 Day 1 · How Google thinks about crawl budgetGoogle's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
repeatsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
repeatsStage D1-C355 Day 1 · How crawling errors affect SearchA soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.
- Stage D2-C847 Day 2 · Welcome to indexing day!
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
repeatsStage D1-C317 Day 1 · How crawling worksGoogle's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
- Stage D2-C848 Day 2 · Welcome to indexing day!
Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.
repeatsStage D1-C329 Day 1 · How crawling worksApart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.
- Stage D2-C850 Day 2 · Controlling indexing
John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
repeatsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D2-C851 Day 2 · Controlling indexing
When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.
repeatsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D3-C624 Day 3 · How long does it take to..?
Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.
repeats