Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Area

Crawling

How Google discovers and fetches URLs, and what gets in the way.

18 topics · 420 claims

Topics in this area 18

Across days 107

  1. Stage D2-C393 Day 2 · Handling web duplication

    Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

    contradicts
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  2. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Slide D1-C079 Day 1 · How Google interprets robots.txt

    The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

  3. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  4. Slide D2-C020 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  5. Slide D2-C021 Day 2 · Welcome to indexing day!

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

    extends
    Slide D1-C105 Day 1 · How Google thinks about crawl budget

    URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.

  6. Docs D2-C022 Day 2 · Welcome to indexing day!

    Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  7. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  8. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  9. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C092 Day 1 · How Google thinks about crawl budget

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

  10. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  11. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  12. Stage D2-C038 Day 2 · How is HTML interpreted

    Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  13. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  14. Stage D2-C058 Day 2 · Controlling indexing

    John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  15. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Docs D1-C127 Day 1 · How Google thinks about crawl budget

    Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

  16. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Stage D1-C482 Day 1 · Q&A

    Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

  17. Docs D2-C069 Day 2 · Controlling indexing

    Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  18. Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    extends
    Stage D1-C434 Day 1 · Q&A

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

  19. Slide D2-C127 Day 2 · Lightning session D: Rendering and JavaScript

    A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  20. Slide D2-C167 Day 2 · Lightning session D: Rendering and JavaScript

    The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  21. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  22. Docs D2-C169 Day 2 · Lightning session D: Rendering and JavaScript

    Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  23. Slide D2-C184 Day 2 · Lightning session D: Rendering and JavaScript

    Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  24. Slide D2-C187 Day 2 · Lightning session D: Rendering and JavaScript

    A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  25. Analysis D2-C188 Day 2 · Lightning session D: Rendering and JavaScript

    Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.

    extends
    Docs D1-C101 Day 1 · How Google thinks about crawl budget

    Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

  26. Analysis D2-C189 Day 2 · Lightning session D: Rendering and JavaScript

    Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.

    extends
    Analysis D1-C067 Day 1 · How crawling works

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

  27. Stage D2-C190 Day 2 · Lightning session D: Rendering and JavaScript

    In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  28. Slide D2-C195 Day 2 · Lightning session D: Rendering and JavaScript

    Content that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  29. Slide D2-C197 Day 2 · Lightning session D: Rendering and JavaScript

    Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.

    extends
    Stage D1-C114 Day 1 · Q&A

    A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

  30. Slide D2-C201 Day 2 · Lightning session D: Rendering and JavaScript

    Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () => fetch('/api/specs'), does not exist for Google until someone clicks.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  31. Slide D2-C204 Day 2 · Lightning session D: Rendering and JavaScript

    Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  32. Analysis D2-C208 Day 2 · Lightning session D: Rendering and JavaScript

    A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.

    extends
    Docs D1-C081 Day 1 · How Google interprets robots.txt

    When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

  33. Analysis D2-C210 Day 2 · Lightning session D: Rendering and JavaScript

    An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.

    extends
    Docs D1-C074 Day 1 · How crawling errors affect Search

    If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.

  34. Stage D2-C213 Day 2 · Lightning session D: Rendering and JavaScript

    A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  35. Stage D2-C268 Day 2 · What is Google friendly JavaScript

    Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  36. Stage D2-C271 Day 2 · What is Google friendly JavaScript

    Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.

    extends
    Stage D1-C114 Day 1 · Q&A

    A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

  37. Stage D2-C271 Day 2 · What is Google friendly JavaScript

    Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.

    extends
    Stage D1-C491 Day 1 · Q&A

    If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.

  38. Docs D2-C276 Day 2 · What is Google friendly JavaScript

    To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each chunk its own persistent, unique URL, link sequentially to those URLs, and update the displayed URL with the History API when a new chunk becomes the main visible element.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  39. Stage D2-C287 Day 2 · What is Google friendly JavaScript

    A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  40. Stage D2-C291 Day 2 · What is Google friendly JavaScript

    In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  41. Stage D2-C296 Day 2 · What is Google friendly JavaScript

    If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  42. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    extends
    Stage D1-C070 Day 1 · How crawling errors affect Search

    Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.

  43. Slide D2-C336 Day 2 · Understanding what's on a page

    Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  44. Stage D2-C338 Day 2 · Understanding what's on a page

    Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

    extends
    Slide D1-C042 Day 1 · How Search works and where's AI?

    BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.

  45. Stage D2-C339 Day 2 · Understanding what's on a page

    For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  46. Stage D2-C367 Day 2 · Handling web duplication

    Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  47. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  48. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  49. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C069 Day 1 · How crawling errors affect Search

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

  50. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  51. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  52. Stage D2-C375 Day 2 · Handling web duplication

    Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  53. Docs D2-C376 Day 2 · Handling web duplication

    Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  54. Stage D2-C377 Day 2 · Handling web duplication

    AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

    extends
    Stage D1-C436 Day 1 · Q&A

    A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

  55. Docs D2-C395 Day 2 · Handling web duplication

    Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  56. Stage D2-C427 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.

    extends
    Analysis D1-C067 Day 1 · How crawling works

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

  57. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  58. Stage D2-C685 Day 2 · Deciding what goes in the index?

    Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  59. Analysis D2-C686 Day 2 · Deciding what goes in the index?

    Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

    extends
    Analysis D1-C095 Day 1 · How Google thinks about crawl budget

    New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.

  60. Stage D2-C695 Day 2 · Deciding what goes in the index?

    Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  61. Stage D2-C697 Day 2 · Deciding what goes in the index?

    Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  62. Stage D2-C700 Day 2 · Deciding what goes in the index?

    Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  63. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  64. Stage D2-C706 Day 2 · Deciding what goes in the index?

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    extends
    Docs D1-C097 Day 1 · How Google thinks about crawl budget

    Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.

  65. Analysis D2-C708 Day 2 · Deciding what goes in the index?

    The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  66. Stage D2-C709 Day 2 · Deciding what goes in the index?

    Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  67. Stage D2-C714 Day 2 · Deciding what goes in the index?

    'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  68. Analysis D2-C716 Day 2 · Deciding what goes in the index?

    Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

    extends
    Analysis D1-C098 Day 1 · How Google thinks about crawl budget

    On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

  69. Stage D2-C717 Day 2 · Deciding what goes in the index?

    Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

    extends
    Stage D1-C368 Day 1 · How crawling errors affect Search

    Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.

  70. Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  71. Slide D2-C820 Day 2 · Welcome to indexing day!

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

    extends
    Stage D1-C202 Day 1 · How Search works and where's AI?

    Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.

  72. Stage D2-C842 Day 2 · Welcome to indexing day!

    Google said listing the sitemap in robots.txt is fine, as many websites do.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  73. Stage D2-C846 Day 2 · Welcome to indexing day!

    Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  74. Analysis D2-C849 Day 2 · Welcome to indexing day!

    Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  75. Stage D2-C884 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.

    extends
    Stage D1-C540 Day 1 · Q&A

    Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

  76. Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  77. Stage D3-C385 Day 3 · Shopping on Search: Beyond the blue links

    Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.

    extends
    Slide D1-C065 Day 1 · How crawling works

    The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

  78. Stage D3-C602 Day 3 · How long does it take to..?

    Google said it knows hundreds of trillions of URLs (as of October 2026).

    extends
    Stage D1-C201 Day 1 · How Search works and where's AI?

    Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.

  79. Stage D3-C606 Day 3 · How long does it take to..?

    For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).

    extends
    Stage D1-C466 Day 1 · Q&A

    Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

  80. Slide D3-C610 Day 3 · How long does it take to..?

    Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.

    extends
    Stage D1-C337 Day 1 · How crawling works

    An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

  81. Slide D3-C611 Day 3 · How long does it take to..?

    If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.

    extends
    Stage D1-C337 Day 1 · How crawling works

    An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

  82. Slide D3-C612 Day 3 · How long does it take to..?

    Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.

    extends
    Stage D1-C331 Day 1 · How crawling works

    Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.

  83. Slide D3-C613 Day 3 · How long does it take to..?

    Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.

    extends
    Docs D1-C085 Day 1 · How Google interprets robots.txt

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

  84. Stage D3-C614 Day 3 · How long does it take to..?

    Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

    extends
    Docs D1-C085 Day 1 · How Google interprets robots.txt

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

  85. Stage D3-C615 Day 3 · How long does it take to..?

    A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

    extends
    Docs D1-C140 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

  86. Slide D3-C617 Day 3 · How long does it take to..?

    Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2 weeks, and 1-3 weeks in recovery.

    extends
    Stage D1-C493 Day 1 · Q&A

    There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.

  87. Stage D3-C618 Day 3 · How long does it take to..?

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    extends
    Docs D1-C126 Day 1 · How crawling errors affect Search

    Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.

  88. Stage D3-C618 Day 3 · How long does it take to..?

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    extends
    Stage D1-C354 Day 1 · How crawling errors affect Search

    Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.

  89. Stage D3-C618 Day 3 · How long does it take to..?

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    extends
    Stage D1-C440 Day 1 · Q&A

    If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.

  90. Stage D3-C623 Day 3 · How long does it take to..?

    When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  91. Stage D3-C623 Day 3 · How long does it take to..?

    When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.

    extends
    Stage D1-C377 Day 1 · How Google thinks about crawl budget

    The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.

  92. Stage D3-C623 Day 3 · How long does it take to..?

    When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.

    extends
    Stage D1-C439 Day 1 · Q&A

    Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.

  93. Stage D3-C624 Day 3 · How long does it take to..?

    Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.

    extends
    Stage D1-C375 Day 1 · How Google thinks about crawl budget

    Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

  94. Docs D3-C667 Day 3 · How long does it take to..?

    Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.

    extends
    Stage D1-C528 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

  95. Slide D2-C040 Day 2 · How is HTML interpreted

    Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.

    repeats
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  96. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    repeats
    Stage D1-C328 Day 1 · How crawling works

    During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

  97. Stage D2-C050 Day 2 · Controlling indexing

    John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

    repeats
    Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

  98. Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    repeats
    Stage D1-C273 Day 1 · Lightning session A: Automation and AI

    A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

  99. Slide D2-C264 Day 2 · What is Google friendly JavaScript

    Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but returns a 200 HTTP status code.

    repeats
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  100. Slide D2-C284 Day 2 · What is Google friendly JavaScript

    URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.

    repeats
    Docs D1-C101 Day 1 · How Google thinks about crawl budget

    Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

  101. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    repeats
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  102. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    repeats
    Stage D1-C355 Day 1 · How crawling errors affect Search

    A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.

  103. Stage D2-C847 Day 2 · Welcome to indexing day!

    Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.

    repeats
    Stage D1-C317 Day 1 · How crawling works

    Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.

  104. Stage D2-C848 Day 2 · Welcome to indexing day!

    Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.

    repeats
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

  105. Stage D2-C850 Day 2 · Controlling indexing

    John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.

    repeats
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  106. Stage D2-C851 Day 2 · Controlling indexing

    When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

    repeats
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  107. Stage D3-C624 Day 3 · How long does it take to..?

    Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.

    repeats
    Stage D1-C538 Day 1 · Q&A

    A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.