Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Area

Indexing

How Google processes a crawled page: parsing and extraction, main content, tokens, duplicates, canonicals, redirects and site moves.

8 topics · 282 claims

Topics in this area 8

Across days 55

  1. Stage D2-C325 Day 2 · Understanding what's on a page

    Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

    contradicts
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  2. Stage D2-C393 Day 2 · Handling web duplication

    Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

    contradicts
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  3. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  4. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  5. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C092 Day 1 · How Google thinks about crawl budget

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

  6. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  7. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  8. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

  9. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C036 Day 1 · How Search works and where's AI?

    Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.

  10. Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  11. Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  12. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  13. Slide D2-C127 Day 2 · Lightning session D: Rendering and JavaScript

    A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  14. Slide D2-C167 Day 2 · Lightning session D: Rendering and JavaScript

    The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  15. Slide D2-C168 Day 2 · Lightning session D: Rendering and JavaScript

    In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  16. Docs D2-C169 Day 2 · Lightning session D: Rendering and JavaScript

    Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.

    extends
    Slide D1-C040 Day 1 · How Search works and where's AI?

    URL discovery works through links: a homepage links to section pages, which link to further pages.

  17. Stage D2-C338 Day 2 · Understanding what's on a page

    Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

    extends
    Slide D1-C042 Day 1 · How Search works and where's AI?

    BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.

  18. Stage D2-C339 Day 2 · Understanding what's on a page

    For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  19. Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  20. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  21. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  22. Stage D2-C346 Day 2 · Handling web duplication

    For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  23. Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

    extends
    Stage D1-C111 Day 1 · session not recorded

    Gary Illyes said there is no such thing as a duplicate content penalty.

  24. Stage D2-C367 Day 2 · Handling web duplication

    Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  25. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  26. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  27. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C069 Day 1 · How crawling errors affect Search

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

  28. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  29. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  30. Stage D2-C375 Day 2 · Handling web duplication

    Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  31. Docs D2-C376 Day 2 · Handling web duplication

    Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  32. Stage D2-C380 Day 2 · Handling web duplication

    Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

    extends
    Analysis D1-C113 Day 1 · session not recorded

    The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

  33. Docs D2-C395 Day 2 · Handling web duplication

    Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  34. Stage D2-C401 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    Broken canonical tags can make the wrong pages of a site show up in search results.

    extends
    Analysis D1-C113 Day 1 · session not recorded

    The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

  35. Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  36. Stage D2-C427 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.

    extends
    Analysis D1-C067 Day 1 · How crawling works

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

  37. Slide D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

    extends
    Slide D1-C063 Day 1 · How crawling works

    The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

  38. Stage D2-C688 Day 2 · Deciding what goes in the index?

    Index selection is the last step before documents enter Google's index.

    extends
    Stage D1-C211 Day 1 · How Search works and where's AI?

    Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

  39. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  40. Stage D2-C884 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.

    extends
    Stage D1-C540 Day 1 · Q&A

    Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

  41. Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  42. Stage D3-C011 Day 3 · Making sense of users' queries

    Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.

    extends
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

  43. Stage D3-C013 Day 3 · Making sense of users' queries

    Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

    extends
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  44. Stage D3-C013 Day 3 · Making sense of users' queries

    Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

    extends
    Stage D2-C737 Day 2 · How does the index look like?

    A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

  45. Stage D3-C075 Day 3 · Making sense of users' queries

    At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.

    extends
    Stage D2-C738 Day 2 · How does the index look like?

    At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.

  46. Analysis D3-C109 Day 3 · Lightning session K: Facets of quality

    The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.

    extends
    Stage D2-C318 Day 2 · Understanding what's on a page

    Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

  47. Stage D3-C170 Day 3 · How Google thinks about Quality

    Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.

    extends
    Stage D2-C311 Day 2 · Understanding what's on a page

    Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.

  48. Stage D3-C316 Day 3 · How Search results are born

    Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.

    extends
    Stage D2-C724 Day 2 · How does the index look like?

    The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.

  49. Stage D3-C323 Day 3 · How Search results are born

    Most of Google's search features need nothing extra from the site owner; Google generates them from what it extracted from the page during indexing.

    extends
    Stage D2-C446 Day 2 · Finding the gold nuggets: structured data, media, and more!

    The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.

  50. Stage D3-C626 Day 3 · How long does it take to..?

    Google renders pages in two ways: immediately after crawling, or later through a queue-based process that runs elsewhere.

    extends
    Slide D2-C170 Day 2 · Lightning session D: Rendering and JavaScript

    After processing, an indexable page is placed in Google's render queue to wait for rendering.

  51. Stage D3-C642 Day 3 · How long does it take to..?

    Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.

    extends
    Stage D2-C351 Day 2 · Handling web duplication

    Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

  52. Analysis D3-C676 Day 3 · How long does it take to..?

    The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.

    extends
    Docs D2-C131 Day 2 · Lightning session D: Rendering and JavaScript

    Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.

  53. Slide D2-C048 Day 2 · How is HTML interpreted

    Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

    repeats
    Stage D1-C328 Day 1 · How crawling works

    During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

  54. Slide D3-C070 Day 3 · Making sense of users' queries

    Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.

    repeats
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

  55. Stage D3-C074 Day 3 · Making sense of users' queries

    Google's index uses posting lists: for each word, a list of the URLs associated with that word.

    repeats
    Stage D2-C733 Day 2 · How does the index look like?

    For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.