Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Not in docs 5
Storage is a second reason for deduplication: Google's storage has many competing uses and storage prices have risen sharply, so the space for any one use is limited and Google has to draw a line somewhere.
Day 2 · Handling web duplication
City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.
Day 2 · Handling web duplication
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
Day 2 · Handling web duplication
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
Gary Illyes · Day 2 · Finding the gold nuggets: structured data, media, and more!
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
Gary Illyes · Day 2 · Finding the gold nuggets: structured data, media, and more!
Consistent with docs 12
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
Day 1 · How Search works and where's AI?
When a page is selected for Google's index, its whole duplicate cluster goes into the index with it.
Cherry Prommawin · Day 1 · How Search works and where's AI?
A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.
Tobias Schwarz · Day 1 · Lightning session C: Crawling
Gary Illyes said there is no such thing as a duplicate content penalty.
Gary Illyes · Day 1 · session not recorded
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
Day 2 · How is HTML interpreted
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
Day 2 · Handling web duplication
For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.
Day 2 · Handling web duplication
Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.
Day 2 · Handling web duplication
Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.
Day 2 · Handling web duplication
Google builds duplicate clusters from four kinds of input: redirects, content, rel=canonical, and a 'magic bucket' of other things.
Day 2 · Handling web duplication
To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.
Day 2 · Handling web duplication
Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.
Day 2 · Handling web duplication
Confirmed by docs 6
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Day 2 · How is HTML interpreted
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
Day 2 · Handling web duplication
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
Day 2 · Handling web duplication
The first reason Google deduplicates is that users do not want to see the same page repeated in the search results, even if site owners would like it to rank ten times on page one.
Day 2 · Handling web duplication
Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.
Day 2 · Handling web duplication
Google's slide said there is not one single ranking system and named spam detection systems, the reviews system, BERT, MUM, RankBrain, freshness systems, deduplication systems, crisis information systems and link analysis systems (PageRank).
Day 3 · How Google thinks about Quality