Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Thing · Concept

Duplicate content

The same or nearly the same content at several URLs, which Google deduplicates: it clusters the URLs and serves one canonical; not a penalty.

ConceptAlso calledduplicate cluster, duplicate clusters, deduplication, deduplicates, deduping

Open in Reef mapOpen in Graph

Narrative see the topic Duplicate content

Claims
28
In Google’s docs
4
Said at the event
23
Not in docs
5
Kit items
10

Glossary · Duplicate cluster

A group of pages Google considers essentially equivalent: it stores one in the index and keeps track of the other URLs.

Google’s documentation 4

Documented in

DocsSourceD1-C128

Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

Google Search Central · Day 1 · session not recorded

DocsSourceD2-C373

Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.

Google Search Central · Day 2 · Handling web duplication

DocsSourceD2-C702

Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.

Google Search Central · Day 2 · Deciding what goes in the index?

DocsSourceD3-C669

Google's guide to fixing canonicalization issues says that even after content issues are fixed, Google might hold pages in a duplicate cluster for up to two weeks, and that pages split out faster when they differ clearly and significantly.

Google Search Central · Day 3 · How long does it take to..?

Said at the event 23

Slide and stage claims that name it, the ones Google’s documentation does not cover first.

Not in docs 5

StageNot in docsD2-C350

Storage is a second reason for deduplication: Google's storage has many competing uses and storage prices have risen sharply, so the space for any one use is limited and Google has to draw a line somewhere.

Day 2 · Handling web duplication

SlideNot in docsD2-C370

City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.

Day 2 · Handling web duplication

Consistent with docs 12

StageConsistent with docsD2-C346

For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

Day 2 · Handling web duplication

StageConsistent with docsD2-C351

Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

Day 2 · Handling web duplication

SlideConsistent with docsD2-C371

To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.

Day 2 · Handling web duplication

SlideConsistent with docsD2-C381

Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.

Day 2 · Handling web duplication

Confirmed by docs 6

SlideConfirmed by docsD2-C345

Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

Day 2 · Handling web duplication

SlideConfirmed by docsD2-C348

Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

Day 2 · Handling web duplication

StageConfirmed by docsD2-C349

The first reason Google deduplicates is that users do not want to see the same page repeated in the search results, even if site owners would like it to rank ten times on page one.

Day 2 · Handling web duplication

SlideConfirmed by docsD2-C352

Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.

Day 2 · Handling web duplication

SlideConfirmed by docsD3-C133

Google's slide said there is not one single ranking system and named spam detection systems, the reviews system, BERT, MUM, RankBrain, freshness systems, deduplication systems, crisis information systems and link analysis systems (PageRank).

Day 3 · How Google thinks about Quality

Press and analysis 1

Built on these claims 10

Kit items about Duplicate content: their own words name it, or several of the claims they rest on do.

Developer requirements 6

1 more

Also inglossary terms Duplicate cluster, Alternate names, Canonical, Similar URLs

Connected things 10

Most often named with it

Things named in the same claim, with the number of claims they share.

Topics that feature it