Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Indexing

Canonical selection

Some duplicate content is normal and is no spam violation or penalty: Google clusters duplicate pages and picks one as canonical. For Google the canonical is the representative of a duplicate cluster, the URL it would ideally show, and Google makes its own choice because rel=canonical is often wrong, for example, apparently, a placeholder left in place of a URL. Google named three considerations: protection against hijacking across pages or sites, user experience (whether the page loads, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps); it said machine learning sets how much each criterion weighs and that the weighting changes over time, which is not in its docs. When the signals agree Google follows the site owner, and when they point in different directions it cannot tell what the owner wants; Google also said it forwards the signals of every duplicate to the URL it picks, which sits uneasily with a community speaker's advice to make the strongest page of a canonical group its leader. On stage rel=canonical was said to 'also help a bit' next to redirects and sitemaps, but Google's canonical guide rates redirects and rel=canonical as strong signals and sitemap inclusion as weak, and says Google prefers URLs that are part of hreflang clusters. A permanent redirect matters only for which URL becomes canonical, not for clustering, and with a 302 Google shows the source URL; Google also said pointing paginated pages' canonical to page 1 can sometimes make sense, which its pagination guidance still advises against. Day 3 gave timings: a change of canonical URL usually shows within one to three weeks, though it can happen in seconds, and Google's guide to canonicalization issues says Google may hold pages in a duplicate cluster for up to two weeks after content issues are fixed. Google also described a site move as a complex canonicalization in which every signal of the old site is recalculated and moved to the new one. Day 1's second recording added Cherry Prommawin's definition: Google clusters duplicates and selects one page per cluster as its representative, the canonical, so users are not shown duplicates. A community speaker advised that a web application compute the expected URL for every request and redirect, or return an error, when the requested URL differs; on Day 2 the same speaker added that under the canonical link specification an improperly declared canonical can be ignored completely, not only replaced by the processor's own heuristic. Author’s view: of the two answers, Google's canonicalization guide favours the redirect, a strong canonical signal, while an error page throws away the links pointing at the variant.

What to do

  • Align every canonical signal for the preferred URL: redirects, rel=canonical, internal links, sitemap entries and hreflang.
  • Make sure the preferred URL loads reliably over HTTPS, without meta refresh, certificate errors or failing JavaScript.
  • Declare one rel=canonical per page, in the head of the server-sent HTML, with a real URL rather than a template placeholder.
  • Use 301 or 308 redirects for moves you want reflected in Search; a 302 keeps the source URL as the canonical.
  • Audit rel=canonical links regularly with a crawler and the URL Inspection tool.
  • Report canonicals that look hijacked or plainly wrong in Google's forums.

Day 1: Crawling 4

Said on stage 2

StageConfirmed by docsD1-C207

Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript

  • Repeats D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
  • Extended by D2-C346 Day 2: For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of…
StageConsistent with docsD1-C403

A community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.

Speaker Tobias SchwarzIn Day 1, 16:20 · Lightning session C: CrawlingEvidence transcript

Things

Used byrequirement DEV-URL-11

  • Repeated by D1-C445 Day 1: One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it…

What Google's documentation says 1

DocsSourceD1-C128

Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

“Some duplicate content on a site is normal and it's not a violation of Google's spam policies.”

Publisher Google Search CentralAnnotates Day 1 · session not recorded

Used byglossary term Duplicate cluster

  • Repeated by D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
  • Extended by D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
  • Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
  • Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
  • Extended by D2-C408 Day 2: A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying…
  • Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…

Analysis by the author 1

AnalysisD1-C421

Of the two answers a community speaker offered for a request to a URL variant (redirect, or an error page), Google's canonicalization guide favours the redirect: a redirect is a strong signal that its target should become canonical, while an error page throws away any links pointing at the variant.

Author Ibrahim AnjroAnnotates Day 1, 16:20 · Lightning session C: Crawling

Day 2: Indexing 47

Shown on screen 12

SlideConfirmed by docsD2-C031

Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript

Used byrequirement DEV-CAN-03

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
SlideConfirmed by docsD2-C345

Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
  • Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
SlideConfirmed by docsD2-C348

Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byglossary term Canonical

  • Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
  • Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
SlideNot in docsD2-C368

Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideNot in docsD2-C370

City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-08

SlideConsistent with docsD2-C371

To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.

“Misleading site structure (use clear signals!)”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideConsistent with docsD2-C387

Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirements DEV-CAN-01, DEV-CAN-02, DEV-CAN-07

SlideConsistent with docsD2-C388

Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-07

SlideConsistent with docsD2-C389

Whether a page can load is really important for canonical selection: a broken certificate, failing JavaScript or a page that cannot be loaded counts against a URL, and the slide also listed meta refresh and security.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-02

SlideConsistent with docsD2-C390

Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirements DEV-CAN-05, DEV-URL-05

SlideConsistent with docsD2-C396

Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-02

  • Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…

Said on stage 18

StageConsistent with docsD2-C359

Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

StageConsistent with docsD2-C362

Whether a redirect is permanent or temporary matters only for choosing the canonical, not for clustering the URLs together.

“the permanent redirect really only matters for canonicalization, not for clustering.”

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirement DEV-CAN-01

StageConsistent with docsD2-C379

rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirement DEV-CAN-03

  • Extends D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
StageConfirmed by docsD2-C380

Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byglossary term rel=canonical

  • Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
  • Repeated by D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
StageNot in docsD2-C393

Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirement DEV-URL-06

  • Contradicts D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
StageConsistent with docsD2-C399

When all canonical signals point to the same URL, Google follows what the site owner says; when they point in different directions, Google cannot tell what the owner wants.

“if there are multiple things in different directions, we don't know what to do.”

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirement DEV-CAN-05

StageConfirmed by docsD2-C401

Broken canonical tags can make the wrong pages of a site show up in search results.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

Used byrequirement DEV-MON-06

  • Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
StageConsistent with docsD2-C402

According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

  • Repeats D2-C380 Day 2: Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google…
  • Extended by D2-C873 Day 2: A community speaker said that, under the canonical link specification, an improperly declared canonical tag…
StageConsistent with docsD2-C417

Multiple canonical declarations on one page conflict and are invalid, so the search engine applies its own heuristic and picks the canonical for the site owner; a page should declare only one canonical.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

Used byrequirement DEV-CAN-03

StageNot in docsD2-C426

A community speaker said that at best a search engine's heuristic would treat a canonical loop through redirects as self-referencing canonicals, and doubted that this is often done when the loop runs through a client-side redirect.

“But regarding the client-side redirect, I highly doubt that this is often done.”

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

StageConsistent with docsD2-C428

For a canonical leader that users cannot reach through links, a community speaker suggested revisiting the canonical graph and probably making the internally linked page the leader instead.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

StageNot in docsD2-C429

A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

  • Contradicts D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…
StageConfirmed by docsD2-C873

A community speaker said that, under the canonical link specification, an improperly declared canonical tag can also be ignored completely by the application that processes it, not only replaced by that application's own heuristic.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

Used byrequirement DEV-CAN-03

  • Extends D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
StageConsistent with docsD2-C701

When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…

What Google's documentation says 8

DocsSourceD2-C373

Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.

Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirement DEV-CAN-08

DocsSourceD2-C385

Google's canonical guide says that for canonicalization Google prefers URLs that are part of hreflang clusters: if German pages for Germany and Switzerland point to each other with hreflang but not to the Austrian page, the German and Swiss pages are preferred as canonicals.

Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirement DEV-INT-08

DocsSourceD2-C391

Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.

Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirements DEV-CAN-05, DEV-URL-05glossary term rel=canonical

DocsSourceD2-C395

Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

Publisher Search Central blog (8 April 2013)Annotates Day 2, 11:55 · Handling web duplication

Used byrequirement DEV-URL-06

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
DocsSourceD2-C403

Google's guide to specifying canonical URLs says Google supports explicit rel=canonical link annotations as described in RFC 6596, the canonical link relation specification.

Publisher Google Search CentralAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

DocsSourceD2-C418

Google's 2013 Search Central blog post on rel=canonical mistakes says to specify no more than one rel=canonical per page: when a page has more than one, Google will likely ignore all of them, and any benefit of a legitimate canonical is lost.

Publisher Search Central blog (8 April 2013)Annotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

Used byrequirement DEV-CAN-03

DocsSourceD2-C430

Google's guide to specifying canonical URLs says a canonical helps consolidate signals for duplicate pages: links to a duplicate URL are consolidated with links to the preferred URL once the preferred URL becomes canonical.

Publisher Google Search CentralAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

Used byrequirement DEV-CAN-05

DocsSourceD2-C702

Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.

Publisher Google Search CentralAnnotates Day 2, 15:40 · Deciding what goes in the index?

Analysis by the author 9

AnalysisD2-C392

Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.

Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication

AnalysisD2-C394

Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.

Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirement DEV-URL-06

AnalysisD2-C404

A broken canonical setup does not fail visibly: pages still load normally while the choice of URL passes to the search engine's own heuristics, so canonical tags need a regular audit with a crawler and the URL Inspection tool.

Author Ibrahim AnjroAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

Used byrequirement DEV-MON-06

AnalysisD2-C440

Google says canonicalization consolidates link signals from duplicates into the chosen canonical, so a leader with few links of its own is not necessarily weaker once Google accepts it; the practical risk of a weakly linked leader is that Google, treating rel=canonical as a hint, picks a different page as canonical.

Author Ibrahim AnjroAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

Day 3: Serving: Ranking, Search Console, and Performance 5

Said on stage 3

StageConsistent with docsD3-C642

Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.

Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript

  • Extends D2-C351 Day 2: Google treats a site migration as deduplication across sites, in which the site owner says the old and the…

What Google's documentation says 1

DocsSourceD3-C669

Google's guide to fixing canonicalization issues says that even after content issues are fixed, Google might hold pages in a duplicate cluster for up to two weeks, and that pages split out faster when they differ clearly and significantly.

Publisher Google Search CentralAnnotates Day 3, 15:45 · How long does it take to..?

Used byrequirement DEV-MON-10

Analysis by the author 1

AnalysisD3-C670

The canonicalization documentation update mentioned on stage appears in Google's documentation changelog on 10 July 2026 (re-evaluation time, page last updated 21 August 2026), earlier than 'about a month' before the event; its 'up to two weeks' sits at the short end of the one to three weeks said on stage.

Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?

Across days and sessions 22

  1. Stage D2-C393 Day 2 · Handling web duplication

    Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

    contradicts
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  2. Stage D2-C429 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.

    contradicts
    Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

  3. Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  4. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  5. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  6. Stage D2-C346 Day 2 · Handling web duplication

    For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  7. Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

    extends
    Stage D1-C111 Day 1 · session not recorded

    Gary Illyes said there is no such thing as a duplicate content penalty.

  8. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  9. Stage D2-C379 Day 2 · Handling web duplication

    rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

    extends
    Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

  10. Stage D2-C380 Day 2 · Handling web duplication

    Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

    extends
    Analysis D1-C113 Day 1 · session not recorded

    The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

  11. Docs D2-C395 Day 2 · Handling web duplication

    Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  12. Stage D2-C401 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    Broken canonical tags can make the wrong pages of a site show up in search results.

    extends
    Analysis D1-C113 Day 1 · session not recorded

    The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

  13. Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  14. Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

    extends
    Stage D2-C398 Day 2 · Handling web duplication

    Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.

  15. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  16. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

  17. Stage D2-C873 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker said that, under the canonical link specification, an improperly declared canonical tag can also be ignored completely by the application that processes it, not only replaced by that application's own heuristic.

    extends
    Stage D2-C402 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.

  18. Stage D2-C882 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs with no same-intent match were removed with an error status instead of being redirected (the exact code is unclear in the recording).

    extends
    Slide D2-C396 Day 2 · Handling web duplication

    Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

  19. Stage D3-C642 Day 3 · How long does it take to..?

    Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.

    extends
    Stage D2-C351 Day 2 · Handling web duplication

    Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

  20. Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

    repeats
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  21. Stage D1-C445 Day 1 · Q&A

    One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

    repeats
    Stage D1-C403 Day 1 · Lightning session C: Crawling

    A community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.

  22. Stage D2-C402 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.

    repeats
    Stage D2-C380 Day 2 · Handling web duplication

    Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

Built on these claims 14

Developer requirements 14

Sources 15