Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 2 · Thursday 1 October 2026 · 11:55

Handling web duplication

Speaker John Mueller, Search Relations Team Lead

TalkCoverageTranscriptSlides

Speaker from the next talk's back-reference ('pick up where John left off with the canonical links'), heard in a second attendee recording.

Shown on screen 20

SlideConfirmed by docsD2-C345

Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

Speaker John MuellerEvidence slide photo, transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
  • Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
SlideConfirmed by docsD2-C348

Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

Speaker John MuellerEvidence slide photo, transcript

Used byglossary term Canonical

  • Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
  • Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
SlideConfirmed by docsD2-C352

Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.

Speaker John MuellerEvidence slide photo, transcript

Used byglossary term Alternate names

SlideConsistent with docsD2-C354

Alternate names are why a site: query for an old domain still shows the old domain's URLs after a site migration, which site owners often misread as a migration that is not working.

“FYI "alternate names" is why you see old domains in site:-queries”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-09

SlideConsistent with docsD2-C356

After a rebrand that changes the domain (the slide's example was johns-bikes to slow-bikes), Google can still show the old domain to people who search for the old brand by name, in navigational and branded queries.

Speaker John MuellerEvidence slide photo, transcript

SlideConfirmed by docsD2-C366

Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.

“When main content is the same, pages may be clustered.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-INT-07

SlideNot in docsD2-C368

Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideNot in docsD2-C369

Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

“Do we even need to crawl /buy/seo-service ?”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo

Used byrequirements DEV-CAN-08, DEV-URL-09

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
SlideNot in docsD2-C370

City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-08

SlideConsistent with docsD2-C371

To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.

“Misleading site structure (use clear signals!)”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideConsistent with docsD2-C381

Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-INT-08

SlideConfirmed by docsD2-C382

When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

“We try to use hreflang alternates.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Things

Used byrequirement DEV-INT-08

  • Extends D2-C033 Day 2: Google extracts hreflang annotations, through which site owners specify the language variants of their…
SlideConsistent with docsD2-C387

Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-01, DEV-CAN-02, DEV-CAN-07

SlideConsistent with docsD2-C388

Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-07

SlideConsistent with docsD2-C389

Whether a page can load is really important for canonical selection: a broken certificate, failing JavaScript or a page that cannot be loaded counts against a URL, and the slide also listed meta refresh and security.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-02

SlideConsistent with docsD2-C390

Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-05, DEV-URL-05

SlideConsistent with docsD2-C396

Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-02

  • Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…
SlideNot in docsD2-C397

Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.

“Don't block agents.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-AIF-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…

Said on stage 23

StageConsistent with docsD2-C344

Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

Speaker John MuellerEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
StageConsistent with docsD2-C346

For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

Speaker John MuellerEvidence transcript

Used byglossary term Duplicate cluster

  • Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
StageConsistent with docsD2-C351

Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-09

  • Extended by D3-C642 Day 3: Google treats a site move as a complex canonicalization process in which every signal of the old site is…
StageConsistent with docsD2-C361

Google trusts redirects very much for clustering, because a redirect is a clear sign that there is one version of the content; Google keeps track of both URLs but stores only one copy of the content.

Speaker John MuellerEvidence transcript

Things

Used byrequirement DEV-CAN-01

StageNot in docsD2-C364

Google clusters duplicate pages by content in four ways: exact matches, near matches, structurally similar content, and soft 404s.

Speaker John MuellerEvidence transcript

Things
StageConsistent with docsD2-C365

Exact-match duplicates, such as the www and non-www versions of the same page, are clustered and Google keeps only one of them.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-02

StageConsistent with docsD2-C367

Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

Speaker John MuellerEvidence transcript

Things

Used byrequirement DEV-ERR-01

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
StageConsistent with docsD2-C374

A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

Speaker John MuellerEvidence transcript

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
StageConsistent with docsD2-C375

Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

Speaker John MuellerEvidence transcript

Things

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
StageNot in docsD2-C377

AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-AIF-04

  • Extends D1-C436 Day 1: A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to…
StageConsistent with docsD2-C379

rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-03

  • Extends D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
StageConfirmed by docsD2-C380

Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

Speaker John MuellerEvidence transcript

Used byglossary term rel=canonical

  • Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
  • Repeated by D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
StageNot in docsD2-C386

Google picks the canonical from a variety of criteria and uses some kind of machine learning to decide how much weight each criterion gets; the weighting changes from time to time.

“we use some kind of machine learning to understand how strong these criteria should be. And this changes from time to time.”

Speaker John MuellerEvidence transcript

StageNot in docsD2-C393

Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-URL-06

  • Contradicts D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
StageConsistent with docsD2-C399

When all canonical signals point to the same URL, Google follows what the site owner says; when they point in different directions, Google cannot tell what the owner wants.

“if there are multiple things in different directions, we don't know what to do.”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-05

What Google's documentation says 7

DocsSourceD2-C357

Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.

“This is normal and as users get used to the new domain name, the alternate names will fade away without you doing anything.”

Publisher Google Search Central

Used byrequirements DEV-CAN-01, DEV-CAN-09glossary term Alternate names

DocsSourceD2-C373

Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.

Publisher Google Search Central

Used byrequirement DEV-CAN-08

DocsSourceD2-C376

Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

Publisher Search Central blog (24 December 2024)

Things

Used byrequirements DEV-ERR-03, DEV-SRV-02

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
DocsSourceD2-C385

Google's canonical guide says that for canonicalization Google prefers URLs that are part of hreflang clusters: if German pages for Germany and Switzerland point to each other with hreflang but not to the Austrian page, the German and Swiss pages are preferred as canonicals.

Publisher Google Search Central

Used byrequirement DEV-INT-08

DocsSourceD2-C391

Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.

Publisher Google Search Central

Used byrequirements DEV-CAN-05, DEV-URL-05glossary term rel=canonical

DocsSourceD2-C395

Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

Publisher Search Central blog (8 April 2013)

Used byrequirement DEV-URL-06

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…

Analysis by the author 8

AnalysisD2-C378

Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.

Author Ibrahim Anjro

Used byrequirements DEV-AIF-04, DEV-SRV-02

AnalysisD2-C392

Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.

Author Ibrahim Anjro

AnalysisD2-C394

Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.

Author Ibrahim Anjro

Used byrequirement DEV-URL-06

  1. Stage D2-C393 Day 2 · Handling web duplication

    Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

    contradicts
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  2. Stage D2-C429 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.

    contradicts
    Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

  3. Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  4. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  5. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  6. Stage D2-C346 Day 2 · Handling web duplication

    For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  7. Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

    extends
    Stage D1-C111 Day 1 · session not recorded

    Gary Illyes said there is no such thing as a duplicate content penalty.

  8. Stage D2-C367 Day 2 · Handling web duplication

    Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  9. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  10. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  11. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C069 Day 1 · How crawling errors affect Search

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

  12. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  13. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  14. Stage D2-C375 Day 2 · Handling web duplication

    Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  15. Docs D2-C376 Day 2 · Handling web duplication

    Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  16. Stage D2-C377 Day 2 · Handling web duplication

    AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

    extends
    Stage D1-C436 Day 1 · Q&A

    A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

  17. Stage D2-C379 Day 2 · Handling web duplication

    rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

    extends
    Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

  18. Stage D2-C380 Day 2 · Handling web duplication

    Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

    extends
    Analysis D1-C113 Day 1 · session not recorded

    The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

  19. Slide D2-C382 Day 2 · Handling web duplication

    When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

    extends
    Slide D2-C033 Day 2 · How is HTML interpreted

    Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.

  20. Docs D2-C395 Day 2 · Handling web duplication

    Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

    extends
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  21. Slide D2-C397 Day 2 · Handling web duplication

    Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.

    extends
    Docs D1-C131 Day 1 · session not recorded

    Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

  22. Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

    extends
    Stage D2-C398 Day 2 · Handling web duplication

    Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.

  23. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

  24. Stage D2-C882 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs with no same-intent match were removed with an error status instead of being redirected (the exact code is unclear in the recording).

    extends
    Slide D2-C396 Day 2 · Handling web duplication

    Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

  25. Stage D3-C642 Day 3 · How long does it take to..?

    Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.

    extends
    Stage D2-C351 Day 2 · Handling web duplication

    Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

  26. Stage D2-C402 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.

    repeats
    Stage D2-C380 Day 2 · Handling web duplication

    Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

  27. Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

    repeats
    Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.