Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Indexing

Duplicate content

There is no duplicate content penalty: Google's documentation says some duplicate content is normal and not a spam-policy violation, and Day 2 described deduplication as clustering duplicates, indexing one representative URL and forwarding the signals of every URL in the cluster, such as links, to it. Google deduplicates because users do not want repeated results and the index has no room for everything; on stage it added rising storage costs, which is not in its docs. Clusters are built from redirects, content, rel=canonical and other inputs, and Google said content clustering catches exact matches, near matches such as pages with a translated template but untranslated main content, soft 404s and URL patterns, which can fold city or service pages into one canonical without each page being looked at (the four-way split and the pattern behaviour are not in Google's docs). A CDN challenge page served with a 200 status on many URLs can get them clustered as well, and Google's troubleshooting guide says pages split out once they are clearly different, which can take up to two weeks after a fix. Author’s view: the real costs are losing control over which URL is chosen and crawling spent on copies. Day 3 widened the frame: Google's spam slide said sites with mostly scraped content and no added services or content may not provide value to users, and a community speaker said some keyword cannibalization is logical and fine, while for harmful cases the agency groups an article's Search Console queries before deciding whether to redirect the pages or rewrite and split the content. The second recording of Day 1 added the basics: Google clusters duplicates and picks one representative, the canonical, so users are not shown duplicates, and when a page is selected for the index its whole duplicate cluster goes in with it. A community speaker treated similar URLs (protocol, host, case, slashes, encoding, parameter order) as an indicator of duplicate content, and in the Q&A Google suggested redirecting parameter variants to the normalised URL; another community speaker advised against markdown copies of pages for AI agents, a duplicate (which the speaker also called a possible source of cloaking, a word the recordings disagree on). Day 2's second recording completed the deduplication talk: in a migration the site owner says the old and new domains are the same and that Google should pick the new one, and the short advice for same-language, different-country pages was to use hreflang. A Day 2 case study merged two competing loan-comparison sites and folded pages with the same search intent into one strong article each, from over 2,000 URLs to about 100.

What to do

  • Signal the preferred URL consistently with redirects, rel=canonical, internal links and sitemaps.
  • Make location, variant and country pages differ in their main content, and return 404 for empty or invalid combinations, so URL patterns are not folded into one canonical.
  • Serve bot challenges to crawlers with a 503 status, never as a 200 page.
  • Use hreflang with region codes for same-language country versions, such as de-DE, de-AT and de-CH.
  • Put the same structured data on duplicates as on the canonical.
  • After fixing a wrong cluster, allow up to two weeks before judging the result.

Day 1: Crawling 13

Said on stage 10

StageConfirmed by docsD1-C207

Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript

  • Repeats D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
  • Extended by D2-C346 Day 2: For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of…
StageConsistent with docsD1-C272

A community speaker advised against publishing markdown copies of HTML pages for AI agents: the copy is a duplicate (which the speaker also called a possible source of cloaking, an uncertain word in the recordings), and the models are trained to read HTML, CSS and JavaScript.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-AIF-02

StageConfirmed by docsD1-C379

The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

  • Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
StageConsistent with docsD1-C397

A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.

Speaker Tobias SchwarzIn Day 1, 16:20 · Lightning session C: CrawlingEvidence transcript

Used byrequirement DEV-URL-11glossary term Similar URLs

  • Extended by D1-C444 Day 1: Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily…
StageConsistent with docsD1-C399

A community speaker showed five kinds of similar URLs on well-known brand sites: a different protocol or host (HTTP vs HTTPS, www vs non-www), different capitalisation, a different number of delimiters such as slashes, a space encoded as %20 in one URL and + in another, and parameters in a different order.

Speaker Tobias SchwarzIn Day 1, 16:20 · Lightning session C: CrawlingEvidence transcript

Used byrequirement DEV-URL-11glossary term Similar URLs

StageConsistent with docsD1-C445

One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-11

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
  • Repeats D1-C403 Day 1: A community speaker advised that a web application compute the expected URL for every request, for example…
StageConsistent with docsD1-C479

Where content was duplicated across languages, Google's own site consolidation redirected two language versions into one, giving both old URLs one target path to redirect to.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Things

Used byrequirement DEV-CAN-10

  • Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
  • Extended by D1-C541 Day 1: A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in…
StageConsistent with docsD1-C111

Gary Illyes said there is no such thing as a duplicate content penalty.

Speaker Gary IllyesIn Day 1 · session not recordedEvidence notes

  • Extended by D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…

What Google's documentation says 1

DocsSourceD1-C128

Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

“Some duplicate content on a site is normal and it's not a violation of Google's spam policies.”

Publisher Google Search CentralAnnotates Day 1 · session not recorded

Used byglossary term Duplicate cluster

  • Repeated by D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
  • Extended by D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
  • Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
  • Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
  • Extended by D2-C408 Day 2: A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying…
  • Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…

What the press reported 1

PressSourceD1-C112

As reported by Search Engine Journal, John Mueller said in April 2026 that there is no penalty or ranking demotion for having multiple URLs with the same content; Google picks one to keep.

“There's no penalty or ranking demotion if you have multiple URLs going to the same content.”

Reported by Search Engine Journal (8 April 2026)Annotates Day 1 · session not recorded

Analysis by the author 1

AnalysisD1-C113

The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

Author Ibrahim AnjroAnnotates Day 1 · session not recorded

  • Extended by D2-C380 Day 2: Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google…
  • Extended by D2-C401 Day 2: Broken canonical tags can make the wrong pages of a site show up in search results.

Day 2: Indexing 48

Shown on screen 16

SlideConfirmed by docsD2-C031

Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript

Used byrequirement DEV-CAN-03

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
SlideConfirmed by docsD2-C345

Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
  • Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
SlideConfirmed by docsD2-C348

Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byglossary term Canonical

  • Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
  • Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
SlideConfirmed by docsD2-C352

Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byglossary term Alternate names

SlideConsistent with docsD2-C354

Alternate names are why a site: query for an old domain still shows the old domain's URLs after a site migration, which site owners often misread as a migration that is not working.

“FYI "alternate names" is why you see old domains in site:-queries”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-09

SlideConfirmed by docsD2-C366

Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.

“When main content is the same, pages may be clustered.”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-INT-07

SlideNot in docsD2-C368

Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideNot in docsD2-C369

Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

“Do we even need to crawl /buy/seo-service ?”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo

Used byrequirements DEV-CAN-08, DEV-URL-09

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
SlideNot in docsD2-C370

City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-08

SlideConsistent with docsD2-C371

To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.

“Misleading site structure (use clear signals!)”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideConsistent with docsD2-C381

Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-INT-08

SlideConfirmed by docsD2-C382

When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

“We try to use hreflang alternates.”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Things

Used byrequirement DEV-INT-08

  • Extends D2-C033 Day 2: Google extracts hreflang annotations, through which site owners specify the language variants of their…
SlideConsistent with docsD2-C388

Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-07

SlideConsistent with docsD2-C396

Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-CAN-02

  • Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…

Said on stage 21

StageConsistent with docsD2-C344

Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
StageConsistent with docsD2-C346

For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byglossary term Duplicate cluster

  • Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
StageConsistent with docsD2-C351

Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirement DEV-CAN-09

  • Extended by D3-C642 Day 3: Google treats a site move as a complex canonicalization process in which every signal of the old site is…
StageConsistent with docsD2-C359

Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

StageConsistent with docsD2-C361

Google trusts redirects very much for clustering, because a redirect is a clear sign that there is one version of the content; Google keeps track of both URLs but stores only one copy of the content.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Things

Used byrequirement DEV-CAN-01

StageConsistent with docsD2-C362

Whether a redirect is permanent or temporary matters only for choosing the canonical, not for clustering the URLs together.

“the permanent redirect really only matters for canonicalization, not for clustering.”

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirement DEV-CAN-01

StageConsistent with docsD2-C367

Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Things

Used byrequirement DEV-ERR-01

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
StageConsistent with docsD2-C374

A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
StageConsistent with docsD2-C375

Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript

Things

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
StageD2-C408

A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

Used byrequirement DEV-MON-06

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D2-C398 Day 2: Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure…
StageNot in docsD2-C444

Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.

Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript, slide photo

Used byrequirement DEV-SDA-10

  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
StageConsistent with docsD2-C701

When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…

What Google's documentation says 7

DocsSourceD2-C310

Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.

“When Google indexes a page, it determines the primary content (or centerpiece) of each page.”

Publisher Google Search CentralAnnotates Day 2, 11:30 · Understanding what's on a page

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

DocsSourceD2-C357

Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.

“This is normal and as users get used to the new domain name, the alternate names will fade away without you doing anything.”

Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirements DEV-CAN-01, DEV-CAN-09glossary term Alternate names

DocsSourceD2-C373

Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.

Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirement DEV-CAN-08

DocsSourceD2-C376

Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

Publisher Search Central blog (24 December 2024)Annotates Day 2, 11:55 · Handling web duplication

Things

Used byrequirements DEV-ERR-03, DEV-SRV-02

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
DocsSourceD2-C430

Google's guide to specifying canonical URLs says a canonical helps consolidate signals for duplicate pages: links to a duplicate URL are consolidated with links to the preferred URL once the preferred URL becomes canonical.

Publisher Google Search CentralAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves

Used byrequirement DEV-CAN-05

DocsSourceD2-C451

Google's general structured data guidelines recommend placing the same structured data on all duplicate pages of the same content, not just on the canonical page.

“we recommend placing the same structured data on all page duplicates, not just on the canonical page”

Publisher Google Search CentralAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!

Used byrequirement DEV-SDA-10

DocsSourceD2-C702

Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.

Publisher Google Search CentralAnnotates Day 2, 15:40 · Deciding what goes in the index?

Analysis by the author 4

AnalysisD2-C450

Because Google extracts structured data, images and videos only after deduplication, put markup and media on the URL you want as canonical and keep them identical on its duplicates; markup that exists only on a duplicate that loses canonical selection may never be extracted.

Author Ibrahim AnjroAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!

Used byrequirement DEV-SDA-10

Day 3: Serving: Ranking, Search Console, and Performance 2

Shown on screen 1

Said on stage 1

Across days and sessions 32

  1. Stage D2-C429 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.

    contradicts
    Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

  2. Stage D1-C379 Day 1 · How Google thinks about crawl budget

    The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.

    extends
    Slide D1-C103 Day 1 · How Google thinks about crawl budget

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

  3. Stage D1-C444 Day 1 · Q&A

    Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

    extends
    Stage D1-C397 Day 1 · Lightning session C: Crawling

    A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.

  4. Stage D1-C541 Day 1 · Q&A

    A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in Google's own site migration, because it was the only option available to the person doing it (the recording does not make fully clear whether the JavaScript performed the redirects or built the mapping).

    extends
    Stage D1-C479 Day 1 · Q&A

    Where content was duplicated across languages, Google's own site consolidation redirected two language versions into one, giving both old URLs one target path to redirect to.

  5. Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  6. Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

    extends
    Slide D1-C037 Day 1 · How Search works and where's AI?

    For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

  7. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  8. Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  9. Stage D2-C346 Day 2 · Handling web duplication

    For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

    extends
    Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

  10. Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

    extends
    Stage D1-C111 Day 1 · session not recorded

    Gary Illyes said there is no such thing as a duplicate content penalty.

  11. Stage D2-C367 Day 2 · Handling web duplication

    Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  12. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Slide D1-C094 Day 1 · How Google thinks about crawl budget

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

  13. Slide D2-C369 Day 2 · Handling web duplication

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  14. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C069 Day 1 · How crawling errors affect Search

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

  15. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  16. Stage D2-C374 Day 2 · Handling web duplication

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  17. Stage D2-C375 Day 2 · Handling web duplication

    Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

    extends
    Stage D1-C367 Day 1 · How crawling errors affect Search

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

  18. Docs D2-C376 Day 2 · Handling web duplication

    Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  19. Stage D2-C379 Day 2 · Handling web duplication

    rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

    extends
    Slide D2-C031 Day 2 · How is HTML interpreted

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

  20. Stage D2-C380 Day 2 · Handling web duplication

    Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

    extends
    Analysis D1-C113 Day 1 · session not recorded

    The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

  21. Slide D2-C382 Day 2 · Handling web duplication

    When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

    extends
    Slide D2-C033 Day 2 · How is HTML interpreted

    Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.

  22. Stage D2-C401 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    Broken canonical tags can make the wrong pages of a site show up in search results.

    extends
    Analysis D1-C113 Day 1 · session not recorded

    The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

  23. Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  24. Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

    extends
    Stage D2-C398 Day 2 · Handling web duplication

    Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.

  25. Stage D2-C444 Day 2 · Finding the gold nuggets: structured data, media, and more!

    Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.

    extends
    Slide D2-C026 Day 2 · How is HTML interpreted

    A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

  26. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  27. Stage D2-C701 Day 2 · Deciding what goes in the index?

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    extends
    Slide D2-C345 Day 2 · Handling web duplication

    Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

  28. Stage D2-C882 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs with no same-intent match were removed with an error status instead of being redirected (the exact code is unclear in the recording).

    extends
    Slide D2-C396 Day 2 · Handling web duplication

    Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

  29. Stage D3-C642 Day 3 · How long does it take to..?

    Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.

    extends
    Stage D2-C351 Day 2 · Handling web duplication

    Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

  30. Stage D1-C207 Day 1 · How Search works and where's AI?

    Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

    repeats
    Docs D1-C128 Day 1 · session not recorded

    Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

  31. Stage D1-C445 Day 1 · Q&A

    One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

    repeats
    Stage D1-C403 Day 1 · Lightning session C: Crawling

    A community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.

  32. Stage D2-C680 Day 2 · Deciding what goes in the index?

    Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

    repeats
    Stage D2-C344 Day 2 · Handling web duplication

    Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

Built on these claims 21

Developer requirements 21

Sources 21