Day 1: Crawling 13
Said on stage 10
Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
- Repeats D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
- Extended by D2-C346 Day 2: For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of…
When a page is selected for Google's index, its whole duplicate cluster goes into the index with it.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
For a content audit of a client blog with over 3,000 posts, an agency defined the parameters, scoring and weightings manually and had AI apply them, producing posts grouped by category, possible cannibalisation flags, a priority list and a list of posts to check by hand.
Speaker James PowleyIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
A community speaker advised against publishing markdown copies of HTML pages for AI agents: the copy is a duplicate (which the speaker also called a possible source of cloaking, an uncertain word in the recordings), and the models are trained to read HTML, CSS and JavaScript.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-AIF-02
The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.
Speaker Tobias SchwarzIn Day 1, 16:20 · Lightning session C: CrawlingEvidence transcript
Used byrequirement DEV-URL-11glossary term Similar URLs
- Extended by D1-C444 Day 1: Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily…
A community speaker showed five kinds of similar URLs on well-known brand sites: a different protocol or host (HTTP vs HTTPS, www vs non-www), different capitalisation, a different number of delimiters such as slashes, a space encoded as %20 in one URL and + in another, and parameters in a different order.
Speaker Tobias SchwarzIn Day 1, 16:20 · Lightning session C: CrawlingEvidence transcript
Used byrequirement DEV-URL-11glossary term Similar URLs
One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-11
- Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
- Repeats D1-C403 Day 1: A community speaker advised that a web application compute the expected URL for every request, for example…
Where content was duplicated across languages, Google's own site consolidation redirected two language versions into one, giving both old URLs one target path to redirect to.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-CAN-10
- Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
- Extended by D1-C541 Day 1: A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in…
Gary Illyes said there is no such thing as a duplicate content penalty.
Speaker Gary IllyesIn Day 1 · session not recordedEvidence notes
- Extended by D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…
What Google's documentation says 1
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
“Some duplicate content on a site is normal and it's not a violation of Google's spam policies.”
Publisher Google Search CentralAnnotates Day 1 · session not recorded
Used byglossary term Duplicate cluster
- Repeated by D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
- Extended by D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
- Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
- Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
- Extended by D2-C408 Day 2: A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying…
- Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
What the press reported 1
As reported by Search Engine Journal, John Mueller said in April 2026 that there is no penalty or ranking demotion for having multiple URLs with the same content; Google picks one to keep.
“There's no penalty or ranking demotion if you have multiple URLs going to the same content.”
Reported by Search Engine Journal (8 April 2026)Annotates Day 1 · session not recorded
Analysis by the author 1
The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
Author Ibrahim AnjroAnnotates Day 1 · session not recorded
- Extended by D2-C380 Day 2: Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google…
- Extended by D2-C401 Day 2: Broken canonical tags can make the wrong pages of a site show up in search results.
Day 2: Indexing 48
Shown on screen 16
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript
Used byrequirement DEV-CAN-03
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
- Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byglossary term Canonical
- Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
- Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byglossary term Alternate names
Alternate names also serve localization: if Google knows that country versions such as you.de and you.at are equivalent, it can pick the right version to show using hreflang.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo
Alternate names are why a site: query for an old domain still shows the old domain's URLs after a site migration, which site owners often misread as a migration that is not working.
“FYI "alternate names" is why you see old domains in site:-queries”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-09
Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.
“When main content is the same, pages may be clustered.”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-INT-07
Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
“Do we even need to crawl /buy/seo-service ?”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo
Used byrequirements DEV-CAN-08, DEV-URL-09
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-08
To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.
“Misleading site structure (use clear signals!)”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-INT-08
When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.
“We try to use hreflang alternates.”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-INT-08
- Extends D2-C033 Day 2: Google extracts hreflang annotations, through which site owners specify the language variants of their…
Google's speaker advised against 'clever' geo-redirecting, because it very often goes wrong.
“clever" geo-redirecting (is often not so clever)”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-INT-02
Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-07
Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-02
- Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…
Said on stage 21
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
- Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
- Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byglossary term Duplicate cluster
- Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
Google's speaker said what SEOs call the canonical is, for Google, the representative of a cluster of duplicate pages: the URL Google would ideally show.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byglossary term Canonical
The first reason Google deduplicates is that users do not want to see the same page repeated in the search results, even if site owners would like it to rank ten times on page one.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Storage is a second reason for deduplication: Google's storage has many competing uses and storage prices have risen sharply, so the space for any one use is limited and Google has to draw a line somewhere.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-CAN-09
- Extended by D3-C642 Day 3: Google treats a site move as a complex canonicalization process in which every signal of the old site is…
Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Google builds duplicate clusters from four kinds of input: redirects, content, rel=canonical, and a 'magic bucket' of other things.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Google trusts redirects very much for clustering, because a redirect is a clear sign that there is one version of the content; Google keeps track of both URLs but stores only one copy of the content.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-CAN-01
Whether a redirect is permanent or temporary matters only for choosing the canonical, not for clustering the URLs together.
“the permanent redirect really only matters for canonicalization, not for clustering.”
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-CAN-01
Google clusters duplicate pages by content in four ways: exact matches, near matches, structurally similar content, and soft 404s.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Exact-match duplicates, such as the www and non-www versions of the same page, are clustered and Google keeps only one of them.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-CAN-02
Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-ERR-01
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirements DEV-SRV-01, DEV-SRV-02
- Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
- Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
- Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirements DEV-SRV-01, DEV-SRV-02
- Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
In a community speaker's terms, a canonical group is the set of all URLs connected through canonical links, and its canonical leader is the page that all the group's canonical links ultimately point to.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-MON-06
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D2-C398 Day 2: Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure…
In a community case study, two of the first loan-comparison websites in Poland did exactly the same thing, so they competed with each other for the same users, the same keywords and the same space in search results.
Speaker Martyna AğanoğluIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
In a community case study, the kept pages that shared the same search intent were merged into one strong article each, taking the merged site from over 2,000 URLs to about 100, a cut of about 95% of URLs.
Speaker Martyna AğanoğluIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
Speaker Gary IllyesIn Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!Evidence transcript, slide photo
Used byrequirement DEV-SDA-10
- Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
What Google's documentation says 7
Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.
“When Google indexes a page, it determines the primary content (or centerpiece) of each page.”
Publisher Google Search CentralAnnotates Day 2, 11:30 · Understanding what's on a page
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.
“This is normal and as users get used to the new domain name, the alternate names will fade away without you doing anything.”
Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirements DEV-CAN-01, DEV-CAN-09glossary term Alternate names
Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.
Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-CAN-08
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
Publisher Search Central blog (24 December 2024)Annotates Day 2, 11:55 · Handling web duplication
Used byrequirements DEV-ERR-03, DEV-SRV-02
- Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
Google's guide to specifying canonical URLs says a canonical helps consolidate signals for duplicate pages: links to a duplicate URL are consolidated with links to the preferred URL once the preferred URL becomes canonical.
Publisher Google Search CentralAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves
Used byrequirement DEV-CAN-05
Google's general structured data guidelines recommend placing the same structured data on all duplicate pages of the same content, not just on the canonical page.
“we recommend placing the same structured data on all page duplicates, not just on the canonical page”
Publisher Google Search CentralAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!
Used byrequirement DEV-SDA-10
Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.
Publisher Google Search CentralAnnotates Day 2, 15:40 · Deciding what goes in the index?
Analysis by the author 4
Make location and variant pages differ in their main content (local stock, staff, prices, addresses) and return 404 for empty or invalid combinations; otherwise every URL that fits the pattern can be folded into one canonical.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirements DEV-CAN-08, DEV-URL-09
For German-language sites serving Germany, Austria and Switzerland, add hreflang with region codes (de-DE, de-AT, de-CH) and make the country pages differ in more than boilerplate (prices, shipping, legal details), or expect Google to cluster them and show one.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-INT-08
Because Google extracts structured data, images and videos only after deduplication, put markup and media on the URL you want as canonical and keep them identical on its duplicates; markup that exists only on a duplicate that loses canonical selection may never be extracted.
Author Ibrahim AnjroAnnotates Day 2, 13:30 · Finding the gold nuggets: structured data, media, and more!
Used byrequirement DEV-SDA-10
'Only canonicals end up in search results' as said on stage is a simplification: non-canonical duplicates are dropped from the index, but Google's documentation says an alternate from the same cluster can still be shown in some contexts, such as a mobile version to a mobile user.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?