Speaker from the next talk's back-reference ('pick up where John left off with the canonical links'), heard in a second attendee recording.
Shown on screen 20
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
Speaker John MuellerEvidence slide photo, transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
- Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
Speaker John MuellerEvidence slide photo, transcript
Used byglossary term Canonical
- Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
- Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.
Speaker John MuellerEvidence slide photo, transcript
Used byglossary term Alternate names
Alternate names also serve localization: if Google knows that country versions such as you.de and you.at are equivalent, it can pick the right version to show using hreflang.
Speaker John MuellerEvidence slide photo
Alternate names are why a site: query for an old domain still shows the old domain's URLs after a site migration, which site owners often misread as a migration that is not working.
“FYI "alternate names" is why you see old domains in site:-queries”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-09
After a rebrand that changes the domain (the slide's example was johns-bikes to slow-bikes), Google can still show the old domain to people who search for the old brand by name, in navigational and branded queries.
Speaker John MuellerEvidence slide photo, transcript
Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.
“When main content is the same, pages may be clustered.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-07
Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
“Do we even need to crawl /buy/seo-service ?”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo
Used byrequirements DEV-CAN-08, DEV-URL-09
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-08
To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.
“Misleading site structure (use clear signals!)”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-08
When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.
“We try to use hreflang alternates.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-08
- Extends D2-C033 Day 2: Google extracts hreflang annotations, through which site owners specify the language variants of their…
Google's speaker advised against 'clever' geo-redirecting, because it very often goes wrong.
“clever" geo-redirecting (is often not so clever)”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-INT-02
Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-01, DEV-CAN-02, DEV-CAN-07
Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-07
Whether a page can load is really important for canonical selection: a broken certificate, failing JavaScript or a page that cannot be loaded counts against a URL, and the slide also listed meta refresh and security.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-02
Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirements DEV-CAN-05, DEV-URL-05
Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-CAN-02
- Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…
Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.
“Don't block agents.”
Wording checked against the slide or recording
Speaker John MuellerEvidence slide photo, transcript
Used byrequirement DEV-AIF-04
- Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
Said on stage 23
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
Speaker John MuellerEvidence transcript
- Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
- Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.
Speaker John MuellerEvidence transcript
Used byglossary term Duplicate cluster
- Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
Google's speaker said what SEOs call the canonical is, for Google, the representative of a cluster of duplicate pages: the URL Google would ideally show.
Speaker John MuellerEvidence transcript
Used byglossary term Canonical
The first reason Google deduplicates is that users do not want to see the same page repeated in the search results, even if site owners would like it to rank ten times on page one.
Speaker John MuellerEvidence transcript
Storage is a second reason for deduplication: Google's storage has many competing uses and storage prices have risen sharply, so the space for any one use is limited and Google has to draw a line somewhere.
Speaker John MuellerEvidence transcript
Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-09
- Extended by D3-C642 Day 3: Google treats a site move as a complex canonicalization process in which every signal of the old site is…
When someone searches for an old domain after a migration, Google shows the old URL as an alternate version because that is what was searched for, and relies on the redirect to take the user to the new site.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-09
Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.
Speaker John MuellerEvidence transcript
Google builds duplicate clusters from four kinds of input: redirects, content, rel=canonical, and a 'magic bucket' of other things.
Speaker John MuellerEvidence transcript
Google trusts redirects very much for clustering, because a redirect is a clear sign that there is one version of the content; Google keeps track of both URLs but stores only one copy of the content.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-01
Whether a redirect is permanent or temporary matters only for choosing the canonical, not for clustering the URLs together.
“the permanent redirect really only matters for canonicalization, not for clustering.”
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-01
Google clusters duplicate pages by content in four ways: exact matches, near matches, structurally similar content, and soft 404s.
Speaker John MuellerEvidence transcript
Exact-match duplicates, such as the www and non-www versions of the same page, are clustered and Google keeps only one of them.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-02
Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-ERR-01
- Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
Speaker John MuellerEvidence transcript
Used byrequirements DEV-SRV-01, DEV-SRV-02
- Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
- Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
- Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
Speaker John MuellerEvidence transcript
Used byrequirements DEV-SRV-01, DEV-SRV-02
- Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-AIF-04
- Extends D1-C436 Day 1: A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to…
rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-03
- Extends D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.
Speaker John MuellerEvidence transcript
Used byglossary term rel=canonical
- Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
- Repeated by D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
Google picks the canonical from a variety of criteria and uses some kind of machine learning to decide how much weight each criterion gets; the weighting changes from time to time.
“we use some kind of machine learning to understand how strong these criteria should be. And this changes from time to time.”
Speaker John MuellerEvidence transcript
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-URL-06
- Contradicts D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.
Speaker John MuellerEvidence transcript
Used byrequirement DEV-MON-06
- Extended by D2-C408 Day 2: A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying…
When all canonical signals point to the same URL, Google follows what the site owner says; when they point in different directions, Google cannot tell what the owner wants.
“if there are multiple things in different directions, we don't know what to do.”
Speaker John MuellerEvidence transcript
Used byrequirement DEV-CAN-05
What Google's documentation says 7
Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.
“This is normal and as users get used to the new domain name, the alternate names will fade away without you doing anything.”
Publisher Google Search Central
Used byrequirements DEV-CAN-01, DEV-CAN-09glossary term Alternate names
Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.
Publisher Google Search Central
Used byrequirement DEV-CAN-08
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
Publisher Search Central blog (24 December 2024)
Used byrequirements DEV-ERR-03, DEV-SRV-02
- Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
Google's canonical guide says that for canonicalization Google prefers URLs that are part of hreflang clusters: if German pages for Germany and Switzerland point to each other with hreflang but not to the Austrian page, the German and Swiss pages are preferred as canonicals.
Publisher Google Search Central
Used byrequirement DEV-INT-08
Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.
Publisher Google Search Central
Used byrequirements DEV-CAN-05, DEV-URL-05glossary term rel=canonical
Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.
Publisher Search Central blog (8 April 2013)
Used byrequirement DEV-URL-06
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Google's pagination guide says the pages of a paginated sequence may share the same title and description, and suggests linking every page of the sequence back to the first page.
Publisher Google Search Central
Used byrequirement DEV-URL-06
Analysis by the author 8
After a migration, judge success by the new domain's indexing and traffic in Search Console rather than by a site: query on the old domain, and keep the old domain's redirects in place long term so searches for the old brand still reach the new site.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-09
Use permanent redirects (301 or 308) for moves you want reflected in Search: a temporary redirect still groups the URLs, but it changes which URL Google is likely to pick as the canonical.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-01
Make location and variant pages differ in their main content (local stock, staff, prices, addresses) and return 404 for empty or invalid combinations; otherwise every URL that fits the pattern can be folded into one canonical.
Author Ibrahim Anjro
Used byrequirements DEV-CAN-08, DEV-URL-09
Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.
Author Ibrahim Anjro
Used byrequirements DEV-AIF-04, DEV-SRV-02
For German-language sites serving Germany, Austria and Switzerland, add hreflang with region codes (de-DE, de-AT, de-CH) and make the country pages differ in more than boilerplate (prices, shipping, legal details), or expect Google to cluster them and show one.
Author Ibrahim Anjro
Used byrequirement DEV-INT-08
Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.
Author Ibrahim Anjro
Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.
Author Ibrahim Anjro
Used byrequirement DEV-URL-06
Before a migration or template change, align every canonical signal for the preferred URL (redirects, rel=canonical, internal links, sitemap entries, hreflang and working HTTPS); Google says it follows the site owner only when the signals agree.
Author Ibrahim Anjro
Used byrequirement DEV-CAN-05