Day 2: Indexing 47
Shown on screen 12
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript
Used byrequirement DEV-CAN-03
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
The rel=canonical link element is placed in the head section of the HTML and tells Google that one page is the representative of another.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo, transcript
Used byrequirement DEV-CAN-03glossary term rel=canonical
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
- Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byglossary term Canonical
- Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
- Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-08
To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.
“Misleading site structure (use clear signals!)”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirements DEV-CAN-08, DEV-URL-09
Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirements DEV-CAN-01, DEV-CAN-02, DEV-CAN-07
Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-07
Whether a page can load is really important for canonical selection: a broken certificate, failing JavaScript or a page that cannot be loaded counts against a URL, and the slide also listed meta refresh and security.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-02
Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirements DEV-CAN-05, DEV-URL-05
Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-CAN-02
- Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…
Said on stage 18
Google's speaker said what SEOs call the canonical is, for Google, the representative of a cluster of duplicate pages: the URL Google would ideally show.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byglossary term Canonical
Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Whether a redirect is permanent or temporary matters only for choosing the canonical, not for clustering the URLs together.
“the permanent redirect really only matters for canonicalization, not for clustering.”
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-CAN-01
rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-CAN-03
- Extends D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byglossary term rel=canonical
- Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
- Repeated by D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
Google picks the canonical from a variety of criteria and uses some kind of machine learning to decide how much weight each criterion gets; the weighting changes from time to time.
“we use some kind of machine learning to understand how strong these criteria should be. And this changes from time to time.”
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-URL-06
- Contradicts D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-MON-06
- Extended by D2-C408 Day 2: A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying…
When all canonical signals point to the same URL, Google follows what the site owner says; when they point in different directions, Google cannot tell what the owner wants.
“if there are multiple things in different directions, we don't know what to do.”
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-CAN-05
Broken canonical tags can make the wrong pages of a site show up in search results.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-MON-06
- Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
- Repeats D2-C380 Day 2: Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google…
- Extended by D2-C873 Day 2: A community speaker said that, under the canonical link specification, an improperly declared canonical tag…
Multiple canonical declarations on one page conflict and are invalid, so the search engine applies its own heuristic and picks the canonical for the site owner; a page should declare only one canonical.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-CAN-03
Canonical loops, such as two HTML pages whose canonical tags point to each other, are a structural conflict: the group has no canonical leader, so its canonical information cannot be used.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-CAN-04glossary term Canonical chain and canonical loop
A community speaker said that at best a search engine's heuristic would treat a canonical loop through redirects as self-referencing canonicals, and doubted that this is often done when the loop runs through a client-side redirect.
“But regarding the client-side redirect, I highly doubt that this is often done.”
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
For a canonical leader that users cannot reach through links, a community speaker suggested revisiting the canonical graph and probably making the internally linked page the leader instead.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
- Contradicts D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…
A community speaker said that, under the canonical link specification, an improperly declared canonical tag can also be ignored completely by the application that processes it, not only replaced by that application's own heuristic.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-CAN-03
- Extends D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
- Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
- Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
What Google's documentation says 8
Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.
Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-CAN-08
Google's canonical guide says that for canonicalization Google prefers URLs that are part of hreflang clusters: if German pages for Germany and Switzerland point to each other with hreflang but not to the Austrian page, the German and Swiss pages are preferred as canonicals.
Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-INT-08
Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.
Publisher Google Search CentralAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirements DEV-CAN-05, DEV-URL-05glossary term rel=canonical
Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.
Publisher Search Central blog (8 April 2013)Annotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-URL-06
- Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
Google's guide to specifying canonical URLs says Google supports explicit rel=canonical link annotations as described in RFC 6596, the canonical link relation specification.
Publisher Google Search CentralAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves
Google's 2013 Search Central blog post on rel=canonical mistakes says to specify no more than one rel=canonical per page: when a page has more than one, Google will likely ignore all of them, and any benefit of a legitimate canonical is lost.
Publisher Search Central blog (8 April 2013)Annotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves
Used byrequirement DEV-CAN-03
Google's guide to specifying canonical URLs says a canonical helps consolidate signals for duplicate pages: links to a duplicate URL are consolidated with links to the preferred URL once the preferred URL becomes canonical.
Publisher Google Search CentralAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves
Used byrequirement DEV-CAN-05
Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.
Publisher Google Search CentralAnnotates Day 2, 15:40 · Deciding what goes in the index?
Analysis by the author 9
Put rel=canonical and hreflang link elements in the head of the HTML the server sends, not only in JavaScript-rendered HTML, so Google can read them during HTML parsing without depending on rendering.
Author Ibrahim AnjroAnnotates Day 2, 10:25 · How is HTML interpreted
Used byrequirement DEV-CAN-03
Use permanent redirects (301 or 308) for moves you want reflected in Search: a temporary redirect still groups the URLs, but it changes which URL Google is likely to pick as the canonical.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-CAN-01
Make location and variant pages differ in their main content (local stock, staff, prices, addresses) and return 404 for empty or invalid combinations; otherwise every URL that fits the pattern can be folded into one canonical.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirements DEV-CAN-08, DEV-URL-09
Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-URL-06
Before a migration or template change, align every canonical signal for the preferred URL (redirects, rel=canonical, internal links, sitemap entries, hreflang and working HTTPS); Google says it follows the site owner only when the signals agree.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirement DEV-CAN-05
A broken canonical setup does not fail visibly: pages still load normally while the choice of URL passes to the search engine's own heuristics, so canonical tags need a regular audit with a crawler and the URL Inspection tool.
Author Ibrahim AnjroAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves
Used byrequirement DEV-MON-06
Google says canonicalization consolidates link signals from duplicates into the chosen canonical, so a leader with few links of its own is not necessarily weaker once Google accepts it; the practical risk of a weakly linked leader is that Google, treating rel=canonical as a hint, picks a different page as canonical.
Author Ibrahim AnjroAnnotates Day 2, 12:05 · Lightning session E: Managing Duplicates and Site Moves
'Only canonicals end up in search results' as said on stage is a simplification: non-canonical duplicates are dropped from the index, but Google's documentation says an alternate from the same cluster can still be shown in some contexts, such as a mobile version to a mobile user.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?