Gary Illyes said there is no such thing as a duplicate content penalty.
Speaker Gary IllyesEvidence notes
- Extended by D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Day 1 · Wednesday 30 September 2026
Points from my notes where I did not write down which session they came from.
Gary Illyes said there is no such thing as a duplicate content penalty.
Speaker Gary IllyesEvidence notes
Images were stressed as important.
Evidence notes
Attendees were told to check ARIA and accessibility.
Evidence notes, transcript
Used byrequirement DEV-HTM-04
SEO is not only content; it has many parts.
Evidence notes
Check and focus on rich results for Google.
Evidence notes
Speakers at the event were openly dismissive of llms.txt.
Evidence notes, transcript
Googlebot does not support HTTP/3 today.
Evidence notes
Used byrequirement DEV-SRV-07
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
“Some duplicate content on a site is normal and it's not a violation of Google's spam policies.”
Publisher Google Search Central
Used byglossary term Duplicate cluster
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
Publisher Google Search Central
Used byrequirement DEV-URL-06
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
Publisher Google Search Central
Used byrequirement DEV-HTM-04
Google's crawlers support HTTP/1.1 and HTTP/2, use whichever gives the best crawling performance, and may switch between sessions. HTTP/2 can save server resources but gives no ranking benefit.
Publisher Google
Used byrequirement DEV-SRV-07
Google's crawler overview says HTTP/1.1 is the default protocol of Google's crawlers and that a site can opt out of crawling over HTTP/2 by answering Google's HTTP/2 requests with a 421 status code.
Publisher Google
Used byrequirement DEV-SRV-07
As reported by Search Engine Journal, John Mueller said in April 2026 that there is no penalty or ranking demotion for having multiple URLs with the same content; Google picks one to keep.
“There's no penalty or ranking demotion if you have multiple URLs going to the same content.”
Reported by Search Engine Journal (8 April 2026)
The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
Author Ibrahim Anjro
With one in six AI Mode searches being multimodal, original images with descriptive file names, alt text and captions feed AI answers as well as image search.
Author Ibrahim Anjro
Google has not called ARIA a ranking signal. The case for it is that AI agents and assistive tools read pages through the same structure: real buttons, labelled forms and semantic headings.
Author Ibrahim Anjro
Used byrequirement DEV-HTM-04
This does not mean avoiding HTTP/3, which benefits browsers. The rule is never to run a host that only answers over HTTP/3, and to check that CDN or firewall rules written for HTTP/3 traffic do not break HTTP/1.1 and HTTP/2.
Author Ibrahim Anjro
Used byrequirement DEV-SRV-07
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard as Ahrefs') that looked at almost 40,000 sites over one month: 97% had no AI agent hits on the file, and those that had any got about two hits a month.
A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
Many SEOs still treat everything beyond the raw HTML as the developers' business, but developers often do not handle rendering problems, a community speaker warned.
Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
Content that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () => fetch('/api/specs'), does not exist for Google until someone clicks.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each chunk its own persistent, unique URL, link sequentially to those URLs, and update the displayed URL with the History API when a new chunk becomes the main visible element.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
Gary Illyes said there is no such thing as a duplicate content penalty.
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.
The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
Broken canonical tags can make the wrong pages of a site show up in search results.
The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.
Gary Illyes said images and videos drive a large amount of traffic to publishers.
An image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.
An image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.
With one in six AI Mode searches being multimodal, original images with descriptive file names, alt text and captions feed AI answers as well as image search.
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.
Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.