Area
Indexing
How Google processes a crawled page: parsing and extraction, main content, tokens, duplicates, canonicals, redirects and site moves.
Topics in this area 8
63 claims · 12 sessions
Duplicate content
There is no duplicate content penalty: Google's documentation says some duplicate content is normal and not a spam-policy violation, and Day 2 described deduplication as clustering duplicates, indexing one representative URL and forwarding the signals of every URL in the cluster, such as links, to it. Google deduplicates because users do not want repeated results and the index has no room for everything; on stage it added rising storage costs, which is not in its docs. Clusters are built from redirects, content, rel=canonical and other inputs, and Google said content clustering catches exact matches, near matches such as pages with a translated template but untranslated main content, soft 404s and URL patterns, which can fold city or service pages into one canonical without each page being looked at (the four-way split and the pattern behaviour are not in Google's docs). A CDN challenge page served with a 200 status on many URLs can get them clustered as well, and Google's troubleshooting guide says pages split out once they are clearly different, which can take up to two weeks after a fix. Author’s view: the real costs are losing control over which URL is chosen and crawling spent on copies. Day 3 widened the frame: Google's spam slide said sites with mostly scraped content and no added services or content may not provide value to users, and a community speaker said some keyword cannibalization is logical and fine, while for harmful cases the agency groups an article's Search Console queries before deciding whether to redirect the pages or rewrite and split the content. The second recording of Day 1 added the basics: Google clusters duplicates and picks one representative, the canonical, so users are not shown duplicates, and when a page is selected for the index its whole duplicate cluster goes in with it. A community speaker treated similar URLs (protocol, host, case, slashes, encoding, parameter order) as an indicator of duplicate content, and in the Q&A Google suggested redirecting parameter variants to the normalised URL; another community speaker advised against markdown copies of pages for AI agents, a duplicate (which the speaker also called a possible source of cloaking, a word the recordings disagree on). Day 2's second recording completed the deduplication talk: in a migration the site owner says the old and new domains are the same and that Google should pick the new one, and the short advice for same-language, different-country pages was to use hreflang. A Day 2 case study merged two competing loan-comparison sites and folded pages with the same search intent into one strong article each, from over 2,000 URLs to about 100.
Day 1Day 2Day 3
40 claims · 9 sessions
Processing: from fetched page to index
Day 2 opened the stage that Day 1's crawl pipeline hands its fetches to: the crawler passes on a fetch record, shown on a slide as an example (fetch result, connect time, time to first byte, the robots policies that apply, such as a Google-Extended opt-out, and the raw HTTP response), and processing runs HTML parsing, rendering, deduplication, feature extraction, signal extraction and finally index selection. HTML parsing builds a DOM from which Google extracts every element to separate header, navigation and main content, plus the robots meta tag, rel=canonical (used for deduplication and canonical selection), hreflang and links, which go back to the crawl queue. Links are extracted again from the rendered HTML; rendering itself sits among the processing steps on the slides, is a detached system because it is so expensive, Erin Sparling said, and happens during the crawl according to Google's guide to how Search works. Gary Illyes said deduplication runs before feature extraction, so the costly extraction of structured data, images and videos is spent on fewer documents, and images and videos go on to dedicated indexing services. The fetch record, the DOM step and that order were shown or said at the event and are not in Google's documentation, which describes indexing more broadly. Day 3 timed the stage: Google estimated that indexing a document end to end, from entering indexing until it reaches the serving index tokenized and ready to serve, takes about 1.5 hours on average, and that meta annotations such as robots meta tags are typically processed in 45 to 90 minutes, a critical step without which indexing cannot continue (said at the event, not in Google's docs). Rendering happens either right after crawling or later through a queue. Day 1's second recording added that indexing starts with parsing the fetched HTML so that elements such as the title can be extracted, and a second recording of Day 2 that the media indexer attaches each image and video to the URL of the page that hosts it (said at the event).
Day 1Day 2Day 3
36 claims · 6 sessions
Main content and page sections
Google extracts every element of a page to tell header, navigation and main content apart; its slides marked the main content very important and the header and navigation not so important, and its canonicalization guide says it determines each page's primary content, or centerpiece, when indexing. Gary Illyes said words are weighted by where they sit: footer text is unlikely to contribute much to ranking, and moving a term into the main content is the simplest way to make it count; Google's Search Essentials supports this only in general terms, advising that search words go in prominent places such as the title and main heading, and describes no weighting. Page structure also drives soft 404 detection, extending Day 1's soft 404 discussion: a BERT-like model trained on page layout ignores navigation and footer, so an error message alone in the main content makes a soft 404 while one in the navigation need not. Pages that translate only the menu and footer around the same main content are clustered as duplicates, as Google's hreflang guide also says. A Google software engineer working on data ingestion said extraction sometimes misjudges the main content or picks up extraneous data such as related products' prices, which structured data helps prevent, and Google said a 'Crawled – currently not indexed' page on a site of even quality may use a template that hides where its content is. A second recording captured Gary Illyes's part on main content in Day 2: main content is any part of a page that directly helps it achieve its purpose, not only text but images, videos or a tool, user-generated content on a UGC site, a comment section, content in tabs, and all headings and the visible title, matching Google's Search Quality Rater Guidelines, which add that tabs and comments can be main or supplementary content depending on the page's purpose. He said the main content is what Google considers when ranking a page, and that whatever a site puts in its navigation or header tells Google it does not particularly care about that content. Both recordings have the token metadata marking a word as in the header, the main content, bold or a heading (a third item, heard as title in one recording and italics in the other, is left out). Author’s view: a logical heading hierarchy is not a Google Search requirement, since Google's SEO Starter Guide says out-of-order headings do not matter to Search; the case for it is accessibility and agents that read the accessibility tree.
Day 1Day 2
23 claims · 4 sessions
Tokenization: how text is stored
Google said, in detail not in its documentation, that the Search index holds neither full pages nor sentences but tokens, the smallest searchable units (words, for languages written with spaces), each stored with its position and metadata, with posting lists of the URLs that contain most tokens. The metadata records where a word appeared (header, main content or 'centerpiece', bold, a heading; a further item, heard as title in one recording and italics in another, is left open) and spam signals such as white-on-white text, and snippets are rebuilt from the stored token positions. Thai, Chinese and other languages written without spaces are segmented with statistical models built from web content in that language, and queries go through the same segmenter so they match the index; retrieval then looks up only the query's important words. Day 2 partly contradicts Day 1's slide that Gemini shares tokenization with Search: Gary Illyes said AI-model tokenization is different ('or mostly'), and his slides showed Search keeping 'robots.txt' and 'tl;dr' whole where the AI tokenizer split them into sub-word pieces, as Google's Gemini documentation describes for long words (the slide also showed each token's numeric ID). Author’s view: read it as a shared step with different outputs, and treat leftover hidden keyword blocks as a liability, since spam metadata is stored with the tokens. Day 3 confirmed the mirror on the query side: Google transforms a query into something that can be matched against the index, removing stop words as part of that, while a phrase whose stop words matter is recognised as a whole and its words are indexed together (said at the event, not in Google's docs). Author’s view: a community speaker's picture of Googlebot seeing a page as ones and zeros blends two steps, since crawling fetches the page and tokenization happens later, when it is processed.
Day 2Day 3
56 claims · 7 sessions
Canonical selection
Some duplicate content is normal and is no spam violation or penalty: Google clusters duplicate pages and picks one as canonical. For Google the canonical is the representative of a duplicate cluster, the URL it would ideally show, and Google makes its own choice because rel=canonical is often wrong, for example, apparently, a placeholder left in place of a URL. Google named three considerations: protection against hijacking across pages or sites, user experience (whether the page loads, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps); it said machine learning sets how much each criterion weighs and that the weighting changes over time, which is not in its docs. When the signals agree Google follows the site owner, and when they point in different directions it cannot tell what the owner wants; Google also said it forwards the signals of every duplicate to the URL it picks, which sits uneasily with a community speaker's advice to make the strongest page of a canonical group its leader. On stage rel=canonical was said to 'also help a bit' next to redirects and sitemaps, but Google's canonical guide rates redirects and rel=canonical as strong signals and sitemap inclusion as weak, and says Google prefers URLs that are part of hreflang clusters. A permanent redirect matters only for which URL becomes canonical, not for clustering, and with a 302 Google shows the source URL; Google also said pointing paginated pages' canonical to page 1 can sometimes make sense, which its pagination guidance still advises against. Day 3 gave timings: a change of canonical URL usually shows within one to three weeks, though it can happen in seconds, and Google's guide to canonicalization issues says Google may hold pages in a duplicate cluster for up to two weeks after content issues are fixed. Google also described a site move as a complex canonicalization in which every signal of the old site is recalculated and moved to the new one. Day 1's second recording added Cherry Prommawin's definition: Google clusters duplicates and selects one page per cluster as its representative, the canonical, so users are not shown duplicates. A community speaker advised that a web application compute the expected URL for every request and redirect, or return an error, when the requested URL differs; on Day 2 the same speaker added that under the canonical link specification an improperly declared canonical can be ignored completely, not only replaced by the processor's own heuristic. Author’s view: of the two answers, Google's canonicalization guide favours the redirect, a strong canonical signal, while an error page throws away the links pointing at the variant.
Day 1Day 2Day 3
41 claims · 2 sessions
Canonical graphs: chains, loops and leaders
A community talk on Day 2 argued for auditing canonicals as whole groups rather than page pairs: a canonical group is every URL connected through canonical links, and its leader is the URL they ultimately point to. The speaker's problem cases were chains (also through server-side or client-side redirects), loops with no leader, several canonicals on one page, a leader that is noindexed, blocked or returns an error, a leader reachable only through a canonical link, and canonicals between language versions, which hreflang should connect instead. Google's documentation is more lenient on chains: a 2009 post says a canonical may point to a URL that redirects and that Google can follow chains, while strongly recommending links to a single canonical, and a 2013 post says Google will likely ignore all canonicals on a page that declares more than one. Google's own advice on the day was to check rel=canonical links with a crawler and keep canonical signals clear, and its guide says that on hreflang pages the canonical should be in the same language. The speaker's view that the leader should be the group's strongest page sits uneasily with Google's statement that it forwards the signals of all duplicates to the representative URL. In a second recording the speaker added that, under the canonical link specification, an improperly declared canonical can also be ignored completely by the application that processes it, not only replaced by its own heuristic.
Day 2
85 claims · 9 sessions
Redirects, site moves and alternate names
Google treats a site migration as deduplication across sites: redirects are trusted strongly for clustering, and whether a redirect is permanent or temporary only decides which URL becomes canonical, so with a 302 Google keeps showing the source URL. The URLs that lose are kept as 'alternate names', which is why a site: query for an old domain still lists old URLs after a migration and why people searching for an old brand can still see the old domain; Google's redirects guide calls this normal and says it fades. Google's closing advice on duplication included using redirects for site migrations and keeping canonical signals clear, and a community speaker showed a canonical chain running through a 301 to another subdomain. For a product out of stock for months, a Google Q&A slide said keeping or redirecting the page depends on how important it is to users, who may wait or pre-order; Google's guide to pausing a business recommends staying online and updating Product structured data with current availability. Google also said a site should not move to country-code domains just because a ccTLD is a strong country signal. Day 3 put a site move in time: Google treats it as a complex canonicalization in which every signal of the old site is recalculated and moved, so a move takes one to three months on average, a few weeks for a small site and up to about a year in the worst case, because Google's slowest signal is recalculated about once a year (said at the event, in a partly uncertain passage of the recording; Google's site move guide says a few weeks for most pages of a small to medium site and to keep redirects for at least a year). The second recordings added two migration case studies and Google's own example. In the Day 1 Q&A Google described consolidating a dozen or so language blogs, a Help Center and its old developers site into one site: the first priority was to identify the popular URLs and protect them, and duplicated language versions were merged by giving both old URLs one target path to redirect to. Asked how to plan a migration that does not leave many URLs unindexed, a panelist said the answer is probably not sitemaps but deciding what matters to the business, such as whether to consolidate languages; listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool. As a fun fact the panel said JavaScript was used for the language-consolidation redirects in Google's own migration, because it was the only option available to the person doing it (said at the event; the recording does not make fully clear how it was used). Author’s view: Google's redirects guide says Google Search follows JavaScript redirects only after rendering and recommends them only when server-side or meta refresh redirects are impossible, and its site move guide asks for server-side 301 or 308 redirects, so Google's own JavaScript was a fallback, not a pattern to copy. Author’s view: this matches Google's November 2020 move to Google Search Central, about six years before the event rather than the four or five said on stage. Cherry Prommawin said Google follows permanent and temporary redirects and indexes whatever the target returns (Google's docs: up to 10 hops by default), and Day 2's opening Q&A added that several sloppy migrations on one domain cause many short-term effects on indexing as well as crawling, and contrasted out-of-stock products buyers will wait for with easily replaced ones they simply swap. In Day 2's Lightning session E a community speaker merged two competing loan-comparison sites into one brand: every URL was given a keep, merge or remove decision, an old URL got a 301 only when a new page served the same intent, removed URLs were never redirected to the homepage, and new sitemaps, a Change of Address in Search Console and updated internal links followed, as Google's site move guide advises; the speaker reported traffic growth from the first day and revenue up 100% month over month (the speaker's figures). A second community speaker ran a post-acquisition migration around one redirect map giving every old URL an approved destination and an owner, froze the inventory (export, crawl, sitemap, search data, logs) before launch, and after it checked each URL's route, content and lead rather than its 200 status; he cited an unnamed study's median recovery time of 304 days but said broken routes need action at once. Author’s view: return 404 or 410 for removed content rather than redirecting it to an unrelated page, and report indexing and traffic recovery separately, since Google's guide measures the switch of indexed URLs in weeks while traffic can take far longer.
Day 1Day 2Day 3
10 claims · 1 session
Pruning and consolidating content
In Day 2's lightning session on duplicates and site moves, a community speaker presented the consolidation of two competing loan-comparison sites into one brand. Every URL of both domains was compared on traffic, conversion, backlinks, revenue and rankings and given one decision: keep, merge or remove; pages that served the same search intent were merged into one strong article, taking the site from over 2,000 URLs to about 100. An old URL got a 301 redirect only when a new page served the same intent, and removed URLs were never redirected to the homepage. The speaker reported traffic growth from the first day and revenue up 100% month over month (the speaker's figures). A second speaker advised agreeing what to keep and retire with an owner for every decision, reviewing it with data and prioritising by business value, not only traffic. Author’s view: return 404 or 410 for removed URLs, as Google's site move guide says; for a site of about 2,000 URLs the crawl-budget gain is likely minor, and the benefit of pruning more plausibly comes from one strong URL per intent and consolidated signals.
Day 2
Across days 55
- Stage D2-C325 Day 2 · Understanding what's on a page
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
contradictsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C393 Day 2 · Handling web duplication
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
contradictsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C092 Day 1 · How Google thinks about crawl budgetHostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
- D2-C025 Day 2 · How is HTML interpreted
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- D2-C025 Day 2 · How is HTML interpreted
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
extendsDocs D1-C086 Day 1 · How Google interprets robots.txtGoogle-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
- D2-C025 Day 2 · How is HTML interpreted
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
extendsStage D1-C522 Day 1 · How Google interprets robots.txtGoogle said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
- D2-C026 Day 2 · How is HTML interpreted
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
extendsD1-C036 Day 1 · How Search works and where's AI?Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- D2-C026 Day 2 · How is HTML interpreted
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C031 Day 2 · How is HTML interpreted
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- D2-C048 Day 2 · How is HTML interpreted
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C127 Day 2 · Lightning session D: Rendering and JavaScript
A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C167 Day 2 · Lightning session D: Rendering and JavaScript
The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C168 Day 2 · Lightning session D: Rendering and JavaScript
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Docs D2-C169 Day 2 · Lightning session D: Rendering and JavaScript
Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Stage D2-C338 Day 2 · Understanding what's on a page
Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.
extendsD1-C042 Day 1 · How Search works and where's AI?BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.
- Stage D2-C339 Day 2 · Understanding what's on a page
For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C344 Day 2 · Handling web duplication
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- D2-C345 Day 2 · Handling web duplication
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- D2-C345 Day 2 · Handling web duplication
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
extendsStage D1-C207 Day 1 · How Search works and where's AI?Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.
- Stage D2-C346 Day 2 · Handling web duplication
For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.
extendsStage D1-C207 Day 1 · How Search works and where's AI?Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.
- D2-C348 Day 2 · Handling web duplication
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
extendsStage D1-C111 Day 1 · session not recordedGary Illyes said there is no such thing as a duplicate content penalty.
- Stage D2-C367 Day 2 · Handling web duplication
Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C369 Day 2 · Handling web duplication
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
extendsD1-C094 Day 1 · How Google thinks about crawl budgetIf the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
- D2-C369 Day 2 · Handling web duplication
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsStage D1-C069 Day 1 · How crawling errors affect SearchDNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsStage D1-C367 Day 1 · How crawling errors affect SearchCDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
- Stage D2-C375 Day 2 · Handling web duplication
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
extendsStage D1-C367 Day 1 · How crawling errors affect SearchCDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
- Docs D2-C376 Day 2 · Handling web duplication
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C380 Day 2 · Handling web duplication
Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.
extendsAnalysis D1-C113 Day 1 · session not recordedThe real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
- Docs D2-C395 Day 2 · Handling web duplication
Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Stage D2-C401 Day 2 · Lightning session E: Managing Duplicates and Site Moves
Broken canonical tags can make the wrong pages of a site show up in search results.
extendsAnalysis D1-C113 Day 1 · session not recordedThe real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
- Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C427 Day 2 · Lightning session E: Managing Duplicates and Site Moves
When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.
extendsAnalysis D1-C067 Day 1 · How crawling worksA page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
- D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C688 Day 2 · Deciding what goes in the index?
Index selection is the last step before documents enter Google's index.
extendsStage D1-C211 Day 1 · How Search works and where's AI?Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
- Stage D2-C701 Day 2 · Deciding what goes in the index?
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C884 Day 2 · Lightning session E: Managing Duplicates and Site Moves
The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.
extendsStage D1-C540 Day 1 · Q&AAsked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.
- Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.
extendsD1-C103 Day 1 · How Google thinks about crawl budgetFour ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.
- Stage D3-C011 Day 3 · Making sense of users' queries
Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.
extendsStage D2-C320 Day 2 · Understanding what's on a pageText in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
- Stage D3-C013 Day 3 · Making sense of users' queries
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
extendsStage D2-C321 Day 2 · Understanding what's on a pageFor languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
- Stage D3-C013 Day 3 · Making sense of users' queries
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
extendsStage D2-C737 Day 2 · How does the index look like?A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
- Stage D3-C075 Day 3 · Making sense of users' queries
At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.
extendsStage D2-C738 Day 2 · How does the index look like?At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.
- Analysis D3-C109 Day 3 · Lightning session K: Facets of quality
The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.
extendsStage D2-C318 Day 2 · Understanding what's on a pageGoogle does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
- Stage D3-C170 Day 3 · How Google thinks about Quality
Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.
extendsStage D2-C311 Day 2 · Understanding what's on a pageGary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.
- Stage D3-C316 Day 3 · How Search results are born
Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.
extendsStage D2-C724 Day 2 · How does the index look like?The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.
- Stage D3-C323 Day 3 · How Search results are born
Most of Google's search features need nothing extra from the site owner; Google generates them from what it extracted from the page during indexing.
extendsStage D2-C446 Day 2 · Finding the gold nuggets: structured data, media, and more!The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.
- Stage D3-C626 Day 3 · How long does it take to..?
Google renders pages in two ways: immediately after crawling, or later through a queue-based process that runs elsewhere.
extendsD2-C170 Day 2 · Lightning session D: Rendering and JavaScriptAfter processing, an indexable page is placed in Google's render queue to wait for rendering.
- Stage D3-C642 Day 3 · How long does it take to..?
Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.
extendsStage D2-C351 Day 2 · Handling web duplicationGoogle treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.
- Analysis D3-C676 Day 3 · How long does it take to..?
The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.
extendsDocs D2-C131 Day 2 · Lightning session D: Rendering and JavaScriptGoogle's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.
- D2-C048 Day 2 · How is HTML interpreted
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
repeatsStage D1-C328 Day 1 · How crawling worksDuring indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.
- D3-C070 Day 3 · Making sense of users' queries
Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.
repeatsStage D2-C320 Day 2 · Understanding what's on a pageText in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
- Stage D3-C074 Day 3 · Making sense of users' queries
Google's index uses posting lists: for each word, a list of the URLs associated with that word.
repeatsStage D2-C733 Day 2 · How does the index look like?For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.