Connections
Across days
When a later talk repeats, extends, contradicts or updates an earlier one, the two claims are linked. Questions from the audience are linked to their answers.
Contradicts 5 Updates 1 Extends 393 Repeats 85 Answers 70
Contradicts 5
A claim that disagrees with an earlier one. The later claim is on the left.
- Stage D1-C296 Day 1 · Lightning session A: Automation and AI
A community speaker argued for adopting the GEO label as the industry's chance to leave behind the bad reputation SEO built, unlike Google's view earlier the same day that the new name is not needed.
contradictsStage D1-C049 Day 1 · What's new in the world of SearchGary Illyes argued that GEO is a label invented to create a new field and is not needed. Understanding how SEO works is enough.
- Stage D1-C296 Day 1 · Lightning session A: Automation and AI
A community speaker argued for adopting the GEO label as the industry's chance to leave behind the bad reputation SEO built, unlike Google's view earlier the same day that the new name is not needed.
contradictsD1-C050 Day 1 · How Search works and where's AI?Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
- Stage D2-C325 Day 2 · Understanding what's on a page
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
contradictsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C393 Day 2 · Handling web duplication
Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.
contradictsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Stage D2-C429 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.
contradictsD2-C348 Day 2 · Handling web duplicationGoogle forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
Updates 1
A claim that replaces or revises an earlier one. The later claim is on the left.
- Docs D1-C125 Day 1 · What's new in the world of Search
Since 16 September 2026, US publishers and creators qualify for a Search profile with 10,000 followers across YouTube, Instagram, X or TikTok, and media organisations can claim and manage profiles for all their sub-brands from one login.
updatesDocs D1-C026 Day 1 · What's new in the world of SearchSearch profiles launched on 4 June 2026, in the US first, for creators and publishers with a sizable following on at least one major social or video platform. They appear in knowledge panels and Discover.
Extends 393
A claim that adds detail to an earlier one. The later claim is on the left.
- D1-C066 Day 1 · How Google thinks about crawl budget
Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.
extendsStage D1-C326 Day 1 · How crawling worksGoogle does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.
- Stage D1-C182 Day 1 · What's new in the world of Search
Gary Illyes said that in about 30 years of web search, helping users complete their journeys has made them search more, not less, unlike a product whose problem is solved once and for all.
extendsD1-C001 Day 1 · Welcome and opening keynotesThe keynote showed a slide quoting Elizabeth Reid (VP, Search) saying Search is never a solved problem because the internet and the world keep changing.
- Stage D1-C184 Day 1 · What's new in the world of Search
Google said AI Mode is growing almost exponentially and that people are searching more because of it, not less.
extendsD1-C004 Day 1 · Welcome and opening keynotesGoogle says users like AI Overviews and search more because of them, with early feedback overwhelmingly positive and satisfaction highest among 18–24 year olds.
- Stage D1-C184 Day 1 · What's new in the world of Search
Google said AI Mode is growing almost exponentially and that people are searching more because of it, not less.
extendsD1-C016 Day 1 · What's new in the world of SearchAI Mode queries have doubled every quarter since launch.
- Stage D1-C185 Day 1 · What's new in the world of Search
Google said a multi-part query, such as a white three-row SUV for two teenagers and a car-sick dog, cannot be satisfied by a list of links because no single article covers every aspect, while AI Overviews and AI Mode can synthesise an answer that helps users decide where to go.
extendsD1-C021 Day 1 · What's new in the world of SearchQueries of five or more words are growing in volume 1.5 times faster than shorter queries. The slide footnote cites Google internal data on global English-language queries, for a comparison period that starts in November 2022 (the remaining dates are only partly legible).
- Stage D1-C187 Day 1 · What's new in the world of Search
Google presented the intelligent search box as one of its recent Search launches; it accepts long queries and keeps expanding to make room for them.
extendsStage D1-C150 Day 1 · Welcome and opening keynotesGoogle said it has made the biggest change to its search box in the last 25 years.
- Stage D1-C189 Day 1 · What's new in the world of Search
AI Overviews and AI Mode work together: a user can start with an AI Overview and continue the conversation in AI Mode.
extendsStage D1-C152 Day 1 · Welcome and opening keynotesGoogle said AI Overviews and AI Mode are integrated into one seamless AI experience in Search.
- Stage D1-C192 Day 1 · What's new in the world of Search
Google said Search agents in AI Mode had launched globally a day or two before 30 September 2026.
extendsD1-C022 Day 1 · What's new in the world of SearchSearch Agents in AI Mode let users set standing requests for updates, such as being told when a product goes on sale, a team wins or a price drops below a level.
- Stage D1-C194 Day 1 · What's new in the world of Search
A Search agent is set up by typing a natural-language request, which AI Mode interprets to create an alert; Google called this an information agent.
extendsD1-C022 Day 1 · What's new in the world of SearchSearch Agents in AI Mode let users set standing requests for updates, such as being told when a product goes on sale, a team wins or a price drops below a level.
- Stage D1-C198 Day 1 · What's new in the world of Search
Gary Illyes said SEO is not dead and that the new AI features only create more opportunities for site owners to take advantage of.
extendsD1-C050 Day 1 · How Search works and where's AI?Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
- Stage D1-C199 Day 1 · What's new in the world of Search
Gary Illyes said searching by image or by video, as Search offers it today, was not possible before transformers were invented.
extendsStage D1-C161 Day 1 · Welcome and opening keynotesGoogle said it created the Transformer architecture (the paper 'Attention Is All You Need', the T in ChatGPT) and published it to move the industry forward, competitors included.
- Stage D1-C202 Day 1 · How Search works and where's AI?
Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Stage D1-C205 Day 1 · How Search works and where's AI?
Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D1-C208 Day 1 · How Search works and where's AI?
AI Overviews and AI Mode may have extra processes of their own, like any other search feature, but the bulk of their processing is the same as for normal Search.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D1-C209 Day 1 · How Search works and where's AI?
Google launched the 'Did you mean' feature around 2001-2002 using a statistical model, which Gary Illyes counted as AI because it is a form of machine learning.
extendsD1-C041 Day 1 · How Search works and where's AI?Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.
- Stage D1-C211 Day 1 · How Search works and where's AI?
Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D1-C214 Day 1 · How Search works and where's AI?
Crawling and indexing happen before anyone searches, while serving and ranking happen in real time when someone types a query.
extendsD1-C036 Day 1 · How Search works and where's AI?Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- Stage D1-C217 Day 1 · How Search works and where's AI?
Google began talking publicly about its use of AI in Search around 2015-2016, and RankBrain was the first such system it publicised widely.
extendsD1-C043 Day 1 · How Search works and where's AI?RankBrain is used in serving to interpret the intent behind queries, especially new or unusual ones, and match them to relevant results.
- Stage D1-C218 Day 1 · How Search works and where's AI?
Google said MUM helps it understand the context of the words in a query: a search for hiking shoes that mentions Mount Everest rather than Kilimanjaro gets results suited to that climb.
extendsD1-C044 Day 1 · How Search works and where's AI?MUM (Multitask Unified Model) understands information across text, images, audio and video, and processes information in more than 75 languages.
- Stage D1-C220 Day 1 · How Search works and where's AI?
Google said the more important documents in its index naturally tend to be documents created by humans.
extendsD1-C047 Day 1 · How Search works and where's AI?Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so they understand and promote natural content better.
- Stage D1-C221 Day 1 · How Search works and where's AI?
Google said the natural content its ranking algorithms understand and promote better includes content created by humans or at least edited and reviewed by them.
extendsD1-C047 Day 1 · How Search works and where's AI?Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so they understand and promote natural content better.
- Stage D1-C225 Day 1 · How Search works and where's AI?
Google said query fan-out is nothing new: it fires, for example, ten different searches in the background, each a normal search on the same systems.
extendsD1-C051 Day 1 · How Search works and where's AI?Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.
- Stage D1-C226 Day 1 · How Search works and where's AI?
Google said it does not own third-party AI chatbots such as ChatGPT and has no insight into how they work or are built.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D1-C271 Day 1 · Lightning session A: Automation and AI
A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard as Ahrefs') that looked at almost 40,000 sites over one month: 97% had no AI agent hits on the file, and those that had any got about two hits a month.
extends - Stage D1-C278 Day 1 · Lightning session A: Automation and AI
A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.
extendsDocs D1-C131 Day 1 · session not recordedGoogle's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
- Stage D1-C279 Day 1 · Lightning session A: Automation and AI
A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.
extends - Analysis D1-C314 Day 1 · Lightning session A: Automation and AI
A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.
extendsStage D1-C277 Day 1 · Lightning session A: Automation and AIA community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.
- Analysis D1-C315 Day 1 · Lightning session A: Automation and AI
Google's AI optimisation guide says structured data is not required for generative AI search and needs no special schema.org markup, and on Day 2 Google said raw schema.org is generally not put into model context (D2-C477); use markup for rich-result eligibility and clear data, not as an AI-visibility lever.
extendsStage D1-C276 Day 1 · Lightning session A: Automation and AIA community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which number is the price saves the agent from guessing.
- Docs D1-C316 Day 1 · Lightning session A: Automation and AI
Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.
extendsStage D1-C282 Day 1 · Lightning session A: Automation and AIA community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.
- Stage D1-C325 Day 1 · How crawling works
Googlebot is the crawler Google uses for web search, including Search's AI features.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D1-C335 Day 1 · How crawling works
Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D1-C336 Day 1 · How crawling works
Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Stage D1-C359 Day 1 · How crawling errors affect Search
Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.
extendsStage D1-C069 Day 1 · How crawling errors affect SearchDNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
- Stage D1-C373 Day 1 · How Google thinks about crawl budget
Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it determines how many of the site's pages are discovered and how often they are revisited.
extendsStage D1-C096 Day 1 · How Google thinks about crawl budgetCrawl budget was described as the attention span Google gives a website, and better performance increases it.
- Stage D1-C374 Day 1 · How Google thinks about crawl budget
Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.
extendsD1-C091 Day 1 · How Google thinks about crawl budgetCrawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
- Stage D1-C375 Day 1 · How Google thinks about crawl budget
Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.
extendsD1-C066 Day 1 · How Google thinks about crawl budgetGoogle Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.
- Stage D1-C376 Day 1 · How Google thinks about crawl budget
Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.
extendsD1-C091 Day 1 · How Google thinks about crawl budgetCrawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
- Stage D1-C377 Day 1 · How Google thinks about crawl budget
The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D1-C378 Day 1 · How Google thinks about crawl budget
Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.
extendsD1-C099 Day 1 · How Google thinks about crawl budgetThree things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.
- Stage D1-C379 Day 1 · How Google thinks about crawl budget
The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.
extendsD1-C103 Day 1 · How Google thinks about crawl budgetFour ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.
- Stage D1-C388 Day 1 · How Google thinks about crawl budget
AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D1-C404 Day 1 · Lightning session C: Crawling
For query parameters, a community speaker advised checking that every parameter in a request is actually used and in the expected order, and otherwise redirecting to the expected URL with only the used parameters in the correct order.
extendsDocs D1-C102 Day 1 · How Google thinks about crawl budgetIf faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return a 404 when a filter combination has no results.
- Stage D1-C412 Day 1 · Lightning session C: Crawling
A community speaker described AI search as turning one query into many related searches, including searches in other languages, before the information is retrieved.
extendsDocs D1-C053 Day 1 · How Search works and where's AI?Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.
- Analysis D1-C433 Day 1 · Q&A
The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.
extendsStage D1-C432 Day 1 · Q&AA Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).
- Stage D1-C439 Day 1 · Q&A
Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D1-C444 Day 1 · Q&A
Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.
extendsStage D1-C397 Day 1 · Lightning session C: CrawlingA community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.
- Analysis D1-C470 Day 1 · Q&A
Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.
extends - Analysis D1-C480 Day 1 · Q&A
The consolidation described on stage matches Google's November 2020 move from Google Webmasters to Google Search Central, which moved its main blog, 13 localized blogs and the Search help content of the Search Console Help Center to its developer site: about six years before the event, not the four or five said on stage.
extends - Analysis D1-C498 Day 1 · Welcome and opening keynotes
Cite YouTube's lead as US streaming watch time measured by Nielsen: that is the figure YouTube publishes, and no published Google source backs the keynote's 'around the world'.
extendsStage D1-C497 Day 1 · Welcome and opening keynotesGoogle said YouTube has surpassed Netflix as the number one streaming service, in the US and around the world.
- Stage D1-C499 Day 1 · Welcome and opening keynotes
The keynote added that people who also use a standalone LLM keep certain types of use anchored on Google (the example given is unclear in both recordings).
extendsD1-C007 Day 1 · Welcome and opening keynotesGoogle says 94% of frequent LLM users are also frequent users of Google Search, as evidence that AI is growing search use rather than replacing it.
- Stage D1-C500 Day 1 · Lightning session A: Automation and AI
A community speaker applied Pascal's line about makers of false windows built for symmetry, whose rule is to make pleasing figures rather than to speak accurately, to today's LLMs: sometimes they are not trying to produce the most accurate figures or the most accurate interpretations of them.
extendsStage D1-C242 Day 1 · Lightning session A: Automation and AIA community speaker reported that a code refactor of his company's AI analysis product left data fetching correct but made the analysis shallow while the answers still read well, with the same model, so polished output is no proof of correct analysis.
- Stage D1-C501 Day 1 · Lightning session A: Automation and AI
A community demo's script used both input modes of Google's Rich Results Test: the URL mode for public pages, which Google fetches itself, and the code mode, into which the script pasted the page's HTML, for private pages and pages Google cannot fetch (best reading of a largely unintelligible recording; the second kind of page was heard as 'dead', possibly 'dev').
extendsStage D1-C249 Day 1 · Lightning session A: Automation and AIA community demo showed a script that opens Chrome with a saved session, pastes a URL into Google's Rich Results Test, waits about 15 seconds, then screenshots and saves the result and retries on failure (the demo's recording is largely unintelligible; the tool name is a best reading).
- Stage D1-C502 Day 1 · Lightning session A: Automation and AI
In a community demo, the structured-data problems Google's Rich Results Test reported on the test page included an empty name, a breadcrumb problem and a price of zero, with errors shown in pink and warnings in orange (best readings of a largely unintelligible recording; each item is heard in only one of the two recordings).
extendsStage D1-C251 Day 1 · Lightning session A: Automation and AIA community demo treated warnings in Google's structured-data test as not critical, in an example result of 22 items with two errors and two warnings.
- Stage D1-C503 Day 1 · Lightning session A: Automation and AI
In a community demo, the small model choosing keywords (Claude Haiku) followed fixed rules, no brand terms, no navigation terms and at most two keywords per page; the Spanish example was 'zapatos de mujer' and 'comprar zapatos de mujer' (women's shoes, buy women's shoes).
extendsStage D1-C252 Day 1 · Lightning session A: Automation and AIA community demo matched model size to the task: a small, cheap model (Claude Haiku) chose keywords from a page's URL, title, H1, description and schema types, while a large, expensive model (Claude Opus) fixed code and wrote the summary.
- Analysis D1-C504 Day 1 · Lightning session A: Automation and AI
The study cited on stage matches Peec AI's analysis (reported in February 2026) of over 10 million ChatGPT prompts and 20 million fan-outs: 43% of the fan-out searches for non-English prompts ran in English, and nearly 78% of non-English prompt runs had at least one English fan-out (from 66% for Spanish to 94% for Turkish), so the 43% is a share of fan-out searches, not of prompts.
extendsStage D1-C294 Day 1 · Lightning session A: Automation and AIA community speaker cited a third-party study (heard as Peec AI's) finding that 43% of prompts produced fan-out queries in a language other than the one the user searched in.
- Analysis D1-C505 Day 1 · Lightning session A: Automation and AI
The llms.txt study cited on stage matches Ahrefs' June 2026 study of 137,210 domains: 28% (about 38,000, the 'almost 40,000' sites heard on stage) published an llms.txt file and 97% of those files received no requests at all in May 2026; the files that were fetched got about 22,000 requests across some 1,100 domains, around 20 each, not the 'two hits a month' heard on stage.
extendsStage D1-C271 Day 1 · Lightning session A: Automation and AIA community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard as Ahrefs') that looked at almost 40,000 sites over one month: 97% had no AI agent hits on the file, and those that had any got about two hits a month.
- Docs D1-C506 Day 1 · Lightning session A: Automation and AI
Google's Rich Results Test help says the tool tests either a page's full URL or a pasted code snippet (Code instead of URL); all page resources must be reachable by an anonymous user on the internet, so resources behind a firewall or a password are not available to the test unless exposed, for example through a tunnel.
extendsStage D1-C249 Day 1 · Lightning session A: Automation and AIA community demo showed a script that opens Chrome with a saved session, pastes a URL into Google's Rich Results Test, waits about 15 seconds, then screenshots and saves the result and retries on failure (the demo's recording is largely unintelligible; the tool name is a best reading).
- Stage D1-C508 Day 1 · How crawling errors affect Search
Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.
extendsStage D1-C367 Day 1 · How crawling errors affect SearchCDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
- Stage D1-C508 Day 1 · How crawling errors affect Search
Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.
extendsStage D1-C368 Day 1 · How crawling errors affect SearchSearch Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
- Stage D1-C509 Day 1 · How crawling errors affect Search
Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.
extendsStage D1-C365 Day 1 · How crawling errors affect SearchGoogle has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.
- Stage D1-C509 Day 1 · How crawling errors affect Search
Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.
extendsStage D1-C366 Day 1 · How crawling errors affect SearchGoogle treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
- Stage D1-C512 Day 1 · How Google interprets robots.txt
Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
extendsStage D1-C329 Day 1 · How crawling worksApart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.
- Stage D1-C513 Day 1 · How Google interprets robots.txt
The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.
extendsDocs D1-C086 Day 1 · How Google interprets robots.txtGoogle-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
- Stage D1-C516 Day 1 · How Google interprets robots.txt
Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.
extendsDocs D1-C084 Day 1 · How Google interprets robots.txtGoogle supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
- Stage D1-C519 Day 1 · How Google interprets robots.txt
A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.
extendsDocs D1-C080 Day 1 · How Google interprets robots.txtRules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.
- Stage D1-C520 Day 1 · How Google interprets robots.txt
In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.
extendsD1-C079 Day 1 · How Google interprets robots.txtThe robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
- Stage D1-C520 Day 1 · How Google interprets robots.txt
In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.
extendsDocs D1-C081 Day 1 · How Google interprets robots.txtWhen matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.
- Stage D1-C521 Day 1 · How Google interprets robots.txt
To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.
extendsD1-C079 Day 1 · How Google interprets robots.txtThe robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
- Stage D1-C521 Day 1 · How Google interprets robots.txt
To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.
extendsDocs D1-C082 Day 1 · How Google interprets robots.txtIn robots.txt, * matches zero or more of any character and $ marks the end of the URL.
- Stage D1-C522 Day 1 · How Google interprets robots.txt
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
extendsDocs D1-C086 Day 1 · How Google interprets robots.txtGoogle-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
- Analysis D1-C523 Day 1 · How Google interprets robots.txt
The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.
extendsDocs D1-C086 Day 1 · How Google interprets robots.txtGoogle-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
- Stage D1-C526 Day 1 · How Google interprets robots.txt
The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).
extendsD1-C087 Day 1 · How Google interprets robots.txtGooglebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.
- Stage D1-C526 Day 1 · How Google interprets robots.txt
The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).
extendsD1-C123 Day 1 · How Google interprets robots.txtBing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.
- Stage D1-C527 Day 1 · How Google interprets robots.txt
The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.
extendsDocs D1-C080 Day 1 · How Google interprets robots.txtRules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.
- Stage D1-C528 Day 1 · How Google interprets robots.txt
Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.
extendsDocs D1-C140 Day 1 · How Google interprets robots.txtSearch Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.
- Stage D1-C530 Day 1 · How Google interprets robots.txt
Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
extendsDocs D1-C140 Day 1 · How Google interprets robots.txtSearch Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.
- Stage D1-C531 Day 1 · How Google interprets robots.txt
In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.
extendsDocs D1-C080 Day 1 · How Google interprets robots.txtRules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.
- Analysis D1-C532 Day 1 · How Google interprets robots.txt
A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.
extendsStage D1-C525 Day 1 · How Google interprets robots.txtGoogle called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.
- Stage D1-C533 Day 1 · Lightning session B: Robots.txt
Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.
extendsDocs D1-C080 Day 1 · How Google interprets robots.txtRules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.
- Stage D1-C538 Day 1 · Q&A
A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.
extendsStage D1-C334 Day 1 · How crawling worksDifferent Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.
- Stage D1-C538 Day 1 · Q&A
A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.
extends - Stage D1-C539 Day 1 · Q&A
A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).
extends - Stage D1-C539 Day 1 · Q&A
A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).
extends - Stage D1-C541 Day 1 · Q&A
A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in Google's own site migration, because it was the only option available to the person doing it (the recording does not make fully clear whether the JavaScript performed the redirects or built the mapping).
extends - Analysis D1-C542 Day 1 · Q&A
Google's redirects guide says Google Search follows JavaScript location redirects only after rendering, may never see one if rendering fails, and should be used only when server-side or meta refresh redirects are impossible; its site move guide asks for server-side permanent redirects (301 or 308) where technically possible. The JavaScript in Google's own migration was a fallback ('my only option'), not a pattern to copy.
extendsStage D1-C541 Day 1 · Q&AA Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in Google's own site migration, because it was the only option available to the person doing it (the recording does not make fully clear whether the JavaScript performed the redirects or built the mapping).
- Stage D1-C543 Day 1 · Q&A
Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).
extends - Stage D1-C544 Day 1 · Q&A
John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering new pages, that is a lot of discovery crawling, which suggests that crawling of new pages is not the problem.
extends - Stage D1-C546 Day 1 · Q&A
A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).
extendsStage D1-C374 Day 1 · How Google thinks about crawl budgetCrawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.
- Analysis D1-C547 Day 1 · Q&A
Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the total time a server spends holding connections open for Google, counting both the number of parallel connections and their duration. That fits the panel's advice to watch how many connections Googlebot opens (D1-C494): fewer connections is the documented form of a lower capacity limit.
extends - D2-C017 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
extendsD1-C079 Day 1 · How Google interprets robots.txtThe robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
- D2-C017 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
extendsDocs D1-C084 Day 1 · How Google interprets robots.txtGoogle supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
- D2-C020 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- D2-C021 Day 2 · Welcome to indexing day!
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
extendsD1-C105 Day 1 · How Google thinks about crawl budgetURLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
- Docs D2-C022 Day 2 · Welcome to indexing day!
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
extendsD1-C106 Day 1 · How Google thinks about crawl budgetThe noindex rule consumes crawl budget, because Google must fetch the page to see it.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- D2-C024 Day 2 · How is HTML interpreted
A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.
extendsD1-C092 Day 1 · How Google thinks about crawl budgetHostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.
- D2-C025 Day 2 · How is HTML interpreted
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- D2-C025 Day 2 · How is HTML interpreted
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
extendsDocs D1-C086 Day 1 · How Google interprets robots.txtGoogle-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
- D2-C025 Day 2 · How is HTML interpreted
In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.
extendsStage D1-C522 Day 1 · How Google interprets robots.txtGoogle said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
- D2-C026 Day 2 · How is HTML interpreted
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
extendsD1-C036 Day 1 · How Search works and where's AI?Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- D2-C026 Day 2 · How is HTML interpreted
A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C031 Day 2 · How is HTML interpreted
Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C038 Day 2 · How is HTML interpreted
Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- D2-C048 Day 2 · How is HTML interpreted
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C058 Day 2 · Controlling indexing
John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Analysis D2-C067 Day 2 · Controlling indexing
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
extendsDocs D1-C127 Day 1 · How Google thinks about crawl budgetGoogle treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.
- Analysis D2-C067 Day 2 · Controlling indexing
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
extends - Docs D2-C069 Day 2 · Controlling indexing
Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.
extendsD1-C106 Day 1 · How Google thinks about crawl budgetThe noindex rule consumes crawl budget, because Google must fetch the page to see it.
- D2-C072 Day 2 · Controlling indexing
The nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI Mode in Search.
extendsAnalysis D1-C034 Day 1 · What's new in the world of SearchLead-generation, local-service and e-commerce sites should normally stay included, because AI answers cite and link sources. Sites selling paywalled or licensed content should test exclusion on a child property first. Use nosnippet only if you also want out of regular snippets.
- Stage D2-C073 Day 2 · Controlling indexing
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Docs D2-C074 Day 2 · Controlling indexing
Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- D2-C103 Day 2 · Controlling indexing
Setting the Search generative AI control in Search Console to exclude keeps a site's links and content out of Search generative AI features such as AI Overviews and AI Mode, so the site gets no traffic or impressions from them; include is the default.
extendsD1-C029 Day 1 · What's new in the world of SearchSearch Console has a property setting called Search generative AI that gives direct control over AI Overviews and AI Mode without affecting web rankings. Its states are inherit from parent, include and exclude.
- Stage D2-C118 Day 2 · Lightning session D: Rendering and JavaScript
To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C120 Day 2 · Lightning session D: Rendering and JavaScript
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Docs D2-C121 Day 2 · Lightning session D: Rendering and JavaScript
Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.
extendsDocs D1-C086 Day 1 · How Google interprets robots.txtGoogle-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
- Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
extends - Stage D2-C124 Day 2 · Lightning session D: Rendering and JavaScript
Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.
extendsD1-C056 Day 1 · How Search works and where's AI?Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.
- D2-C127 Day 2 · Lightning session D: Rendering and JavaScript
A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C128 Day 2 · Lightning session D: Rendering and JavaScript
Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached system, kept separate because rendering is time-consuming and computationally expensive.
extendsD2-C026 Day 2 · How is HTML interpretedA Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
- Stage D2-C134 Day 2 · Lightning session D: Rendering and JavaScript
Google urged sites to make sure their JavaScript content can be crawled, rendered and indexed, calling this important today and also tomorrow, as AI systems increasingly ground answers to user requests.
extendsD1-C056 Day 1 · How Search works and where's AI?Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.
- Stage D2-C136 Day 2 · Lightning session D: Rendering and JavaScript
Many SEOs still treat everything beyond the raw HTML as the developers' business, but developers often do not handle rendering problems, a community speaker warned.
extends - Stage D2-C139 Day 2 · Lightning session D: Rendering and JavaScript
A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.
extendsD1-C056 Day 1 · How Search works and where's AI?Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.
- D2-C167 Day 2 · Lightning session D: Rendering and JavaScript
The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- D2-C168 Day 2 · Lightning session D: Rendering and JavaScript
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- D2-C168 Day 2 · Lightning session D: Rendering and JavaScript
In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.
extendsD2-C048 Day 2 · How is HTML interpretedLinks extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
- Docs D2-C169 Day 2 · Lightning session D: Rendering and JavaScript
Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Stage D2-C173 Day 2 · Lightning session D: Rendering and JavaScript
Google's renderer is a headless Chromium that runs the page's JavaScript.
extendsStage D1-C206 Day 1 · How Search works and where's AI?Google renders JavaScript-heavy pages from their HTML, CSS and JavaScript as a browser would, using the latest version of Chromium.
- D2-C184 Day 2 · Lightning session D: Rendering and JavaScript
Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C184 Day 2 · Lightning session D: Rendering and JavaScript
Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.
extendsD2-C040 Day 2 · How is HTML interpretedGoogle cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
- D2-C187 Day 2 · Lightning session D: Rendering and JavaScript
A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Analysis D2-C188 Day 2 · Lightning session D: Rendering and JavaScript
Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.
extendsDocs D1-C101 Day 1 · How Google thinks about crawl budgetGoogle's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.
- Analysis D2-C189 Day 2 · Lightning session D: Rendering and JavaScript
Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.
extendsAnalysis D1-C067 Day 1 · How crawling worksA page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
- Stage D2-C190 Day 2 · Lightning session D: Rendering and JavaScript
In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C195 Day 2 · Lightning session D: Rendering and JavaScript
Content that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C197 Day 2 · Lightning session D: Rendering and JavaScript
Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.
extends - D2-C201 Day 2 · Lightning session D: Rendering and JavaScript
Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () => fetch('/api/specs'), does not exist for Google until someone clicks.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C204 Day 2 · Lightning session D: Rendering and JavaScript
Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- Analysis D2-C208 Day 2 · Lightning session D: Rendering and JavaScript
A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.
extendsDocs D1-C081 Day 1 · How Google interprets robots.txtWhen matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.
- Analysis D2-C210 Day 2 · Lightning session D: Rendering and JavaScript
An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.
extendsDocs D1-C074 Day 1 · How crawling errors affect SearchIf robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.
- Stage D2-C213 Day 2 · Lightning session D: Rendering and JavaScript
A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C259 Day 2 · What is Google friendly JavaScript
Erin Sparling named server-side or hybrid rendering, fallbacks and not leaving placeholders in the DOM, among other measures, as ways to guard against failed JavaScript rendering.
extendsStage D2-C149 Day 2 · Lightning session D: Rendering and JavaScriptA search engine's renderer does not wait forever before taking its snapshot, so content from slow internal systems or APIs can be missing and unresolved placeholders can end up in the final snapshot, a community speaker said.
- Stage D2-C263 Day 2 · What is Google friendly JavaScript
Content is missing from the rendered HTML either because the server does not serve it or because the content has still not appeared after some period of time.
extendsStage D2-C149 Day 2 · Lightning session D: Rendering and JavaScriptA search engine's renderer does not wait forever before taking its snapshot, so content from slow internal systems or APIs can be missing and unresolved placeholders can end up in the final snapshot, a community speaker said.
- Stage D2-C268 Day 2 · What is Google friendly JavaScript
Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Stage D2-C271 Day 2 · What is Google friendly JavaScript
Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.
extends - Stage D2-C271 Day 2 · What is Google friendly JavaScript
Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.
extends - Stage D2-C272 Day 2 · What is Google friendly JavaScript
Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.
extendsD2-C199 Day 2 · Lightning session D: Rendering and JavaScriptAn Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders with is tall, so content lazy-loaded this way can load during rendering.
- Docs D2-C276 Day 2 · What is Google friendly JavaScript
To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each chunk its own persistent, unique URL, link sequentially to those URLs, and update the displayed URL with the History API when a new chunk becomes the main visible element.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Stage D2-C281 Day 2 · What is Google friendly JavaScript
Polyfills work in two directions: they patch backwards where a browser lacks support, or patch forward to build now for a feature expected in the future, as Web Fragments does.
extendsStage D2-C254 Day 2 · Lightning session D: Rendering and JavaScriptNatalia Venditto said the ShadowRealm proposal would do much of what Web Fragments does today with patches: containerizing and isolating execution.
- Stage D2-C287 Day 2 · What is Google friendly JavaScript
A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- Stage D2-C287 Day 2 · What is Google friendly JavaScript
A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.
extendsD2-C040 Day 2 · How is HTML interpretedGoogle cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
- Stage D2-C288 Day 2 · What is Google friendly JavaScript
With the History API, a single-page app can use real links and attach event listeners that intercept the click, rewrite the URL with pushState and load the new content, so Google can follow the links while users avoid full page reloads.
extendsD2-C185 Day 2 · Lightning session D: Rendering and JavaScriptThe crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>, combined with History API routing (window.history.pushState) instead of hash routes.
- Stage D2-C291 Day 2 · What is Google friendly JavaScript
In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C292 Day 2 · What is Google friendly JavaScript
The fix for a client-side soft 404 is to serve a real 404 error page where appropriate; how to detect URLs that do not exist depends on how the app works and where it is hosted.
extendsStage D2-C191 Day 2 · Lightning session D: Rendering and JavaScriptThe fix for a client-side soft 404 is to make a missing page end with a real HTTP 404 status code instead of 200.
- Stage D2-C296 Day 2 · What is Google friendly JavaScript
If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.
extendsD1-C064 Day 1 · How crawling worksThe crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
- Stage D2-C303 Day 2 · What is Google friendly JavaScript
Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.
extendsStage D1-C281 Day 1 · Lightning session A: Automation and AIA community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.
- Stage D2-C330 Day 2 · Understanding what's on a page
Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real, but it matters at the level of an AI model's context window.
extendsD1-C054 Day 1 · How Search works and where's AI?Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.
- Stage D2-C332 Day 2 · Understanding what's on a page
Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.
extendsD1-C054 Day 1 · How Search works and where's AI?Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
extendsStage D1-C070 Day 1 · How crawling errors affect SearchSoft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.
- D2-C336 Day 2 · Understanding what's on a page
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C336 Day 2 · Understanding what's on a page
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
extendsStage D2-C213 Day 2 · Lightning session D: Rendering and JavaScriptA page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.
- Stage D2-C338 Day 2 · Understanding what's on a page
Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.
extendsD1-C042 Day 1 · How Search works and where's AI?BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.
- Stage D2-C339 Day 2 · Understanding what's on a page
For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C344 Day 2 · Handling web duplication
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- D2-C345 Day 2 · Handling web duplication
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- D2-C345 Day 2 · Handling web duplication
Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
extendsStage D1-C207 Day 1 · How Search works and where's AI?Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.
- Stage D2-C346 Day 2 · Handling web duplication
For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.
extendsStage D1-C207 Day 1 · How Search works and where's AI?Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.
- D2-C348 Day 2 · Handling web duplication
Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.
extendsStage D1-C111 Day 1 · session not recordedGary Illyes said there is no such thing as a duplicate content penalty.
- Stage D2-C367 Day 2 · Handling web duplication
Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C369 Day 2 · Handling web duplication
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
extendsD1-C094 Day 1 · How Google thinks about crawl budgetIf the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
- D2-C369 Day 2 · Handling web duplication
Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsStage D1-C069 Day 1 · How crawling errors affect SearchDNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C374 Day 2 · Handling web duplication
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
extendsStage D1-C367 Day 1 · How crawling errors affect SearchCDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
- Stage D2-C375 Day 2 · Handling web duplication
Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.
extendsStage D1-C367 Day 1 · How crawling errors affect SearchCDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
- Docs D2-C376 Day 2 · Handling web duplication
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
extendsAnalysis D1-C078 Day 1 · How crawling errors affect SearchFor temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
- Stage D2-C377 Day 2 · Handling web duplication
AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.
extends - Stage D2-C379 Day 2 · Handling web duplication
rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.
extendsD2-C031 Day 2 · How is HTML interpretedGoogle extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.
- Stage D2-C380 Day 2 · Handling web duplication
Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.
extendsAnalysis D1-C113 Day 1 · session not recordedThe real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
- D2-C382 Day 2 · Handling web duplication
When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.
extendsD2-C033 Day 2 · How is HTML interpretedGoogle extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.
- Docs D2-C395 Day 2 · Handling web duplication
Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C397 Day 2 · Handling web duplication
Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.
extendsDocs D1-C131 Day 1 · session not recordedGoogle's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
- Stage D2-C401 Day 2 · Lightning session E: Managing Duplicates and Site Moves
Broken canonical tags can make the wrong pages of a site show up in search results.
extendsAnalysis D1-C113 Day 1 · session not recordedThe real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.
- Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C408 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.
extendsStage D2-C398 Day 2 · Handling web duplicationGoogle's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.
- Stage D2-C427 Day 2 · Lightning session E: Managing Duplicates and Site Moves
When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.
extendsAnalysis D1-C067 Day 1 · How crawling worksA page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
- D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
extendsD1-C063 Day 1 · How crawling worksThe crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.
- Stage D2-C444 Day 2 · Finding the gold nuggets: structured data, media, and more!
Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.
extendsD2-C026 Day 2 · How is HTML interpretedA Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
- Stage D2-C452 Day 2 · What is Structured Data and why we need it on the internet.
Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
extendsD1-C011 Day 1 · Welcome and opening keynotesEcosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such as videos and cards.
- Stage D2-C452 Day 2 · What is Structured Data and why we need it on the internet.
Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
extendsStage D1-C158 Day 1 · Welcome and opening keynotesGoogle said that even the classic ten-blue-links layout was settled only after millions of experiments, as part of using as much data as possible for product decisions.
- Stage D2-C453 Day 2 · What is Structured Data and why we need it on the internet.
AI Overviews and AI Mode launched as fairly text-heavy answers with little image content and few tables, and they have become more structured over time because that is what users want.
extendsD1-C011 Day 1 · Welcome and opening keynotesEcosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such as videos and cards.
- Docs D2-C462 Day 2 · What is Structured Data and why we need it on the internet.
Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
extendsStage D1-C250 Day 1 · Lightning session A: Automation and AIIn a community demo, a fix loop gave a large LLM (Claude Opus) the current markup, the errors Google's test reported and Google's documentation, had it write new JSON-LD and re-ran the test, with at most three attempts; the demo's page passed on the second.
- Stage D2-C465 Day 2 · What is Structured Data and why we need it on the internet.
Gemini in Chrome relies heavily on the screenshot it takes of a page.
extendsDocs D1-C131 Day 1 · session not recordedGoogle's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
- Stage D2-C474 Day 2 · What is Structured Data and why we need it on the internet.
Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.
extendsD2-C028 Day 2 · How is HTML interpretedGoogle extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
- Stage D2-C476 Day 2 · What is Structured Data and why we need it on the internet.
The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C477 Day 2 · What is Structured Data and why we need it on the internet.
As far as the speaker knows, in most of Google's main AI uses a page's schema.org markup is not turned into text and put directly into the model's context; the data is first sorted out, checked for quality and indexed before it is passed on as grounding context.
extendsDocs D1-C061 Day 1 · How Search works and where's AI?Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in Search, and that such files neither harm nor help visibility because Google Search ignores them.
- Docs D2-C480 Day 2 · What is Structured Data and why we need it on the internet.
Google's guide to optimizing for generative AI features lists 'overfocusing on structured data' among the things site owners don't need to do: structured data is not required for generative AI search and no special schema.org markup is needed, though it remains worth using because it helps pages become eligible for rich results.
extendsDocs D1-C061 Day 1 · How Search works and where's AI?Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in Search, and that such files neither harm nor help visibility because Google Search ignores them.
- Stage D2-C511 Day 2 · What is Structured Data and why we need it on the internet.
Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a chicken-and-egg problem that the schema.org usage statistics are meant to help break.
extends - Stage D2-C521 Day 2 · What is Structured Data and why we need it on the internet.
Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.
extends - Stage D2-C523 Day 2 · Using images to your advantage and Engaging Search users with videos
Gary Illyes said images and videos drive a large amount of traffic to publishers.
extends - Stage D2-C525 Day 2 · Using images to your advantage and Engaging Search users with videos
An image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.
extends - Stage D2-C525 Day 2 · Using images to your advantage and Engaging Search users with videos
An image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.
extendsAnalysis D1-C117 Day 1 · session not recordedWith one in six AI Mode searches being multimodal, original images with descriptive file names, alt text and captions feed AI answers as well as image search.
- Stage D2-C601 Day 2 · Focusing on Internationalisation and Localisation
Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results, so what works for web results also works for the AI features.
extendsStage D2-C116 Day 2 · Lightning session D: Rendering and JavaScriptAI Overviews and AI Mode are built on top of Search results: they are a different experience of the same content Google already has.
- Stage D2-C603 Day 2 · Focusing on Internationalisation and Localisation
Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Analysis D2-C607 Day 2 · Focusing on Internationalisation and Localisation
Check AI Overviews and AI Mode with native-language queries in each target market, not with translated English keywords, and compare with the country breakdown of Search Console's generative AI performance report; topics where competitors are cited and you are not point to missing or weak local content.
extendsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Analysis D2-C612 Day 2 · Focusing on Internationalisation and Localisation
Treat machine translation as a first draft: have a native speaker review it and adapt dates, calendars and units before publishing, because unreviewed bulk translation that adds little value can also fall under Google's scaled content abuse policy.
extendsAnalysis D1-C048 Day 1 · How Search works and where's AI?The answer is not a claim that Google detects AI text. It says ranking favours text that reads as natural to people. The risk with AI content is scale without value, which falls under Google's scaled content abuse policy, not the tool itself.
- Stage D2-C627 Day 2 · Lightning session G: Internationalisation
A community speaker said language versions need adjusting for regional variants of one language: 90 is 'quatre-vingt-dix' in France but 'nonante' in Belgium, and Swiss German differs from German, for example in whether the letter ß (Eszett) is used.
extendsD2-C562 Day 2 · Focusing on Internationalisation and LocalisationLanguage versions can also be language-plus-region variants: Google's example site had a generic English page (en), a UK English page (en-gb) and a Spanish page (es), each on its own URL.
- Stage D2-C648 Day 2 · Calculating (some) signals
Among the many signals Google calculates during indexing, the ones singled out as having large effects on search results were country, language, freshness, SafeSearch and spam.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C651 Day 2 · Calculating (some) signals
Google detects the language of each document and weights language in index selection so that the index is not dominated by one or two languages.
extendsStage D2-C585 Day 2 · Focusing on Internationalisation and LocalisationGoogle determines a page's language for indexing from the page content, not from a language code in the URL.
- Stage D2-C656 Day 2 · Calculating (some) signals
Freshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking event hits a city, such as possible closure of Barcelona's airport, users want really fresh results, not results from two weeks ago.
extendsD1-C045 Day 1 · How Search works and where's AI?Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).
- Stage D2-C668 Day 2 · Calculating (some) signals
Google uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is highly accurate.
extendsD1-C041 Day 1 · How Search works and where's AI?Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.
- Stage D2-C669 Day 2 · Calculating (some) signals
SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C669 Day 2 · Calculating (some) signals
SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.
extendsD1-C041 Day 1 · How Search works and where's AI?Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.
- Stage D2-C682 Day 2 · Deciding what goes in the index?
Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.
extendsStage D1-C211 Day 1 · How Search works and where's AI?Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
- Stage D2-C684 Day 2 · Deciding what goes in the index?
Index selection is a predictive AI system that relies heavily on machine learning.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C685 Day 2 · Deciding what goes in the index?
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
extendsD1-C094 Day 1 · How Google thinks about crawl budgetIf the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
- Analysis D2-C686 Day 2 · Deciding what goes in the index?
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
extendsAnalysis D1-C095 Day 1 · How Google thinks about crawl budgetNew content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.
- Stage D2-C687 Day 2 · Deciding what goes in the index?
Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
extendsStage D1-C205 Day 1 · How Search works and where's AI?Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.
- Stage D2-C688 Day 2 · Deciding what goes in the index?
Index selection is the last step before documents enter Google's index.
extendsStage D1-C211 Day 1 · How Search works and where's AI?Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
- Stage D2-C688 Day 2 · Deciding what goes in the index?
Index selection is the last step before documents enter Google's index.
extendsD2-C026 Day 2 · How is HTML interpretedA Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
- Stage D2-C695 Day 2 · Deciding what goes in the index?
Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D2-C697 Day 2 · Deciding what goes in the index?
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
extendsD1-C106 Day 1 · How Google thinks about crawl budgetThe noindex rule consumes crawl budget, because Google must fetch the page to see it.
- Stage D2-C698 Day 2 · Deciding what goes in the index?
Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.
extendsStage D2-C100 Day 2 · Controlling indexingThe unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.
- Stage D2-C700 Day 2 · Deciding what goes in the index?
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C700 Day 2 · Deciding what goes in the index?
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
extendsStage D2-C334 Day 2 · Understanding what's on a pageA soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
- Stage D2-C701 Day 2 · Deciding what goes in the index?
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C701 Day 2 · Deciding what goes in the index?
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
extendsD2-C345 Day 2 · Handling web duplicationGoogle's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.
- Stage D2-C706 Day 2 · Deciding what goes in the index?
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
extendsD1-C065 Day 1 · How crawling worksThe scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.
- Stage D2-C706 Day 2 · Deciding what goes in the index?
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
extendsDocs D1-C097 Day 1 · How Google thinks about crawl budgetGoogle's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.
- Analysis D2-C708 Day 2 · Deciding what goes in the index?
The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Stage D2-C709 Day 2 · Deciding what goes in the index?
Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D2-C714 Day 2 · Deciding what goes in the index?
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Analysis D2-C716 Day 2 · Deciding what goes in the index?
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Stage D2-C717 Day 2 · Deciding what goes in the index?
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
extendsStage D1-C368 Day 1 · How crawling errors affect SearchSearch Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
- Stage D2-C722 Day 2 · How does the index look like?
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C726 Day 2 · How does the index look like?
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C727 Day 2 · How does the index look like?
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
extendsD1-C051 Day 1 · How Search works and where's AI?Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.
- Stage D2-C727 Day 2 · How does the index look like?
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
extendsDocs D1-C053 Day 1 · How Search works and where's AI?Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.
- Stage D2-C727 Day 2 · How does the index look like?
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
extendsStage D1-C172 Day 1 · Welcome and opening keynotesGoogle said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.
- Stage D2-C727 Day 2 · How does the index look like?
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
extendsStage D2-C073 Day 2 · Controlling indexingJohn Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
- Docs D2-C729 Day 2 · How does the index look like?
Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several data sources, which for AI Mode include the Knowledge Graph and shopping data as well as web content (AI Mode launch post, March 2025).
extendsDocs D1-C053 Day 1 · How Search works and where's AI?Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.
- Stage D2-C741 Day 2 · How does the index look like?
In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the user's query decides which documents are returned.
extendsDocs D1-C129 Day 1 · How Search works and where's AI?Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.
- Stage D2-C770 Day 2 · Google Trends
Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of Nano Banana (the image editing model in the Gemini app), as people searched for both Gemini and Nano Banana.
extendsStage D1-C162 Day 1 · Welcome and opening keynotesGoogle said that after the launch of Nano Banana, its image generation model, the Gemini app became the number one app in the app store in 2025.
- D2-C820 Day 2 · Welcome to indexing day!
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
extendsD1-C040 Day 1 · How Search works and where's AI?URL discovery works through links: a homepage links to section pages, which link to further pages.
- D2-C820 Day 2 · Welcome to indexing day!
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
extendsStage D1-C202 Day 1 · How Search works and where's AI?Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.
- Stage D2-C821 Day 2 · Controlling indexing
John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would probably not be part of them.
extendsStage D2-C063 Day 2 · Controlling indexingThe page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page, which John Mueller called a weird and very broad rule that makes the page stand on its own.
- Stage D2-C822 Day 2 · Focusing on Internationalisation and Localisation
Of the URL structures for country versions in a table on Google's slide (not photographed), the speaker called only one a wrong choice: the last option, which the speaker did not name but said they really do not recommend.
extendsStage D2-C594 Day 2 · Focusing on Internationalisation and LocalisationFor country versions of a site, a ccTLD, subdomains or subdirectories are all acceptable choices; the right one depends on the site's needs, goals and resources.
- Stage D2-C824 Day 2 · How does the index look like?
Google said posting lists, which Google's serving system uses to find the pages that contain a query's words, are not new: they are at least 60 years old (as of 2026).
extendsStage D2-C732 Day 2 · How does the index look like?To find relevant pages, Google's serving system relies on posting lists, a long-established information retrieval structure taught in computer science courses, because simply asking for every page that contains a word would not work.
- Stage D2-C825 Day 2 · Understanding what's on a page
Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.
extendsStage D2-C332 Day 2 · Understanding what's on a pageGemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.
- Stage D2-C838 Day 2 · Welcome to indexing day!
Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.
extendsD2-C820 Day 2 · Welcome to indexing day!Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
- Stage D2-C841 Day 2 · Welcome to indexing day!
Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to them, with easily replaced products, such as a particular cheese, for which the buyer simply picks a similar product instead of waiting.
extendsD2-C010 Day 2 · Welcome to indexing day!Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it depends on how important that product page is to users.
- Stage D2-C842 Day 2 · Welcome to indexing day!
Google said listing the sitemap in robots.txt is fine, as many websites do.
extendsDocs D1-C084 Day 1 · How Google interprets robots.txtGoogle supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
- Stage D2-C843 Day 2 · Welcome to indexing day!
Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.
extendsD2-C017 Day 2 · Welcome to indexing day!Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
- Stage D2-C844 Day 2 · Welcome to indexing day!
Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.
extendsD2-C020 Day 2 · Welcome to indexing day!Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
- Stage D2-C846 Day 2 · Welcome to indexing day!
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D2-C846 Day 2 · Welcome to indexing day!
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
extendsD2-C021 Day 2 · Welcome to indexing day!Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
- Analysis D2-C849 Day 2 · Welcome to indexing day!
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D2-C852 Day 2 · Controlling indexing
When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript, so a script cannot switch the page back to indexable, John Mueller said.
extendsD2-C108 Day 2 · Controlling indexingRemoving a robots restriction such as noindex with JavaScript does not work, a slide said.
- Stage D2-C853 Day 2 · Controlling indexing
John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding them with JavaScript only where that is not possible, as in a JavaScript web app.
extendsD2-C106 Day 2 · Controlling indexingA slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.
- Analysis D2-C855 Day 2 · Controlling indexing
On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex), while Google's guide says only that it may skip rendering; either way, a noindex in the served HTML must never be one that JavaScript is expected to lift.
extendsAnalysis D2-C109 Day 2 · Controlling indexingShip robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may skip rendering a page that arrives with noindex, so the page can stay out of the index even if a script removes the tag later.
- Stage D2-C858 Day 2 · Lightning session D: Rendering and JavaScript
Search engines such as Google and Bing would ignore HTML tags written inside a meta description, a community speaker said.
extendsStage D2-C142 Day 2 · Lightning session D: Rendering and JavaScriptOne site that builds its whole content by rendering also rendered its meta description with HTML tags inside it, which makes no sense and points to a flawed process.
- Stage D2-C859 Day 2 · Lightning session D: Rendering and JavaScript
A rendering failure such as a Content Security Policy blocking a video on a landing page that explains how to open a business account could have a drastic impact for a purely online business, a community speaker warned.
extendsStage D2-C156 Day 2 · Lightning session D: Rendering and JavaScriptContent Security Policy is often misunderstood: on a page of a large German banking group, a video that should have been shown was blocked by a CSP violation, which also appeared in the browser console.
- Stage D2-C860 Day 2 · Lightning session D: Rendering and JavaScript
Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a modern crawler.
extendsStage D2-C162 Day 2 · Lightning session D: Rendering and JavaScriptA community speaker strongly advised SEOs to use Chrome DevTools more often to understand what happens when their pages render.
- Stage D2-C861 Day 2 · Understanding what's on a page
Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.
extendsD2-C309 Day 2 · Understanding what's on a pageA Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
- Stage D2-C866 Day 2 · Understanding what's on a page
Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.
extendsD2-C202 Day 2 · Lightning session D: Rendering and JavaScriptTab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden attribute; Google indexes such hidden content.
- Stage D2-C868 Day 2 · Understanding what's on a page
Gary Illyes said the main content is what Google considers when ranking a page.
extendsD2-C028 Day 2 · How is HTML interpretedGoogle extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
- Stage D2-C869 Day 2 · Understanding what's on a page
Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.
extendsStage D2-C331 Day 2 · Understanding what's on a pageGary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
- Stage D2-C873 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker said that, under the canonical link specification, an improperly declared canonical tag can also be ignored completely by the application that processes it, not only replaced by that application's own heuristic.
extendsStage D2-C402 Day 2 · Lightning session E: Managing Duplicates and Site MovesAccording to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.
- Stage D2-C882 Day 2 · Lightning session E: Managing Duplicates and Site Moves
In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs with no same-intent match were removed with an error status instead of being redirected (the exact code is unclear in the recording).
extendsD2-C396 Day 2 · Handling web duplicationGoogle's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.
- Stage D2-C884 Day 2 · Lightning session E: Managing Duplicates and Site Moves
The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.
extendsStage D1-C540 Day 1 · Q&AAsked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.
- Stage D2-C886 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.
extendsD1-C103 Day 1 · How Google thinks about crawl budgetFour ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.
- Stage D2-C903 Day 2 · Lightning session E: Managing Duplicates and Site Moves
A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.
extendsStage D2-C334 Day 2 · Understanding what's on a pageA soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
- Stage D2-C930 Day 2 · Using images to your advantage and Engaging Search users with videos
Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.
extendsStage D2-C448 Day 2 · Finding the gold nuggets: structured data, media, and more!For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.
- Stage D2-C933 Day 2 · Using images to your advantage and Engaging Search users with videos
Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.
extendsD2-C021 Day 2 · Welcome to indexing day!Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
- Stage D2-C947 Day 2 · Lightning session F: Media
In a community speaker's case, an LLM asked to write alt text for a product image of a veterinary anxiety medicine for cats described only what it could see, a cat and a veterinarian, and dropped the stress, anxiety and medical-treatment intent of the page.
extendsDocs D2-C462 Day 2 · What is Structured Data and why we need it on the internet.Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
- Stage D2-C950 Day 2 · Lightning session F: Media
Because not every generated alt text can be checked by hand, a community speaker's safer script adds the page title to the LLM's image description, so the alt text carries the page's intent and not only what the image shows.
extendsStage D2-C537 Day 2 · Using images to your advantage and Engaging Search users with videosThe text around an image is critical: Google uses it as context to understand the image and to rank it, so an alt attribute alone is not enough.
- Stage D2-C952 Day 2 · Lightning session F: Media
A community speaker warned that articles still recommend rewriting alt text with AI, called AI a tool, and said that adopting such new AI implementations without applying existing SEO knowledge can damage ranking performance.
extendsDocs D2-C462 Day 2 · What is Structured Data and why we need it on the internet.Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
- Stage D2-C962 Day 2 · Focusing on Internationalisation and Localisation
Google said incorrect hreflang language or region codes are a mistake it finds very often, with the same wrong codes coming up again and again, especially in Europe, because site owners assume the codes are easy.
extendsD2-C575 Day 2 · Focusing on Internationalisation and Localisationhreflang language codes must be ISO 639-1 codes: Google's list of common mistakes marked se, dk and cz (the country codes of Sweden, Denmark and Czechia) as wrong and sv, da and cs (Swedish, Danish, Czech) as right.
- Stage D2-C965 Day 2 · Lightning session G: Internationalisation
Users of non-Latin-script languages do not always search in their own script: the same Persian query may be typed in Persian script or in Latin letters, with the same intent and the same expected results.
extendsStage D2-C599 Day 2 · Focusing on Internationalisation and LocalisationGoogle usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.
- Stage D2-C972 Day 2 · Lightning session G: Internationalisation
In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English sources; the presenter explained that Persian content on a topic is often thinner, of lower quality or less relevant, so the systems retrieve from languages with better content, which the presenter called cross-lingual retrieval.
extendsStage D2-C603 Day 2 · Focusing on Internationalisation and LocalisationUsers often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.
- Analysis D2-C982 Day 2 · Focusing on Internationalisation and Localisation
At the point where one recording names the ccTLD as the most important country-targeting signal, a second attendee recording heard 'IP address' instead; Google's multi-regional sites documentation names the ccTLD as a strong signal of a site's target country and server location only as a possible, not definitive one, so the ccTLD is the better-supported reading.
extendsStage D2-C591 Day 2 · Focusing on Internationalisation and LocalisationThe ccTLD (country-code top-level domain) is one of the most important and strongest signals Google uses to decide which country a site targets; Google also considers other signals, which matter less.
- Stage D3-C006 Day 3 · Making sense of users' queries
Google's first step in understanding almost any query is to detect its language, which tells Google roughly what content the user wants: a query in German suggests German content, a query in English English content.
extendsStage D2-C654 Day 2 · Calculating (some) signalsIn ranking, country and language signals help Google serve users the right content for their country and language.
- Stage D3-C007 Day 3 · Making sense of users' queries
Query language detection works poorly when someone searches only for a brand name, such as Facebook or Google, because the query does not show which language the user wants results in.
extendsStage D1-C215 Day 1 · How Search works and where's AI?Serving starts with interpreting the query, which includes cleaning it up, detecting its language and expanding it.
- Stage D3-C011 Day 3 · Making sense of users' queries
Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.
extendsStage D2-C320 Day 2 · Understanding what's on a pageText in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
- Stage D3-C013 Day 3 · Making sense of users' queries
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
extendsStage D2-C321 Day 2 · Understanding what's on a pageFor languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
- Stage D3-C013 Day 3 · Making sense of users' queries
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
extendsStage D2-C737 Day 2 · How does the index look like?A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
- Stage D3-C015 Day 3 · Making sense of users' queries
Google said the difference between words and entities can be seen in Google Trends, where a term can be searched as words or as an entity (Trends calls these a search term and a topic).
extendsStage D2-C761 Day 2 · Google TrendsThe Google Trends Explore page, which Google called the heart of Trends, shows search interest in a query or topic and how it changes over time.
- Stage D3-C048 Day 3 · Making sense of users' queries
Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a German 'ü' written as 'ü', as 'ue' or left out.
extendsStage D2-C600 Day 2 · Focusing on Internationalisation and LocalisationGoogle usually understands a query word whether it is written with or without diacritics (accents).
- Stage D3-C051 Day 3 · Making sense of users' queries
Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.
extendsStage D2-C599 Day 2 · Focusing on Internationalisation and LocalisationGoogle usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.
- D3-C057 Day 3 · Making sense of users' queries
Google's slide on how LLM features with grounding generally work showed a query going to both the search engine and an LLM, the search engine's results going to the LLM, the LLM generating fan-out queries that go back to the search engine, and the LLM returning answers with links.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D3-C059 Day 3 · Making sense of users' queries
Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding how normal queries work explains fan-out queries too.
extendsStage D1-C172 Day 1 · Welcome and opening keynotesGoogle said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.
- Stage D3-C059 Day 3 · Making sense of users' queries
Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding how normal queries work explains fan-out queries too.
extendsStage D2-C727 Day 2 · How does the index look like?Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
- Stage D3-C061 Day 3 · Making sense of users' queries
Google tries to make fan-out queries distinct from each other for better coverage, avoiding asking the same question several times, which would return the same answers.
extendsStage D1-C225 Day 1 · How Search works and where's AI?Google said query fan-out is nothing new: it fires, for example, ten different searches in the background, each a normal search on the same systems.
- Stage D3-C061 Day 3 · Making sense of users' queries
Google tries to make fan-out queries distinct from each other for better coverage, avoiding asking the same question several times, which would return the same answers.
extendsDocs D2-C729 Day 2 · How does the index look like?Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several data sources, which for AI Mode include the Knowledge Graph and shopping data as well as web content (AI Mode launch post, March 2025).
- Stage D3-C063 Day 3 · Making sense of users' queries
Google said fan-out queries are not added to Search Console, because Google considers them part of its infrastructure.
extendsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Stage D3-C065 Day 3 · Making sense of users' queries
Because every system runs fan-out differently, Google advised understanding that fan-out happens but not overfocusing on individual fan-out queries or on how to rank for them.
extendsDocs D1-C129 Day 1 · How Search works and where's AI?Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.
- Stage D3-C075 Day 3 · Making sense of users' queries
At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.
extendsStage D2-C736 Day 2 · How does the index look like?In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.
- Stage D3-C075 Day 3 · Making sense of users' queries
At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.
extendsStage D2-C738 Day 2 · How does the index look like?At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.
- Stage D3-C077 Day 3 · Making sense of users' queries
The first condition for retrieving a document is that the query's words, or its concepts in the case of vectors or embeddings, are in the document or related to it.
extendsStage D2-C740 Day 2 · How does the index look like?Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.
- Stage D3-C079 Day 3 · Making sense of users' queries
To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.
extendsStage D2-C649 Day 2 · Calculating (some) signalsCountry and language are among Google's most important signals and have been used since Google's early days.
- Stage D3-C079 Day 3 · Making sense of users' queries
To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.
extendsStage D2-C722 Day 2 · How does the index look like?Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
- Stage D3-C080 Day 3 · Making sense of users' queries
At retrieval, Google tries to match results to the user's language wherever possible: someone searching in Spanish does not necessarily want results in Italian.
extendsStage D2-C654 Day 2 · Calculating (some) signalsIn ranking, country and language signals help Google serve users the right content for their country and language.
- Stage D3-C082 Day 3 · Making sense of users' queries
Country is the second retrieval signal: a user searching from Switzerland wants cheese from Switzerland, not from Germany, and a user in Spain is poorly served by results targeting a South American country.
extendsStage D2-C649 Day 2 · Calculating (some) signalsCountry and language are among Google's most important signals and have been used since Google's early days.
- Stage D3-C083 Day 3 · Making sense of users' queries
Google called quality the most important of the signals used to order candidates at retrieval: a URL of high quality is more likely to be retrieved from the index for specific queries.
extendsStage D2-C722 Day 2 · How does the index look like?Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
- Stage D3-C104 Day 3 · Lightning session K: Facets of quality
To choose what to cover in depth, a community speaker's team used Google Trends to find what people in their market were interested in and then gave them that content.
extendsStage D2-C789 Day 2 · Google TrendsGoogle presented three uses of Google Trends for content creation, marketing and SEO: keyword selection, scheduling and ideation.
- Analysis D3-C109 Day 3 · Lightning session K: Facets of quality
The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.
extendsStage D2-C318 Day 2 · Understanding what's on a pageGoogle does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
- Stage D3-C131 Day 3 · How Google thinks about Quality
Google's quality talk said quality is one of the most important ranking signals, although Google uses hundreds of ranking signals (the word 'quality' is a repaired speech-to-text reading).
extendsStage D3-C084 Day 3 · Making sense of users' queriesThe same quality signal that Google uses at retrieval is also used in ranking when results are served.
- Analysis D3-C135 Day 3 · How Google thinks about Quality
The quality talk's slide listed MUM among Google's ranking systems, while Google's ranking systems guide says MUM is not currently used for general ranking in Search; read the slide as a list of systems Google runs, not as proof that each one ranks every query.
extendsDocs D1-C130 Day 1 · How Search works and where's AI?Google's ranking systems guide says MUM is not currently used for general ranking in Search, only for specific applications such as COVID-19 vaccine searches and featured snippet callouts.
- Stage D3-C147 Day 3 · How Google thinks about Quality
Google's quality talk said that at any moment thousands of Search experiments are probably running.
extendsStage D1-C158 Day 1 · Welcome and opening keynotesGoogle said that even the classic ten-blue-links layout was settled only after millions of experiments, as part of using as much data as possible for product decisions.
- Stage D3-C170 Day 3 · How Google thinks about Quality
Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.
extendsStage D2-C311 Day 2 · Understanding what's on a pageGary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.
- Stage D3-C179 Day 3 · How Google thinks about Quality
Google's quality talk said non-commodity content offers a unique, experienced take, and advised writing content that readers will find very helpful and reliable.
extendsD1-C055 Day 1 · How Search works and where's AI?Myth: build content for every possible consumer need. Google's answer: prioritise unique perspectives, expertise and in-depth experience that go beyond common knowledge.
- Stage D3-C205 Day 3 · How Google thinks about Quality
Google's quality talk said AI has fundamentally changed how Google builds spam updates, letting it evaluate many more candidates and drastically increasing its velocity, so it launches faster with more impact.
extendsStage D2-C668 Day 2 · Calculating (some) signalsGoogle uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is highly accurate.
- Stage D3-C206 Day 3 · How Google thinks about Quality
Google's quality talk said AI lets Google catch new types of spam and catch more of it.
extendsStage D2-C668 Day 2 · Calculating (some) signalsGoogle uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is highly accurate.
- Stage D3-C212 Day 3 · How Google thinks about Quality
Google's quality talk said Google clarified that traditional spam techniques aimed at manipulating AI responses also violate its spam policies, and that Google can take action against them.
extendsDocs D1-C129 Day 1 · How Search works and where's AI?Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.
- D3-C224 Day 3 · Uncovering Trustworthy Experiences on Discover
Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels and a 16x9 aspect ratio, enabled by the max-image-preview:large setting.
extendsStage D2-C089 Day 2 · Controlling indexingmax-image-preview:large matters mainly in Discover, where it allows a large image that draws people's attention, so the rule can make a page more visible than leaving it out, John Mueller said.
- D3-C239 Day 3 · What are quality updates
Google's slide, headed 'In 2023, there were...', said 40 billion spammy pages are detected every day.
extendsStage D3-C197 Day 3 · How Google thinks about QualityGoogle's quality talk said Google discovers tens of billions of spam pages every day.
- Stage D3-C265 Day 3 · What are quality updates
Google sometimes couples a core update with an update to its Search Quality Rater Guidelines.
extendsStage D3-C170 Day 3 · How Google thinks about QualityGoogle's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.
- Stage D3-C284 Day 3 · What are quality updates
Reviewing the spam slide, the speaker said link spam would probably be replaced with scaled content abuse as the spam type worth talking about today.
extendsStage D3-C207 Day 3 · How Google thinks about QualityGoogle's quality talk listed recent updates to Google's spam policies: back button hijacking, scaled content abuse and manipulating AI responses.
- Docs D3-C287 Day 3 · What are quality updates
Google's spam updates page says its automated spam detection systems run constantly, and a notable improvement to them, such as to the AI-based SpamBrain system, is called a spam update and listed with Google's ranking updates.
extendsStage D2-C670 Day 2 · Calculating (some) signalsSpamBrain is central to Google's spam-fighting efforts and has been improved many times since its launch.
- Stage D3-C304 Day 3 · How Search results are born
Which kinds of results Google shows for a query is decided by query understanding, which tries to predict the intent behind the query.
extendsStage D3-C004 Day 3 · Making sense of users' queriesGoogle placed query understanding in the traditional part of Search, the part where answers are looked up in an index and then presented and ranked.
- Stage D3-C311 Day 3 · How Search results are born
Google handles an image used as a search query much like a text query interpreted as an embedding: the image is broken down into vectors (embeddings) that are then searched for in the index.
extendsStage D2-C740 Day 2 · How does the index look like?Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.
- Stage D3-C316 Day 3 · How Search results are born
Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.
extendsStage D2-C724 Day 2 · How does the index look like?The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.
- Stage D3-C317 Day 3 · How Search results are born
Google can add elements to a text result, such as a 'highly cited' badge or a preferred-source badge.
extendsD1-C027 Day 1 · What's new in the world of SearchSites that users mark as a preferred source are more visible in Top Stories and are labelled in AI Mode and AI Overviews.
- Stage D3-C319 Day 3 · How Search results are born
Image results shown among web results come from Google's image index and are roughly the same images that Google Images shows for the same query.
extendsStage D2-C525 Day 2 · Using images to your advantage and Engaging Search users with videosAn image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.
- Stage D3-C323 Day 3 · How Search results are born
Most of Google's search features need nothing extra from the site owner; Google generates them from what it extracted from the page during indexing.
extendsStage D2-C446 Day 2 · Finding the gold nuggets: structured data, media, and more!The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.
- Stage D3-C324 Day 3 · How Search results are born
Rich results differ from other search features because Google builds them from extra data that site owners provide, usually structured data and usually in JSON-LD format.
extendsStage D2-C521 Day 2 · What is Structured Data and why we need it on the internet.Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.
- Stage D3-C325 Day 3 · How Search results are born
AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to function and work with the normal text results from Google's index.
extendsStage D2-C726 Day 2 · How does the index look like?AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
- Stage D3-C337 Day 3 · How Search results are born
Review structured data lets a site specify how its users rated something; the review snippet shows an average star rating and often the number of reviews of a product, service or piece of content.
extendsStage D2-C454 Day 2 · What is Structured Data and why we need it on the internet.Structured data turns the loosely structured web into structured information that powers visual search features such as review stars and recipe filters (for example by preparation time).
- Stage D3-C341 Day 3 · How Search results are born
Some countries rely a lot on social proof such as reviews, Google's speaker said, recalling a point from Google's Day 2 internationalisation talk.
extendsStage D2-C614 Day 2 · Focusing on Internationalisation and LocalisationAccording to consumer data shown on a slide in the talk (source not captured), US consumers judge product quality more by user feedback and reviews, while European shoppers seem to look more at brand reputation.
- Stage D3-C365 Day 3 · Shopping on Search: Beyond the blue links
Google added six Merchant Center feed attributes for AI shopping experiences: question and answer, documents, related products, item group title and variant option for variants, and popularity rank.
extendsStage D2-C509 Day 2 · What is Structured Data and why we need it on the internet.More shopping structured data news was left for a Day 3 talk by a Google colleague, Alex.
- Stage D3-C378 Day 3 · Shopping on Search: Beyond the blue links
Google Shopping called rich data with light structure the AI sweet spot: a little structure that helps the AI system understand the data, not deep, highly nested, complex structures.
extendsStage D2-C489 Day 2 · What is Structured Data and why we need it on the internet.Describing every semantic detail of a page in markup is probably not worth the effort; focus on the structured data that Google or other consumers actually use.
- Stage D3-C383 Day 3 · Shopping on Search: Beyond the blue links
In product markup, an offer's shipping and return information can point through a JSON-LD identifier (@id) to shipping and return data defined elsewhere.
extendsStage D2-C505 Day 2 · What is Structured Data and why we need it on the internet.Google's shopping structured data launches of the previous year (2025) added support for merchant loyalty programs and shipping policies, letting merchants define a policy at organisation level and specify details at product level.
- Stage D3-C385 Day 3 · Shopping on Search: Beyond the blue links
Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.
extendsD1-C065 Day 1 · How crawling worksThe scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.
- Stage D3-C396 Day 3 · Shopping on Search: Beyond the blue links
Google works to make sure that what merchants can express in its shopping feeds can also be expressed in schema.org, adding vocabulary where schema.org lacks it, for example for product details.
extendsStage D2-C506 Day 2 · What is Structured Data and why we need it on the internet.Google's structured data speaker said that in the months before the event Google added support for validity dates on sale prices in product structured data, so merchants no longer need to rush to remove a sale price when the sale ends for fear it shows wrongly in snippets.
- Stage D3-C440 Day 3 · Inside Search Console: What’s New & How to Use It
Google called AI reporting in Search Console an evolving space and expects the generative AI report to get richer, with more information in the future.
extends - Stage D3-C442 Day 3 · Inside Search Console: What’s New & How to Use It
Search Console launched platform properties, which bring social platforms into Search Console, about three or four months before the October 2026 event.
extendsStage D2-C960 Day 2 · Lightning session F: MediaA community speaker said publishers can also verify their social accounts in Search Console, which makes it easier to track the performance of content repurposed for social platforms.
- Docs D3-C468 Day 3 · Inside Search Console: What’s New & How to Use It
Google's social and video performance guide says that if you already claimed your Search profile, all of its verified accounts are added automatically as platform properties in Search Console.
extendsD1-C025 Day 1 · What's new in the world of SearchPublishers and creators can claim a dedicated Search profile with a name, profile photo, cover image, bio, connected social and website accounts, pinned posts and links, and users can follow it.
- Docs D3-C468 Day 3 · Inside Search Console: What’s New & How to Use It
Google's social and video performance guide says that if you already claimed your Search profile, all of its verified accounts are added automatically as platform properties in Search Console.
extendsStage D2-C960 Day 2 · Lightning session F: MediaA community speaker said publishers can also verify their social accounts in Search Console, which makes it easier to track the performance of content repurposed for social platforms.
- Stage D3-C469 Day 3 · Lightning session L: Understanding SERPs and your users
In a community lightning talk, agency founder Nik Vujic presented how his agency feeds Search Console and GA4 data to an LLM agent to support decisions for its clients.
extendsStage D1-C255 Day 1 · Lightning session A: Automation and AIA community speaker who is not a developer automated monthly SEO reports with Python in Visual Studio Code, using Claude and ChatGPT as coding partners and Google Cloud for access to the Search Console API.
- Docs D3-C473 Day 3 · Lightning session L: Understanding SERPs and your users
Google's launch post for Search Console's branded queries filter (20 November 2025) says the branded versus non-branded split is made by an internal AI-assisted system, not by a regular expression, and that some queries may occasionally be misidentified.
extendsStage D3-C417 Day 3 · Inside Search Console: What’s New & How to Use ItGoogle said query groups are not built by matching query text alone: deciding which queries belong in which group goes beyond simple text filtering.
- D3-C477 Day 3 · Lightning session L: Understanding SERPs and your users
The community workflow's third step feeds the exported Search Console dataset into the agency's LLM agent, which queries the dataset and returns the evidence rather than a sample.
extendsStage D1-C256 Day 1 · Lightning session A: Automation and AIAn agency's automated monthly SEO report follows four simple steps: the data comes in, a script processes it, AI summarises it and the result goes onto a dashboard.
- D3-C478 Day 3 · Lightning session L: Understanding SERPs and your users
Nik Vujic's slide said no agent is needed to start: pull the Search Console data, import it into any LLM and query it in conversation; his agency built its agent to have everything in one place.
extendsStage D1-C256 Day 1 · Lightning session A: Automation and AIAn agency's automated monthly SEO report follows four simple steps: the data comes in, a script processes it, AI summarises it and the result goes onto a dashboard.
- Analysis D3-C487 Day 3 · Lightning session L: Understanding SERPs and your users
Google's ranking systems guide describes 'query deserves freshness' systems that show fresher content where it would be expected, which is narrower than a general preference for fresh content; refresh pages whose queries expect current information, and judge other refreshes by quality.
extendsStage D2-C656 Day 2 · Calculating (some) signalsFreshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking event hits a city, such as possible closure of Barcelona's airport, users want really fresh results, not results from two weeks ago.
- Stage D3-C492 Day 3 · Lightning session L: Understanding SERPs and your users
Nik Vujic said GA4 setups that separate out traffic from LLMs do not capture visits from AI Overviews and AI Mode.
extendsStage D3-C438 Day 3 · Inside Search Console: What’s New & How to Use ItThe regular Search Console Performance report still shows all traffic including AI features, while the generative AI view shows only impressions from AI surfaces.
- Stage D3-C495 Day 3 · Lightning session L: Understanding SERPs and your users
Nik Vujic said he hopes Search Console's generative AI report will get query data.
extendsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Stage D3-C495 Day 3 · Lightning session L: Understanding SERPs and your users
Nik Vujic said he hopes Search Console's generative AI report will get query data.
extendsStage D3-C440 Day 3 · Inside Search Console: What’s New & How to Use ItGoogle called AI reporting in Search Console an evolving space and expects the generative AI report to get richer, with more information in the future.
- Analysis D3-C497 Day 3 · Lightning session L: Understanding SERPs and your users
Pages that do better in AI answers than in classic results do not refute Google's line that optimising for people is optimising for generative AI Search, since the comparison mixed Google's AI features with third-party LLMs; measure each surface separately, per page, before deciding what to change.
extendsD1-C058 Day 1 · How Search works and where's AI?Optimising for people is optimising for generative AI Search.
- Docs D3-C498 Day 3 · Lightning session L: Understanding SERPs and your users
Google's AI features guide says traffic from sites appearing in AI features such as AI Overviews and AI Mode is included in the overall search traffic in Search Console and reported in the Performance report within the Web search type.
extendsStage D3-C438 Day 3 · Inside Search Console: What’s New & How to Use ItThe regular Search Console Performance report still shows all traffic including AI features, while the generative AI view shows only impressions from AI surfaces.
- Analysis D3-C501 Day 3 · Lightning session L: Understanding SERPs and your users
Before feeding a branded versus non-branded split to an LLM, spot-check a sample of queries in each group, since Google's AI-assisted classification can misidentify queries; on URL-path or subdomain properties, where the filter is unavailable, a regex query filter is the fallback.
extendsAnalysis D3-C435 Day 3 · Inside Search Console: What’s New & How to Use ItReview the filters and regexes that Search Console's AI-powered configuration proposes before trusting the numbers: Google's launch post calls the feature experimental, warns that AI can misinterpret requests, and limits it to the Search results Performance report (not Discover or News) and to configuration, not sorting or exporting.
- Stage D3-C518 Day 3 · Lightning session L: Understanding SERPs and your users
A community speaker concluded that rankings and clicks no longer equal real business outcomes.
extendsD1-C057 Day 1 · How Search works and where's AI?Myth: the old metrics don't work in the AI era. Google's answer: measure success through metrics that matter to your business.
- Docs D3-C535 Day 3 · Lightning session L: Understanding SERPs and your users
Beyond Search Console, Google's AI features guide suggests tracking conversions and time spent on the site in tools such as Google Analytics to understand the value of traffic from AI features.
extendsDocs D1-C062 Day 1 · How Search works and where's AI?Google's guide for generative AI features recommends the Generative AI performance report in Search Console for measuring how content performs in generative AI features on Google Search and Discover.
- Docs D3-C536 Day 3 · Lightning session L: Understanding SERPs and your users
A May 2025 Search Central blog post advised site owners to look at the overall value of visits from Search rather than focusing too much on clicks, using indicators of conversion such as sales, sign-ups, a more engaged audience or information lookups about the business.
extendsD1-C057 Day 1 · How Search works and where's AI?Myth: the old metrics don't work in the AI era. Google's answer: measure success through metrics that matter to your business.
- Analysis D3-C542 Day 3 · Lightning session L: Understanding SERPs and your users
Google's own line on Day 1 was to measure success by metrics that matter to the business, and its Generative AI performance report shows impressions but no clicks; together with the community talk this argues for a report that puts organic next to direct, branded paid search, conversions and revenue.
extendsD1-C057 Day 1 · How Search works and where's AI?Myth: the old metrics don't work in the AI era. Google's answer: measure success through metrics that matter to your business.
- Analysis D3-C542 Day 3 · Lightning session L: Understanding SERPs and your users
Google's own line on Day 1 was to measure success by metrics that matter to the business, and its Generative AI performance report shows impressions but no clicks; together with the community talk this argues for a report that puts organic next to direct, branded paid search, conversions and revenue.
extendsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Stage D3-C562 Day 3 · Mastering the messy middle
Google's research also counted easy access as part of AI's ease benefit: multimodal input, such as taking pictures, and low friction lower the barrier to entry.
extendsD1-C018 Day 1 · What's new in the world of SearchOne in six AI Mode searches is multimodal, using voice or images.
- Stage D3-C569 Day 3 · Mastering the messy middle
Google's research found that people who use AI run more searches than people who do not.
extendsD1-C004 Day 1 · Welcome and opening keynotesGoogle says users like AI Overviews and search more because of them, with early feedback overwhelmingly positive and satisfaction highest among 18–24 year olds.
- Stage D3-C571 Day 3 · Mastering the messy middle
Google's research found that people who use AI also visit more brand websites, because it is simply easier.
extendsD1-C005 Day 1 · Welcome and opening keynotesGoogle says AI Overviews help with new types of questions and lead users to visit a greater diversity of sites, with more total sites shown on the results page.
- Stage D3-C602 Day 3 · How long does it take to..?
Google said it knows hundreds of trillions of URLs (as of October 2026).
extendsStage D1-C201 Day 1 · How Search works and where's AI?Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.
- Stage D3-C606 Day 3 · How long does it take to..?
For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).
extends - D3-C610 Day 3 · How long does it take to..?
Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.
extendsStage D1-C337 Day 1 · How crawling worksAn XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.
- D3-C611 Day 3 · How long does it take to..?
If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.
extendsStage D1-C337 Day 1 · How crawling worksAn XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.
- D3-C612 Day 3 · How long does it take to..?
Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.
extendsStage D1-C331 Day 1 · How crawling worksGoogle's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.
- D3-C613 Day 3 · How long does it take to..?
Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.
extendsDocs D1-C085 Day 1 · How Google interprets robots.txtGoogle generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
- Stage D3-C614 Day 3 · How long does it take to..?
Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.
extendsDocs D1-C085 Day 1 · How Google interprets robots.txtGoogle generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
- Stage D3-C615 Day 3 · How long does it take to..?
A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.
extendsDocs D1-C140 Day 1 · How Google interprets robots.txtSearch Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.
- D3-C617 Day 3 · How long does it take to..?
Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2 weeks, and 1-3 weeks in recovery.
extends - Stage D3-C618 Day 3 · How long does it take to..?
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
extendsDocs D1-C126 Day 1 · How crawling errors affect SearchGoogle treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.
- Stage D3-C618 Day 3 · How long does it take to..?
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
extendsStage D1-C354 Day 1 · How crawling errors affect SearchGoogle slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.
- Stage D3-C618 Day 3 · How long does it take to..?
When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.
extends - Stage D3-C623 Day 3 · How long does it take to..?
When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D3-C623 Day 3 · How long does it take to..?
When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.
extendsStage D1-C377 Day 1 · How Google thinks about crawl budgetThe quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.
- Stage D3-C623 Day 3 · How long does it take to..?
When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.
extends - Stage D3-C624 Day 3 · How long does it take to..?
Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.
extendsStage D1-C375 Day 1 · How Google thinks about crawl budgetGooglebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.
- Stage D3-C626 Day 3 · How long does it take to..?
Google renders pages in two ways: immediately after crawling, or later through a queue-based process that runs elsewhere.
extendsD2-C170 Day 2 · Lightning session D: Rendering and JavaScriptAfter processing, an indexable page is placed in Google's render queue to wait for rendering.
- Stage D3-C627 Day 3 · How long does it take to..?
Content that JavaScript adds to a page is typically seen by Google's indexing system within a few hours, and at worst within weeks.
extendsDocs D2-C172 Day 2 · Lightning session D: Rendering and JavaScriptGoogle's JavaScript SEO basics guide says a page may wait in the render queue for a few seconds but that it can take longer, and it gives no upper limit.
- Stage D3-C628 Day 3 · How long does it take to..?
Gary Illyes said Google keeps saying it renders every URL on the internet and that he would say this is not true, though it is what he was told; he went on to say that Google's logs show the rendering queue cleared within weeks.
extendsStage D2-C256 Day 2 · What is Google friendly JavaScriptGoogle renders nearly all of the web by replicating what a browser does, using a real browser's rendering engine.
- Stage D3-C642 Day 3 · How long does it take to..?
Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.
extendsStage D2-C351 Day 2 · Handling web duplicationGoogle treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.
- Stage D3-C647 Day 3 · How long does it take to..?
Google may never use structured data from a site it does not trust: once it sees markup it does not trust, it does not touch it.
extendsStage D2-C500 Day 2 · What is Structured Data and why we need it on the internet.Structured data that is not relevant to the page's content can be treated as abusive: Google's filters make it ineffective, and egregious cases can lead to a manual action.
- D3-C662 Day 3 · How long does it take to..?
Recovery after a core update typically takes 3 to 6 months once the site has put in work to regain traffic, and up to 6 months to a year, until the next core update.
extendsStage D3-C294 Day 3 · What are quality updatesAfter a core update a site can win back some of its rankings by improving the site as a whole.
- D3-C663 Day 3 · How long does it take to..?
Google estimated that a spam update affects sites within 1-2 days of the rollout, 1-2 weeks for continuous processing (said on stage as 2 weeks on average) and months for batch refreshes.
extendsStage D3-C286 Day 3 · What are quality updatesWhen spammers find loopholes that Google's systems miss, Google releases spam updates that change its systems to catch them.
- Docs D3-C667 Day 3 · How long does it take to..?
Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.
extendsStage D1-C528 Day 1 · How Google interprets robots.txtSearch Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.
- Analysis D3-C676 Day 3 · How long does it take to..?
The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.
extendsDocs D2-C131 Day 2 · Lightning session D: Rendering and JavaScriptGoogle's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.
- Stage D3-C680 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Gary Illyes said machine learning is about 50 years old and that Google has been using it for 'probably 30 years', starting with statistical models that made the 'Did you mean' feature possible.
extendsD1-C041 Day 1 · How Search works and where's AI?Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.
- Stage D3-C684 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Predictive language models have been used a lot in Search, and BERT is one: technically, in the purest sense, an LLM.
extendsD1-C042 Day 1 · How Search works and where's AI?BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.
- Stage D3-C686 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Hallucinations can happen with any AI model, and with current training methods there is no way to get rid of them.
extendsDocs D2-C462 Day 2 · What is Structured Data and why we need it on the internet.Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
- Stage D3-C688 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Generative models, including image diffusion models, make things up when they lack information or because of issues in their training.
extendsStage D2-C939 Day 2 · Using images to your advantage and Engaging Search users with videosThe diffusion models that generate images were built to generate images, not text, so they are typically poor at rendering text inside an image.
- Stage D3-C690 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Quality raters cannot give a site a penalty or a manual action; their ratings are converted into labels that Google uses to improve its algorithms.
extendsD3-C160 Day 3 · How Google thinks about QualityGoogle's slide said search quality raters cannot affect the rankings of individual sites.
- Stage D3-C694 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Scaled content abuse is becoming a problem again: in the early 2000s pages were churned out with Perl or PHP scripts, and now the same is done with LLMs.
extendsStage D3-C284 Day 3 · What are quality updatesReviewing the spam slide, the speaker said link spam would probably be replaced with scaled content abuse as the spam type worth talking about today.
- Stage D3-C695 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
The cheaper tokens become, the more AI slop is created, and Google counts AI slop as scaled content abuse.
extendsDocs D2-C611 Day 2 · Focusing on Internationalisation and LocalisationGoogle's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.
- Stage D3-C697 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem, but how it was used and what for.
extendsDocs D2-C611 Day 2 · Focusing on Internationalisation and LocalisationGoogle's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.
- D3-C700 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's closing slide said to use AI responsibly because AI hallucinates, and, especially when creating content briefs with AI, to make sure not to add to the sea of AI slop already flooding the internet.
extendsStage D2-C941 Day 2 · Using images to your advantage and Engaging Search users with videosSites that use AI-generated images or videos should make sure they work for users, check them for hallucinations and regenerate them where needed.
- D3-C701 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's closing slide said AI on Google is just SEO: AI features on Google Search use exactly the same processes as traditional results, so no new acronym is needed, as none was for mobile-first indexing or structured data.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D3-C708 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google does not tell lightning-talk and poster speakers what to talk about; they bring their own ideas.
extendsStage D1-C227 Day 1 · Lightning session A: Automation and AIGoogle received close to 200 submissions for the community lightning talks of the Deep Dive, reviewed them over about a month, and gave each selected community speaker seven minutes on stage.
Repeats 85
The same point made again, often on a later day. The later claim is on the left.
- Stage D1-C196 Day 1 · What's new in the world of Search
Google said Search Console's generative AI reporting launched alongside the generative AI control, starting with impressions only, and may expand later.
repeatsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Stage D1-C207 Day 1 · How Search works and where's AI?
Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.
repeatsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D1-C210 Day 1 · How Search works and where's AI?
Google said it is not really trying to tell AI-written from human-written content, because it cares more about the quality of content than about how it was created.
repeatsStage D1-C175 Day 1 · Welcome and opening keynotesExplaining the principle of incentivising high-quality content, Google said the content may be created by humans or by AI: what matters is that it is high quality and made for users, not for search.
- Stage D1-C271 Day 1 · Lightning session A: Automation and AI
A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard as Ahrefs') that looked at almost 40,000 sites over one month: 97% had no AI agent hits on the file, and those that had any got about two hits a month.
repeatsD1-C054 Day 1 · How Search works and where's AI?Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.
- Stage D1-C274 Day 1 · Lightning session A: Automation and AI
A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.
repeatsDocs D1-C131 Day 1 · session not recordedGoogle's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
- Stage D1-C385 Day 1 · How Google thinks about crawl budget
4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.
repeatsDocs D1-C072 Day 1 · How crawling errors affect Search4xx status codes other than 429 have no effect on crawl rate.
- Stage D1-C434 Day 1 · Q&A
User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.
repeatsStage D1-C273 Day 1 · Lightning session A: Automation and AIA community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.
- Stage D1-C445 Day 1 · Q&A
One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.
repeatsStage D1-C403 Day 1 · Lightning session C: CrawlingA community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.
- Stage D1-C466 Day 1 · Q&A
Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.
repeatsStage D1-C381 Day 1 · How Google thinks about crawl budgetNot every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.
- Stage D1-C522 Day 1 · How Google interprets robots.txt
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
repeats - Stage D1-C524 Day 1 · How Google interprets robots.txt
Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.
repeats - D2-C040 Day 2 · How is HTML interpreted
Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.
repeatsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- D2-C048 Day 2 · How is HTML interpreted
Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
repeatsStage D1-C328 Day 1 · How crawling worksDuring indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.
- Stage D2-C050 Day 2 · Controlling indexing
John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.
repeatsStage D1-C512 Day 1 · How Google interprets robots.txtRobots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
- D2-C103 Day 2 · Controlling indexing
Setting the Search generative AI control in Search Console to exclude keeps a site's links and content out of Search generative AI features such as AI Overviews and AI Mode, so the site gets no traffic or impressions from them; include is the default.
repeatsDocs D1-C030 Day 1 · What's new in the world of SearchThe Search generative AI control covers AI Overviews, AI Mode and generative AI features in Discover. Excluding removes both links to the site and use of its content to ground answers.
- Stage D2-C116 Day 2 · Lightning session D: Rendering and JavaScript
AI Overviews and AI Mode are built on top of Search results: they are a different experience of the same content Google already has.
repeatsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
repeatsStage D1-C273 Day 1 · Lightning session A: Automation and AIA community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.
- Stage D2-C126 Day 2 · Lightning session D: Rendering and JavaScript
Google renders pages with Chromium, the browser technology that also underlies Chrome, Edge and other Chromium-based browsers.
repeatsStage D1-C206 Day 1 · How Search works and where's AI?Google renders JavaScript-heavy pages from their HTML, CSS and JavaScript as a browser would, using the latest version of Chromium.
- Stage D2-C217 Day 2 · Lightning session D: Rendering and JavaScript
Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard, and that this is what the Web Fragments work is trying to do.
repeatsStage D2-C052 Day 2 · Controlling indexingJohn Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.
- D2-C262 Day 2 · What is Google friendly JavaScript
Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.
repeatsD2-C174 Day 2 · Lightning session D: Rendering and JavaScriptWhatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.
- D2-C264 Day 2 · What is Google friendly JavaScript
Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but returns a 200 HTTP status code.
repeatsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- D2-C265 Day 2 · What is Google friendly JavaScript
Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the crawling of critical JavaScript files or API endpoints.
repeatsD2-C204 Day 2 · Lightning session D: Rendering and JavaScriptBlocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.
- Stage D2-C268 Day 2 · What is Google friendly JavaScript
Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.
repeatsD2-C195 Day 2 · Lightning session D: Rendering and JavaScriptContent that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.
- Stage D2-C271 Day 2 · What is Google friendly JavaScript
Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.
repeatsD2-C197 Day 2 · Lightning session D: Rendering and JavaScriptLazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.
- Stage D2-C274 Day 2 · What is Google friendly JavaScript
Content loaded as elements enter the viewport, for example with an Intersection Observer, does load when Google renders a page, because Google's rendering viewport is very tall.
repeatsD2-C199 Day 2 · Lightning session D: Rendering and JavaScriptAn Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders with is tall, so content lazy-loaded this way can load during rendering.
- D2-C284 Day 2 · What is Google friendly JavaScript
URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.
repeatsDocs D1-C101 Day 1 · How Google thinks about crawl budgetGoogle's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.
- D2-C285 Day 2 · What is Google friendly JavaScript
Google recommends the History API to give single-page apps clean URLs instead of fragment-based routes.
repeatsD2-C185 Day 2 · Lightning session D: Rendering and JavaScriptThe crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>, combined with History API routing (window.history.pushState) instead of hash routes.
- Stage D2-C291 Day 2 · What is Google friendly JavaScript
In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.
repeatsStage D2-C190 Day 2 · Lightning session D: Rendering and JavaScriptIn single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.
- Stage D2-C296 Day 2 · What is Google friendly JavaScript
If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.
repeatsD2-C204 Day 2 · Lightning session D: Rendering and JavaScriptBlocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.
- D2-C309 Day 2 · Understanding what's on a page
A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
repeatsD2-C028 Day 2 · How is HTML interpretedGoogle extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
repeatsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
repeatsStage D1-C355 Day 1 · How crawling errors affect SearchA soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.
- Stage D2-C402 Day 2 · Lightning session E: Managing Duplicates and Site Moves
According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.
repeatsStage D2-C380 Day 2 · Handling web duplicationMany SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.
- D2-C441 Day 2 · Finding the gold nuggets: structured data, media, and more!
Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.
repeatsD2-C026 Day 2 · How is HTML interpretedA Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.
- D2-C442 Day 2 · Finding the gold nuggets: structured data, media, and more!
Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
repeatsD2-C048 Day 2 · How is HTML interpretedLinks extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.
- Stage D2-C524 Day 2 · Using images to your advantage and Engaging Search users with videos
Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.
repeatsStage D2-C447 Day 2 · Finding the gold nuggets: structured data, media, and more!For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.
- Stage D2-C537 Day 2 · Using images to your advantage and Engaging Search users with videos
The text around an image is critical: Google uses it as context to understand the image and to rank it, so an alt attribute alone is not enough.
repeatsD1-C045 Day 1 · How Search works and where's AI?Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).
- Stage D2-C601 Day 2 · Focusing on Internationalisation and Localisation
Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results, so what works for web results also works for the AI features.
repeatsD1-C050 Day 1 · How Search works and where's AI?Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
- Stage D2-C601 Day 2 · Focusing on Internationalisation and Localisation
Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results, so what works for web results also works for the AI features.
repeatsD1-C051 Day 1 · How Search works and where's AI?Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.
- Stage D2-C679 Day 2 · Calculating (some) signals
Combing Google's documentation for signals is not the best use of an SEO's time; creating content that users will like is a better one.
repeatsStage D1-C059 Day 1 · Welcome and opening keynotesThe opening keynote closed with the advice to think about UEO, user engine optimisation, next to SEO and GEO: focus on the user and the rest will follow.
- Stage D2-C680 Day 2 · Deciding what goes in the index?
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
repeatsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C680 Day 2 · Deciding what goes in the index?
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
repeatsStage D2-C344 Day 2 · Handling web duplicationGoogle deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
- Stage D2-C719 Day 2 · How does the index look like?
The speaker recapped Google's pipeline up to the index: Google crawls pages, processes the fetched documents and then stores them in its index.
repeatsD1-C036 Day 1 · How Search works and where's AI?Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- Stage D2-C720 Day 2 · How does the index look like?
Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.
repeatsStage D2-C323 Day 2 · Understanding what's on a pageWhen tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.
- Stage D2-C721 Day 2 · How does the index look like?
Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.
repeatsStage D2-C318 Day 2 · Understanding what's on a pageGoogle does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
- Stage D2-C737 Day 2 · How does the index look like?
A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
repeatsStage D2-C321 Day 2 · Understanding what's on a pageFor languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
- Stage D2-C752 Day 2 · Google Trends
Google handles more than five trillion searches a year, a figure Google announced publicly at the beginning of 2025.
repeats - Stage D2-C847 Day 2 · Welcome to indexing day!
Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.
repeatsStage D1-C317 Day 1 · How crawling worksGoogle's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
- Stage D2-C848 Day 2 · Welcome to indexing day!
Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.
repeatsStage D1-C329 Day 1 · How crawling worksApart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.
- Stage D2-C850 Day 2 · Controlling indexing
John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
repeatsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D2-C851 Day 2 · Controlling indexing
When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.
repeatsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D2-C851 Day 2 · Controlling indexing
When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.
repeatsD2-C021 Day 2 · Welcome to indexing day!Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
- Stage D2-C934 Day 2 · Using images to your advantage and Engaging Search users with videos
The noimageindex robots meta tag tells Google not to index the images on the page.
repeatsStage D2-C099 Day 2 · Controlling indexingThe noimageindex rule tells Google not to index any of the images on the page, and John Mueller said he could not see why a site would want that.
- Stage D2-C936 Day 2 · Using images to your advantage and Engaging Search users with videos
Setting the max-image-preview robots meta tag to large can make content perform surprisingly well in Discover, Gary Illyes said.
repeatsStage D2-C089 Day 2 · Controlling indexingmax-image-preview:large matters mainly in Discover, where it allows a large image that draws people's attention, so the rule can make a page more visible than leaving it out, John Mueller said.
- Stage D3-C056 Day 3 · Making sense of users' queries
Google's generative AI features in Search build on the traditional ways of searching, so query understanding also flows into AI Overviews and AI Mode.
repeatsD1-C051 Day 1 · How Search works and where's AI?Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.
- Stage D3-C058 Day 3 · Making sense of users' queries
In Google's AI features, the query and the search results are additionally sent to an LLM, which returns new queries for Google to run.
repeatsDocs D1-C053 Day 1 · How Search works and where's AI?Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.
- D3-C070 Day 3 · Making sense of users' queries
Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.
repeatsStage D2-C320 Day 2 · Understanding what's on a pageText in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
- Stage D3-C074 Day 3 · Making sense of users' queries
Google's index uses posting lists: for each word, a list of the URLs associated with that word.
repeatsStage D2-C733 Day 2 · How does the index look like?For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.
- Stage D3-C076 Day 3 · Making sense of users' queries
For retrieval, Google uses signals attached individually to each document in the index.
repeatsStage D2-C722 Day 2 · How does the index look like?Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
- Stage D3-C116 Day 3 · Lightning session K: Facets of quality
Google's long-standing advice to write for people and give them what they want has become true in practice because Googlebot has become more and more human, a community speaker argued.
repeatsD1-C047 Day 1 · How Search works and where's AI?Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so they understand and promote natural content better.
- Stage D3-C129 Day 3 · Lightning session K: Facets of quality
If you cannot say what problems your pages solve for users, or the answer is 'it depends', you are optimising for the search engine rather than for people, and that will not work for long, a community speaker warned.
repeatsD1-C013 Day 1 · Welcome and opening keynotesEcosystem principle 4, incentivise high-quality content: content made for Search will not be successful.
- Stage D3-C142 Day 3 · How Google thinks about Quality
Google's quality talk said ranking signals differ by result type: for web pages they include the text on the page, links and passages, while for news, probably, freshness, diversity and originality become more important.
repeatsD1-C045 Day 1 · How Search works and where's AI?Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).
- Stage D3-C190 Day 3 · How Google thinks about Quality
Google's quality talk said quality problems should be treated as quality issues, not as AI versus human content, because a lot of good AI-assisted or AI-written content exists.
repeatsStage D1-C210 Day 1 · How Search works and where's AI?Google said it is not really trying to tell AI-written from human-written content, because it cares more about the quality of content than about how it was created.
- D3-C224 Day 3 · Uncovering Trustworthy Experiences on Discover
Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels and a 16x9 aspect ratio, enabled by the max-image-preview:large setting.
repeatsStage D2-C936 Day 2 · Using images to your advantage and Engaging Search users with videosSetting the max-image-preview robots meta tag to large can make content perform surprisingly well in Discover, Gary Illyes said.
- D3-C237 Day 3 · What are quality updates
Google's slide said Google made more than 4,700 launches to Search in 2023; the speaker rounded this to close to 5,000.
repeatsD3-C149 Day 3 · How Google thinks about QualityGoogle's slide said that in 2023 Google ran 719,326 search quality tests, 124,942 side-by-side experiments and 16,871 live traffic experiments, and made 4,781 launches to Search.
- Stage D3-C256 Day 3 · What are quality updates
Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.
repeatsStage D2-C680 Day 2 · Deciding what goes in the index?Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
- Stage D3-C275 Day 3 · What are quality updates
Google aims to keep more than 99% of search results free from spam and said it already achieves this, thanks to advances in AI.
repeatsStage D2-C676 Day 2 · Calculating (some) signalsGoogle's testing shows that, thanks to SpamBrain, more than 99% of visits from Search are now spam-free.
- Stage D3-C311 Day 3 · How Search results are born
Google handles an image used as a search query much like a text query interpreted as an embedding: the image is broken down into vectors (embeddings) that are then searched for in the index.
repeatsStage D1-C224 Day 1 · How Search works and where's AI?For visual search, Google breaks an image down into vectors, sends the vectors to the index and returns results based on them.
- D3-C321 Day 3 · How Search results are born
Google's serving diagram shows the query passing through query understanding and retrieval to the index, then back through ranking and search features to the user, with the Search Features step highlighted (example query 'Where to eat orange').
repeatsD3-C001 Day 3 · Welcome to serving and ranking day!Google's opening slide for serving day showed serving as the third stage after crawling and indexing, and drew Google's serving infrastructure as query understanding and retrieval leading into the index, then ranking and search features leading back to the user, for the example query 'Where to eat jamon'.
- Stage D3-C329 Day 3 · How Search results are born
Google's structured data feature guide lists the kinds of structured data Google supports with a search feature and what each can do to a site's search results.
repeatsD2-C490 Day 2 · What is Structured Data and why we need it on the internet.Google recommends using the Search gallery in its developer documentation to find the structured data features that suit a site; the gallery shows each feature and how Google uses the markup.
- Stage D3-C405 Day 3 · Shopping on Search: Beyond the blue links
Web markup is an efficient and unambiguous way for sites to share product data with Google, Google Shopping said, repeating the Day 2 structured data talk.
repeatsD2-C469 Day 2 · What is Structured Data and why we need it on the internet.Parsing structured data is significantly cheaper and more efficient than relying on LLMs for every complex extraction task.
- Analysis D3-C435 Day 3 · Inside Search Console: What’s New & How to Use It
Review the filters and regexes that Search Console's AI-powered configuration proposes before trusting the numbers: Google's launch post calls the feature experimental, warns that AI can misinterpret requests, and limits it to the Search results Performance report (not Discover or News) and to configuration, not sorting or exporting.
repeatsStage D1-C243 Day 1 · Lightning session A: Automation and AIThe first rule for using LLMs on search data, according to a community speaker, is never to let the model check its own output.
- Stage D3-C436 Day 3 · Inside Search Console: What’s New & How to Use It
Search Console has a tool with which every property owner can tell Google whether the site's content may be used in AI search results, because Google wants content owners to control their content.
repeatsD1-C029 Day 1 · What's new in the world of SearchSearch Console has a property setting called Search generative AI that gives direct control over AI Overviews and AI Mode without affecting web rankings. Its states are inherit from parent, include and exclude.
- Stage D3-C437 Day 3 · Inside Search Console: What’s New & How to Use It
Search Console added reporting of the impressions a site's content gets in the AI surfaces of Google Search, found in the left navigation nested under Search results.
repeatsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Stage D3-C439 Day 3 · Inside Search Console: What’s New & How to Use It
Search Console's generative AI report shows impressions broken down by pages, countries and devices.
repeatsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Stage D3-C440 Day 3 · Inside Search Console: What’s New & How to Use It
Google called AI reporting in Search Console an evolving space and expects the generative AI report to get richer, with more information in the future.
repeatsStage D1-C196 Day 1 · What's new in the world of SearchGoogle said Search Console's generative AI reporting launched alongside the generative AI control, starting with impressions only, and may expand later.
- D3-C586 Day 3 · Mastering the messy middle
Google's marketing research talk reached the same advice as Search: optimise for people to win in generative AI search.
repeatsD1-C058 Day 1 · How Search works and where's AI?Optimising for people is optimising for generative AI Search.
- D3-C588 Day 3 · Mastering the messy middle
Google's slide advised producing unique, helpful, human-centric content.
repeatsD1-C055 Day 1 · How Search works and where's AI?Myth: build content for every possible consumer need. Google's answer: prioritise unique perspectives, expertise and in-depth experience that go beyond common knowledge.
- Stage D3-C624 Day 3 · How long does it take to..?
Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.
repeats - Stage D3-C680 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Gary Illyes said machine learning is about 50 years old and that Google has been using it for 'probably 30 years', starting with statistical models that made the 'Did you mean' feature possible.
repeatsStage D1-C209 Day 1 · How Search works and where's AI?Google launched the 'Did you mean' feature around 2001-2002 using a statistical model, which Gary Illyes counted as AI because it is a form of machine learning.
- Stage D3-C697 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem, but how it was used and what for.
repeatsStage D3-C190 Day 3 · How Google thinks about QualityGoogle's quality talk said quality problems should be treated as quality issues, not as AI versus human content, because a lot of good AI-assisted or AI-written content exists.
- D3-C701 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's closing slide said AI on Google is just SEO: AI features on Google Search use exactly the same processes as traditional results, so no new acronym is needed, as none was for mobile-first indexing or structured data.
repeatsD1-C050 Day 1 · How Search works and where's AI?Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
- Stage D3-C702 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.
repeatsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D3-C702 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.
repeatsStage D2-C726 Day 2 · How does the index look like?AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
- Stage D3-C704 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google warned that shiny new things, such as trying to work out fan-out queries, distract from the real mission of creating helpful content, since Search is just piping that connects users to information.
repeatsStage D3-C065 Day 3 · Making sense of users' queriesBecause every system runs fan-out differently, Google advised understanding that fan-out happens but not overfocusing on individual fan-out queries or on how to rank for them.
Answers 70
An answer to a question from the audience. The later claim is on the left.
- D1-C047 Day 1 · How Search works and where's AI?
Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so they understand and promote natural content better.
answersD1-C046 Day 1 · How Search works and where's AI?A Google slide headed 'Your Question' raised whether Google can distinguish between AI-written and human-written content, presented on stage as the question this topic usually prompts rather than one asked live.
- Stage D1-C114 Day 1 · Q&A
A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.
answers - Stage D1-C210 Day 1 · How Search works and where's AI?
Google said it is not really trying to tell AI-written from human-written content, because it cares more about the quality of content than about how it was created.
answersD1-C046 Day 1 · How Search works and where's AI?A Google slide headed 'Your Question' raised whether Google can distinguish between AI-written and human-written content, presented on stage as the question this topic usually prompts rather than one asked live.
- Stage D1-C381 Day 1 · How Google thinks about crawl budget
Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.
answersStage D1-C380 Day 1 · How Google thinks about crawl budgetAn audience member asked, in a question submitted before the event, whether crawl budget is still an SEO priority in 2026 or only relevant for very large sites.
- Stage D1-C388 Day 1 · How Google thinks about crawl budget
AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.
answersStage D1-C387 Day 1 · How Google thinks about crawl budgetAn audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.
- Stage D1-C389 Day 1 · How Google thinks about crawl budget
Google expects sites to see more crawling overall, because many other services, including AI services, now crawl the web besides Google.
answersStage D1-C387 Day 1 · How Google thinks about crawl budgetAn audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.
- Stage D1-C391 Day 1 · How Google thinks about crawl budget
A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.
answersStage D1-C390 Day 1 · How Google thinks about crawl budgetAn audience member asked, in a question submitted before the event, whether crawl frequency affects the ranking position of a URL.
- Stage D1-C392 Day 1 · How Google thinks about crawl budget
Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal that people are interested in the site and that its content is of high quality.
answersStage D1-C390 Day 1 · How Google thinks about crawl budgetAn audience member asked, in a question submitted before the event, whether crawl frequency affects the ranking position of a URL.
- Stage D1-C423 Day 1 · Q&A
A Google panelist said many AI crawlers are less sophisticated than search crawlers: in server logs, search crawlers tend to follow where a site changes and which pages are valuable, while AI crawlers may simply work through a site in order, which he put down to less crawling experience and different priorities.
answers - Stage D1-C424 Day 1 · Q&A
Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas many AI systems crawl a site with no understanding of it and take everything.
answers - Stage D1-C425 Day 1 · Q&A
Crawling for AI model training differs from search crawling because training mainly needs a very large number of tokens, and it matters little which pages they come from.
answers - Stage D1-C428 Day 1 · Q&A
Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.
answers - Stage D1-C429 Day 1 · Q&A
A Google panelist said crawlers that ignore robots.txt and cause a nuisance are better treated as a scraping problem than as a crawling problem.
answers - Stage D1-C439 Day 1 · Q&A
Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.
answers - Stage D1-C440 Day 1 · Q&A
If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.
answers - Stage D1-C442 Day 1 · Q&A
Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the site lies; many similar URLs being crawled may not matter if the site is nowhere near that limit.
answers - Stage D1-C443 Day 1 · Q&A
To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued over time, and whether the average response time shows the server is fast enough or is limiting Googlebot.
answers - Stage D1-C444 Day 1 · Q&A
Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.
answers - Stage D1-C445 Day 1 · Q&A
One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.
answers - Stage D1-C449 Day 1 · Q&A
A Google panelist said he expected more AI reporting to launch in Search Console, without a timeline or a promise, because many people at Google care about giving site owners the information to understand and adapt to what happens in Search.
answers - Stage D1-C450 Day 1 · Q&A
Google said reporting on AI features should not simply mirror classic search results, because users get more information before they click and interact with pages differently, so it first has to work out which data would be useful and actionable.
answers - Stage D1-C452 Day 1 · Q&A
Query data for AI features is hard to provide because people ask AI very different kinds of questions that do not map back to keywords as in classic search, and the data would have to be grouped to protect privacy while staying useful.
answers - Stage D1-C454 Day 1 · Q&A
Gary Illyes said llms.txt does not matter for Google Search right now and that he does not expect it to, while adding that he had been wrong before.
answers - Stage D1-C455 Day 1 · Q&A
Gary Illyes said llms.txt matters to some people, which is why tools such as Lighthouse add checks for it, so both sides of the debate are right in their own context.
answers - Stage D1-C456 Day 1 · Q&A
A Google panelist said he knew of no plan for Google to use llms.txt, though it could happen, and that a possible future change is no reason to implement it now.
answers - Stage D1-C457 Day 1 · Q&A
A Google panelist said most llms.txt files he had seen are generated automatically by a setting in SEO plugins, so if Google Search ever made llms.txt matter, a site could add one by ticking a checkbox.
answers - Stage D1-C458 Day 1 · Q&A
A Google panelist objected that llms.txt is designed for supposedly intelligent systems that should be able to parse a website.
answers - Stage D1-C459 Day 1 · Q&A
A Google panelist compared llms.txt to the old meta keywords debate: an AI agent should not blindly trust what a site says about its own authority, just as a site calling itself the best car insurance site is no reason to stop looking at others.
answers - Stage D1-C462 Day 1 · Q&A
Google said it knew of no plans to add crawl stats to the Search Console API, because that would be a completely new part of the API.
answers - Stage D1-C463 Day 1 · Q&A
A Google panelist said adding a report that is a variation of the Performance report to the Search Console API would be easier than adding crawl stats, because the Performance report is already in the API.
answers - Stage D1-C464 Day 1 · Q&A
Google's Search Relations team talks with the Search Console team about the API more than before, partly because with so many vibe-coded systems around it is very easy to just say 'use the Search Console API to do something', and it is annoying when Search Console has no API for that.
answers - Stage D1-C466 Day 1 · Q&A
Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.
answers - Stage D1-C467 Day 1 · Q&A
To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.
answers - Stage D1-C468 Day 1 · Q&A
When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google crawls and recrawls a lot and judging whether they are useful.
answers - Stage D1-C469 Day 1 · Q&A
Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.
answers - Stage D1-C472 Day 1 · Q&A
cats.txt was created by an SEO as a satirical take on llms.txt, to show that crawlers fetch whatever files they are given, so requests for llms.txt in server logs do not mean the file is important.
answers - Stage D1-C473 Day 1 · Q&A
Files such as cats.txt have no importance for Google Search: if a site links to one, Googlebot will find and crawl it, but it has no other effect.
answers - Stage D1-C475 Day 1 · Q&A
There is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that Google can find each of them.
answers - Stage D1-C477 Day 1 · Q&A
In a migration about four or five years before the event, Google consolidated 12 or 14 of its blogs in different languages, a Help Center and its old developers site into one site.
answers - Stage D1-C478 Day 1 · Q&A
The first priority in Google's own site consolidation was to identify the popular URLs people care about and make sure the migration did not damage them.
answers - Stage D1-C479 Day 1 · Q&A
Where content was duplicated across languages, Google's own site consolidation redirected two language versions into one, giving both old URLs one target path to redirect to.
answers - Stage D1-C482 Day 1 · Q&A
Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.
answers - Stage D1-C483 Day 1 · Q&A
A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense yet, because nobody knows how these systems will develop.
answers - Stage D1-C484 Day 1 · Q&A
A Google panelist called blocking all AI crawlers while allowing search crawlers such as Googlebot, Bingbot and Applebot a personal, philosophical decision that every site owner can make.
answers - Stage D1-C485 Day 1 · Q&A
To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
answers - Stage D1-C487 Day 1 · Q&A
A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.
answers - Stage D1-C488 Day 1 · Q&A
When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.
answers - Stage D1-C489 Day 1 · Q&A
A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.
answers - Stage D1-C491 Day 1 · Q&A
If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.
answers - Stage D1-C493 Day 1 · Q&A
There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.
answers - Stage D1-C494 Day 1 · Q&A
To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot opens to the site: if it has dropped, the capacity limit changed.
answers - Stage D1-C540 Day 1 · Q&A
Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.
answers - Stage D1-C546 Day 1 · Q&A
A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).
answers - D2-C002 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
answersD2-C001 Day 2 · Welcome to indexing day!An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
- D2-C003 Day 2 · Welcome to indexing day!
Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have the extra benefit that they may also be useful to users.
answersD2-C001 Day 2 · Welcome to indexing day!An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
- D2-C010 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it depends on how important that product page is to users.
answersD2-C009 Day 2 · Welcome to indexing day!An audience member asked whether a product detail page that is out of stock for two to three months should keep returning 200 with links to similar products, or be 302-redirected to a similar product or to its parent product listing page.
- D2-C011 Day 2 · Welcome to indexing day!
Google's Q&A slide on out-of-stock product pages said users might wait months for some products, or even pre-order them if the site offers pre-ordering.
answersD2-C009 Day 2 · Welcome to indexing day!An audience member asked whether a product detail page that is out of stock for two to three months should keep returning 200 with links to similar products, or be 302-redirected to a similar product or to its parent product listing page.
- D2-C017 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
answersD2-C015 Day 2 · Welcome to indexing day!An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.
- D2-C018 Day 2 · Welcome to indexing day!
Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in Search Console.
answersD2-C015 Day 2 · Welcome to indexing day!An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.
- D2-C020 Day 2 · Welcome to indexing day!
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
answersD2-C019 Day 2 · Welcome to indexing day!An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
- D2-C021 Day 2 · Welcome to indexing day!
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
answersD2-C019 Day 2 · Welcome to indexing day!An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
- D2-C820 Day 2 · Welcome to indexing day!
Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.
answersD2-C001 Day 2 · Welcome to indexing day!An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
- Stage D2-C838 Day 2 · Welcome to indexing day!
Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.
answersD2-C001 Day 2 · Welcome to indexing day!An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.
- Stage D2-C840 Day 2 · Welcome to indexing day!
Google answered that several sloppy migrations on one domain can cause many effects in the short term, and noted that migrations concern indexing as well as crawling, a subject a later Day 2 talk would cover.
answersStage D2-C839 Day 2 · Welcome to indexing day!An audience member asked how several sloppy migrations on the same domain can affect Googlebot's crawling.
- Stage D2-C841 Day 2 · Welcome to indexing day!
Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to them, with easily replaced products, such as a particular cheese, for which the buyer simply picks a similar product instead of waiting.
answersD2-C009 Day 2 · Welcome to indexing day!An audience member asked whether a product detail page that is out of stock for two to three months should keep returning 200 with links to similar products, or be 302-redirected to a similar product or to its parent product listing page.
- Stage D2-C842 Day 2 · Welcome to indexing day!
Google said listing the sitemap in robots.txt is fine, as many websites do.
answersD2-C015 Day 2 · Welcome to indexing day!An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.
- Stage D2-C843 Day 2 · Welcome to indexing day!
Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.
answersD2-C015 Day 2 · Welcome to indexing day!An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.
- Stage D2-C844 Day 2 · Welcome to indexing day!
Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.
answersD2-C019 Day 2 · Welcome to indexing day!An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
- Stage D2-C845 Day 2 · Welcome to indexing day!
Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no figure given).
answersD2-C019 Day 2 · Welcome to indexing day!An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
- Stage D2-C846 Day 2 · Welcome to indexing day!
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
answersD2-C019 Day 2 · Welcome to indexing day!An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.