Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Every claim

All claims

2228 claims in event order. Each one has a permanent link: use its ID, such as D1-C001, to cite it.

Day 1: Crawling 545

11:00 · Welcome and opening keynotes 60

SlideD1-C001

The keynote showed a slide quoting Elizabeth Reid (VP, Search) saying Search is never a solved problem because the internet and the world keep changing.

“Search is never a solved problem. Old challenges evolve and new challenges are constantly popping up — because the internet and the world are always changing.”

Wording checked against the slide or recording

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D1-C182 Day 1: Gary Illyes said that in about 30 years of web search, helping users complete their journeys has made them…
SlideD1-C002

Google named three forces driving the evolution of Search: new platforms (closed-garden apps taking information needs from the open web), new user preferences (personalised, authentic, visual, easy-to-parse content) and new technologies (AI and ML).

Speaker Lino CattaruzziEvidence slide photo, transcript

SlideD1-C003

The stated theme guiding the future of Search is making it effortless to find helpful information.

“Making it effortless to find helpful information”

Wording checked against the slide or recording

Speaker Lino CattaruzziEvidence slide photo, transcript

SlideNot in docsD1-C004

Google says users like AI Overviews and search more because of them, with early feedback overwhelmingly positive and satisfaction highest among 18–24 year olds.

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D1-C184 Day 1: Google said AI Mode is growing almost exponentially and that people are searching more because of it, not…
  • Extended by D3-C569 Day 3: Google's research found that people who use AI run more searches than people who do not.
SlideNot in docsD1-C005

Google says AI Overviews help with new types of questions and lead users to visit a greater diversity of sites, with more total sites shown on the results page.

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D3-C571 Day 3: Google's research found that people who use AI also visit more brand websites, because it is simply easier.
SlideNot in docsD1-C007

Google says 94% of frequent LLM users are also frequent users of Google Search, as evidence that AI is growing search use rather than replacing it.

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D1-C499 Day 1: The keynote added that people who also use a standalone LLM keep certain types of use anchored on Google (the…
SlideNot in docsD1-C008

Google says 71% of shoppers coming to Google Search agree they are open to trying new brands or products (the slide footnote cites a Google-commissioned Ipsos global consumer survey), and that people use Search to figure out what they want, not only to find what they already know.

Speaker Lino CattaruzziEvidence slide photo, transcript

SlideD1-C009

Google quoted Larry Page from 2000 on the ultimate search engine understanding exactly what you mean, adding that it has never been closer to making this a reality.

Speaker Lino CattaruzziEvidence slide photo, transcript

SlideD1-C011

Ecosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such as videos and cards.

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D2-C452 Day 2: Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
  • Extended by D2-C453 Day 2: AI Overviews and AI Mode launched as fairly text-heavy answers with little image content and few tables, and…
SlideD1-C012

Ecosystem principle 3, multiple user needs: some users want links, some quick answers, some depth; Google says it will keep prioritising traffic to the ecosystem.

Speaker Lino CattaruzziEvidence slide photo, transcript

SlideConsistent with docsD1-C013

Ecosystem principle 4, incentivise high-quality content: content made for Search will not be successful.

“Made for Search" content will not be successful, as we'll work to incentivise high-quality content.”

Wording checked against the slide or recording

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Repeated by D3-C129 Day 3: If you cannot say what problems your pages solve for users, or the answer is 'it depends', you are optimising…
SlideD1-C014

Ecosystem principle 5, traffic patterns may fluctuate: long-held traffic patterns are likely to change, which Google frames as new opportunities for all sites.

Speaker Lino CattaruzziEvidence slide photo, transcript

StageD1-C059

The opening keynote closed with the advice to think about UEO, user engine optimisation, next to SEO and GEO: focus on the user and the rest will follow.

Speaker Lino CattaruzziEvidence notes, transcript

Things
  • Repeated by D2-C679 Day 2: Combing Google's documentation for signals is not the best use of an SEO's time; creating content that users…
StageD1-C142

The event host stressed that Search Central Live is about organic Google Search, which is separate from Google Ads.

Speaker not identifiedEvidence transcript

StageD1-C145

The opening keynote said this is probably the most exciting time ever in the history of the web and of search.

Speaker Lino CattaruzziEvidence transcript

StageD1-C146

The keynote named AI as the first of three drivers of change in Search: like mobile and social media before it, AI will change the way users search, and the keynote called it a profound inflection point for the industry, to be handled responsibly together with the ecosystem.

Speaker Lino CattaruzziEvidence transcript

StageD1-C147

The keynote's second driver was changing content consumption: users expect content that is instant, engaging and interactive, and user-generated content is the way young audiences engage the most.

Speaker Lino CattaruzziEvidence transcript

StageD1-C148

The keynote's third driver was that information quality remains a core focus for Google, which wants high-quality sources to serve across Google Search, its AI experiences in Search, YouTube and Discover.

Speaker Lino CattaruzziEvidence transcript

StageConfirmed by docsD1-C150

Google said it has made the biggest change to its search box in the last 25 years.

Speaker Lino CattaruzziEvidence transcript

  • Extended by D1-C187 Day 1: Google presented the intelligent search box as one of its recent Search launches; it accepts long queries and…
StageD1-C153

The keynote said AI answers change the dynamic of a search, because a response leads to follow-up after follow-up.

Speaker Lino CattaruzziEvidence transcript

StageNot in docsD1-C155

The keynote cited Google CEO Sundar Pichai as saying Google is the biggest contributor of clicks to the open web and wants to remain so.

Speaker Lino CattaruzziEvidence transcript

StageD1-C157

The keynote described a growing serving challenge: new content and creators keep appearing and trending, and Google has to integrate this new content with existing content to answer each query, with the volume changing faster than ever.

Speaker Lino CattaruzziEvidence transcript

StageNot in docsD1-C158

Google said that even the classic ten-blue-links layout was settled only after millions of experiments, as part of using as much data as possible for product decisions.

Speaker Lino CattaruzziEvidence transcript

  • Extended by D2-C452 Day 2: Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
  • Extended by D3-C147 Day 3: Google's quality talk said that at any moment thousands of Search experiments are probably running.
StageD1-C159

The keynote's personal favourites in the evolution of Search were Google Translate, voice search, the Knowledge Graph (relating entities to give context) and Google Lens.

Speaker Lino CattaruzziEvidence transcript

StageConsistent with docsD1-C161

Google said it created the Transformer architecture (the paper 'Attention Is All You Need', the T in ChatGPT) and published it to move the industry forward, competitors included.

Speaker Lino CattaruzziEvidence transcript

Used byglossary term Transformer

  • Extended by D1-C199 Day 1: Gary Illyes said searching by image or by video, as Search offers it today, was not possible before…
StageNot in docsD1-C162

Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the number one app in the app store in 2025.

Speaker Lino CattaruzziEvidence transcript

Things
  • Extended by D2-C770 Day 2: Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of…
StageNot in docsD1-C166

Google said 18 to 24 year olds are more engaged than ever with the AI-powered features in Search: they requery and ask more, and more complex, questions, which leads to more searches.

Speaker Lino CattaruzziEvidence transcript

StageNot in docsD1-C167

Google said its searches are growing year on year in both commercial and non-commercial queries.

Speaker Lino CattaruzziEvidence transcript

StageD1-C170

The keynote said monetisation, such as shoppers being open to new brands, is key to the sustainability of the open web.

Speaker Lino CattaruzziEvidence transcript

StageD1-C171

Google's examples of brainstorming queries were a request for birthday gift ideas for a son turning 18 and, instead of asking for the nearest pizza, describing a party of fifteen guests that needs healthy, warm food delivered in under twenty minutes.

Speaker Lino CattaruzziEvidence transcript

StageConsistent with docsD1-C172

Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.

Speaker Lino CattaruzziEvidence transcript

  • Extended by D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
  • Extended by D3-C059 Day 3: Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding…
StageConsistent with docsD1-C175

Explaining the principle of incentivising high-quality content, Google said the content may be created by humans or by AI: what matters is that it is high quality and made for users, not for search.

“we want to incentivise high-quality content, created by humans or by AI. But the key is that it's high quality”

Wording checked against the slide or recording

Speaker Lino CattaruzziEvidence transcript

  • Repeated by D1-C210 Day 1: Google said it is not really trying to tell AI-written from human-written content, because it cares more…
StageD1-C177

The event host said the event would not reveal a lot of secrets of the Google Search formula; its goal was to put publicly available information into context and show how the pieces fit together.

Speaker not identifiedEvidence transcript

StageConfirmed by docsD1-C178

Under Google's Honest Results Policy, no website gets an advantage in Search because it is a Google client, a Google partner or has a personal relationship with Google.

“No website should have an advantage in search because it is a Google client, a Google partner, or because of a personal relationship.”

Speaker not identifiedEvidence transcript

Used byglossary term Honest Results Policy

StageD1-C180

Because of the Honest Results Policy, the information shared at Search Central Live is information available to everyone, giving everyone the same opportunity to understand how Search works.

Speaker not identifiedEvidence transcript

AnalysisD1-C498

Cite YouTube's lead as US streaming watch time measured by Nielsen: that is the figure YouTube publishes, and no published Google source backs the keynote's 'around the world'.

Author Ibrahim Anjro

Things
  • Extends D1-C497 Day 1: Google said YouTube has surpassed Netflix as the number one streaming service, in the US and around the world.
StageNot in docsD1-C499

The keynote added that people who also use a standalone LLM keep certain types of use anchored on Google (the example given is unclear in both recordings).

Speaker Lino CattaruzziEvidence transcript

  • Extends D1-C007 Day 1: Google says 94% of frequent LLM users are also frequent users of Google Search, as evidence that AI is…

11:35 · What's new in the world of Search 41

SlideNot in docsD1-C016

AI Mode queries have doubled every quarter since launch.

Speaker Gary IllyesEvidence slide photo

Things
  • Extended by D1-C184 Day 1: Google said AI Mode is growing almost exponentially and that people are searching more because of it, not…
SlideNot in docsD1-C017

The average AI Mode query is about 3 times the length of a traditional search (the opening keynote's slide and its speaker said two to three times).

Speaker Gary IllyesEvidence 2 slide photos, transcript

Things

Used byglossary term AI Mode

SlideNot in docsD1-C018

One in six AI Mode searches is multimodal, using voice or images.

Speaker Gary IllyesEvidence slide photo

Things
  • Extended by D3-C562 Day 3: Google's research also counted easy access as part of AI's ease benefit: multimodal input, such as taking…
SlideNot in docsD1-C020

Google describes four new ways people search in AI Mode: Explore (brainstorming queries, grown 30% faster than AI Mode queries overall in the US since launch), Learn (study guides, deep dives), Decide (searches beginning with 'which', up 40% faster than AI Mode queries overall in the past six months) and Do (planning queries, up 80% faster than AI Mode queries overall in the past six months).

Speaker Gary IllyesEvidence 2 slide photos, transcript

Things
SlideNot in docsD1-C021

Queries of five or more words are growing in volume 1.5 times faster than shorter queries. The slide footnote cites Google internal data on global English-language queries, for a comparison period that starts in November 2022 (the remaining dates are only partly legible).

Speaker Gary IllyesEvidence slide photo

  • Extended by D1-C185 Day 1: Google said a multi-part query, such as a white three-row SUV for two teenagers and a car-sick dog, cannot be…
SlideNot in docsD1-C022

Search Agents in AI Mode let users set standing requests for updates, such as being told when a product goes on sale, a team wins or a price drops below a level.

Speaker Gary IllyesEvidence slide photo, transcript

Things
  • Extended by D1-C192 Day 1: Google said Search agents in AI Mode had launched globally a day or two before 30 September 2026.
  • Extended by D1-C194 Day 1: A Search agent is set up by typing a natural-language request, which AI Mode interprets to create an alert…
SlideNot in docsD1-C023

Google listed eight ways AI Search helps users find and visit websites: preferred sources expansion, fresh perspectives and prominent links, subscription content, in-line quotes in AI Mode and AI Overviews, rich linking and embedded web, context on content sources, 'Highly Cited' labels and Search profiles.

Speaker Gary IllyesEvidence slide photo, transcript

SlideNot in docsD1-C024

Users can link their subscriptions so answers highlight paywalled content they already pay for; The Indian Express saw 34% more engagement among linked subscribers.

Speaker Gary IllyesEvidence slide photo, transcript

SlideConfirmed by docsD1-C025

Publishers and creators can claim a dedicated Search profile with a name, profile photo, cover image, bio, connected social and website accounts, pinned posts and links, and users can follow it.

Speaker Gary IllyesEvidence slide photo

Used byglossary term Search profile

  • Extended by D3-C468 Day 3: Google's social and video performance guide says that if you already claimed your Search profile, all of its…
DocsSourceD1-C026

Search profiles launched on 4 June 2026, in the US first, for creators and publishers with a sizable following on at least one major social or video platform. They appear in knowledge panels and Discover.

Publisher Google blog (4 June 2026)

  • Updated by D1-C125 Day 1: Since 16 September 2026, US publishers and creators qualify for a Search profile with 10,000 followers across…
DocsSourceD1-C125

Since 16 September 2026, US publishers and creators qualify for a Search profile with 10,000 followers across YouTube, Instagram, X or TikTok, and media organisations can claim and manage profiles for all their sub-brands from one login.

“We've lowered the eligibility threshold to 10,000 followers across YouTube, Instagram, X, or TikTok.”

Publisher Google blog (16 September 2026)

Used byglossary term Search profile

  • Updates D1-C026 Day 1: Search profiles launched on 4 June 2026, in the US first, for creators and publishers with a sizable…
SlideNot in docsD1-C027

Sites that users mark as a preferred source are more visible in Top Stories and are labelled in AI Mode and AI Overviews.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extended by D3-C317 Day 3: Google can add elements to a text result, such as a 'highly cited' badge or a preferred-source badge.
SlideConfirmed by docsD1-C029

Search Console has a property setting called Search generative AI that gives direct control over AI Overviews and AI Mode without affecting web rankings. Its states are inherit from parent, include and exclude.

“Direct control over AI Overviews and AI Mode without impacting web rankings.”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-IDX-10fact F-012glossary term Search generative AI control

  • Extended by D2-C103 Day 2: Setting the Search generative AI control in Search Console to exclude keeps a site's links and content out of…
  • Repeated by D3-C436 Day 3: Search Console has a tool with which every property owner can tell Google whether the site's content may be…
DocsSourceD1-C030

The Search generative AI control covers AI Overviews, AI Mode and generative AI features in Discover. Excluding removes both links to the site and use of its content to ground answers.

Publisher Google Search Console Help

Used byrequirement DEV-IDX-10glossary term Search generative AI control

  • Repeated by D2-C103 Day 2: Setting the Search generative AI control in Search Console to exclude keeps a site's links and content out of…
AnalysisD1-C034

Lead-generation, local-service and e-commerce sites should normally stay included, because AI answers cite and link sources. Sites selling paywalled or licensed content should test exclusion on a child property first. Use nosnippet only if you also want out of regular snippets.

Author Ibrahim Anjro

Used byrequirement DEV-IDX-10

  • Extended by D2-C072 Day 2: The nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI…
StageConsistent with docsD1-C049

Gary Illyes argued that GEO is a label invented to create a new field and is not needed. Understanding how SEO works is enough.

Speaker Gary IllyesEvidence notes, transcript

Things

Used bystory angle A-001

  • Contradicted by D1-C296 Day 1: A community speaker argued for adopting the GEO label as the industry's chance to leave behind the bad…
StageD1-C182

Gary Illyes said that in about 30 years of web search, helping users complete their journeys has made them search more, not less, unlike a product whose problem is solved once and for all.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C001 Day 1: The keynote showed a slide quoting Elizabeth Reid (VP, Search) saying Search is never a solved problem…
StageNot in docsD1-C183

Google said Gen Z users are a very large part of Google Search's user base and search differently from older users, for example with images.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD1-C184

Google said AI Mode is growing almost exponentially and that people are searching more because of it, not less.

Speaker Gary IllyesEvidence transcript

Things
  • Extends D1-C016 Day 1: AI Mode queries have doubled every quarter since launch.
  • Extends D1-C004 Day 1: Google says users like AI Overviews and search more because of them, with early feedback overwhelmingly…
StageConsistent with docsD1-C185

Google said a multi-part query, such as a white three-row SUV for two teenagers and a car-sick dog, cannot be satisfied by a list of links because no single article covers every aspect, while AI Overviews and AI Mode can synthesise an answer that helps users decide where to go.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C021 Day 1: Queries of five or more words are growing in volume 1.5 times faster than shorter queries. The slide footnote…
StageD1-C186

Gary Illyes said that in the mid-1990s searchers had to speak machine, typing queries like 'weather Barcelona rain', whereas today machines understand how people naturally speak.

Speaker Gary IllyesEvidence transcript

StageConfirmed by docsD1-C187

Google presented the intelligent search box as one of its recent Search launches; it accepts long queries and keeps expanding to make room for them.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C150 Day 1: Google said it has made the biggest change to its search box in the last 25 years.
StageNot in docsD1-C188

Google said it noticed young users uploading pictures to reverse image search and expecting results that answer something about the picture, and that Google Lens was its answer to that behaviour.

Speaker Gary IllyesEvidence transcript

Used byglossary term Google Lens

StageNot in docsD1-C192

Google said Search agents in AI Mode had launched globally a day or two before 30 September 2026.

Speaker Gary IllyesEvidence transcript

Things

Used byglossary term Search agents (information agents)

  • Extends D1-C022 Day 1: Search Agents in AI Mode let users set standing requests for updates, such as being told when a product goes…
PressSourceD1-C193

A search marketing agency reported on 29 September 2026 that Google's VP of Product for Search announced on X, on 28 September 2026, that AI Mode information monitoring was rolling out to everyone globally in the Google app; Google's I/O 2026 post had launched information agents first for AI Pro and Ultra subscribers in summer 2026.

Reported by Relevant Audience (29 September 2026), Google blog (19 May 2026)

Things
StageConsistent with docsD1-C194

A Search agent is set up by typing a natural-language request, which AI Mode interprets to create an alert; Google called this an information agent.

Speaker Gary IllyesEvidence transcript

Things

Used byglossary term Search agents (information agents)

  • Extends D1-C022 Day 1: Search Agents in AI Mode let users set standing requests for updates, such as being told when a product goes…
StageConsistent with docsD1-C195

Google said it would roll out agentic capabilities that allow bookings later in 2026, in the US only at first.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD1-C196

Google said Search Console's generative AI reporting launched alongside the generative AI control, starting with impressions only, and may expand later.

Speaker Gary IllyesEvidence transcript

  • Repeats D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
  • Repeated by D3-C440 Day 3: Google called AI reporting in Search Console an evolving space and expects the generative AI report to get…
StageD1-C197

Gary Illyes traced 'SEO is dead' predictions back to a post titled 'Search engine R.I.P.' on an online advertising forum on 10 November 1997, and said such posts still appear every month or even every week.

Speaker Gary IllyesEvidence transcript

Used bystory angle A-001

StageD1-C198

Gary Illyes said SEO is not dead and that the new AI features only create more opportunities for site owners to take advantage of.

Speaker Gary IllyesEvidence transcript

Used bystory angle A-001

  • Extends D1-C050 Day 1: Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
StageNot in docsD1-C199

Gary Illyes said searching by image or by video, as Search offers it today, was not possible before transformers were invented.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C161 Day 1: Google said it created the Transformer architecture (the paper 'Attention Is All You Need', the T in ChatGPT)…
AnalysisD1-C200

Visual search predates transformers in Google's own history: its Search timeline dates Google Images to 2001 and Google Lens to 2017, the year Google Research published the Transformer, so read the stage remark as being about today's multimodal AI search.

Author Ibrahim Anjro

11:45 · How Search works and where's AI? 55

SlideConfirmed by docsD1-C036

Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.

Speaker Cherry Prommawin, Gary IllyesEvidence 2 slide photos, transcript

  • Extended by D1-C214 Day 1: Crawling and indexing happen before anyone searches, while serving and ranking happen in real time when…
  • Extended by D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Repeated by D2-C719 Day 2: The speaker recapped Google's pipeline up to the index: Google crawls pages, processes the fetched documents…
SlideConsistent with docsD1-C037

For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Extended by D1-C205 Day 1: Signals calculated for a page during indexing are stored in the index and used both to decide whether the…
  • Extended by D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
  • Extended by D2-C344 Day 2: Google deduplicates pages because many sites have very many pages and Google's index does not have room for…
  • Extended by D2-C648 Day 2: Among the many signals Google calculates during indexing, the ones singled out as having large effects on…
  • Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
  • Extended by D2-C684 Day 2: Index selection is a predictive AI system that relies heavily on machine learning.
  • Extended by D2-C722 Day 2: Each document in Google's index has pretty much all the signals calculated for it attached, for example…
SlideConsistent with docsD1-C038

AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-AIF-01story angle A-001glossary terms AI Mode, Grounding

  • Extended by D1-C208 Day 1: AI Overviews and AI Mode may have extra processes of their own, like any other search feature, but the bulk…
  • Extended by D1-C325 Day 1: Googlebot is the crawler Google uses for web search, including Search's AI features.
  • Extended by D1-C388 Day 1: AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl…
  • Extended by D2-C073 Day 2: John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page…
  • Extended by D2-C074 Day 2: Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a…
  • Repeated by D2-C116 Day 2: AI Overviews and AI Mode are built on top of Search results: they are a different experience of the same…
  • Extended by D2-C476 Day 2: The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning…
  • Extended by D2-C603 Day 2: Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is…
  • Extended by D2-C726 Day 2: AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a…
  • Extended by D3-C057 Day 3: Google's slide on how LLM features with grounding generally work showed a query going to both the search…
  • Extended by D3-C701 Day 3: Google's closing slide said AI on Google is just SEO: AI features on Google Search use exactly the same…
  • Repeated by D3-C702 Day 3: AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have…
SlideConsistent with docsD1-C039

Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Extended by D1-C226 Day 1: Google said it does not own third-party AI chatbots such as ChatGPT and has no insight into how they work or…
  • Extended by D1-C335 Day 1: Crawling for Gemini may be set to care less about quality and more about the amount of content, because for…
  • Extended by D2-C118 Day 2: To train Gemini models, Google renders every page just as it does for Search, so a page that renders…
  • Extended by D2-C120 Day 2: When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML…
  • Contradicted by D2-C325 Day 2: Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for…
  • Extended by D2-C669 Day 2: SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically…
SlideConfirmed by docsD1-C040

URL discovery works through links: a homepage links to section pages, which link to further pages.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-URL-04

  • Extended by D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
  • Extended by D1-C336 Day 1: Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from…
  • Extended by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
  • Extended by D2-C038 Day 2: Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure…
  • Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
  • Extended by D2-C169 Day 2: Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before…
  • Extended by D2-C287 Day 2: A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google…
SlideNot in docsD1-C041

Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Extended by D1-C209 Day 1: Google launched the 'Did you mean' feature around 2001-2002 using a statistical model, which Gary Illyes…
  • Extended by D2-C668 Day 2: Google uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is…
  • Extended by D2-C669 Day 2: SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically…
  • Extended by D3-C680 Day 3: Gary Illyes said machine learning is about 50 years old and that Google has been using it for 'probably 30…
SlideConsistent with docsD1-C042

BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

Used byglossary term BERT, RankBrain and MUM

  • Extended by D2-C338 Day 2: Google detects soft 404s with a language model, described as something like BERT, that is trained to…
  • Extended by D3-C684 Day 3: Predictive language models have been used a lot in Search, and BERT is one: technically, in the purest sense…
DocsSourceD1-C133

Google's ranking systems guide lists BERT among its ranking systems, as an AI system that helps Google understand how combinations of words express different meanings and intent.

“an AI system Google uses that allows us to understand how combinations of words express different meanings and intent”

Publisher Google Search Central

SlideConsistent with docsD1-C043

RankBrain is used in serving to interpret the intent behind queries, especially new or unusual ones, and match them to relevant results.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo

Used byglossary term BERT, RankBrain and MUM

  • Extended by D1-C217 Day 1: Google began talking publicly about its use of AI in Search around 2015-2016, and RankBrain was the first…
SlideNot in docsD1-C044

MUM (Multitask Unified Model) understands information across text, images, audio and video, and processes information in more than 75 languages.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo

  • Extended by D1-C218 Day 1: Google said MUM helps it understand the context of the words in a query: a search for hiking shoes that…
DocsSourceD1-C134

Google's May 2021 announcement of MUM says it is trained across 75 languages and understands information across text and images, and that it could expand to other modalities such as video and audio in the future.

“MUM is multimodal, so it understands information across text and images and, in the future, can expand to more modalities like video and audio.”

Publisher Google blog (18 May 2021)

DocsSourceD1-C130

Google's ranking systems guide says MUM is not currently used for general ranking in Search, only for specific applications such as COVID-19 vaccine searches and featured snippet callouts.

“It's not currently used for general ranking in Search”

Publisher Google Search Central

Used byglossary term BERT, RankBrain and MUM

  • Extended by D3-C135 Day 3: The quality talk's slide listed MUM among Google's ranking systems, while Google's ranking systems guide says…
SlideNot in docsD1-C045

Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Repeated by D2-C537 Day 2: The text around an image is critical: Google uses it as context to understand the image and to rank it, so an…
  • Extended by D2-C656 Day 2: Freshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking…
  • Repeated by D3-C142 Day 3: Google's quality talk said ranking signals differ by result type: for web pages they include the text on the…
SlideD1-C046

A Google slide headed 'Your Question' raised whether Google can distinguish between AI-written and human-written content, presented on stage as the question this topic usually prompts rather than one asked live.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Answered by D1-C047 Day 1: Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so…
  • Answered by D1-C210 Day 1: Google said it is not really trying to tell AI-written from human-written content, because it cares more…
SlideNot in docsD1-C047

Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so they understand and promote natural content better.

“ML based ranking algorithms and signals are trained on content by humans for humans. They "understand" and promote natural content better.”

Wording checked against the slide or recording

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Answers D1-C046 Day 1: A Google slide headed 'Your Question' raised whether Google can distinguish between AI-written and…
  • Extended by D1-C220 Day 1: Google said the more important documents in its index naturally tend to be documents created by humans.
  • Extended by D1-C221 Day 1: Google said the natural content its ranking algorithms understand and promote better includes content created…
  • Repeated by D3-C116 Day 3: Google's long-standing advice to write for people and give them what they want has become true in practice…
AnalysisD1-C048

The answer is not a claim that Google detects AI text. It says ranking favours text that reads as natural to people. The risk with AI content is scale without value, which falls under Google's scaled content abuse policy, not the tool itself.

Author Ibrahim Anjro

  • Extended by D2-C612 Day 2: Treat machine translation as a first draft: have a native speaker review it and adapt dates, calendars and…
SlideConfirmed by docsD1-C050

Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.

“Don't worry about what to call it. Good SEO is good GEO, AEO.”

Wording checked against the slide or recording

Speaker Cherry Prommawin, Gary IllyesEvidence 2 slide photos, transcript

Things

Used bymyth M-001quote Q-001story angle A-001glossary term GEO and AEO

  • Extended by D1-C198 Day 1: Gary Illyes said SEO is not dead and that the new AI features only create more opportunities for site owners…
  • Contradicted by D1-C296 Day 1: A community speaker argued for adopting the GEO label as the industry's chance to leave behind the bad…
  • Repeated by D2-C601 Day 2: Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results…
  • Repeated by D3-C701 Day 3: Google's closing slide said AI on Google is just SEO: AI features on Google Search use exactly the same…
SlideConfirmed by docsD1-C051

Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-AIF-01myth M-001story angle A-001glossary term AI Overviews

  • Extended by D1-C225 Day 1: Google said query fan-out is nothing new: it fires, for example, ten different searches in the background…
  • Repeated by D2-C601 Day 2: Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results…
  • Extended by D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
  • Repeated by D3-C056 Day 3: Google's generative AI features in Search build on the traditional ways of searching, so query understanding…
DocsSourceD1-C053

Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.

Publisher Google Search Central

Used byglossary term Query fan-out

  • Extended by D1-C412 Day 1: A community speaker described AI search as turning one query into many related searches, including searches…
  • Extended by D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
  • Extended by D2-C729 Day 2: Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several…
  • Repeated by D3-C058 Day 3: In Google's AI features, the query and the search results are additionally sent to an LLM, which returns new…
DocsSourceD1-C129

Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.

Publisher Google Search Central

Used byrequirement DEV-AIF-03glossary term Scaled content abuse

  • Extended by D2-C741 Day 2: In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the…
  • Extended by D3-C065 Day 3: Because every system runs fan-out differently, Google advised understanding that fan-out happens but not…
  • Extended by D3-C212 Day 3: Google's quality talk said Google clarified that traditional spam techniques aimed at manipulating AI…
SlideConfirmed by docsD1-C054

Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.

“LLMS.txt isn't necessary.”

Wording checked against the slide or recording

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

Things

Used byrequirements DEV-AIF-02, DEV-AIF-03myth M-002story angle A-001glossary term llms.txt

  • Repeated by D1-C271 Day 1: A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard…
  • Extended by D2-C330 Day 2: Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real…
  • Extended by D2-C332 Day 2: Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller…
SlideConfirmed by docsD1-C055

Myth: build content for every possible consumer need. Google's answer: prioritise unique perspectives, expertise and in-depth experience that go beyond common knowledge.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Extended by D3-C179 Day 3: Google's quality talk said non-commodity content offers a unique, experienced take, and advised writing…
  • Repeated by D3-C588 Day 3: Google's slide advised producing unique, helpful, human-centric content.
SlideConsistent with docsD1-C056

Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Extended by D2-C124 Day 2: Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render…
  • Extended by D2-C134 Day 2: Google urged sites to make sure their JavaScript content can be crawled, rendered and indexed, calling this…
  • Extended by D2-C139 Day 2: A community speaker strongly advised putting everything you want cited into the raw, server-side rendered…
SlideD1-C057

Myth: the old metrics don't work in the AI era. Google's answer: measure success through metrics that matter to your business.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Extended by D3-C518 Day 3: A community speaker concluded that rankings and clicks no longer equal real business outcomes.
  • Extended by D3-C536 Day 3: A May 2025 Search Central blog post advised site owners to look at the overall value of visits from Search…
  • Extended by D3-C542 Day 3: Google's own line on Day 1 was to measure success by metrics that matter to the business, and its Generative…
SlideConsistent with docsD1-C058

Optimising for people is optimising for generative AI Search.

“Optimizing for people is optimizing for Generative AI Search”

Wording checked against the slide or recording

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

Used byquote Q-002

  • Extended by D3-C497 Day 3: Pages that do better in AI answers than in classic results do not refute Google's line that optimising for…
  • Repeated by D3-C586 Day 3: Google's marketing research talk reached the same advice as Search: optimise for people to win in generative…
DocsSourceD1-C061

Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in Search, and that such files neither harm nor help visibility because Google Search ignores them.

Publisher Google Search Central

Used byrequirement DEV-AIF-02story angle A-001glossary term llms.txt

  • Extended by D2-C477 Day 2: As far as the speaker knows, in most of Google's main AI uses a page's schema.org markup is not turned into…
  • Extended by D2-C480 Day 2: Google's guide to optimizing for generative AI features lists 'overfocusing on structured data' among the…
DocsSourceD1-C062

Google's guide for generative AI features recommends the Generative AI performance report in Search Console for measuring how content performs in generative AI features on Google Search and Discover.

Publisher Google Search Central

Used byrequirement DEV-MON-07glossary term Generative AI performance report

  • Extended by D3-C535 Day 3: Beyond Search Console, Google's AI features guide suggests tracking conversions and time spent on the site in…
DocsSourceD1-C124

Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.

Publisher Search Central blog (3 June 2026)

Used byrequirement DEV-MON-07glossary term Generative AI performance report

  • Repeated by D1-C196 Day 1: Google said Search Console's generative AI reporting launched alongside the generative AI control, starting…
  • Extended by D2-C607 Day 2: Check AI Overviews and AI Mode with native-language queries in each target market, not with translated…
  • Extended by D3-C063 Day 3: Google said fan-out queries are not added to Search Console, because Google considers them part of its…
  • Repeated by D3-C437 Day 3: Search Console added reporting of the impressions a site's content gets in the AI surfaces of Google Search…
  • Repeated by D3-C439 Day 3: Search Console's generative AI report shows impressions broken down by pages, countries and devices.
  • Extended by D3-C495 Day 3: Nik Vujic said he hopes Search Console's generative AI report will get query data.
  • Extended by D3-C542 Day 3: Google's own line on Day 1 was to measure success by metrics that matter to the business, and its Generative…
StageNot in docsD1-C201

Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.

Speaker Cherry PrommawinEvidence transcript

  • Extended by D3-C602 Day 3: Google said it knows hundreds of trillions of URLs (as of October 2026).
StageConsistent with docsD1-C202

Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-URL-04glossary term Hub pages

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extended by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
StageConsistent with docsD1-C205

Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.

Speaker Cherry PrommawinEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Extended by D2-C687 Day 2: Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
StageConfirmed by docsD1-C206

Google renders JavaScript-heavy pages from their HTML, CSS and JavaScript as a browser would, using the latest version of Chromium.

Speaker Cherry PrommawinEvidence transcript

  • Repeated by D2-C126 Day 2: Google renders pages with Chromium, the browser technology that also underlies Chrome, Edge and other…
  • Extended by D2-C173 Day 2: Google's renderer is a headless Chromium that runs the page's JavaScript.
StageConfirmed by docsD1-C207

Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so that users are not shown duplicates.

Speaker Cherry PrommawinEvidence transcript

  • Repeats D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
  • Extended by D2-C346 Day 2: For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of…
StageConsistent with docsD1-C208

AI Overviews and AI Mode may have extra processes of their own, like any other search feature, but the bulk of their processing is the same as for normal Search.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageConfirmed by docsD1-C209

Google launched the 'Did you mean' feature around 2001-2002 using a statistical model, which Gary Illyes counted as AI because it is a form of machine learning.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
  • Repeated by D3-C680 Day 3: Gary Illyes said machine learning is about 50 years old and that Google has been using it for 'probably 30…
StageConfirmed by docsD1-C210

Google said it is not really trying to tell AI-written from human-written content, because it cares more about the quality of content than about how it was created.

“We care more about the quality of the content than how it was created.”

Speaker Gary IllyesEvidence transcript

  • Answers D1-C046 Day 1: A Google slide headed 'Your Question' raised whether Google can distinguish between AI-written and…
  • Repeats D1-C175 Day 1: Explaining the principle of incentivising high-quality content, Google said the content may be created by…
  • Repeated by D3-C190 Day 3: Google's quality talk said quality problems should be treated as quality issues, not as AI versus human…
StageConsistent with docsD1-C211

Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.

Speaker Cherry PrommawinEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Extended by D2-C682 Day 2: Google's index selection system calculates thresholds and decides which documents are kept and which are…
  • Extended by D2-C688 Day 2: Index selection is the last step before documents enter Google's index.
StageNot in docsD1-C213

Google said its index, printed on paper, would reach the Moon and back twelve times.

Speaker Cherry PrommawinEvidence transcript

StageConsistent with docsD1-C215

Serving starts with interpreting the query, which includes cleaning it up, detecting its language and expanding it.

Speaker Cherry PrommawinEvidence transcript

  • Extended by D3-C007 Day 3: Query language detection works poorly when someone searches only for a brand name, such as Facebook or…
StageConsistent with docsD1-C217

Google began talking publicly about its use of AI in Search around 2015-2016, and RankBrain was the first such system it publicised widely.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C043 Day 1: RankBrain is used in serving to interpret the intent behind queries, especially new or unusual ones, and…
StageConsistent with docsD1-C218

Google said MUM helps it understand the context of the words in a query: a search for hiking shoes that mentions Mount Everest rather than Kilimanjaro gets results suited to that climb.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C044 Day 1: MUM (Multitask Unified Model) understands information across text, images, audio and video, and processes…
StageNot in docsD1-C220

Google said the more important documents in its index naturally tend to be documents created by humans.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C047 Day 1: Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so…
StageNot in docsD1-C221

Google said the natural content its ranking algorithms understand and promote better includes content created by humans or at least edited and reviewed by them.

“content that was created by humans, or at least edited and reviewed”

Speaker Gary IllyesEvidence transcript

  • Extends D1-C047 Day 1: Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so…
StageNot in docsD1-C224

For visual search, Google breaks an image down into vectors, sends the vectors to the index and returns results based on them.

Speaker Gary IllyesEvidence transcript

  • Repeated by D3-C311 Day 3: Google handles an image used as a search query much like a text query interpreted as an embedding: the image…
StageConsistent with docsD1-C225

Google said query fan-out is nothing new: it fires, for example, ten different searches in the background, each a normal search on the same systems.

“it is just ten different searches that fire in the background”

Speaker Gary IllyesEvidence transcript

  • Extends D1-C051 Day 1: Three reasons were given: generative AI features are built directly on the core ranking systems, query…
  • Extended by D3-C061 Day 3: Google tries to make fan-out queries distinct from each other for better coverage, avoiding asking the same…
StageD1-C226

Google said it does not own third-party AI chatbots such as ChatGPT and has no insight into how they work or are built.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…

13:10 · Lightning session A: Automation and AI 98

StageD1-C227

Google received close to 200 submissions for the community lightning talks of the Deep Dive, reviewed them over about a month, and gave each selected community speaker seven minutes on stage.

Speaker not identifiedEvidence transcript

  • Extended by D3-C708 Day 3: Google does not tell lightning-talk and poster speakers what to talk about; they bring their own ideas.
StageD1-C228

Google's hosts said the lightning talks exist because the SEO community holds a lot of knowledge, some of it untapped, and highlighting the people who have it serves the community best.

Speaker not identifiedEvidence transcript

StageD1-C229

A community speaker's approach to AI in agency SEO work is to use it where it meaningfully elevates and enriches existing work rather than replacing it, keeping full strategic oversight and control of the output.

“Human judgment and strategy is the default, and AI simply pushes it at scale.”

Wording checked against the slide or recording

Speaker James PowleyEvidence transcript

StageD1-C230

Asking AI to produce whole SEO deliverables from scratch, such as creating content or auditing a website, is where people get stuck; break a task into its elements and decide where AI can and should play a part, a community speaker advised.

Speaker James PowleyEvidence transcript

StageD1-C231

A community speaker designs AI-assisted deliverables like a flowchart, keeping guardrails in place and keeping the prompts succinct and measured, so the team controls the process end to end.

Speaker James PowleyEvidence transcript

StageD1-C232

In an agency's competitor keyword analysis, a deliverable defined before AI, the team set the strategy buckets and the keyword scoring metrics itself and used AI only to pull in the data, never to make a judgment.

Speaker James PowleyEvidence transcript

StageD1-C233

With AI pulling in the data, a competitor keyword analysis covering 26 competitors and 79 data points per keyword went from three days of work to a few hours, a community speaker reported for his agency.

Speaker James PowleyEvidence transcript

StageD1-C234

For a content audit of a client blog with over 3,000 posts, an agency defined the parameters, scoring and weightings manually and had AI apply them, producing posts grouped by category, possible cannibalisation flags, a priority list and a list of posts to check by hand.

Speaker James PowleyEvidence transcript

StageD1-C235

The key lesson of a community talk on AI automation: AI should not decide what good looks like; people set the rules for what is good and AI applies them.

“it didn't decide what good looked like. We did that.”

Wording checked against the slide or recording

Speaker James PowleyEvidence transcript

StageD1-C236

In a post-migration analysis of over 200,000 lines of search traffic data and more than 45,000 pages, where many URLs had been merged into one, an agency set the rules for what counted as the same page and used AI to match old and new addresses and roll the traffic up per page.

Speaker James PowleyEvidence transcript

StageD1-C237

A community speaker builds SEO recommendations without AI and then uses AI to turn them into mock-ups and visuals, which help win buy-in from clients and internal teams.

Speaker James PowleyEvidence transcript

StageD1-C240

How often an LLM gives a wrong answer about data depends on the model, but how often the mistake is noticed depends on the person using it, a community speaker said.

“how often can you notice the mistake? And that depends on you.”

Wording checked against the slide or recording

Speaker Rafael KovashikawaEvidence transcript

StageD1-C241

In agency marketing-data pipelines, as in financial forecasting, one misplaced figure travels downstream and compounds into later analysis and decisions, which is why LLM errors on search data must be caught early, a community speaker warned.

Speaker Rafael KovashikawaEvidence transcript

StageD1-C242

A community speaker reported that a code refactor of his company's AI analysis product left data fetching correct but made the analysis shallow while the answers still read well, with the same model, so polished output is no proof of correct analysis.

Speaker Rafael KovashikawaEvidence transcript

  • Extended by D1-C500 Day 1: A community speaker applied Pascal's line about makers of false windows built for symmetry, whose rule is to…
StageD1-C243

The first rule for using LLMs on search data, according to a community speaker, is never to let the model check its own output.

“not let the model grade its own homework”

Wording checked against the slide or recording

Speaker Rafael KovashikawaEvidence transcript

  • Repeated by D3-C435 Day 3: Review the filters and regexes that Search Console's AI-powered configuration proposes before trusting the…
StageD1-C244

A community speaker's first guardrail for LLM answers about data is a receipt: every figure the model gives must be traceable to the data request behind it, so the answer can be verified.

“If there is no receipt, there is no reimbursement.”

Wording checked against the slide or recording

Speaker Rafael KovashikawaEvidence transcript

StageD1-C245

In a community speaker's verification system, each claim in an LLM's data answer is marked verified, flagged (a hallucination), or skipped or not found.

Speaker Rafael KovashikawaEvidence transcript

StageD1-C246

Claims that cannot be verified go to a judge, models from other vendors that score them; a failing claim is struck through with the reason shown, and the system annotates and proposes corrections rather than rewriting the answer, a community speaker described.

Speaker Rafael KovashikawaEvidence transcript

StageD1-C248

Before using AI-generated numbers, for example in a client meeting, always ask the AI for the receipts and the reason behind every number, because AI hallucinates; LLMs present their reasoning as a forward chain, so the user has to check it backwards, a community speaker said.

Speaker Rafael KovashikawaEvidence transcript

StageD1-C249

A community demo showed a script that opens Chrome with a saved session, pastes a URL into Google's Rich Results Test, waits about 15 seconds, then screenshots and saves the result and retries on failure (the demo's recording is largely unintelligible; the tool name is a best reading).

Speaker not identifiedEvidence transcript

  • Extended by D1-C501 Day 1: A community demo's script used both input modes of Google's Rich Results Test: the URL mode for public pages…
  • Extended by D1-C506 Day 1: Google's Rich Results Test help says the tool tests either a page's full URL or a pasted code snippet (Code…
StageD1-C250

In a community demo, a fix loop gave a large LLM (Claude Opus) the current markup, the errors Google's test reported and Google's documentation, had it write new JSON-LD and re-ran the test, with at most three attempts; the demo's page passed on the second.

Speaker not identifiedEvidence transcript

Things
  • Extended by D2-C462 Day 2: Google's guidance on generative AI content warns that AI output can contain hallucinations and says…
StageConsistent with docsD1-C251

A community demo treated warnings in Google's structured-data test as not critical, in an example result of 22 items with two errors and two warnings.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-SDA-11

  • Extended by D1-C502 Day 1: In a community demo, the structured-data problems Google's Rich Results Test reported on the test page…
StageD1-C252

A community demo matched model size to the task: a small, cheap model (Claude Haiku) chose keywords from a page's URL, title, H1, description and schema types, while a large, expensive model (Claude Opus) fixed code and wrote the summary.

Speaker not identifiedEvidence transcript

  • Extended by D1-C503 Day 1: In a community demo, the small model choosing keywords (Claude Haiku) followed fixed rules, no brand terms…
StageD1-C254

To automate SEO reporting with AI-assisted coding, a community speaker advised against aiming for an app or a perfect dashboard: pick one repetitive task, define the smallest useful output, ask the AI what is possible, then build and debug one step at a time.

Speaker Kira BreuerEvidence transcript

StageD1-C255

A community speaker who is not a developer automated monthly SEO reports with Python in Visual Studio Code, using Claude and ChatGPT as coding partners and Google Cloud for access to the Search Console API.

Speaker Kira BreuerEvidence transcript

  • Extended by D3-C469 Day 3: In a community lightning talk, agency founder Nik Vujic presented how his agency feeds Search Console and GA4…
StageD1-C256

An agency's automated monthly SEO report follows four simple steps: the data comes in, a script processes it, AI summarises it and the result goes onto a dashboard.

Speaker Kira BreuerEvidence transcript

  • Extended by D3-C477 Day 3: The community workflow's third step feeds the exported Search Console dataset into the agency's LLM agent…
  • Extended by D3-C478 Day 3: Nik Vujic's slide said no agent is needed to start: pull the Search Console data, import it into any LLM and…
StageConsistent with docsD1-C258

Pulling Search Console data through the API required setting up a Google Cloud account to get API access, a community speaker said.

Speaker Kira BreuerEvidence transcript

  • Prerequisites Google Search Console API documentation · checked 3 October 2026
StageD1-C259

A community speaker gave the AI a detailed description of the tech stack and requirements and started tiny, with one script, one client and one output, before having the AI write the code.

Speaker Kira BreuerEvidence transcript

StageD1-C260

When debugging AI-written code with AI, always read the code and understand what went wrong, so you can fix it yourself next time, a community speaker advised.

“Don't trust it blindly, and try to understand what is happening in your code”

Wording checked against the slide or recording

Speaker Kira BreuerEvidence transcript

StageD1-C263

To start automating SEO work with AI, a community speaker said you need no expertise, only one painful task, a clear outcome and the patience to iterate.

“you don't need to be an expert; you just need to try.”

Wording checked against the slide or recording

Speaker Kira BreuerEvidence transcript

DocsSourceD1-C265

Search Console Help separates critical issues, which make a structured-data item invalid and stop it from appearing as a rich result, from non-critical issues, listed under 'Improve item appearance'; the Rich Results Test reports valid items that have warnings.

Publisher Google Search Console Help

Used byrequirement DEV-SDA-11

DocsSourceD1-C266

To use the Search Console API you need a Google Account with Search Console permission on the property, a project in the Google API Console and OAuth2 credentials (all Search Console APIs except the Testing Tools API need OAuth2).

Publisher Google Search Console API documentation

  • Prerequisites Google Search Console API documentation · checked 3 October 2026
StageConsistent with docsD1-C271

A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard as Ahrefs') that looked at almost 40,000 sites over one month: 97% had no AI agent hits on the file, and those that had any got about two hits a month.

Speaker Carlos OrtegaEvidence transcript

Things

Used byrequirement DEV-AIF-02

  • Repeats D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
  • Extends D1-C060 Day 1: Speakers at the event were openly dismissive of llms.txt.
  • Extended by D1-C505 Day 1: The llms.txt study cited on stage matches Ahrefs' June 2026 study of 137,210 domains: 28% (about 38,000, the…
StageConsistent with docsD1-C272

A community speaker advised against publishing markdown copies of HTML pages for AI agents: the copy is a duplicate (which the speaker also called a possible source of cloaking, an uncertain word in the recordings), and the models are trained to read HTML, CSS and JavaScript.

Speaker Carlos OrtegaEvidence transcript

Used byrequirement DEV-AIF-02

StageConfirmed by docsD1-C273

A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

Speaker Carlos OrtegaEvidence transcript

Used byrequirement DEV-AIF-05glossary term User-triggered fetchers

  • Repeated by D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…
  • Repeated by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
StageConfirmed by docsD1-C274

A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.

Speaker Carlos OrtegaEvidence transcript

Used byrequirement DEV-HTM-04

  • Repeats D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
StageConsistent with docsD1-C275

A community speaker said layout shifts, such as a button or image popping in after the first load, annoy users and may confuse AI agents, and recommended watching Cumulative Layout Shift (CLS), the Core Web Vitals metric for visual stability.

Speaker Carlos OrtegaEvidence transcript

Used byglossary term Cumulative Layout Shift (CLS)

StageNot in docsD1-C276

A community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which number is the price saves the agent from guessing.

Speaker Carlos OrtegaEvidence transcript

  • Extended by D1-C315 Day 1: Google's AI optimisation guide says structured data is not required for generative AI search and needs no…
StageConsistent with docsD1-C277

A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.

Speaker Carlos OrtegaEvidence transcript

Used byrequirement DEV-HTM-04

  • Extended by D1-C314 Day 1: A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do…
StageConsistent with docsD1-C278

A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.

Speaker Carlos OrtegaEvidence transcript

Used byrequirement DEV-HTM-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
StageConfirmed by docsD1-C279

A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.

Speaker Carlos OrtegaEvidence transcript

Used byrequirement DEV-HTM-04glossary term ARIA

  • Extends D1-C118 Day 1: Attendees were told to check ARIA and accessibility.
StageConsistent with docsD1-C280

A community speaker warned that no ARIA is better than bad ARIA: wrong or confusing ARIA attributes do more harm than leaving them out.

Speaker Carlos OrtegaEvidence transcript

Used byrequirement DEV-HTM-04glossary term ARIA

StageConfirmed by docsD1-C281

A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.

Speaker Carlos OrtegaEvidence transcript

Things

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extended by D2-C303 Day 2: Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents…
StageConfirmed by docsD1-C282

A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.

Speaker Carlos OrtegaEvidence transcript

Things

Used byrequirement DEV-AIF-06

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extended by D1-C316 Day 1: Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through…
StageConfirmed by docsD1-C283

A community speaker said WebMCP has two kinds of tools: declarative ones, mostly HTML annotations such as the fields of a contact form, and imperative ones for other actions such as booking, filtering a catalogue, adding products to a cart, getting product specs or reordering.

Speaker Carlos OrtegaEvidence transcript

Things

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
StageD1-C285

A community speaker contrasted real audience understanding (reading logs, tracking how the brand is mentioned and perceived, letting customer complaints change pricing pages and products) with keyword-tool research that turns People Also Ask questions into H2s and blog posts.

Speaker Thiago PojdaEvidence transcript

StageD1-C286

A community speaker said the SEO industry turned everything into a metric because budgets are approved on numbers, not ideas.

“an idea doesn't get you budget approval; a number does.”

Wording checked against the slide or recording

Speaker Thiago PojdaEvidence transcript

StageD1-C287

A community speaker said not all SEO work was wrong: a site still has to be crawlable and visible before it is even considered, and much of that work is plumbing, not gaming.

“There's a lot of plumbing that's not gaming.”

Speaker Thiago PojdaEvidence transcript

StageD1-C288

A community speaker said Google's AMP project, meant to get people building fast websites and addressed directly to engineers and marketers, was smart, but the industry reduced it to a metric and a checklist ('now I have an AMP website').

Speaker Thiago PojdaEvidence transcript

StageD1-C291

A community speaker said listicles do earn more mentions in AI answers, not because they are trusted more but because LLMs mainly check that the thing being talked about exists, not whether the content is honest.

Speaker Thiago PojdaEvidence transcript

StageNot in docsD1-C295

A community speaker showed a Spanish-language AI prompt about places to eat in Barcelona whose fan-out queries were all in English, so the content used to answer may be written by and for a different audience than the one asking.

Speaker Thiago PojdaEvidence transcript

StageD1-C296

A community speaker argued for adopting the GEO label as the industry's chance to leave behind the bad reputation SEO built, unlike Google's view earlier the same day that the new name is not needed.

Speaker Thiago PojdaEvidence transcript

Things

Used bystory angle A-001

  • Contradicts D1-C049 Day 1: Gary Illyes argued that GEO is a label invented to create a new field and is not needed. Understanding how…
  • Contradicts D1-C050 Day 1: Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
StageD1-C300

A community speaker said most AI-visibility measurement covers only the selection stage, whether a brand is mentioned or selected, without showing why or why not.

Speaker not identifiedEvidence transcript

StageD1-C301

A community speaker named two signals that in their view strongly influence whether an AI selects a brand: the bias already in the model's memory before it searches, and the brand's presence in the search results the model is grounded on.

Speaker not identifiedEvidence transcript

StageD1-C304

A community speaker described the association stage as whether an AI already connects a brand with the topics that matter to it before running any search, and said problems there are mostly brand problems.

Speaker not identifiedEvidence transcript

StageD1-C305

A community speaker's 'brand association rate' asks a model, with web browsing switched off so that only its memory answers, to name ten topics it associates with the brand, or the other way round, starting from a topic.

Speaker not identifiedEvidence transcript

StageD1-C306

A community speaker advised repeating brand-association prompts over time: association in every run is good, in about half the runs shows some association with work to do, and little or none shows a clear gap.

Speaker not identifiedEvidence transcript

StageD1-C307

A community speaker proposed a 'brand co-occurrence rate' for the search stage: the share of an AI's fan-out queries that contain the brand (one in three is 33%), which also hints at the associations the model already has; problems there are mostly relevance problems.

Speaker not identifiedEvidence transcript

StageD1-C310

A community speaker described the selection stage as whether a brand is chosen in the answer, measured with mentions and similar metrics; failure there is mostly a trust or relevance issue.

Speaker not identifiedEvidence transcript

StageD1-C311

Closing the lightning talks on automation and AI, one of the session's Google hosts called AI a tool that can be used well or badly: it can even be used for content generation, which none of the talks featured, and it is fantastic for coding and analysis but less so for other tasks.

Speaker not identifiedEvidence transcript

AnalysisD1-C314

A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.

Author Ibrahim Anjro

  • Extends D1-C277 Day 1: A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section…
AnalysisD1-C315

Google's AI optimisation guide says structured data is not required for generative AI search and needs no special schema.org markup, and on Day 2 Google said raw schema.org is generally not put into model context (D2-C477); use markup for rich-result eligibility and clear data, not as an AI-visibility lever.

Author Ibrahim Anjro

  • Extends D1-C276 Day 1: A community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which…
DocsSourceD1-C316

Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.

Publisher Chrome for Developers

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extends D1-C282 Day 1: A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all…
StageD1-C500

A community speaker applied Pascal's line about makers of false windows built for symmetry, whose rule is to make pleasing figures rather than to speak accurately, to today's LLMs: sometimes they are not trying to produce the most accurate figures or the most accurate interpretations of them.

Speaker Rafael KovashikawaEvidence transcript

  • Extends D1-C242 Day 1: A community speaker reported that a code refactor of his company's AI analysis product left data fetching…
StageConsistent with docsD1-C501

A community demo's script used both input modes of Google's Rich Results Test: the URL mode for public pages, which Google fetches itself, and the code mode, into which the script pasted the page's HTML, for private pages and pages Google cannot fetch (best reading of a largely unintelligible recording; the second kind of page was heard as 'dead', possibly 'dev').

Speaker not identifiedEvidence transcript

Used byrequirement DEV-SDA-11

  • Extends D1-C249 Day 1: A community demo showed a script that opens Chrome with a saved session, pastes a URL into Google's Rich…
StageConsistent with docsD1-C502

In a community demo, the structured-data problems Google's Rich Results Test reported on the test page included an empty name, a breadcrumb problem and a price of zero, with errors shown in pink and warnings in orange (best readings of a largely unintelligible recording; each item is heard in only one of the two recordings).

Speaker not identifiedEvidence transcript

Used byrequirement DEV-SDA-08

  • Extends D1-C251 Day 1: A community demo treated warnings in Google's structured-data test as not critical, in an example result of…
StageD1-C503

In a community demo, the small model choosing keywords (Claude Haiku) followed fixed rules, no brand terms, no navigation terms and at most two keywords per page; the Spanish example was 'zapatos de mujer' and 'comprar zapatos de mujer' (women's shoes, buy women's shoes).

Speaker not identifiedEvidence transcript

  • Extends D1-C252 Day 1: A community demo matched model size to the task: a small, cheap model (Claude Haiku) chose keywords from a…
AnalysisD1-C504

The study cited on stage matches Peec AI's analysis (reported in February 2026) of over 10 million ChatGPT prompts and 20 million fan-outs: 43% of the fan-out searches for non-English prompts ran in English, and nearly 78% of non-English prompt runs had at least one English fan-out (from 66% for Spanish to 94% for Turkish), so the 43% is a share of fan-out searches, not of prompts.

Author Ibrahim Anjro

  • Extends D1-C294 Day 1: A community speaker cited a third-party study (heard as Peec AI's) finding that 43% of prompts produced…
AnalysisD1-C505

The llms.txt study cited on stage matches Ahrefs' June 2026 study of 137,210 domains: 28% (about 38,000, the 'almost 40,000' sites heard on stage) published an llms.txt file and 97% of those files received no requests at all in May 2026; the files that were fetched got about 22,000 requests across some 1,100 domains, around 20 each, not the 'two hits a month' heard on stage.

Author Ibrahim Anjro

Things

Used byrequirement DEV-AIF-02

  • Extends D1-C271 Day 1: A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard…
DocsSourceD1-C506

Google's Rich Results Test help says the tool tests either a page's full URL or a pasted code snippet (Code instead of URL); all page resources must be reachable by an anonymous user on the internet, so resources behind a firewall or a password are not available to the test unless exposed, for example through a tunnel.

Publisher Google Search Console Help

Used byrequirement DEV-SDA-11glossary term Rich Results Test

  • Extends D1-C249 Day 1: A community demo showed a script that opens Chrome with a saved session, pastes a URL into Google's Rich…

14:05 · How crawling works 34

SlideConsistent with docsD1-C063

The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and passes the fetch reply to indexing.

Speaker Gary IllyesEvidence 3 slide photos, transcript

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
  • Extended by D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extended by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
  • Extended by D2-C127 Day 2: A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing…
  • Extended by D2-C167 Day 2: The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler…
  • Extended by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
SlideConsistent with docsD1-C064

The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

“Ensure we don't break the internet”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
  • Extended by D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
  • Extended by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
SlideConsistent with docsD1-C065

The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
  • Extended by D3-C385 Day 3: Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh…
AnalysisD1-C067

A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

Author Ibrahim Anjro

Used byrequirements DEV-URL-04, DEV-URL-05

  • Extended by D2-C189 Day 2: Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script…
  • Extended by D2-C427 Day 2: When internal links point only to page A and the canonical leader is reached only through A's canonical link…
DocsSourceD1-C136

Google's Inside Googlebot post (March 2026) says Googlebot currently fetches only the first 2MB of each URL, HTTP headers included (64MB for PDFs); bytes past that cutoff are not fetched, rendered or indexed, and each resource the page loads has its own separate limit.

Publisher Search Central blog (31 March 2026)

Things

Used byrequirement DEV-PRF-04

DocsSourceD1-C137

Google's Inside Googlebot post warns that bloated inline base64 images, large blocks of inline CSS or JavaScript, or megabytes of menus can push a page's text or structured data past Googlebot's 2MB cutoff, and advises moving heavy CSS and JavaScript to external files and placing meta tags, the title, the canonical and essential structured data high in the HTML.

Publisher Search Central blog (31 March 2026)

Used byrequirement DEV-PRF-04

StageD1-C317

Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.

Speaker Cherry PrommawinEvidence transcript

  • Repeated by D2-C847 Day 2: Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including…
StageNot in docsD1-C318

A URL states how a resource is requested (the protocol, HTTP or HTTPS), where (the host, meaning which computer on the network) and what (the path to the exact page or file).

Speaker Cherry PrommawinEvidence transcript

StageConsistent with docsD1-C320

An HTTP 200 OK status only means that the server believes it managed to do what the client asked for.

“just means that the server believes that it managed to accomplish whatever the user was asking for”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-ERR-03

StageConsistent with docsD1-C322

Googlebot is an ordinary HTTP client with nothing special about it: like a browser, it fetches a URL it was given and returns the fetched bytes to Google's servers.

“Googlebot is just a client. It is an HTTP client. There's nothing all that much special about it.”

Speaker Gary IllyesEvidence transcript

Things
StageNot in docsD1-C323

AI agents are, technically, the same thing as crawlers: HTTP clients that accomplish something on behalf of a user or a service.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD1-C326

Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure, because every crawler must accomplish a few specific tasks and obey Google's internal crawling policies.

Speaker Gary IllyesEvidence transcript

  • Extended by D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
StageD1-C327

Google said it does not usually talk publicly about crawl components such as the scheduler and the crawl queue, because the details get confusing and taken out of context.

Speaker Gary IllyesEvidence transcript

Used byglossary term Crawl scheduler

StageConsistent with docsD1-C328

During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the scheduler to be crawled.

Speaker Gary IllyesEvidence transcript

  • Repeated by D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
StageConsistent with docsD1-C329

Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

Speaker Gary IllyesEvidence transcript

  • Extended by D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
  • Repeated by D2-C848 Day 2: Google said it strongly believes site owners should be able to opt out of crawling and control how their site…
StageConsistent with docsD1-C330

The only Google-owned crawlers that do not obey robots.txt are contractual crawlers, which crawl a site whose owner has agreed that Google may crawl it however it likes.

Speaker Gary IllyesEvidence transcript

Used byglossary term Contractual crawlers (special-case crawlers)

StageConsistent with docsD1-C331

Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C612 Day 3: Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower…
StageConsistent with docsD1-C332

The crawl scheduler logs how often each page changes and crawls frequently changing pages first: a news site's homepage, which changes very often, is prioritised over its terms of service page, which may change once a year.

Speaker Gary IllyesEvidence transcript

Used byglossary term Crawl scheduler

StageNot in docsD1-C333

The scheduler hands the crawler an ordered list of URLs from the crawl queue, and the crawler works through the list from top to bottom.

Speaker Gary IllyesEvidence transcript

Used byglossary term Crawl scheduler

StageConsistent with docsD1-C334

Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.

Speaker Gary IllyesEvidence transcript

  • Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
StageNot in docsD1-C335

Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

“the number of tokens is actually more important”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

Things
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageConfirmed by docsD1-C336

Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
StageConsistent with docsD1-C337

An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Extended by D3-C610 Day 3: Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.
  • Extended by D3-C611 Day 3: If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within…
StageNot in docsD1-C338

Google described its crawling as a large-scale distributed swarm of simple HTTP clients, roughly what one would get by deploying many wget or curl libraries on cloud compute instances.

“a large-scale distributed swarm of simple HTTP clients”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

StageD1-C341

Gary Illyes called the Crawl Stats report the most valuable tool for debugging crawl errors, together with server logs, which he called the real thing for those who know how to read them.

Speaker Gary IllyesEvidence transcript

DocsSourceD1-C342

Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.

“Googlebot is just a user of something that resembles a centralized crawling platform”

Publisher Search Central blog (31 March 2026)

DocsSourceD1-C343

Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.

Publisher Google

Used byglossary term Contractual crawlers (special-case crawlers)

AnalysisD1-C344

On stage Google spoke of probably hundreds, if not thousands, of crawlers, while its Inside Googlebot post speaks of dozens of other clients; both agree that only the larger crawlers are documented, so a Google user agent missing from the public lists is not proof of a fake request, and reverse DNS or Google's published IP ranges are the test.

Author Ibrahim Anjro

Things

14:35 · How crawling errors affect Search 39

StageConsistent with docsD1-C069

DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

Speaker Gary IllyesEvidence notes, transcript

Used byrequirements DEV-MON-04, DEV-SRV-01

  • Extended by D1-C359 Day 1: Most network errors happen somewhere between the site's origin server and Google's data centers, where they…
  • Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
DocsSourceD1-C126

Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.

“Google treats network timeouts, connection reset, and DNS errors similarly to 5xx server errors.”

Publisher Google

Things

Used byrequirements DEV-MON-04, DEV-SRV-01, DEV-SRV-03fact F-017

  • Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
StageConsistent with docsD1-C070

Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.

Speaker Gary IllyesEvidence notes, transcript

Things
  • Extended by D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
DocsSourceD1-C073

A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

Publisher Google, Google Search Central

Used byrequirement DEV-ERR-01glossary term Soft 404

  • Extended by D2-C190 Day 2: In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because…
  • Extended by D2-C213 Day 2: A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as…
  • Repeated by D2-C264 Day 2: Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but…
  • Extended by D2-C291 Day 2: In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200…
  • Repeated by D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
  • Extended by D2-C336 Day 2: Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty…
  • Extended by D2-C367 Day 2: Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.
  • Extended by D2-C700 Day 2: Index selection drops soft 404 pages that were not dropped earlier, for example when a document is…
AnalysisD1-C078

For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

Author Ibrahim Anjro

Used byrequirements DEV-MON-04, DEV-SRV-02, DEV-SRV-03

  • Extended by D2-C339 Day 2: For soft 404 detection, the position of error text decides: an error in a less important part such as the…
  • Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
  • Extended by D2-C376 Day 2: Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200…
DocsSourceD1-C138

Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.

Publisher Google

Used byrequirements DEV-MON-04, DEV-SRV-01

StageNot in docsD1-C346

1xx informational status codes, which only say that a request was received and more data is coming, have no meaning of their own for crawling.

Speaker Cherry PrommawinEvidence transcript

StageConfirmed by docsD1-C354

Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.

Speaker Cherry PrommawinEvidence transcript

Things

Used byrequirement DEV-SRV-03

  • Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
StageConfirmed by docsD1-C355

A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-ERR-03

  • Repeated by D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
StageD1-C357

Gary Illyes called DNS and network errors a very common issue on the internet nowadays, and DNS problems very pesky to debug.

Speaker Gary IllyesEvidence transcript

StageNot in docsD1-C359

Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-01

  • Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
StageConsistent with docsD1-C363

For DNS and network errors caused by a firewall or CDN, Google advises checking whether new firewall rules were set recently and otherwise asking in the CDN's forum, as Google itself cannot see or help with these errors.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-SRV-01

StageConsistent with docsD1-C365

Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-09

  • Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
StageConsistent with docsD1-C366

Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.

“402 will just mean 404 to us”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-09

  • Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
StageConsistent with docsD1-C367

CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-02glossary term Soft 404

  • Extended by D1-C508 Day 1: Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing…
  • Extended by D2-C374 Day 2: A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows…
  • Extended by D2-C375 Day 2: Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot…
StageConfirmed by docsD1-C368

Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.

Speaker Gary IllyesEvidence transcript

  • Extended by D1-C508 Day 1: Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing…
  • Extended by D2-C717 Day 2: Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed…
DocsSourceD1-C370

Google's crawlers follow up to 10 redirect hops by default (some products' crawlers have other limits); content served by the redirecting URL is ignored and the final target's content is processed instead.

“Any content Google receives from the redirecting URL is ignored, and the final target URL's content is processed instead.”

Publisher Google

Things

Used byrequirement DEV-CAN-01glossary term Redirect chain

AnalysisD1-C371

Day 1 said a CDN captcha page served with HTTP 200 becomes a soft 404, while Day 2 (D2-C374, D2-C375) said such challenge pages are hard to recognise as errors and can be clustered as duplicates; Google's CDN post describes both outcomes, and in both the real pages drop out of Search.

Author Ibrahim Anjro

StageConfirmed by docsD1-C508

Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-MON-03

  • Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…
  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
StageNot in docsD1-C509

Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-SRV-09

  • Extends D1-C366 Day 1: Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
  • Extends D1-C365 Day 1: Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication…

15:10 · How Google interprets robots.txt 36

SlideD1-C079

The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

Speaker GoogleEvidence slide photo, transcript

  • Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
  • Extended by D1-C521 Day 1: To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the…
  • Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
DocsSourceD1-C080

Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

Publisher Google

Used byrequirements DEV-SRV-05, DEV-SRV-10glossary term User-agent group

  • Extended by D1-C519 Day 1: A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several…
  • Extended by D1-C527 Day 1: The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another…
  • Extended by D1-C531 Day 1: In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of…
  • Extended by D1-C533 Day 1: Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that…
DocsSourceD1-C081

When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

“In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”

Publisher Google

Used byrequirement DEV-SRV-05

  • Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
  • Extended by D2-C208 Day 2: A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/…
AnalysisD1-C083

Under the example file, Googlebot may crawl /, /politics/eu-vote and /sports/live/, and is blocked from /?utm_source=x, /index.html, /politics (no trailing slash), /live/ and /sports/live-score. An 'allow: /$' rule does not cover the homepage with tracking parameters.

Author Ibrahim Anjro

DocsSourceD1-C084

Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

Publisher Google

Used byrequirement DEV-SRV-05glossary term robots.txt

  • Extended by D1-C516 Day 1: Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was…
  • Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
  • Extended by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
DocsSourceD1-C085

Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

Publisher Google

Used byrequirement DEV-SRV-05

  • Extended by D3-C613 Day 3: Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an…
  • Extended by D3-C614 Day 3: Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for…
DocsSourceD1-C086

Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

“Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.”

Publisher Google

Used byrequirement DEV-IDX-11glossary term Google-Extended

  • Extended by D1-C513 Day 1: The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in…
  • Extended by D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
  • Extended by D1-C523 Day 1: The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023…
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
  • Extended by D2-C121 Day 2: Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex…
SlideConfirmed by docsD1-C087

Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.

Speaker GoogleEvidence slide photo, transcript

Things

Used byrequirement DEV-SRV-06

  • Extended by D1-C526 Day 1: The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked…
SlideNot in docsD1-C123

Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

“Unpredictability is never a good time.”

Wording checked against the slide or recording

Speaker GoogleEvidence slide photo, transcript

Used byrequirement DEV-SRV-06

  • Extended by D1-C526 Day 1: The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked…
AnalysisD1-C088

Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files. Never place any non-standard directive between user-agent lines; put it after a group's rules and test the file with each search engine's tools, such as Bing's robots.txt tester and Google's open-source robots.txt parser.

Author Ibrahim Anjro

Used byrequirement DEV-SRV-06

StageNot in docsD1-C089

The robots.txt talk pointed to the robots.txt file of thebestfriedchickenever.com to show why comments in robots.txt are useful; the live file is mostly a large ASCII-art drawing written as # comment lines, followed by a single user-agent group (checked 2026-10-04).

Speaker GoogleEvidence notes, transcript

DocsSourceD1-C140

Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

Publisher Google Search Console Help

Used byrequirement DEV-SRV-05

  • Extended by D1-C528 Day 1: Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version…
  • Extended by D1-C530 Day 1: Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
  • Extended by D3-C615 Day 3: A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.
StageConfirmed by docsD1-C512

Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)

  • Extends D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
  • Repeated by D2-C050 Day 2: John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes…
StageConsistent with docsD1-C513

The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.

Speaker GoogleEvidence transcript

Used byglossary term RFC 9309 (Robots Exclusion Protocol)

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
StageConsistent with docsD1-C514

Robots.txt is not a security measure: the file sits at a predictable public address, so anyone can read the paths it disallows. Protect a secret folder with authentication, or do not put it on the internet.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-11glossary term robots.txt

StageConsistent with docsD1-C516

Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
StageConsistent with docsD1-C519

A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConsistent with docsD1-C520

In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.

Speaker GoogleEvidence transcript, slide photo

Used byrequirement DEV-SRV-05

  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
StageConsistent with docsD1-C521

To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.

Speaker GoogleEvidence transcript, slide photo

Used byrequirement DEV-SRV-05

  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extends D1-C082 Day 1: In robots.txt, * matches zero or more of any character and $ marks the end of the URL.
StageConsistent with docsD1-C522

Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

Speaker GoogleEvidence transcript

Used byrequirement DEV-IDX-11glossary term Google-Extended

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
  • Repeats D1-C485 Day 1: To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
AnalysisD1-C523

The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

Author Ibrahim Anjro

Used byrequirement DEV-AIF-05

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
StageNot in docsD1-C525

Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-05

  • Extended by D1-C532 Day 1: A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts…
StageNot in docsD1-C526

The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-06

  • Extends D1-C087 Day 1: Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and…
  • Extends D1-C123 Day 1: Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's…
StageConsistent with docsD1-C527

The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConfirmed by docsD1-C528

Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

Speaker GoogleEvidence transcript

Used byrequirements DEV-SRV-05, DEV-SRV-06glossary term robots.txt report

  • Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
  • Extended by D3-C667 Day 3: Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours…
StageConsistent with docsD1-C531

In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-10

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
AnalysisD1-C532

A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

Author Ibrahim Anjro

Used byrequirement DEV-SRV-05

  • Extends D1-C525 Day 1: Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule…

15:30 · Lightning session B: Robots.txt 5

StageConfirmed by docsD1-C533

Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

Speaker Dave SmartEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageNot in docsD1-C534

Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.

Speaker Dave SmartEvidence transcript

Used byrequirement DEV-CAN-11glossary term Redirect chain

StageNot in docsD1-C536

Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.

Speaker Dave SmartEvidence transcript

Used byrequirement DEV-CAN-11

16:00 · How Google thinks about crawl budget 52

SlideConsistent with docsD1-C066

Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-PRF-01

  • Extends D1-C326 Day 1: Google does not let each team build its own crawler: its many crawlers share one crawler infrastructure…
  • Extended by D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…
SlideConfirmed by docsD1-C091

Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Extended by D1-C374 Day 1: Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can…
  • Extended by D1-C376 Day 1: Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app…
DocsSourceD1-C132

Google's crawl budget guide says each crawler has its own crawl demand, but the crawl capacity limit (hostload) is shared across all crawlers, so high demand from one crawler can reduce the capacity left for others.

“high demand from one crawler can reduce the capacity available for others”

Publisher Google

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

SlideConsistent with docsD1-C092

Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirements DEV-MON-04, DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
SlideConsistent with docsD1-C093

Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byglossary term Crawl demand

  • Extended by D1-C377 Day 1: The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual…
  • Extended by D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
  • Extended by D2-C695 Day 2: Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most…
  • Extended by D2-C709 Day 2: Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and…
  • Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
SlideNot in docsD1-C094

If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

“If quality or popularity is unknown, use parent root's aggregate quality or popularity is used, then, that path's parent's, and so on.”

Wording checked against the slide or recording

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-10

  • Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
  • Extended by D2-C685 Day 2: Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies…
AnalysisD1-C095

New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.

Author Ibrahim Anjro

Used byrequirement DEV-URL-10

  • Extended by D2-C686 Day 2: Launch new pages under sections that Google already indexes well, and improve or remove weak sections…
StageConsistent with docsD1-C096

Crawl budget was described as the attention span Google gives a website, and better performance increases it.

Speaker Cherry PrommawinEvidence notes, transcript

Used byrequirement DEV-PRF-01glossary term Crawl budget

  • Extended by D1-C373 Day 1: Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it…
DocsSourceD1-C097

Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.

Publisher Google

  • Extended by D2-C706 Day 2: 'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL…
AnalysisD1-C098

On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

Author Ibrahim Anjro

  • Extended by D2-C708 Day 2: The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on…
  • Extended by D2-C714 Day 2: 'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages…
  • Extended by D2-C716 Day 2: Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the…
SlideConsistent with docsD1-C100

Six faceted-URL problems were shown: the same facets in a different order, irrelevant or conflicting facet combinations, excessive facet selection, facets on paginated series, facets combined with search queries, and optional facets with default values.

Speaker Cherry PrommawinEvidence slide photo

Used byrequirement DEV-URL-08glossary term Faceted navigation

DocsSourceD1-C101

Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

Publisher Google

Used byrequirement DEV-URL-08glossary term Faceted navigation

  • Extended by D2-C188 Day 2: Hash-fragment links are a problem only where Google should follow them: product, category and language links…
  • Repeated by D2-C284 Day 2: URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot…
SlideConsistent with docsD1-C103

Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirements DEV-SRV-08, DEV-URL-08

  • Extended by D1-C379 Day 1: The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages…
  • Extended by D2-C886 Day 2: A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster…
DocsSourceD1-C104

HTTP caching for crawlers means supporting conditional requests (ETag with If-None-Match, or Last-Modified with If-Modified-Since) and answering 304 Not Modified when nothing changed. Google's crawling team has said it prefers ETag.

Publisher Google, Search Central blog (9 December 2024)

Used byrequirement DEV-SRV-08

SlideConfirmed by docsD1-C106

The noindex rule consumes crawl budget, because Google must fetch the page to see it.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-IDX-02fact F-025

  • Extended by D2-C022 Day 2: Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a…
  • Extended by D2-C069 Day 2: Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a…
  • Extended by D2-C697 Day 2: Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing…
SlideConfirmed by docsD1-C107

The nofollow rule can still consume crawl budget: Google does not crawl through the nofollow link itself, but it still crawls the linked page when it finds that page through other links.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-IDX-02

DocsSourceD1-C127

Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

Publisher Google Search Central, Search Central blog (10 September 2019)

Used byrequirements DEV-IDX-02, DEV-IDX-06glossary term nofollow

  • Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
AnalysisD1-C110

A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

Author Ibrahim Anjro

Used byrequirement DEV-IDX-01

  • Extended by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
  • Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
  • Extended by D2-C849 Day 2: Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can…
  • Extended by D2-C058 Day 2: John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use…
  • Repeated by D2-C850 Day 2: John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
  • Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
StageConsistent with docsD1-C373

Crawl budget is the finite amount of resources Google allocates to crawling a specific website, and it determines how many of the site's pages are discovered and how often they are revisited.

Speaker Cherry PrommawinEvidence transcript

Used byglossary term Crawl budget

  • Extends D1-C096 Day 1: Crawl budget was described as the attention span Google gives a website, and better performance increases it.
StageConsistent with docsD1-C374

Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-PRF-01

  • Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
  • Extended by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
StageConfirmed by docsD1-C375

Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling infrastructure, so a request from a Google user agent in server logs is a request routed through that shared platform.

Speaker Cherry PrommawinEvidence transcript

Things
  • Extends D1-C066 Day 1: Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose…
  • Extended by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
StageConfirmed by docsD1-C376

Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Extends D1-C091 Day 1: Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.
StageNot in docsD1-C377

The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.

Speaker Cherry PrommawinEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
  • Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
StageConsistent with docsD1-C378

Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

Speaker Cherry PrommawinEvidence transcript

Used byrequirements DEV-URL-08, DEV-URL-11

  • Extends D1-C099 Day 1: Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and…
StageConfirmed by docsD1-C379

The useless content worth improving or removing to save crawl budget includes very bad or low-quality pages, spam content, duplicate pages and soft error pages.

Speaker Cherry PrommawinEvidence transcript

  • Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
StageD1-C380

An audience member asked, in a question submitted before the event, whether crawl budget is still an SEO priority in 2026 or only relevant for very large sites.

From the audienceEvidence transcript

  • Answered by D1-C381 Day 1: Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few…
StageConfirmed by docsD1-C381

Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.

Speaker Cherry PrommawinEvidence transcript

Used byglossary term Crawl budget

  • Answers D1-C380 Day 1: An audience member asked, in a question submitted before the event, whether crawl budget is still an SEO…
  • Repeated by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
AnalysisD1-C382

The stage rule of thumb (under a few thousand URLs crawl budget is unlikely to be a problem) and Google's crawl budget guide (D1-C097: sites with over a million pages that change about weekly, or over 10,000 pages that change daily) leave a middle range where a site should check Search Console's Crawl Stats report before blaming crawl budget for slow indexing.

Author Ibrahim Anjro

AnalysisD1-C386

The stage point that 4xx responses do not affect crawl budget matches Google's documentation that 4xx codes have no effect on crawl rate (D1-C072), with one exception: 429 Too Many Requests counts as a server error and slows crawling like a 5xx (D1-C071, D1-C092).

Author Ibrahim Anjro

StageD1-C387

An audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.

From the audienceEvidence transcript

  • Answered by D1-C388 Day 1: AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl…
  • Answered by D1-C389 Day 1: Google expects sites to see more crawling overall, because many other services, including AI services, now…
StageConsistent with docsD1-C388

AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.

Speaker Cherry PrommawinEvidence transcript

Things
  • Answers D1-C387 Day 1: An audience member asked, in questions submitted before the event, how often Googlebot should be expected to…
  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageD1-C389

Google expects sites to see more crawling overall, because many other services, including AI services, now crawl the web besides Google.

Speaker Cherry PrommawinEvidence transcript

  • Answers D1-C387 Day 1: An audience member asked, in questions submitted before the event, how often Googlebot should be expected to…
StageD1-C390

An audience member asked, in a question submitted before the event, whether crawl frequency affects the ranking position of a URL.

From the audienceEvidence transcript

  • Answered by D1-C391 Day 1: A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.
  • Answered by D1-C392 Day 1: Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal…
StageConfirmed by docsD1-C391

A higher crawl rate does not make a URL rank better; it only means Google checks the site more often.

“if your crawl rate is increased, that doesn't mean that you would rank better.”

Speaker Cherry PrommawinEvidence transcript

  • Answers D1-C390 Day 1: An audience member asked, in a question submitted before the event, whether crawl frequency affects the…
StageConsistent with docsD1-C392

Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal that people are interested in the site and that its content is of high quality.

Speaker Cherry PrommawinEvidence transcript

  • Answers D1-C390 Day 1: An audience member asked, in a question submitted before the event, whether crawl frequency affects the…
AnalysisD1-C419

A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.

Author Ibrahim Anjro

AnalysisD1-C420

The stage remark that crawl demand follows the quality of the site as a whole sits beside the same talk's slide (D1-C094), which falls back to the parent path's aggregate only when a URL's own quality is unknown, and Google's crawl budget guide lists page quality among the demand factors; read it as site quality setting the baseline while known URL-level signals still count.

Author Ibrahim Anjro

16:20 · Lightning session C: Crawling 23

StageConsistent with docsD1-C397

A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-URL-11glossary term Similar URLs

  • Extended by D1-C444 Day 1: Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily…
StageConsistent with docsD1-C399

A community speaker showed five kinds of similar URLs on well-known brand sites: a different protocol or host (HTTP vs HTTPS, www vs non-www), different capitalisation, a different number of delimiters such as slashes, a space encoded as %20 in one URL and + in another, and parameters in a different order.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-URL-11glossary term Similar URLs

StageConsistent with docsD1-C400

To find similar URLs in a URL list, a community speaker first applies basic normalisation that keeps the meaning: lowercase the scheme and host, remove default ports such as 80 for HTTP, resolve relative path segments and usually drop the fragment, which matters only to clients.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-URL-11

StageD1-C401

A community speaker then derives a uniform URL that deliberately changes the meaning (one scheme, everything lowercase, www and index files such as index.html removed, duplicate and trailing delimiters removed, parameters sorted alphabetically) and sorts the URL list by it, so that similar URLs group together.

Speaker Tobias SchwarzEvidence transcript

StageNot in docsD1-C402

Similar URLs usually come from programming errors or manually set links, but anyone can link to them from other sites, including malicious actors, so a site should be hardened against them.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-URL-11

StageConsistent with docsD1-C403

A community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.

Speaker Tobias SchwarzEvidence transcript

Things

Used byrequirement DEV-URL-11

  • Repeated by D1-C445 Day 1: One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it…
StageConsistent with docsD1-C404

For query parameters, a community speaker advised checking that every parameter in a request is actually used and in the expected order, and otherwise redirecting to the expected URL with only the used parameters in the correct order.

Speaker Tobias SchwarzEvidence transcript

Things

Used byrequirement DEV-URL-11

  • Extends D1-C102 Day 1: If faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return…
StageD1-C405

A community speaker said most web applications wrongly accept a numeric ID in a URL written with leading zeros or as a decimal with trailing zeros, which creates further similar URLs, so such variants are worth testing.

Speaker Tobias SchwarzEvidence transcript

StageNot in docsD1-C407

A community speaker cited more than 900 million weekly active ChatGPT users and 2.5 billion monthly users of a Google AI feature (the recording is unclear which), adding that these are not comparable market-share figures.

Speaker Jovana AvramovicEvidence transcript

StageD1-C408

A community speaker argued that search no longer happens in one place, so SEO is about how content is found, understood and trusted wherever search occurs.

Speaker Jovana AvramovicEvidence transcript

StageD1-C409

A community speaker argued that the change in search behaviour is not only longer AI Mode queries but conversation: a follow-up such as 'compare these two' relies on the context of earlier turns.

Speaker Jovana AvramovicEvidence transcript

Things
StageD1-C410

A community speaker described a cognitive offloading shift: people increasingly hand remembering, comparing, analysing and even decision-making over to search tools.

Speaker Jovana AvramovicEvidence transcript

StageD1-C411

In a community speaker's view, AI search changes the workflow but not the foundation: where one query once returned ten blue links and left research and comparison to the user, one query can now lead to many different answers.

Speaker Jovana AvramovicEvidence transcript

StageConsistent with docsD1-C412

A community speaker described AI search as turning one query into many related searches, including searches in other languages, before the information is retrieved.

Speaker Jovana AvramovicEvidence transcript

  • Extends D1-C053 Day 1: Query fan-out means running several related searches at once to gather more results; a question about lawn…
StageD1-C413

A community speaker said technical SEO remains the foundation for AI search, with site architecture showing the structure and thematic connections of content and structured data adding clarity, but that a technically sound site is only the beginning.

Speaker Jovana AvramovicEvidence transcript

StageNot in docsD1-C414

A community speaker linked the serial-position effect in human memory (people best remember the first and last items of a list) to the 'lost in the middle' pattern that research has found in large language models.

Speaker Jovana AvramovicEvidence transcript

Used byglossary term Lost in the middle

StageD1-C415

A community speaker advised placing the most important information at the beginning or the end of a piece of content, because information in the middle is less likely to be cited by AI systems.

“It's relevant to put your most important information at the beginning or at the end.”

Speaker Jovana AvramovicEvidence transcript

StageD1-C416

A community speaker reported that on client sites information towards the end of the content got cited, and invited others to check their own examples.

Speaker Jovana AvramovicEvidence transcript

AnalysisD1-C417

Placing key facts first or last is a community heuristic based on research into language models, not something Google has said its systems do; Google's own advice (D1-C054) is to write for people without chopping content, and a short summary at the top serves readers either way.

Author Ibrahim Anjro

Used byglossary term Lost in the middle

StageD1-C418

A community speaker summed up the shift as SEO's target expanding rather than moving: besides asking whether a page can rank, ask whether its information is useful as a source and easy to extract.

“The target is not shifting; it is expanding.”

Speaker Jovana AvramovicEvidence transcript

AnalysisD1-C421

Of the two answers a community speaker offered for a request to a URL variant (redirect, or an error page), Google's canonicalization guide favours the redirect: a redirect is a strong signal that its target should become canonical, while an error page throws away any links pointing at the variant.

Author Ibrahim Anjro

16:35 · Q&A 85

StageD1-C120

To get Google to build something, such as an addition to an API, report and request it publicly and in volume.

Speaker not identifiedEvidence notes, transcript

  • Extended by D2-C511 Day 2: Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a…
StageD1-C422

An audience member asked, in a question submitted at registration, what the biggest difference is between crawling for search and crawling for AI models.

From the audienceEvidence transcript

  • Answered by D1-C423 Day 1: A Google panelist said many AI crawlers are less sophisticated than search crawlers: in server logs, search…
  • Answered by D1-C424 Day 1: Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas…
  • Answered by D1-C425 Day 1: Crawling for AI model training differs from search crawling because training mainly needs a very large number…
StageD1-C423

A Google panelist said many AI crawlers are less sophisticated than search crawlers: in server logs, search crawlers tend to follow where a site changes and which pages are valuable, while AI crawlers may simply work through a site in order, which he put down to less crawling experience and different priorities.

“my feeling is a lot of the AI crawlers are still a bit stupid”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

  • Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
StageConsistent with docsD1-C424

Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas many AI systems crawl a site with no understanding of it and take everything.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
StageNot in docsD1-C425

Crawling for AI model training differs from search crawling because training mainly needs a very large number of tokens, and it matters little which pages they come from.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
StageD1-C426

Asked whether optimising for AI search can hurt classic search or the reverse, a Google panelist said yes, pointing to questionable GEO advice published online that he declined to name.

Speaker not identifiedEvidence transcript

Things
StageD1-C427

An audience question asked how to make sure bots behave well beyond robots.txt.

From the audienceEvidence transcript

  • Answered by D1-C428 Day 1: Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so…
  • Answered by D1-C429 Day 1: A Google panelist said crawlers that ignore robots.txt and cause a nuisance are better treated as a scraping…
StageConfirmed by docsD1-C428

Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C427 Day 1: An audience question asked how to make sure bots behave well beyond robots.txt.
StageNot in docsD1-C432

A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

Speaker not identifiedEvidence transcript

  • Extended by D1-C433 Day 1: The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices'…
AnalysisD1-C433

The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.

Author Ibrahim Anjro

  • Extends D1-C432 Day 1: A Google panelist said they were working on a set of crawler best practices and offering research…
StageConfirmed by docsD1-C434

User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05glossary term User-triggered fetchers

  • Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
  • Extended by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
StageD1-C436

A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

Speaker not identifiedEvidence transcript

  • Extended by D2-C377 Day 2: AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and…
StageD1-C437

A Google panelist said much of robots.txt is written with search engines in mind: a search engine should never add items to a cart or check out, but an agent acting for a user probably should be able to.

Speaker not identifiedEvidence transcript

StageD1-C438

An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as sports sites, whose content changes constantly and is largely near-duplicate across seasons.

From the audienceEvidence transcript

  • Answered by D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
  • Answered by D1-C440 Day 1: If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its…
StageConsistent with docsD1-C439

Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.

Speaker not identifiedEvidence transcript

  • Answers D1-C438 Day 1: An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as…
  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
  • Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
  • Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
StageConfirmed by docsD1-C440

If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-PRF-01

  • Answers D1-C438 Day 1: An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as…
  • Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…
StageD1-C441

An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl budget inefficiently on a site.

From the audienceEvidence transcript

  • Answered by D1-C442 Day 1: Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the…
  • Answered by D1-C443 Day 1: To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued…
  • Answered by D1-C444 Day 1: Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily…
  • Answered by D1-C445 Day 1: One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it…
StageNot in docsD1-C442

Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the site lies; many similar URLs being crawled may not matter if the site is nowhere near that limit.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-MON-04

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
StageConsistent with docsD1-C443

To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued over time, and whether the average response time shows the server is fast enough or is limiting Googlebot.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-MON-04

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
StageConsistent with docsD1-C444

Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
  • Extends D1-C397 Day 1: A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary…
StageConsistent with docsD1-C445

One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-11

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
  • Repeats D1-C403 Day 1: A community speaker advised that a web application compute the expected URL for every request, for example…
StageNot in docsD1-C446

Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
StageConsistent with docsD1-C447

Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
StageD1-C448

An audience member asked what changes to expect in Search Console's reporting on Google's AI features.

From the audienceEvidence transcript

  • Answered by D1-C449 Day 1: A Google panelist said he expected more AI reporting to launch in Search Console, without a timeline or a…
  • Answered by D1-C450 Day 1: Google said reporting on AI features should not simply mirror classic search results, because users get more…
StageD1-C449

A Google panelist said he expected more AI reporting to launch in Search Console, without a timeline or a promise, because many people at Google care about giving site owners the information to understand and adapt to what happens in Search.

Speaker not identifiedEvidence transcript

  • Answers D1-C448 Day 1: An audience member asked what changes to expect in Search Console's reporting on Google's AI features.
  • Extended by D3-C440 Day 3: Google called AI reporting in Search Console an evolving space and expects the generative AI report to get…
StageD1-C450

Google said reporting on AI features should not simply mirror classic search results, because users get more information before they click and interact with pages differently, so it first has to work out which data would be useful and actionable.

Speaker not identifiedEvidence transcript

  • Answers D1-C448 Day 1: An audience member asked what changes to expect in Search Console's reporting on Google's AI features.
StageD1-C451

An audience member asked for data on AI features as a feedback channel for improving content, as an alternative to prompt-tracking tools.

From the audienceEvidence transcript

  • Answered by D1-C452 Day 1: Query data for AI features is hard to provide because people ask AI very different kinds of questions that do…
StageConsistent with docsD1-C452

Query data for AI features is hard to provide because people ask AI very different kinds of questions that do not map back to keywords as in classic search, and the data would have to be grouped to protect privacy while staying useful.

Speaker not identifiedEvidence transcript

  • Answers D1-C451 Day 1: An audience member asked for data on AI features as a feedback channel for improving content, as an…
StageD1-C453

An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed Insights and Lighthouse advocate it and audit tools ask site owners to create the file.

From the audienceEvidence transcript

Things
  • Answered by D1-C454 Day 1: Gary Illyes said llms.txt does not matter for Google Search right now and that he does not expect it to…
  • Answered by D1-C455 Day 1: Gary Illyes said llms.txt matters to some people, which is why tools such as Lighthouse add checks for it, so…
  • Answered by D1-C456 Day 1: A Google panelist said he knew of no plan for Google to use llms.txt, though it could happen, and that a…
  • Answered by D1-C457 Day 1: A Google panelist said most llms.txt files he had seen are generated automatically by a setting in SEO…
  • Answered by D1-C458 Day 1: A Google panelist objected that llms.txt is designed for supposedly intelligent systems that should be able…
  • Answered by D1-C459 Day 1: A Google panelist compared llms.txt to the old meta keywords debate: an AI agent should not blindly trust…
StageConfirmed by docsD1-C454

Gary Illyes said llms.txt does not matter for Google Search right now and that he does not expect it to, while adding that he had been wrong before.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-02glossary term llms.txt

  • Answers D1-C453 Day 1: An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed…
StageConsistent with docsD1-C455

Gary Illyes said llms.txt matters to some people, which is why tools such as Lighthouse add checks for it, so both sides of the debate are right in their own context.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-02

  • llms.txt Chrome for Developers · checked 3 October 2026
  • Answers D1-C453 Day 1: An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed…
StageD1-C456

A Google panelist said he knew of no plan for Google to use llms.txt, though it could happen, and that a possible future change is no reason to implement it now.

“Just because maybe something will change in the future doesn't mean you should take action now.”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Things
  • Answers D1-C453 Day 1: An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed…
StageD1-C457

A Google panelist said most llms.txt files he had seen are generated automatically by a setting in SEO plugins, so if Google Search ever made llms.txt matter, a site could add one by ticking a checkbox.

Speaker not identifiedEvidence transcript

Things
  • Answers D1-C453 Day 1: An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed…
StageD1-C458

A Google panelist objected that llms.txt is designed for supposedly intelligent systems that should be able to parse a website.

“it's designed for allegedly intelligent systems”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Things
  • Answers D1-C453 Day 1: An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed…
StageD1-C459

A Google panelist compared llms.txt to the old meta keywords debate: an AI agent should not blindly trust what a site says about its own authority, just as a site calling itself the best car insurance site is no reason to stop looking at others.

Speaker not identifiedEvidence transcript

Things
  • Answers D1-C453 Day 1: An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed…
DocsSourceD1-C460

Chrome's Lighthouse documentation lists an llms.txt audit among its agentic browsing audits: it flags a server error when llms.txt is fetched and marks the audit not applicable when the file is missing, because providing the file is optional for now.

Publisher Chrome for Developers

Used byrequirement DEV-AIF-02

  • llms.txt Chrome for Developers · checked 3 October 2026
StageD1-C461

An audience member asked whether and when crawl stats will be added to the Search Console API.

From the audienceEvidence transcript

  • Answered by D1-C462 Day 1: Google said it knew of no plans to add crawl stats to the Search Console API, because that would be a…
  • Answered by D1-C463 Day 1: A Google panelist said adding a report that is a variation of the Performance report to the Search Console…
  • Answered by D1-C464 Day 1: Google's Search Relations team talks with the Search Console team about the API more than before, partly…
StageConsistent with docsD1-C462

Google said it knew of no plans to add crawl stats to the Search Console API, because that would be a completely new part of the API.

Speaker not identifiedEvidence transcript

  • API Reference Google Search Console API documentation · checked 3 October 2026
  • Answers D1-C461 Day 1: An audience member asked whether and when crawl stats will be added to the Search Console API.
StageConsistent with docsD1-C463

A Google panelist said adding a report that is a variation of the Performance report to the Search Console API would be easier than adding crawl stats, because the Performance report is already in the API.

Speaker not identifiedEvidence transcript

  • API Reference Google Search Console API documentation · checked 3 October 2026
  • Answers D1-C461 Day 1: An audience member asked whether and when crawl stats will be added to the Search Console API.
StageD1-C464

Google's Search Relations team talks with the Search Console team about the API more than before, partly because with so many vibe-coded systems around it is very easy to just say 'use the Search Console API to do something', and it is annoying when Search Console has no API for that.

Speaker not identifiedEvidence transcript

  • Answers D1-C461 Day 1: An audience member asked whether and when crawl stats will be added to the Search Console API.
StageD1-C465

An audience member asked how a large news site can tell whether crawl budget is limiting how fast new articles are discovered (within minutes), and which statistics in Search Console and the logs show this.

From the audienceEvidence transcript

  • Answered by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
  • Answered by D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
  • Answered by D1-C468 Day 1: When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google…
  • Answered by D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…
StageConsistent with docsD1-C466

Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-MON-11

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Repeats D1-C381 Day 1: Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few…
  • Extended by D3-C606 Day 3: For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading…
StageNot in docsD1-C467

To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Extended by D1-C543 Day 1: Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises…
StageConsistent with docsD1-C468

When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google crawls and recrawls a lot and judging whether they are useful.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-MON-11

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
StageConsistent with docsD1-C469

Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-URL-08

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Extended by D1-C470 Day 1: Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless…
AnalysisD1-C470

Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.

Author Ibrahim Anjro

  • Extends D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…
StageD1-C471

An audience question noted that a cats.txt file received more crawler requests than llms.txt in the server logs and asked how important that is.

From the audienceEvidence transcript

Things
  • Answered by D1-C472 Day 1: cats.txt was created by an SEO as a satirical take on llms.txt, to show that crawlers fetch whatever files…
  • Answered by D1-C473 Day 1: Files such as cats.txt have no importance for Google Search: if a site links to one, Googlebot will find and…
StageD1-C472

cats.txt was created by an SEO as a satirical take on llms.txt, to show that crawlers fetch whatever files they are given, so requests for llms.txt in server logs do not mean the file is important.

Speaker not identifiedEvidence transcript

Things

Used byglossary term cats.txt

  • Answers D1-C471 Day 1: An audience question noted that a cats.txt file received more crawler requests than llms.txt in the server…
StageConsistent with docsD1-C473

Files such as cats.txt have no importance for Google Search: if a site links to one, Googlebot will find and crawl it, but it has no other effect.

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-AIF-02glossary term cats.txt

  • Answers D1-C471 Day 1: An audience question noted that a cats.txt file received more crawler requests than llms.txt in the server…
StageD1-C474

An audience member asked where the sweet spot is for a sitemap that misses nothing but does not overflow Search Console.

From the audienceEvidence transcript

  • Answered by D1-C475 Day 1: There is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that…
StageConfirmed by docsD1-C475

There is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that Google can find each of them.

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-URL-05

  • Answers D1-C474 Day 1: An audience member asked where the sweet spot is for a sitemap that misses nothing but does not overflow…
StageD1-C476

An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not indexed.

From the audienceEvidence transcript

  • Answered by D1-C477 Day 1: In a migration about four or five years before the event, Google consolidated 12 or 14 of its blogs in…
  • Answered by D1-C478 Day 1: The first priority in Google's own site consolidation was to identify the popular URLs people care about and…
  • Answered by D1-C479 Day 1: Where content was duplicated across languages, Google's own site consolidation redirected two language…
  • Answered by D1-C540 Day 1: Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is…
StageConsistent with docsD1-C477

In a migration about four or five years before the event, Google consolidated 12 or 14 of its blogs in different languages, a Help Center and its old developers site into one site.

Speaker not identifiedEvidence transcript

  • Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
  • Extended by D1-C480 Day 1: The consolidation described on stage matches Google's November 2020 move from Google Webmasters to Google…
StageNot in docsD1-C478

The first priority in Google's own site consolidation was to identify the popular URLs people care about and make sure the migration did not damage them.

Speaker not identifiedEvidence transcript

  • Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
StageConsistent with docsD1-C479

Where content was duplicated across languages, Google's own site consolidation redirected two language versions into one, giving both old URLs one target path to redirect to.

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-CAN-10

  • Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
  • Extended by D1-C541 Day 1: A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in…
AnalysisD1-C480

The consolidation described on stage matches Google's November 2020 move from Google Webmasters to Google Search Central, which moved its main blog, 13 localized blogs and the Search help content of the Search Console Help Center to its developer site: about six years before the event, not the four or five said on stage.

Author Ibrahim Anjro

  • Extends D1-C477 Day 1: In a migration about four or five years before the event, Google consolidated 12 or 14 of its blogs in…
StageD1-C481

An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that site owners can control AI access separately from Googlebot and analyse it in their logs.

From the audienceEvidence transcript

Things
  • Answered by D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
  • Answered by D1-C483 Day 1: A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense…
  • Answered by D1-C484 Day 1: A Google panelist called blocking all AI crawlers while allowing search crawlers such as Googlebot, Bingbot…
  • Answered by D1-C485 Day 1: To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
  • Answered by D1-C487 Day 1: A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site…
  • Answered by D1-C488 Day 1: When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt…
  • Answered by D1-C489 Day 1: A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a…
StageConsistent with docsD1-C482

Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
StageD1-C483

A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense yet, because nobody knows how these systems will develop.

Speaker not identifiedEvidence transcript

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
StageConfirmed by docsD1-C485

To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.

Speaker not identifiedEvidence transcript

Used byrequirements DEV-AIF-05, DEV-IDX-11

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Repeated by D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
StageNot in docsD1-C487

A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Repeated by D1-C524 Day 1: Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.
StageNot in docsD1-C488

When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
StageD1-C489

A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.

“Personally, I think that's a bad pattern, because you don't know what you're blocking.”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
StageD1-C490

An audience member asked what happens to thousands of crawled and indexed paginated category and tag pages if they are replaced by a single page with a JavaScript load-more button, and whether the old pages should then redirect to the canonical first page.

From the audienceEvidence transcript

  • Answered by D1-C491 Day 1: If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them…
  • Answered by D1-C114 Day 1: A Google panelist called pagination one of the trickiest things in web development and said switching to…
StageConfirmed by docsD1-C491

If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.

“Googlebot is not running around clicking on buttons”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-URL-07

  • Answers D1-C490 Day 1: An audience member asked what happens to thousands of crawled and indexed paginated category and tag pages if…
  • Extended by D2-C271 Day 2: Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on…
StageConsistent with docsD1-C492

John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a rough view of whether discovery crawling is in a reasonable range, though measuring publish-to-first-crawl time is more accurate.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-MON-11

  • Extended by D1-C544 Day 1: John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering…
StageD1-C545

An audience member asked, for very large sites of about 100 million pages, what signs show that a site is limited by crawl budget, and how to tell a crawl capacity limit problem from a crawl demand problem.

From the audienceEvidence transcript

  • Answered by D1-C493 Day 1: There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower…
  • Answered by D1-C494 Day 1: To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot…
  • Answered by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
StageConsistent with docsD1-C493

There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Answers D1-C545 Day 1: An audience member asked, for very large sites of about 100 million pages, what signs show that a site is…
  • Extended by D3-C617 Day 3: Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2…
StageConsistent with docsD1-C494

To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot opens to the site: if it has dropped, the capacity limit changed.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Answers D1-C545 Day 1: An audience member asked, for very large sites of about 100 million pages, what signs show that a site is…
StageConsistent with docsD1-C114

A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

“the infinite scrolling or load more thing is kind of dangerous, depending on what you want to achieve”

Wording checked against the slide or recording

Speaker not identifiedEvidence notes, transcript

Things

Used byrequirements DEV-URL-06, DEV-URL-07

  • Answers D1-C490 Day 1: An audience member asked what happens to thousands of crawled and indexed paginated category and tag pages if…
  • Extended by D2-C197 Day 2: Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll'…
  • Extended by D2-C271 Day 2: Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on…
StageConsistent with docsD1-C538

A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.

Speaker not identifiedEvidence transcript

Used byglossary term Crawl demand

  • Extends D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
  • Extends D1-C334 Day 1: Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site…
  • Repeated by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
StageConsistent with docsD1-C539

A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extends D1-C447 Day 1: Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its…
  • Extends D1-C446 Day 1: Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a…
StageConsistent with docsD1-C540

Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-CAN-09

  • Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
  • Extended by D2-C884 Day 2: The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a…
StageNot in docsD1-C541

A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in Google's own site migration, because it was the only option available to the person doing it (the recording does not make fully clear whether the JavaScript performed the redirects or built the mapping).

Speaker not identifiedEvidence transcript

Used byrequirement DEV-CAN-01

  • Extends D1-C479 Day 1: Where content was duplicated across languages, Google's own site consolidation redirected two language…
  • Extended by D1-C542 Day 1: Google's redirects guide says Google Search follows JavaScript location redirects only after rendering, may…
AnalysisD1-C542

Google's redirects guide says Google Search follows JavaScript location redirects only after rendering, may never see one if rendering fails, and should be used only when server-side or meta refresh redirects are impossible; its site move guide asks for server-side permanent redirects (301 or 308) where technically possible. The JavaScript in Google's own migration was a fallback ('my only option'), not a pattern to copy.

Author Ibrahim Anjro

Used byrequirement DEV-CAN-01

  • Extends D1-C541 Day 1: A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in…
StageNot in docsD1-C543

Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Extends D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
StageConsistent with docsD1-C544

John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering new pages, that is a lot of discovery crawling, which suggests that crawling of new pages is not the problem.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-MON-11

  • Extends D1-C492 Day 1: John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a…
StageConsistent with docsD1-C546

A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).

Speaker not identifiedEvidence transcript

  • Answers D1-C545 Day 1: An audience member asked, for very large sites of about 100 million pages, what signs show that a site is…
  • Extends D1-C374 Day 1: Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can…
  • Extended by D1-C547 Day 1: Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the…
AnalysisD1-C547

Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the total time a server spends holding connections open for Google, counting both the number of parallel connections and their duration. That fits the panel's advice to watch how many connections Googlebot opens (D1-C494): fewer connections is the documented form of a lower capacity limit.

Author Ibrahim Anjro

Used byrequirement DEV-PRF-01

  • Extends D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…

Day 1, session not recorded 17

PressSourceD1-C112

As reported by Search Engine Journal, John Mueller said in April 2026 that there is no penalty or ranking demotion for having multiple URLs with the same content; Google picks one to keep.

“There's no penalty or ranking demotion if you have multiple URLs going to the same content.”

Reported by Search Engine Journal (8 April 2026)

DocsSourceD1-C128

Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.

“Some duplicate content on a site is normal and it's not a violation of Google's spam policies.”

Publisher Google Search Central

Used byglossary term Duplicate cluster

  • Repeated by D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
  • Extended by D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
  • Extended by D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
  • Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
  • Extended by D2-C408 Day 2: A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying…
  • Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
AnalysisD1-C113

The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and every copy is still crawled. Copied or scraped content is a separate spam-policy issue.

Author Ibrahim Anjro

  • Extended by D2-C380 Day 2: Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google…
  • Extended by D2-C401 Day 2: Broken canonical tags can make the wrong pages of a site show up in search results.
DocsSourceD1-C115

Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

Publisher Google Search Central

Used byrequirement DEV-URL-06

  • Repeated by D2-C040 Day 2: Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not…
  • Extended by D2-C184 Day 2: Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link…
  • Extended by D2-C187 Day 2: A market or language selector built as a button works for users but leaves the whole cluster of…
  • Extended by D2-C195 Day 2: Content that loads only after a user action such as a click or a scroll is not in the DOM while Google…
  • Extended by D2-C201 Day 2: Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () =>…
  • Extended by D2-C268 Day 2: Content that loads only when a user clicks an element is not supported in the way Google renders pages for…
  • Extended by D2-C276 Day 2: To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each…
  • Contradicted by D2-C393 Day 2: Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can…
  • Extended by D2-C395 Day 2: Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated…
StageD1-C116

Images were stressed as important.

Evidence notes

  • Extended by D2-C523 Day 2: Gary Illyes said images and videos drive a large amount of traffic to publishers.
  • Extended by D2-C525 Day 2: An image Google has extracted can appear almost anywhere Google shows results, including Discover, image…
AnalysisD1-C117

With one in six AI Mode searches being multimodal, original images with descriptive file names, alt text and captions feed AI answers as well as image search.

Author Ibrahim Anjro

Things
  • Extended by D2-C525 Day 2: An image Google has extracted can appear almost anywhere Google shows results, including Discover, image…
DocsSourceD1-C131

Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

Publisher Google Search Central

Used byrequirement DEV-HTM-04

  • Repeated by D1-C274 Day 1: A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the…
  • Extended by D1-C278 Day 1: A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an…
  • Extended by D2-C397 Day 2: Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes…
  • Extended by D2-C465 Day 2: Gemini in Chrome relies heavily on the screenshot it takes of a page.
AnalysisD1-C119

Google has not called ARIA a ranking signal. The case for it is that AI agents and assistive tools read pages through the same structure: real buttons, labelled forms and semantic headings.

Author Ibrahim Anjro

Used byrequirement DEV-HTM-04

StageD1-C121

SEO is not only content; it has many parts.

Evidence notes

  • Extended by D2-C136 Day 2: Many SEOs still treat everything beyond the raw HTML as the developers' business, but developers often do not…
StageNot in docsD1-C122

Check and focus on rich results for Google.

Evidence notes

  • Extended by D2-C521 Day 2: Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its…
StageD1-C060

Speakers at the event were openly dismissive of llms.txt.

Evidence notes, transcript

Things
  • Extended by D1-C271 Day 1: A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard…
AnalysisD1-C077

This does not mean avoiding HTTP/3, which benefits browsers. The rule is never to run a host that only answers over HTTP/3, and to check that CDN or firewall rules written for HTTP/3 traffic do not break HTTP/1.1 and HTTP/2.

Author Ibrahim Anjro

Things

Used byrequirement DEV-SRV-07

Day 2: Indexing 970

10:15 · Welcome to indexing day! 36

SlideD2-C001

An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's crawl discovery and whether they pass link equity to deeper pages.

From the audienceEvidence slide photo

Things
  • Answered by D2-C002 Day 2: Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.
  • Answered by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
  • Answered by D2-C003 Day 2: Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have…
  • Answered by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
SlideNot in docsD2-C002

Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.

“It's hard to make good HTML sitemaps for large sites.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
SlideConsistent with docsD2-C820

Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

“Better rely on hubs like category pages that link out to your important pages.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-04

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extends D1-C202 Day 1: Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they…
  • Extended by D2-C838 Day 2: Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one…
SlideD2-C003

Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have the extra benefit that they may also be useful to users.

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo

Used byrequirement DEV-URL-04

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
AnalysisD2-C004

Google's answer did not say whether HTML sitemaps pass link equity. On a large site, check that every important page is linked from a category or hub page in the normal navigation instead of relying on an HTML sitemap page to reach it.

Author Ibrahim Anjro

AnalysisD2-C008

Google's view of HTML sitemaps has moved: a 2005 blog post encouraged them, the current sitemap and ecommerce documentation does not mention them, and the 2026 Q&A slide pointed large sites to category hub pages instead.

Author Ibrahim Anjro

Things
SlideD2-C009

An audience member asked whether a product detail page that is out of stock for two to three months should keep returning 200 with links to similar products, or be 302-redirected to a similar product or to its parent product listing page.

From the audienceEvidence slide photo, transcript

  • Answered by D2-C010 Day 2: Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it…
  • Answered by D2-C011 Day 2: Google's Q&A slide on out-of-stock product pages said users might wait months for some products, or even…
  • Answered by D2-C841 Day 2: Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to…
SlideNot in docsD2-C010

Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it depends on how important that product page is to users.

“It depends on the importance of the PDPs to the users.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Things

Used byrequirement DEV-ERR-04

  • Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
  • Extended by D2-C841 Day 2: Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to…
SlideConsistent with docsD2-C011

Google's Q&A slide on out-of-stock product pages said users might wait months for some products, or even pre-order them if the site offers pre-ordering.

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-ERR-04

  • Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
DocsSourceD2-C012

Google's guide to temporarily pausing an online business recommends that a shop expecting to sell again within weeks or months stays online with limited functionality, such as a disabled cart, and updates its Product structured data to show current availability.

Publisher Google Search Central

Used byrequirements DEV-ERR-04, DEV-SDA-08, DEV-SRV-03

AnalysisD2-C014

Keep a product page that is out of stock for a few months live with a 200 status, show its availability on the page and in Product structured data, and offer pre-ordering where possible when users would wait for the product. Redirect it only when users would rather switch to a similar product than wait.

Author Ibrahim Anjro

Used byrequirement DEV-ERR-04

SlideD2-C015

An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.

From the audienceEvidence slide photo, transcript

  • Answered by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
  • Answered by D2-C018 Day 2: Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in…
  • Answered by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
  • Answered by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
SlideConfirmed by docsD2-C017

Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

“If it's included in the robots.txt file, any crawler can pick your sitemaps up”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extended by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
SlideConsistent with docsD2-C018

Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in Search Console.

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
SlideD2-C019

An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.

From the audienceEvidence slide photo, transcript

  • Answered by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
  • Answered by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
  • Answered by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
  • Answered by D2-C845 Day 2: Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no…
  • Answered by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
SlideNot in docsD2-C020

Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

“Simply put, it's because of sites that are extremely important and like to disallow their most important pages, either accidentally or out of ignorance.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
  • Extended by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
SlideConfirmed by docsD2-C021

Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D1-C105 Day 1: URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
  • Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
  • Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
  • Extended by D2-C933 Day 2: Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt…
DocsSourceD2-C022

Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.

Publisher Google Search Central

Used byrequirement DEV-IDX-01glossary term noindex

  • Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
DocsSourceD2-C833

Google's sitemap guide says Google uses lastmod only when it is consistently and verifiably accurate, and counts a change to the main content, the structured data or the links of a page as significant, but not a changed copyright date.

Publisher Google Search Central

Used byrequirement DEV-URL-05

StageConsistent with docsD2-C838

Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.

“if you want to make one, knock yourself out, but I will focus on some better things like hub pages, category pages”

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-04glossary term Hub pages

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
  • Extends D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
StageD2-C839

An audience member asked how several sloppy migrations on the same domain can affect Googlebot's crawling.

From the audienceEvidence transcript

Things
  • Answered by D2-C840 Day 2: Google answered that several sloppy migrations on one domain can cause many effects in the short term, and…
StageConsistent with docsD2-C840

Google answered that several sloppy migrations on one domain can cause many effects in the short term, and noted that migrations concern indexing as well as crawling, a subject a later Day 2 talk would cover.

Speaker not identifiedEvidence transcript

  • Answers D2-C839 Day 2: An audience member asked how several sloppy migrations on the same domain can affect Googlebot's crawling.
StageNot in docsD2-C841

Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to them, with easily replaced products, such as a particular cheese, for which the buyer simply picks a similar product instead of waiting.

Speaker not identifiedEvidence transcript

  • Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
  • Extends D2-C010 Day 2: Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it…
StageConfirmed by docsD2-C842

Google said listing the sitemap in robots.txt is fine, as many websites do.

“you can include it in robots.txt. No problem whatsoever.”

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
StageConsistent with docsD2-C843

Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
StageNot in docsD2-C844

Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
StageConsistent with docsD2-C846

Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”

Speaker not identifiedEvidence transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
StageD2-C847

Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including basics of how the internet works such as TCP/IP, so that attendees could understand how crawlers work.

Speaker not identifiedEvidence transcript

  • Repeats D1-C317 Day 1: Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses…
StageConsistent with docsD2-C848

Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.

Speaker not identifiedEvidence transcript

  • Repeats D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
AnalysisD2-C849

Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

Author Ibrahim Anjro

  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…

10:25 · How is HTML interpreted 27

SlideNot in docsD2-C024

A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

Speaker Cherry PrommawinEvidence slide photo

Used byrequirement DEV-PRF-01

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
SlideConsistent with docsD2-C025

In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

Speaker Cherry PrommawinEvidence slide photo

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
SlideConsistent with docsD2-C026

A Google pipeline slide placed processing between the crawler and the index and listed six processing steps: HTML parsing, rendering, deduplication, feature extraction, signal extraction and index selection.

Speaker Cherry PrommawinEvidence slide photo

  • Extends D1-C036 Day 1: Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extended by D2-C128 Day 2: Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached…
  • Repeated by D2-C441 Day 2: Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML…
  • Extended by D2-C444 Day 2: Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a…
  • Extended by D2-C688 Day 2: Index selection is the last step before documents enter Google's index.
SlideConsistent with docsD2-C028

Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • Repeated by D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
  • Extended by D2-C868 Day 2: Gary Illyes said the main content is what Google considers when ranking a page.
  • Extended by D2-C474 Day 2: Google's systems sometimes fail to determine a page's main content correctly, and structured data helps…
AnalysisD2-C029

Make the main content of every template easy to separate from the header, navigation and footer, for example as one clearly delimited main area, because Google identifies the main content and treats it as the most important part of the page.

Author Ibrahim Anjro

Used byrequirement DEV-HTM-01

SlideConfirmed by docsD2-C031

Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirement DEV-CAN-03

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extended by D2-C379 Day 2: rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for…
SlideConfirmed by docsD2-C033

Google extracts hreflang annotations, through which site owners specify the language variants of their content, to know whether a page has an equivalent with similar content in another language.

Speaker Cherry PrommawinEvidence slide photo, transcript

Things

Used byglossary term hreflang

  • Extended by D2-C382 Day 2: When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of…
SlideConfirmed by docsD2-C036

Links and anchors are among the things Google extracts from a page's HTML, and the slide card for them simply read 'We like links.'

“We like links.”

Wording checked against the slide or recording

Speaker Cherry PrommawinEvidence 2 slide photos, transcript

Used byrequirement DEV-URL-01

StageD2-C037

The speaker said links are still an extremely important part of the internet and of most major search and AI systems.

Speaker Cherry PrommawinEvidence transcript

StageConsistent with docsD2-C038

Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.

Speaker Cherry PrommawinEvidence transcript

Used byrequirement DEV-URL-01

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
SlideConfirmed by docsD2-C039

Google can extract links written as an a element with an href attribute that holds an absolute or a relative URL, which the speaker called the good old normal way.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-01

SlideConsistent with docsD2-C040

Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-02

  • Repeats D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
  • Extended by D2-C184 Day 2: Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link…
  • Extended by D2-C287 Day 2: A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google…
SlideConsistent with docsD2-C041

Google cannot extract a link from an href attribute placed on an element other than a, such as a span, because that is not a standard way to make a link.

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-02

DocsSourceD2-C043

Google's link best practices say Google can generally crawl a link only if it is an a element with an href attribute, and list routerLink without href, href on a span, onclick-only a elements and javascript: URLs as not recommended, while noting that Google may still attempt to parse them.

Publisher Google Search Central

Used byrequirements DEV-URL-01, DEV-URL-02

AnalysisD2-C044

Google's link-extraction slide put routerLink, href on a span, onclick-only links and javascript: URLs under 'can not extract', which is stricter than Google's link documentation saying Google may still try to parse them; either way they are not dependable links for discovery.

Author Ibrahim Anjro

AnalysisD2-C045

Audit every template and JavaScript component that outputs links, such as navigation, pagination, filters and product tiles: each needs a real a element with an href, because onclick handlers, routerLink without href, href on a span and javascript: URLs leave the target pages without a link Google can reliably extract.

Author Ibrahim Anjro

StageNot in docsD2-C046

The speaker said Google sometimes also extracts URLs that are typed out as plain text on a page without being hyperlinked; the remarks around this point were unclear in the recording.

Speaker Cherry PrommawinEvidence transcript

AnalysisD2-C047

Do not rely on plain-text URLs for discovery: even if Google sometimes picks them up, a proper a href link is what was described as feeding discovery, site structure and ranking.

Author Ibrahim Anjro

SlideConfirmed by docsD2-C048

Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the newly found URLs.

Speaker Cherry PrommawinEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D1-C328 Day 1: During indexing Google extracts the URLs found in fetched content and passes them back to the crawler and the…
  • Extended by D2-C168 Day 2: In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the…
  • Repeated by D2-C442 Day 2: Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.
StageConsistent with docsD2-C049

The robots meta element is also extracted when Google processes a page's HTML, and the speaker called it probably one of the most important extracted elements, or one the audience is probably interested in.

Speaker Cherry PrommawinEvidence transcript

DocsSourceD2-C831

Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.

Publisher Google Search Central

Used byrequirements DEV-CAN-03, DEV-HTM-05

10:30 · Controlling indexing 66

StageConfirmed by docsD2-C050

John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

Speaker John MuellerEvidence transcript

  • Repeats D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
StageD2-C052

John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.

Speaker John MuellerEvidence transcript

  • Repeated by D2-C217 Day 2: Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard…
SlideConsistent with docsD2-C053

John Mueller opened with a true-or-false quiz slide asking whether robots meta tags can make a page more visible in Search results than having none, and later answered that the statement is true.

“You can use robots meta tags to be more visible in Search results than without robots meta tags.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-IDX-04

SlideConfirmed by docsD2-C057

The robots meta tag is written as <meta name="robots" content="rule1,rule2">, or with the name googlebot instead of robots, with several rules separated by commas.

Speaker John MuellerEvidence slide photo

Used byglossary term Robots meta tag

StageConfirmed by docsD2-C058

John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.

Speaker John MuellerEvidence transcript

Used byglossary term robots.txt

  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
SlideConfirmed by docsD2-C059

The robots meta rule all is the default value and does nothing: it places no restrictions on indexing the page or following its links, the same as having no robots meta tag, and it is not an instruction that search engines must index the page.

“This is the default value - it does nothing”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-IDX-06

SlideNot in docsD2-C062

A slide named internal admin pages, temporary landing pages and thin content as typical use cases for the noindex rule.

Speaker John MuellerEvidence slide photo

Things
StageConsistent with docsD2-C063

The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page, which John Mueller called a weird and very broad rule that makes the page stand on its own.

Speaker John MuellerEvidence transcript

Things

Used byrequirement DEV-IDX-06glossary term nofollow

  • Extended by D2-C821 Day 2: John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would…
StageConsistent with docsD2-C064

John Mueller recommends rel=nofollow on individual links instead of the page-level nofollow robots rule, so a site can choose which links are useful and which are not.

Speaker John MuellerEvidence transcript

Things

Used byrequirement DEV-IDX-06

StageNot in docsD2-C066

John Mueller said AI crawlers do not really know what to do with nofollow links, because they look at the content rather than building a link graph; he did not say whether he meant Google's AI systems, other AI crawlers or both.

“AI crawlers don't really know what to do with a nofollow link, because they're looking at the content”

Speaker John MuellerEvidence transcript

Things
AnalysisD2-C067

nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

Author Ibrahim Anjro

  • Extends D1-C127 Day 1: Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed…
  • Extends D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
DocsSourceD2-C069

Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.

Publisher Google Search Central

Used byrequirements DEV-IDX-01, DEV-IDX-03glossary term X-Robots-Tag

  • Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
SlideConfirmed by docsD2-C070

The nosnippet rule stops Google from showing a text snippet or video preview for a page in search results, while the page's title is still shown.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-IDX-05glossary term nosnippet and data-nosnippet

SlideConfirmed by docsD2-C072

The nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI Mode in Search.

“Also prevents the content from being used as a direct input for AI Overviews and AI Mode in Search results.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-IDX-05glossary term nosnippet and data-nosnippet

  • Extends D1-C034 Day 1: Lead-generation, local-service and e-commerce sites should normally stay included, because AI answers cite…
StageConsistent with docsD2-C073

John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.

“we use the snippet as a way of building out the AI Overviews and the AI Mode answers”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-IDX-05

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
  • Extended by D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
DocsSourceD2-C074

Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements.

Publisher Google Search Central

Used byrequirements DEV-AIF-01, DEV-IDX-05glossary term AI Overviews

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageConsistent with docsD2-C076

John Mueller said data-nosnippet is rarely needed but lets a site keep a specific piece of text, such as a business phone number, out of the snippet, so the page can still be found for it while searchers have to visit the page to see it.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-IDX-07glossary term nosnippet and data-nosnippet

DocsSourceD2-C079

Google's robots meta tag specification says data-nosnippet may be extracted both before and after rendering, so the attribute should not be added to or removed from existing elements with JavaScript.

Publisher Google Search Central

Used byrequirements DEV-IDX-07, DEV-REN-05

StageConsistent with docsD2-C080

John Mueller said valid HTML is technically not a ranking factor but does matter for controls such as data-nosnippet.

“valid HTML is technically not an SEO ranking factor, but it does play a role”

Speaker John MuellerEvidence transcript

Used byrequirements DEV-HTM-05, DEV-IDX-07

AnalysisD2-C081

To hide one detail, such as a phone number, a price or a direct answer, instead of the whole snippet, wrap it in data-nosnippet on a span, div or section in the server HTML and validate the HTML so that an unclosed element cannot hide the rest of the page from snippets.

Author Ibrahim Anjro

Used byrequirement DEV-IDX-07

SlideNot in docsD2-C085

max-snippet:-1 removes the length limit and can produce a longer snippet than having no rule, because by default Google keeps snippets to a length it considers reasonable instead of quoting a page at length.

“max-snippet:-1 = No limit. (can be more than without)”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

SlideConfirmed by docsD2-C087

The max-image-preview rule sets the maximum size of a page's image previews: none shows no preview, standard a default-sized one and large the largest possible, for example in Discover.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-IDX-04glossary term max-image-preview

StageNot in docsD2-C088

John Mueller said Google Search does not always show an image thumbnail for a result, and when it does, the thumbnail is usually limited to the size of the result entry.

Speaker John MuellerEvidence transcript

StageConsistent with docsD2-C089

max-image-preview:large matters mainly in Discover, where it allows a large image that draws people's attention, so the rule can make a page more visible than leaving it out, John Mueller said.

Speaker John MuellerEvidence transcript

  • Repeated by D2-C936 Day 2: Setting the max-image-preview robots meta tag to large can make content perform surprisingly well in…
  • Extended by D3-C224 Day 3: Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels…
StageConsistent with docsD2-C094

John Mueller said notranslate can also be set in a meta tag named google, and that in that form it also turns off Chrome's automatic translation of the page.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-IDX-09

DocsSourceD2-C095

Google's Search documentation gives notranslate as a rule for a robots or googlebot meta tag or an X-Robots-Tag header, and says the rule opts a page out of all translation features in Google Search.

Publisher Google Search Central

Used byrequirements DEV-IDX-03, DEV-IDX-09glossary term X-Robots-Tag

StageConfirmed by docsD2-C100

The unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-IDX-08glossary term unavailable_after

  • Extended by D2-C698 Day 2: Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index…
SlideConfirmed by docsD2-C103

Setting the Search generative AI control in Search Console to exclude keeps a site's links and content out of Search generative AI features such as AI Overviews and AI Mode, so the site gets no traffic or impressions from them; include is the default.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-IDX-10

  • Repeats D1-C030 Day 1: The Search generative AI control covers AI Overviews, AI Mode and generative AI features in Discover.…
  • Extends D1-C029 Day 1: Search Console has a property setting called Search generative AI that gives direct control over AI Overviews…
SlideConsistent with docsD2-C104

To be as visible as possible in Google, John Mueller's closing slide recommended the robots rules max-image-preview:large and max-snippet:-1.

“To be as visible as possible in Google, use: max-image-preview:large, max-snippet:-1”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, video, transcript

Used byrequirement DEV-IDX-04

SlideConsistent with docsD2-C106

A slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-REN-05

  • Extended by D2-C853 Day 2: John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding…
SlideConsistent with docsD2-C108

Removing a robots restriction such as noindex with JavaScript does not work, a slide said.

“But... it takes more time, and removing restrictions (like "noindex") doesn't work.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-REN-05

  • Extended by D2-C852 Day 2: When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript…
AnalysisD2-C109

Ship robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may skip rendering a page that arrives with noindex, so the page can stay out of the index even if a script removes the tag later.

Author Ibrahim Anjro

Used byrequirement DEV-REN-05

  • Extended by D2-C855 Day 2: On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex)…
StageConfirmed by docsD2-C850

John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-IDX-01

  • Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
StageConfirmed by docsD2-C851

When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-IDX-01

  • Repeats D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
  • Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
StageConsistent with docsD2-C852

When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript, so a script cannot switch the page back to indexable, John Mueller said.

“we will see the noindex and say, oh, we will get rid of this page; we won't even process the JavaScript”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-REN-05

  • Extends D2-C108 Day 2: Removing a robots restriction such as noindex with JavaScript does not work, a slide said.
StageConfirmed by docsD2-C853

John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding them with JavaScript only where that is not possible, as in a JavaScript web app.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-REN-05

  • Extends D2-C106 Day 2: A slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.
DocsSourceD2-C854

Google's JavaScript SEO basics guide says that when Google encounters a noindex rule it may skip rendering and JavaScript execution, so using JavaScript to change or remove a noindex robots meta tag may not work as expected.

“it may skip rendering and JavaScript execution”

Publisher Google Search Central

Used byrequirement DEV-REN-05

AnalysisD2-C855

On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex), while Google's guide says only that it may skip rendering; either way, a noindex in the served HTML must never be one that JavaScript is expected to lift.

Author Ibrahim Anjro

  • Extends D2-C109 Day 2: Ship robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may…

10:40 · Lightning session D: Rendering and JavaScript 150

StageD2-C110

Google framed rendering historically: when Google started, most of the web was plain, semantic HTML it could extract directly, but the modern web is built with JavaScript and much content is generated by JavaScript, so Google had to adapt.

Speaker Erin SparlingEvidence transcript

SlideConsistent with docsD2-C112

Google renders pages in order to index what users see: its commitment to reflecting the user experience means understanding what people see when a page loads in a browser and surfacing that in Search.

“Rendering Pages to Index What Users See”

Wording checked against the slide or recording

Speaker Erin SparlingEvidence slide photo, transcript

Things

Used byglossary term Rendering

StageNot in docsD2-C113

Compared with other search engines and AI crawlers, Google said its effort to mimic what the user sees is essential to keeping its knowledge of the web up to date and comprehensive; it did not say what the others do.

Speaker Erin SparlingEvidence transcript

StageNot in docsD2-C115

Google described the mission of its rendering as simply executing JavaScript, while the implementation is complex, expensive and difficult.

“our mission is very simple these days: execute JavaScript”

Speaker Erin SparlingEvidence transcript

StageConsistent with docsD2-C116

AI Overviews and AI Mode are built on top of Search results: they are a different experience of the same content Google already has.

Speaker Erin SparlingEvidence transcript

Used bystory angle A-001

  • Repeats D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
  • Extended by D2-C601 Day 2: Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results…
StageConsistent with docsD2-C117

AI Overviews and AI Mode typically do not ground their answers by reading pages live, unlike Gemini when a user asks about a specific page, Google said.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-PRF-02

StageNot in docsD2-C118

To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.

“if it works for Search, it works for Gemini for training”

Speaker Erin SparlingEvidence transcript

Things

Used byrequirement DEV-IDX-11

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageNot in docsD2-C120

When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.

Speaker Erin SparlingEvidence transcript

Things

Used byrequirement DEV-PRF-02

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
DocsSourceD2-C121

Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

Publisher Google

Used byrequirement DEV-IDX-11glossary term Grounding

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
DocsSourceD2-C122

Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

Publisher Google

Used byrequirement DEV-IDX-11

  • Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
  • Extends D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…
StageConsistent with docsD2-C124

Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.

Speaker Erin SparlingEvidence transcript

Things

Used byrequirements DEV-PRF-02, DEV-REN-01

  • Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
StageConfirmed by docsD2-C126

Google renders pages with Chromium, the browser technology that also underlies Chrome, Edge and other Chromium-based browsers.

Speaker Erin SparlingEvidence transcript

Things

Used byglossary term Rendering

  • Repeats D1-C206 Day 1: Google renders JavaScript-heavy pages from their HTML, CSS and JavaScript as a browser would, using the…
SlideConsistent with docsD2-C127

A Google pipeline diagram ran from the crawl queue to the crawler, then through HTML parsing to processing, which passes pages to rendering and gets them back, and then to the index; rendering fetches its JavaScript and CSS through the crawler.

Speaker Erin SparlingEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
StageConsistent with docsD2-C128

Although Google's pipeline diagram shows rendering as part of indexing, Google's rendering is a detached system, kept separate because rendering is time-consuming and computationally expensive.

“it shows rendering as part of indexing, but it's actually a detached system”

Speaker Erin SparlingEvidence transcript

Things
  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
StageNot in docsD2-C129

After the crawler has fetched a page, Google's rendering system executes it, checks that it loads properly and works out what it looks like at different sizes.

Speaker Erin SparlingEvidence transcript

Things
DocsSourceD2-C131

Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.

Publisher Google Search Central

Used byrequirement DEV-REN-01glossary term Rendering

  • Extended by D3-C676 Day 3: The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page…
StageConsistent with docsD2-C134

Google urged sites to make sure their JavaScript content can be crawled, rendered and indexed, calling this important today and also tomorrow, as AI systems increasingly ground answers to user requests.

Speaker Erin SparlingEvidence transcript

  • Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
SlideD2-C137

The first rendering blind spot a community speaker showed was content changes: text, images and recommendations that differ between the raw HTML and the rendered page, illustrated by a shop page whose rendered version had a whole extra section.

Speaker Sören BendigEvidence slide photo, transcript

Things

Used byrequirement DEV-REN-07

StageNot in docsD2-C138

Switching JavaScript off and on in the browser gives a quick first impression of which content on a page depends on rendering.

Speaker Sören BendigEvidence transcript

StageConsistent with docsD2-C139

A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.

“everything you want cited, include it in the raw HTML, server-side rendered”

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-REN-01

  • Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
StageConsistent with docsD2-C140

In a navigation example, the main navigation was in the raw HTML and would survive a rendering failure, but the whole sub-navigation was built by JavaScript, so all of its links would be inaccessible to bots if rendering broke.

Speaker Sören BendigEvidence transcript

StageD2-C142

One site that builds its whole content by rendering also rendered its meta description with HTML tags inside it, which makes no sense and points to a flawed process.

Speaker Sören BendigEvidence transcript

Things

Used byrequirement DEV-HTM-03

  • Extended by D2-C858 Day 2: Search engines such as Google and Bing would ignore HTML tags written inside a meta description, a community…
StageConsistent with docsD2-C144

Browser console messages reveal further problems on a page, Content Security Policy violations among them; few teams analyse console messages at scale, but they should, a community speaker said.

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-REN-09

StageD2-C145

On a travel site, the number of offers differed between the raw HTML and the rendered page, possibly because the two versions drew on different internal databases or applied extra filters.

Speaker Sören BendigEvidence transcript

StageD2-C146

On a retail brand's page, the rendered version promised a bigger discount for a newsletter sign-up than the non-rendered version that was served, a mismatch that can hurt customer satisfaction; the two recordings disagree on the figure the rendered page promised.

Speaker Sören BendigEvidence transcript

StageNot in docsD2-C149

A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal systems or APIs can be missing and unresolved placeholders can end up in the final snapshot, a community speaker said.

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-REN-06

  • Extended by D2-C259 Day 2: Erin Sparling named server-side or hybrid rendering, fallbacks and not leaving placeholders in the DOM, among…
  • Extended by D2-C263 Day 2: Content is missing from the rendered HTML either because the server does not serve it or because the content…
StageD2-C150

In one shop's example, brand names and product descriptions were rendered from a database; the brand name was missing from the meta description, and the shop sometimes appeared with unresolved placeholders and sometimes as intended.

Speaker Sören BendigEvidence transcript

StageD2-C152

Broken titles and snippets caused by unresolved placeholders may hurt business, especially for e-commerce shops and affiliate-heavy sites, a community speaker warned.

Speaker Sören BendigEvidence transcript

Things
StageNot in docsD2-C153

To catch intermittent rendering problems, archive the full HTML and the resources of each page during a site audit and analyse them yourself, a community speaker advised.

Speaker Sören BendigEvidence transcript

Things

Used byrequirement DEV-MON-05

StageNot in docsD2-C155

Content Security Policy lowers the risk of cross-site scripting and clickjacking by defining trusted hosts; a resource from a host that is not trusted is not used.

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-REN-09

StageD2-C156

Content Security Policy is often misunderstood: on a page of a large German banking group, a video that should have been shown was blocked by a CSP violation, which also appeared in the browser console.

Speaker Sören BendigEvidence transcript

  • Extended by D2-C859 Day 2: A rendering failure such as a Content Security Policy blocking a video on a landing page that explains how to…
StageConfirmed by docsD2-C157

Never block a resource that a page needs for rendering with robots.txt, a community speaker said, calling this common sense.

Speaker Sören BendigEvidence transcript

StageNot in docsD2-C162

A community speaker strongly advised SEOs to use Chrome DevTools more often to understand what happens when their pages render.

Speaker Sören BendigEvidence transcript

Things
  • Extended by D2-C860 Day 2: Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a…
AnalysisD2-C165

Compare the session counts that per-session-billed tools (chat, personalisation, testing, session replay) charge for with real user sessions, and load such tools only after a user interaction where they add no content that needs to be indexed.

Author Ibrahim Anjro

Used byrequirement DEV-PRF-03

SlideConfirmed by docsD2-C167

The diagram from Google's JavaScript SEO basics page shows a URL going from the crawl queue to the crawler, the crawled HTML going to processing, then the render queue and the renderer, whose rendered HTML returns to processing before the page reaches the index.

Speaker Rebecca YuEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
SlideConfirmed by docsD2-C168

In Google's processing step, links are extracted from the HTML Google already has, before rendering, and the URLs found go back to the crawl queue.

Speaker Rebecca YuEvidence slide photo, transcript

Things
  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extends D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
DocsSourceD2-C169

Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.

Publisher Google Search Central

Used byrequirement DEV-URL-01

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
SlideNot in docsD2-C171

A community speaker's slide put the usual wait in Google's render queue at seconds to a couple of minutes per page.

“Usually from seconds to a couple minutes.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo

DocsSourceD2-C172

Google's JavaScript SEO basics guide says a page may wait in the render queue for a few seconds but that it can take longer, and it gives no upper limit.

“The page may stay on this queue for a few seconds, but it can take longer than that.”

Publisher Google Search Central

Used byrequirements DEV-PRF-02, DEV-REN-01glossary term Render queue

  • Extended by D3-C627 Day 3: Content that JavaScript adds to a page is typically seen by Google's indexing system within a few hours, and…
SlideConsistent with docsD2-C174

Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.

“Whatever is in the DOM at that moment is what likely gets indexed.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo, transcript

Things
  • Repeated by D2-C262 Day 2: Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.
StageConfirmed by docsD2-C176

To check whether JavaScript content is affected, open the URL Inspection tool in Search Console or the Rich Results Test and look for that content in the rendered HTML tab.

Speaker Rebecca YuEvidence transcript

Used byrequirement DEV-MON-02glossary terms Raw HTML and rendered HTML, URL Inspection tool

StageD2-C177

A community speaker treats the rendered HTML shown in Google's testing tools as the final source of truth, because what Google sees there is what gets indexed.

“I always treat this page as the final source of truth”

Speaker Rebecca YuEvidence transcript

StageNot in docsD2-C180

If content is missing from the rendered HTML, find the script responsible in Chrome DevTools: open the Network tab, filter by Fetch/XHR and reload the page.

Speaker Rebecca YuEvidence transcript

Things
StageNot in docsD2-C182

If the cause of missing JavaScript content is still unclear after checking the network requests, search the source code for a string related to the missing content.

Speaker Rebecca YuEvidence transcript

SlideConfirmed by docsD2-C184

Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.

Speaker Rebecca YuEvidence slide photo, transcript

Used byrequirement DEV-URL-02

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
  • Extends D2-C040 Day 2: Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not…
SlideConfirmed by docsD2-C185

The crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>, combined with History API routing (window.history.pushState) instead of hash routes.

Speaker Rebecca YuEvidence slide photo

Used byrequirement DEV-URL-03

  • Repeated by D2-C285 Day 2: Google recommends the History API to give single-page apps clean URLs instead of fragment-based routes.
  • Extended by D2-C288 Day 2: With the History API, a single-page app can use real links and attach event listeners that intercept the…
StageConsistent with docsD2-C186

Non-crawlable link markup, such as onclick links and hash pseudo-links, is common in single-page web apps, and sites that use faceted navigation should check their links for it.

Speaker Rebecca YuEvidence transcript

Used byrequirement DEV-URL-02glossary term Single-page app (SPA)

SlideConsistent with docsD2-C187

A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.

“The nav works perfectly for users, and the entire alternate-language cluster is orphaned.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo

Used byrequirement DEV-INT-06

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
AnalysisD2-C188

Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.

Author Ibrahim Anjro

Used byrequirement DEV-URL-08

  • Extends D1-C101 Day 1: Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable…
AnalysisD2-C189

Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.

Author Ibrahim Anjro

Things

Used byrequirement DEV-INT-06

  • Extends D1-C067 Day 1: A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts…
StageConfirmed by docsD2-C190

In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.

Speaker Rebecca YuEvidence transcript

Used byrequirement DEV-ERR-02glossary term Single-page app (SPA)

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Repeated by D2-C291 Day 2: In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200…
DocsSourceD2-C192

Google's JavaScript guides say that when client-side routing makes a real 404 status impractical, a single-page app can avoid soft 404s by redirecting with JavaScript to a URL whose server returns 404, or by adding a robots noindex meta tag with JavaScript.

Publisher Google Search Central

Used byrequirement DEV-ERR-02

SlideConfirmed by docsD2-C195

Content that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.

Speaker Rebecca YuEvidence slide photo, transcript

Used byrequirement DEV-REN-02

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
  • Repeated by D2-C268 Day 2: Content that loads only when a user clicks an element is not supported in the way Google renders pages for…
SlideConsistent with docsD2-C196

For rendering, what matters is whether content is present in the DOM, not whether it is visible on screen: hidden content can be indexed, absent content cannot.

“Not visible versus hidden. Present versus absent in the DOM.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo

Things

Used byrequirement DEV-REN-02

SlideConsistent with docsD2-C197

Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.

“Googlebot doesn't scroll. Never runs.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo, transcript

Used byrequirement DEV-REN-03glossary term Lazy loading

  • Extends D1-C114 Day 1: A Google panelist called pagination one of the trickiest things in web development and said switching to…
  • Repeated by D2-C271 Day 2: Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on…
SlideConsistent with docsD2-C199

An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders with is tall, so content lazy-loaded this way can load during rendering.

“Fires on viewport entry. Viewport is tall.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo

Used byrequirement DEV-REN-03

  • Extended by D2-C272 Day 2: Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.
  • Repeated by D2-C274 Day 2: Content loaded as elements enter the viewport, for example with an Intersection Observer, does load when…
DocsSourceD2-C200

In Google's March 2023 SEO office hours, John Mueller said Google handles infinite scroll with viewport expansion, rendering a page like a very long phone, which is not very efficient and can miss content, so pagination links are strongly recommended.

“this is done through a technique called "viewport expansion", where we render a page like a very long phone.”

Publisher Google Search Central (SEO office hours transcript, March 2023)

Things
SlideConfirmed by docsD2-C201

Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () => fetch('/api/specs'), does not exist for Google until someone clicks.

“Doesn't exist until someone clicks”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo

Used byrequirement DEV-REN-02

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
SlideConsistent with docsD2-C202

Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden attribute; Google indexes such hidden content.

“CSS-hidden is fine. Google indexes it.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo, transcript

Used byrequirement DEV-REN-02

  • Extended by D2-C866 Day 2: Content inside tabs, for example separate tabs for a product description and a manufacturer description…
SlideConfirmed by docsD2-C204

Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

“If the renderer can't fetch it, the renderer can't run it.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo, transcript

Used byrequirement DEV-REN-04

  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Repeated by D2-C265 Day 2: Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the…
  • Repeated by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
SlideConsistent with docsD2-C205

Disallowing a script folder such as /static/js/ or a generic /api/ folder in robots.txt can block the endpoints that supply a page's content, leaving blank modules and missing content after rendering.

Speaker Rebecca YuEvidence slide photo, transcript

DocsSourceD2-C206

Google says its Web Rendering Service fetches the resources a page references through Googlebot, including JavaScript, CSS and XHR requests to APIs, but not images or videos.

Publisher Search Central blog (3 December 2024), Search Central blog (31 March 2026)

Used byrequirement DEV-REN-04

SlideConsistent with docsD2-C207

If an API folder must stay blocked in robots.txt, allow the endpoints that rendering needs, for example Disallow: /api/ together with Allow: /api/products/.

“Carve out only what rendering needs.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo, transcript

Used byrequirement DEV-REN-04

AnalysisD2-C208

A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.

Author Ibrahim Anjro

  • Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
StageConfirmed by docsD2-C213

A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.

Speaker Rebecca YuEvidence transcript

Things

Used byrequirements DEV-ERR-03, DEV-REN-04

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extended by D2-C336 Day 2: Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty…
StageD2-C214

A community speaker summed up JavaScript rendering issues as cases where search engines cannot execute, access or trigger the code needed to display a page's core content.

“JavaScript rendering issues happen when search engines cannot execute, access or trigger the code required to display your core content.”

Speaker Rebecca YuEvidence transcript

StageConsistent with docsD2-C215

To find rendering issues, compare a page's raw source code with its rendered versions, including the rendered HTML that Google's testing tools show; the differences point to the problems.

Speaker Rebecca YuEvidence transcript

Things

Used byrequirement DEV-MON-02glossary term Raw HTML and rendered HTML

StageD2-C217

Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard, and that this is what the Web Fragments work is trying to do.

Speaker Natalia VendittoEvidence transcript

  • Repeats D2-C052 Day 2: John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet…
StageD2-C218

Natalia Venditto said teams are increasingly asked to integrate AI-generated content and applications into existing web applications and want to do so without breaking the host application, a situation she called a typical micro-frontend scenario.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C219

Natalia Venditto named three typical failure modes by which an embedded app can break the application hosting it: collisions in the global JavaScript scope and module registry, CSS bleed, and fate sharing.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C220

CSS bleed, as Natalia Venditto described it, means styles leak between an embedded app and its host, so that suddenly everything looks like the embedded app or the embedded app takes on the look of the host.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C221

Fate sharing means a failure in an embedded app spreads to its host, for example an unhandled exception in the embedded app that ends up breaking the host application.

Speaker Natalia VendittoEvidence transcript

AnalysisD2-C222

Pages that embed third-party or AI-generated apps can share their fate, since an unhandled exception in the embedded code can break the host page; check the rendered HTML and JavaScript console output of such pages in Google's testing tools to make sure the host's own content still renders.

Author Ibrahim Anjro

Used byrequirement DEV-REN-09

SlideD2-C223

A slide titled Orchestration approaches compared four ways to combine micro-frontends in one page (iframe, Shadow DOM, rewriting everything, and reframing with Web Fragments) on three criteria (JavaScript isolated, styles isolated, part of the page) and named the catch of each.

Speaker Natalia VendittoEvidence slide photo, transcript

SlideNot in docsD2-C224

An iframe, the long-established way to embed an app, isolates both JavaScript and styles but is not part of the page: it is walled off from the host's DOM, navigation and layout, which Natalia Venditto said brings many problems with accessibility, layout and navigation.

“Walled off from DOM, navigation, layout”

Wording checked against the slide or recording

Speaker Natalia VendittoEvidence slide photo, transcript

DocsSourceD2-C227

Google says its systems generally try to index the content of a page embedded with an iframe as part of the page that embeds it, but this is not guaranteed because both pages are also normal HTML pages on their own.

Publisher Google Search Central (SEO office hours transcript, December 2023), Search Central blog (21 January 2022)

Used byrequirement DEV-REN-10

SlideNot in docsD2-C228

Shadow DOM isolates styles and keeps the embedded content part of the page, but it does not isolate JavaScript: the embedded code still shares the host's JavaScript globals.

“Still shares JS globals”

Wording checked against the slide or recording

Speaker Natalia VendittoEvidence slide photo, transcript

Used byglossary term Shadow DOM

SlideNot in docsD2-C229

Rewriting everything into one application keeps the result part of the page, but JavaScript and style isolation then have to be done by hand.

Speaker Natalia VendittoEvidence slide photo, transcript

SlideD2-C230

The catch the comparison slide gave for rewriting everything was that you cannot rewrite code you did not write; Natalia Venditto explained that AI cannot rewrite code already in place because it does not know its requirements, design system or other constraints.

“Can't rewrite code you didn't write”

Wording checked against the slide or recording

Speaker Natalia VendittoEvidence slide photo, transcript

SlideNot in docsD2-C231

Reframing with Web Fragments was the only approach on the comparison slide ticked for all three criteria (JavaScript isolated, styles isolated, part of the page); its stated catch is that it relies on browser patches today.

“Needs patches today”

Wording checked against the slide or recording

Speaker Natalia VendittoEvidence slide photo, transcript

StageNot in docsD2-C232

Web Fragments is a lightweight library that puts a hidden iframe in the host page and uses it as an isolated sandbox to run the embedded app's JavaScript, not to render the app.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C233

Web Fragments fetches the embedded app's assets from a remote endpoint through middleware that acts as a gateway.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C234

Web Fragments reframes the embedded app's DOM by placing it inside a shadow DOM in the host page, while the app's JavaScript runs in the hidden iframe.

Speaker Natalia VendittoEvidence transcript

DocsSourceD2-C235

Google's documentation says Google supports web components and flattens shadow DOM and light DOM content when it renders a page, and that content not visible in the rendered HTML cannot be indexed.

Publisher Google Search Central

Used byrequirement DEV-REN-10glossary term Shadow DOM

StageNot in docsD2-C236

With Web Fragments the page ends up as a single document in which the host application does not know the embedded app runs inside it and the embedded app does not know it lives in another document.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C237

Web Fragments monkey-patches browser APIs to containerize the browser, recreating inside it an architecture similar to Docker containers on the back end.

“we monkey patch the browser to containerize it”

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C238

The browser APIs Web Fragments patches include document, history and location; Natalia Venditto said they are virtualized rather than hijacked, so controls work exactly as in a normal application.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C239

Web Fragments retargets dispatchEvent calls to the fragment's shadow root and resolves DOM calls such as appendChild and element lookups against the main document.

Speaker Natalia VendittoEvidence transcript

AnalysisD2-C240

Because Web Fragments monkey-patches core browser APIs such as document, history and location, test pages that use it in Google's rendering (the rendered HTML in Search Console's URL Inspection tool or the Rich Results Test), not only in a normal desktop browser.

Author Ibrahim Anjro

Used byrequirement DEV-REN-10

StageNot in docsD2-C241

Natalia Venditto said Web Fragments suits AI-generated apps because a web fragment used as a custom element sandboxes the app's rendering, avoiding collisions and CSS bleed between the app and its host in either direction.

Speaker Natalia VendittoEvidence transcript

SlideNot in docsD2-C242

A slide set out Web Fragments in five steps: an LLM writes the app and adds a custom element such as <web-fragment fragment-id="some-id"> to the HTML, a standalone HTTP endpoint is set up, the library is imported and initialized with initializeWebFragments(), the fragment is registered in the gateway, and step 5 is a celebration emoji.

Speaker Natalia VendittoEvidence slide photo, transcript

StageNot in docsD2-C243

Natalia Venditto said the five Web Fragments steps are all that is needed to run a containerized application fully on the client side.

Speaker Natalia VendittoEvidence transcript

SlideNot in docsD2-C244

A web fragment is served from its own standalone HTTP endpoint, so the embedded app can be deployed anywhere, separately from the host page.

Speaker Natalia VendittoEvidence slide photo, transcript

AnalysisD2-C245

Make sure robots.txt does not block the URLs from which the browser loads a fragment's assets (the gateway paths on the host page's origin, or the endpoint's own host if assets load from there), because Google does not render JavaScript from blocked files and robots.txt rules apply per host.

Author Ibrahim Anjro

StageNot in docsD2-C246

Natalia Venditto said Web Fragments is framework-agnostic and vendor-agnostic and works the same whatever JavaScript framework the embedded app uses.

Speaker Natalia VendittoEvidence transcript

AnalysisD2-C250

Because Google flattens shadow DOM when it renders, content that Web Fragments places in a shadow root can in principle be indexed with the host page, but only if it appears in the rendered HTML; check that in URL Inspection rather than relying on a library's promise of indexability.

Author Ibrahim Anjro

Used byrequirement DEV-REN-10

StageD2-C857

None of the rendering problems a community speaker showed from established brands' sites had been fixed promptly: some were fixed by the time of the talk and some were still live.

Speaker Sören BendigEvidence transcript

Things
StageD2-C859

A rendering failure such as a Content Security Policy blocking a video on a landing page that explains how to open a business account could have a drastic impact for a purely online business, a community speaker warned.

Speaker Sören BendigEvidence transcript

Things
  • Extends D2-C156 Day 2: Content Security Policy is often misunderstood: on a page of a large German banking group, a video that…
StageNot in docsD2-C860

Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a modern crawler.

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-MON-05

  • Extends D2-C162 Day 2: A community speaker strongly advised SEOs to use Chrome DevTools more often to understand what happens when…

11:15 · What is Google friendly JavaScript 55

StageConfirmed by docsD2-C256

Google renders nearly all of the web by replicating what a browser does, using a real browser's rendering engine.

“Google renders nearly all of the internet by replicating browser actions”

Speaker Erin SparlingEvidence transcript

Things
  • Extended by D3-C628 Day 3: Gary Illyes said Google keeps saying it renders every URL on the internet and that he would say this is not…
StageConsistent with docsD2-C259

Erin Sparling named server-side or hybrid rendering, fallbacks and not leaving placeholders in the DOM, among other measures, as ways to guard against failed JavaScript rendering.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-REN-06

  • Extends D2-C149 Day 2: A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal…
SlideConfirmed by docsD2-C262

Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.

“If it's not in the final DOM, Google can't see it.”

Wording checked against the slide or recording

Speaker Erin SparlingEvidence slide photo, transcript

Things

Used byrequirement DEV-REN-02

  • Repeats D2-C174 Day 2: Whatever is in the DOM at the moment Google's rendering finishes is what likely gets indexed.
StageNot in docsD2-C263

Content is missing from the rendered HTML either because the server does not serve it or because the content has still not appeared after some period of time.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-REN-06

  • Extends D2-C149 Day 2: A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal…
SlideConfirmed by docsD2-C264

Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but returns a 200 HTTP status code.

“Your application serves a "Not Found" message but returns a 200 HTTP status code.”

Wording checked against the slide or recording

Speaker Erin SparlingEvidence slide photo, transcript

Used byrequirement DEV-ERR-02

  • Repeats D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
SlideConsistent with docsD2-C265

Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the crawling of critical JavaScript files or API endpoints.

“robots.txt disallowing crawling of critical .js or API endpoints.”

Wording checked against the slide or recording

Speaker Erin SparlingEvidence slide photo, transcript

Used byrequirement DEV-REN-04

  • Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
SlideConsistent with docsD2-C266

Google named three causes of content missing from the rendered HTML: JavaScript inaccessible to Googlebot, JavaScript DOM event triggers and disabled browser APIs.

Speaker Erin SparlingEvidence slide photo, video, transcript

StageConfirmed by docsD2-C268

Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-REN-02glossary term Lazy loading

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
  • Repeats D2-C195 Day 2: Content that loads only after a user action such as a click or a scroll is not in the DOM while Google…
SlideD2-C270

Google said there is no single fix for content missing from the rendered HTML: the solution varies by case, and the advice is to follow best practices.

“Solution: varies by case, but follow best practices.”

Wording checked against the slide or recording

Speaker Erin SparlingEvidence slide photo, transcript

StageConfirmed by docsD2-C271

Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.

Speaker Erin SparlingEvidence slide photo, transcript

Used byrequirements DEV-REN-03, DEV-URL-07

  • Extends D1-C114 Day 1: A Google panelist called pagination one of the trickiest things in web development and said switching to…
  • Repeats D2-C197 Day 2: Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll'…
  • Extends D1-C491 Day 1: If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them…
StageNot in docsD2-C272

Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.

“it renders it in a very tall viewport. Specifically, around 10,000 pixels is what the viewport gets rendered as.”

Speaker Erin SparlingEvidence slide photo, transcript

Used byrequirements DEV-REN-03, DEV-URL-07glossary term Viewport expansion

  • Extends D2-C199 Day 2: An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders…
DocsSourceD2-C273

Google's March 2023 SEO office hours say Google sees infinite-scroll content through viewport expansion, rendering a page like a very long phone, which is not particularly efficient and can miss infinite content.

“a technique called "viewport expansion", where we render a page like a very long phone”

Publisher Google Search Central (SEO office hours transcript, March 2023)

Things

Used byrequirement DEV-URL-07glossary term Viewport expansion

StageConfirmed by docsD2-C274

Content loaded as elements enter the viewport, for example with an Intersection Observer, does load when Google renders a page, because Google's rendering viewport is very tall.

Speaker Erin SparlingEvidence transcript

Things

Used byrequirement DEV-REN-03

  • Repeats D2-C199 Day 2: An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders…
DocsSourceD2-C276

To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each chunk its own persistent, unique URL, link sequentially to those URLs, and update the displayed URL with the History API when a new chunk becomes the main visible element.

Publisher Google Search Central

Used byrequirement DEV-URL-07

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
DocsSourceD2-C280

Google's JavaScript SEO basics recommends differential serving and polyfills when feature detection finds a missing browser API, and warns that some browser features cannot be polyfilled.

“We recommend using differential serving and polyfills if you feature-detect a missing browser API that you need.”

Publisher Google Search Central

Used byrequirement DEV-REN-08

SlideConfirmed by docsD2-C284

URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.

“Fragment identifiers (#) are often ignored by crawlers.”

Wording checked against the slide or recording

Speaker Erin SparlingEvidence slide photo, transcript

Used byrequirement DEV-URL-03

  • Repeats D1-C101 Day 1: Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable…
SlideConfirmed by docsD2-C285

Google recommends the History API to give single-page apps clean URLs instead of fragment-based routes.

“Use the History API for clean URLs in SPAs.”

Wording checked against the slide or recording

Speaker Erin SparlingEvidence slide photo, transcript

Used byrequirement DEV-URL-03glossary term History API

  • Repeats D2-C185 Day 2: The crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>…
StageConfirmed by docsD2-C287

A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-URL-01

  • Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
  • Extends D2-C040 Day 2: Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not…
StageConfirmed by docsD2-C288

With the History API, a single-page app can use real links and attach event listeners that intercept the click, rewrite the URL with pushState and load the new content, so Google can follow the links while users avoid full page reloads.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-URL-03glossary terms History API, Single-page app (SPA)

  • Extends D2-C185 Day 2: The crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>…
StageNot in docsD2-C289

Real URLs also work as deep links: Android, iOS and desktop operating systems accept full URLs as keys to specific content in an app, so clean URLs simplify the cross-platform user experience, not only indexing.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-URL-03

StageConfirmed by docsD2-C291

In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-ERR-02

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Repeats D2-C190 Day 2: In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because…
StageConsistent with docsD2-C292

The fix for a client-side soft 404 is to serve a real 404 error page where appropriate; how to detect URLs that do not exist depends on how the app works and where it is hosted.

Speaker Erin SparlingEvidence transcript

  • Extends D2-C191 Day 2: The fix for a client-side soft 404 is to make a missing page end with a real HTTP 404 status code instead of…
DocsSourceD2-C294

For client-side rendered single-page apps, where meaningful status codes can be impossible or impractical, Google's documentation gives two ways to avoid soft 404s: a JavaScript redirect to a URL that returns a 404 status, or a noindex robots meta tag added with JavaScript.

Publisher Google Search Central

Used byrequirement DEV-ERR-02

AnalysisD2-C295

After moving to real paths, set the server or hosting rewrite rules so a direct request to every valid path returns 200 with the content (ideally server-rendered) and an unknown path returns a 404 status, which removes client-side soft 404s at the source.

Author Ibrahim Anjro

Used byrequirement DEV-ERR-02

StageConfirmed by docsD2-C296

If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-REN-04

  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
StageNot in docsD2-C299

A headless content management system serves its content as an API; even WordPress, often seen as one monolithic application, can be used headless, with only its editor or only its renderer.

Speaker Erin SparlingEvidence transcript

StageD2-C300

Erin Sparling's practice with a headless CMS is to give an AI agent the content's JSON Schema and have it build a throwaway prototype front end, to see the content before deciding how to render it; the prototype is not meant for launch.

Speaker Erin SparlingEvidence transcript

StageD2-C302

When one content schema feeds several front ends (the example was a news UI, then a recipe UI), Erin Sparling had an AI agent add automated browser UI tests and agent controls so the interfaces could be managed and tested.

Speaker Erin SparlingEvidence transcript

StageConsistent with docsD2-C306

To use Search Console to debug what Google can see on a site and why, Erin Sparling said the first step is to get access to the site by verifying ownership (the speaker's words were authorized domains).

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-MON-01

DocsSourceD2-C829

Google's guide to fixing Search-related JavaScript problems says the Web Rendering Service may ignore caching headers and so use outdated JavaScript or CSS, and recommends content fingerprinting, which puts a hash of the content in the file name, as in main.2bb85551.js.

Publisher Google Search Central

Used byrequirement DEV-REN-06

11:30 · Understanding what's on a page 48

SlideConsistent with docsD2-C309

A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

Speaker Gary IllyesEvidence video, transcript

Used byrequirement DEV-HTM-01

  • Repeats D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
  • Extended by D2-C861 Day 2: Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not…
DocsSourceD2-C310

Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.

“When Google indexes a page, it determines the primary content (or centerpiece) of each page.”

Publisher Google Search Central

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

StageConsistent with docsD2-C311

Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C170 Day 3: Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content…
StageNot in docsD2-C313

Words in the footer of a page get a lower weight, so text placed in the footer is unlikely to contribute much to ranking the page.

“if you put something in a footer, it's more likely that it's not going to contribute much to ranking”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-02

SlideConsistent with docsD2-C314

On Google's example blog page, the post title and opening sentence counted as important because they sit in the main content, in front of the user, while the site tagline, the 'Categories' sidebar and category links such as 'Hugo (7)' counted as less important supplementary text.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-HTM-02

StageConsistent with docsD2-C315

To make a word count for ranking a page, Gary Illyes said the simplest step is to move it into the main content, because where text sits on a page already contributes quite a bit to ranking.

“where you position text on a page will already contribute quite a bit to ranking that page”

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-HTM-02

StageNot in docsD2-C316

Gary Illyes said not everything on a page can or should be important: if everything were placed in the main content, nothing would stand out as main content, which he called working as intended.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C318

Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

Speaker Gary IllyesEvidence slide photo, transcript

Used byglossary term Tokenization

  • Repeated by D2-C721 Day 2: Google's Search index does not hold the full content of pages; Google said storing full pages and pulling…
  • Extended by D3-C109 Day 3: The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens…
StageNot in docsD2-C320

Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
  • Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
StageNot in docsD2-C321

For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

Speaker Gary IllyesEvidence transcript

  • Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageD2-C322

Gary Illyes said a colleague, John, would cover how Google interprets the words of a query on the morning of Day 3.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C323

When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-HTM-07

  • Repeated by D2-C720 Day 2: Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the…
StageNot in docsD2-C325

Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

Speaker Gary IllyesEvidence slide photo, transcript

  • Contradicts D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
SlideConsistent with docsD2-C326

In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.

Speaker Gary IllyesEvidence slide photo, transcript

Used byglossary term Tokenization

StageConsistent with docsD2-C327

Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.

Speaker Gary IllyesEvidence slide photo, transcript

Used byglossary term Tokenization

SlideNot in docsD2-C328

Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.

Speaker Gary IllyesEvidence 2 slide photos

AnalysisD2-C329

Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.

Author Ibrahim Anjro

StageConsistent with docsD2-C330

Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real, but it matters at the level of an AI model's context window.

Speaker Gary IllyesEvidence transcript

Used bymyth M-002

  • Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
StageConsistent with docsD2-C331

Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.

Speaker Gary IllyesEvidence transcript

Things

Used bymyth M-002

  • Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
  • Extended by D2-C869 Day 2: Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps…
StageConsistent with docsD2-C332

Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-03myth M-002

  • Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
  • Extended by D2-C825 Day 2: Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window…
StageNot in docsD2-C825

Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-03

  • Extends D2-C332 Day 2: Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller…
AnalysisD2-C333

Do not rewrite pages into short, self-contained chunks for AI systems; Google says Gemini reads context windows of millions of tokens, so structure content for readers, with clear headings and complete explanations.

Author Ibrahim Anjro

Things

Used byrequirement DEV-AIF-03

StageConfirmed by docsD2-C334

A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

Speaker Gary IllyesEvidence 2 slide photos, transcript

Used byrequirements DEV-ERR-01, DEV-ERR-03glossary term Soft 404

  • Repeats D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D1-C070 Day 1: Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the…
  • Repeats D1-C355 Day 1: A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found'…
  • Extended by D2-C903 Day 2: A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that…
  • Extended by D2-C700 Day 2: Index selection drops soft 404 pages that were not dropped earlier, for example when a document is…
StageNot in docsD2-C335

Because error pages are worded in endless variations, of which 'page not found' is only the classic one, Google cannot detect soft 404s with simple error, word or keyword matching.

Speaker Gary IllyesEvidence transcript

Things
SlideConfirmed by docsD2-C336

Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.

Speaker Gary IllyesEvidence slide photo, transcript

Things

Used byrequirement DEV-ERR-03

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D2-C213 Day 2: A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as…
SlideNot in docsD2-C337

Mistakes in Google's own systems are a further cause of soft 404s, and Gary Illyes asked site owners to report such mistakes in Google's forums.

“BONUS: Mistakes in Google's systems (that you should notify us about)”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

Things
StageNot in docsD2-C338

Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

“This is basically an LLM thing, something like BERT, that is specifically trained to understand page structure”

Speaker Gary IllyesEvidence transcript

Things
  • Extends D1-C042 Day 1: BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at…
StageNot in docsD2-C339

For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-ERR-03

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
StageConsistent with docsD2-C341

Gary Illyes said soft 404 detection keeps searchers from clicking into dead ends and avoids wasting site owners' resources on visitors sent to error pages.

Speaker Gary IllyesEvidence transcript

Things
StageConsistent with docsD2-C861

Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
  • Extends D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
StageConfirmed by docsD2-C862

Gary Illyes defined a page's main content as any part of the page that directly helps the page achieve its purpose, what it was built for.

“Main content is any part of the page that directly helps the page achieve its purpose”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConfirmed by docsD2-C864

Content created by other users can be main content: on a user-generated content site, the user-generated content can be the page's main content.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConsistent with docsD2-C865

A comment section below a blog post can still be part of the page's main content and can contribute to Google's understanding of the page.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConfirmed by docsD2-C866

Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-REN-02

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
  • Extends D2-C202 Day 2: Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden…
StageConsistent with docsD2-C868

Gary Illyes said the main content is what Google considers when ranking a page.

“It's the main content that we consider for ranking.”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
StageConsistent with docsD2-C869

Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.

“the context window is perhaps 900,000 or even closer to a million big”

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-03

  • Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
  • Extends D2-C331 Day 2: Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
DocsSourceD2-C870

Google's Search Quality Rater Guidelines define main content as any part of the page that directly helps it achieve its purpose, including text, images, videos, page features such as calculators and content created by users, and they count the title at the top of the page as part of it.

“Main Content is any part of the page that directly helps the page achieve its purpose.”

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
DocsSourceD2-C871

Google's Search Quality Rater Guidelines say navigation links are a common type of supplementary content, and that content behind tabs and user reviews or comments may count as main content on some pages and as supplementary content on others, depending on the page's purpose.

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
AnalysisD2-C872

The talk gave Gemini's context window both as millions of tokens and as roughly 900,000 to a million; Google's long-context docs say Gemini models have context windows of 1 million or more tokens (about eight average novels per million), so plan with about one million tokens as the documented floor rather than several million.

Author Ibrahim Anjro

Things

11:55 · Handling web duplication 58

StageConsistent with docsD2-C344

Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.

Speaker John MuellerEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeated by D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
SlideConfirmed by docsD2-C345

Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs and index the unique pages, and forward signals to the representative URLs.

Speaker John MuellerEvidence slide photo, transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
  • Extended by D2-C701 Day 2: When Google already has duplicate information for a document, for example when reprocessing it, index…
StageConsistent with docsD2-C346

For deduplication, a cluster is a set of pages Google considers essentially equivalent: Google stores one of them in the index and keeps track of the other related URLs.

Speaker John MuellerEvidence transcript

Used byglossary term Duplicate cluster

  • Extends D1-C207 Day 1: Google clusters duplicate pages and selects one page per cluster as its representative, the canonical URL, so…
SlideConfirmed by docsD2-C348

Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

Speaker John MuellerEvidence slide photo, transcript

Used byglossary term Canonical

  • Extends D1-C111 Day 1: Gary Illyes said there is no such thing as a duplicate content penalty.
  • Contradicted by D2-C429 Day 2: A community speaker said pages carry different link equity, and a canonical leader that is not the strongest…
StageConsistent with docsD2-C351

Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-09

  • Extended by D3-C642 Day 3: Google treats a site move as a complex canonicalization process in which every signal of the old site is…
SlideConfirmed by docsD2-C352

Google keeps the other URLs of a duplicate cluster as 'alternate names': equivalent URLs with the same content that Google still tracks as alternate versions of the representative URL.

Speaker John MuellerEvidence slide photo, transcript

Used byglossary term Alternate names

SlideConsistent with docsD2-C354

Alternate names are why a site: query for an old domain still shows the old domain's URLs after a site migration, which site owners often misread as a migration that is not working.

“FYI "alternate names" is why you see old domains in site:-queries”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-09

SlideConsistent with docsD2-C356

After a rebrand that changes the domain (the slide's example was johns-bikes to slow-bikes), Google can still show the old domain to people who search for the old brand by name, in navigational and branded queries.

Speaker John MuellerEvidence slide photo, transcript

DocsSourceD2-C357

Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.

“This is normal and as users get used to the new domain name, the alternate names will fade away without you doing anything.”

Publisher Google Search Central

Used byrequirements DEV-CAN-01, DEV-CAN-09glossary term Alternate names

StageConsistent with docsD2-C361

Google trusts redirects very much for clustering, because a redirect is a clear sign that there is one version of the content; Google keeps track of both URLs but stores only one copy of the content.

Speaker John MuellerEvidence transcript

Things

Used byrequirement DEV-CAN-01

StageNot in docsD2-C364

Google clusters duplicate pages by content in four ways: exact matches, near matches, structurally similar content, and soft 404s.

Speaker John MuellerEvidence transcript

Things
StageConsistent with docsD2-C365

Exact-match duplicates, such as the www and non-www versions of the same page, are clustered and Google keeps only one of them.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-02

SlideConfirmed by docsD2-C366

Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.

“When main content is the same, pages may be clustered.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-INT-07

StageConsistent with docsD2-C367

Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

Speaker John MuellerEvidence transcript

Things

Used byrequirement DEV-ERR-01

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
SlideNot in docsD2-C368

Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideNot in docsD2-C369

Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

“Do we even need to crawl /buy/seo-service ?”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo

Used byrequirements DEV-CAN-08, DEV-URL-09

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
SlideNot in docsD2-C370

City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-08

SlideConsistent with docsD2-C371

To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.

“Misleading site structure (use clear signals!)”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

DocsSourceD2-C373

Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.

Publisher Google Search Central

Used byrequirement DEV-CAN-08

StageConsistent with docsD2-C374

A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

Speaker John MuellerEvidence transcript

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
StageConsistent with docsD2-C375

Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

Speaker John MuellerEvidence transcript

Things

Used byrequirements DEV-SRV-01, DEV-SRV-02

  • Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
DocsSourceD2-C376

Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

Publisher Search Central blog (24 December 2024)

Things

Used byrequirements DEV-ERR-03, DEV-SRV-02

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
StageNot in docsD2-C377

AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-AIF-04

  • Extends D1-C436 Day 1: A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to…
AnalysisD2-C378

Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.

Author Ibrahim Anjro

Used byrequirements DEV-AIF-04, DEV-SRV-02

StageConsistent with docsD2-C379

rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-03

  • Extends D2-C031 Day 2: Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and…
StageConfirmed by docsD2-C380

Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google has to make its own judgment about the canonical.

Speaker John MuellerEvidence transcript

Used byglossary term rel=canonical

  • Extends D1-C113 Day 1: The real cost of duplicates is control and crawling: Google may choose a canonical you did not want, and…
  • Repeated by D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
SlideConsistent with docsD2-C381

Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-INT-08

SlideConfirmed by docsD2-C382

When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

“We try to use hreflang alternates.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Things

Used byrequirement DEV-INT-08

  • Extends D2-C033 Day 2: Google extracts hreflang annotations, through which site owners specify the language variants of their…
DocsSourceD2-C385

Google's canonical guide says that for canonicalization Google prefers URLs that are part of hreflang clusters: if German pages for Germany and Switzerland point to each other with hreflang but not to the Austrian page, the German and Swiss pages are preferred as canonicals.

Publisher Google Search Central

Used byrequirement DEV-INT-08

StageNot in docsD2-C386

Google picks the canonical from a variety of criteria and uses some kind of machine learning to decide how much weight each criterion gets; the weighting changes from time to time.

“we use some kind of machine learning to understand how strong these criteria should be. And this changes from time to time.”

Speaker John MuellerEvidence transcript

SlideConsistent with docsD2-C387

Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-01, DEV-CAN-02, DEV-CAN-07

SlideConsistent with docsD2-C388

Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-07

SlideConsistent with docsD2-C389

Whether a page can load is really important for canonical selection: a broken certificate, failing JavaScript or a page that cannot be loaded counts against a URL, and the slide also listed meta refresh and security.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-02

SlideConsistent with docsD2-C390

Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-05, DEV-URL-05

DocsSourceD2-C391

Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.

Publisher Google Search Central

Used byrequirements DEV-CAN-05, DEV-URL-05glossary term rel=canonical

AnalysisD2-C392

Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.

Author Ibrahim Anjro

StageNot in docsD2-C393

Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-URL-06

  • Contradicts D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
AnalysisD2-C394

Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.

Author Ibrahim Anjro

Used byrequirement DEV-URL-06

DocsSourceD2-C395

Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

Publisher Search Central blog (8 April 2013)

Used byrequirement DEV-URL-06

  • Extends D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
SlideConsistent with docsD2-C396

Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-02

  • Extended by D2-C882 Day 2: In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs…
SlideNot in docsD2-C397

Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.

“Don't block agents.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-AIF-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
StageConsistent with docsD2-C399

When all canonical signals point to the same URL, Google follows what the site owner says; when they point in different directions, Google cannot tell what the owner wants.

“if there are multiple things in different directions, we don't know what to do.”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-CAN-05

12:05 · Lightning session E: Managing Duplicates and Site Moves 82

StageConsistent with docsD2-C402

According to the canonical link specification cited by a community speaker, when a canonical tag is declared improperly, the application that processes it may apply its own heuristic and decide what to do instead.

Speaker Tobias SchwarzEvidence transcript

  • Repeats D2-C380 Day 2: Many SEOs assume Google must follow their rel=canonical, but because people sometimes get it wrong, Google…
  • Extended by D2-C873 Day 2: A community speaker said that, under the canonical link specification, an improperly declared canonical tag…
AnalysisD2-C404

A broken canonical setup does not fail visibly: pages still load normally while the choice of URL passes to the search engine's own heuristics, so canonical tags need a regular audit with a crawler and the URL Inspection tool.

Author Ibrahim Anjro

Used byrequirement DEV-MON-06

StageD2-C406

Checking canonicals pairwise (page A has a canonical to page B, page B a self-referencing canonical) is common but, in a community speaker's experience, misses the bigger picture, such as other links pointing to A or B.

Speaker Tobias SchwarzEvidence transcript

StageD2-C408

A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-MON-06

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D2-C398 Day 2: Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure…
SlideD2-C410

A crawler report shown on screen drew one canonical group of five URLs as a link graph with the leader marked by a crown, and listed each member's HTTP status, indexability, document language, leader flag, incoming and outgoing links, and hints.

Speaker Tobias SchwarzEvidence slide photo, transcript

SlideD2-C411

The crawler report shown on screen ran group-level checks on each canonical group: an overall group status, whether the leader has self links, whether languages are consistent across members, and whether the group contains a chain.

Speaker Tobias SchwarzEvidence slide photo

SlideConsistent with docsD2-C415

The crawler report shown on screen flagged a URL discovered only through a canonical tag, with no internal a href links leading to it, as a phantom document outside the visible site structure.

“a phantom document not part of the visible site structure”

Wording checked against the slide or recording

Speaker Tobias SchwarzEvidence slide photo

StageConsistent with docsD2-C417

Multiple canonical declarations on one page conflict and are invalid, so the search engine applies its own heuristic and picks the canonical for the site owner; a page should declare only one canonical.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-CAN-03

DocsSourceD2-C418

Google's 2013 Search Central blog post on rel=canonical mistakes says to specify no more than one rel=canonical per page: when a page has more than one, Google will likely ignore all of them, and any benefit of a legitimate canonical is lost.

Publisher Search Central blog (8 April 2013)

Used byrequirement DEV-CAN-03

StageConsistent with docsD2-C419

Canonical chains, where a page's canonical target leads on to yet another URL, are named as improper use in the canonical link specification, according to a community speaker.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-CAN-04glossary term Canonical chain and canonical loop

DocsSourceD2-C420

Google's 2009 Search Central blog post introducing rel=canonical says Google's algorithm is lenient and can follow canonical chains, but strongly recommends updating links to point to a single canonical page for optimal canonicalization.

Publisher Search Central blog (12 February 2009)

Used byrequirement DEV-CAN-04glossary term Canonical chain and canonical loop

SlideConsistent with docsD2-C422

The crawler report shown on screen flagged a canonical chain as a problem: canonical links that form a chain instead of all pointing directly to the group's canonical leader.

“Canonical links form a chain rather than all pointing directly to the leader.”

Wording checked against the slide or recording

Speaker Tobias SchwarzEvidence slide photo

StageNot in docsD2-C426

A community speaker said that at best a search engine's heuristic would treat a canonical loop through redirects as self-referencing canonicals, and doubted that this is often done when the loop runs through a client-side redirect.

“But regarding the client-side redirect, I highly doubt that this is often done.”

Speaker Tobias SchwarzEvidence transcript

StageConsistent with docsD2-C427

When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.

“you're telling the search engine to index something that a human can't reach”

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-CAN-04

  • Extends D1-C067 Day 1: A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts…
StageConsistent with docsD2-C428

For a canonical leader that users cannot reach through links, a community speaker suggested revisiting the canonical graph and probably making the internally linked page the leader instead.

Speaker Tobias SchwarzEvidence transcript

StageNot in docsD2-C429

A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.

Speaker Tobias SchwarzEvidence transcript

  • Contradicts D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…
DocsSourceD2-C430

Google's guide to specifying canonical URLs says a canonical helps consolidate signals for duplicate pages: links to a duplicate URL are consolidated with links to the preferred URL once the preferred URL becomes canonical.

Publisher Google Search Central

Used byrequirement DEV-CAN-05

AnalysisD2-C431

A leader reachable only through a canonical and a leader with weak link equity share one fix: point internal links, sitemap entries and redirects at the URL chosen as canonical, so the leader is both reachable for users and the strongest page in its group.

Author Ibrahim Anjro

Used byrequirement DEV-CAN-05

StageConsistent with docsD2-C432

The canonical leader of a group should be indexable: it should carry no noindex robots directive, return no error status code and not be blocked in robots.txt.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-CAN-04

StageConsistent with docsD2-C433

Canonical links between different language versions of a page are a misuse: if a search engine accepted such a canonical group, only one language version would rank.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-CAN-06

StageConfirmed by docsD2-C434

Language versions of a page should be connected with hreflang annotations instead of canonical links.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-CAN-06

DocsSourceD2-C435

Google's guide to specifying canonical URLs says that on pages using hreflang, the canonical should be a page in the same language, or the best possible substitute language if no canonical page exists in the same language.

Publisher Google Search Central

Used byrequirement DEV-CAN-06

AnalysisD2-C436

Check that every canonical group contains a single document language: a group spanning languages means canonicals are joining versions that hreflang should connect, and each language version should keep its own self-referencing canonical.

Author Ibrahim Anjro

Used byrequirement DEV-CAN-06

AnalysisD2-C439

The crawler report shown at the talk rated a redirecting canonical target as an error, while Google's 2009 guidance accepts a canonical URL that redirects and says Google can follow canonical chains; the gap is one of severity, since both recommend pointing every canonical link straight at the final URL.

Author Ibrahim Anjro

AnalysisD2-C440

Google says canonicalization consolidates link signals from duplicates into the chosen canonical, so a leader with few links of its own is not necessarily weaker once Google accepts it; the practical risk of a weakly linked leader is that Google, treating rel=canonical as a hint, picks a different page as canonical.

Author Ibrahim Anjro

StageConfirmed by docsD2-C873

A community speaker said that, under the canonical link specification, an improperly declared canonical tag can also be ignored completely by the application that processes it, not only replaced by that application's own heuristic.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-CAN-03

  • Extends D2-C402 Day 2: According to the canonical link specification cited by a community speaker, when a canonical tag is declared…
StageD2-C874

In a community case study, two of the first loan-comparison websites in Poland did exactly the same thing, so they competed with each other for the same users, the same keywords and the same space in search results.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C875

A community speaker said running two competing sites in one market meant keeping up twice with changes in how people search and in what Google values, instead of putting all resources behind one site.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C876

In a community case study, an acquisition by a group of loan-comparison brands (operating in eleven markets) prompted the merger of two competing portals into one brand, which the speaker framed as creating order rather than only a technical migration.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C877

A community speaker set three goals before any technical work on a two-site consolidation: lower costs (infrastructure, maintenance and content paid for only once), one strong brand focused on one specialization, and a simple structure instead of hundreds of similar URLs competing with each other.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C878

Before choosing which of two domains to keep in a consolidation, a community speaker's team checked both domains for past penalties and compared their traffic, conversion and revenue.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C880

A community speaker's test for keeping a page during a consolidation was whether it had real potential to make money and really fitted what the business does; pages that failed were removed.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C881

In a community case study, the kept pages that shared the same search intent were merged into one strong article each, taking the merged site from over 2,000 URLs to about 100, a cut of about 95% of URLs.

Speaker Martyna AğanoğluEvidence transcript

StageConsistent with docsD2-C882

In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs with no same-intent match were removed with an error status instead of being redirected (the exact code is unclear in the recording).

Speaker Martyna AğanoğluEvidence transcript

Things

Used byrequirement DEV-CAN-10

  • Extends D2-C396 Day 2: Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't…
StageConfirmed by docsD2-C884

The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.

Speaker Martyna AğanoğluEvidence transcript

Used byrequirement DEV-CAN-09glossary term Change of Address tool

  • Extends D1-C540 Day 1: Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is…
StageConsistent with docsD2-C886

A community speaker listed the technical SEO goals of a two-site consolidation as lighter pages, faster loading and no crawl budget spent on content that no longer mattered.

Speaker Martyna AğanoğluEvidence transcript

  • Extends D1-C103 Day 1: Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers'…
StageD2-C887

As part of a consolidation, a community speaker's team introduced the real people behind the site, financial experts visibly responsible for what it publishes, which the speaker counted as important as the technical work.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C889

A community speaker reported that after consolidating two loan-comparison sites into one domain, traffic grew from the first day, with no drop or waiting period.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C891

A community speaker credited a consolidation's success to doing everything at once: one clear topic, visible expertise instead of an anonymous content factory, and crawling spent only on pages that matter, growing through quality rather than more content.

Speaker Martyna AğanoğluEvidence transcript

StageD2-C892

A community talk on a still unfinished migration, after a client bought a company whose site had to fit into one product line, presented the decisions behind the migration rather than its results.

Speaker David Carrasco PamiesEvidence transcript

StageD2-C893

A community speaker argued that migrations are hard because of people rather than technology: two companies, two teams, a dozen stakeholders and a board that set a deadline before anyone opened the CMS.

“I think there are no complex migrations. There are complex people.”

Speaker David Carrasco PamiesEvidence transcript

StageD2-C894

A community speaker named four kinds of post-acquisition site migration: absorb (one site integrates into the other), keep apart, merge (two or three sites into a new one), and plug in (a full site becomes a section of the other site).

Speaker David Carrasco PamiesEvidence transcript

StageD2-C898

A community speaker said most migration failures look technical but trace back to a decision nobody made; in the migration presented, more than 7,000 URLs still had no approved destination at the time of the talk.

Speaker David Carrasco PamiesEvidence transcript

StageD2-C899

A community speaker recorded each migration decision per URL: old URL, new URL, what must survive (for example a product answer and a demo request), who approves, and the test that proves it works.

Speaker David Carrasco PamiesEvidence transcript

Used byglossary term Redirect map

StageD2-C900

A community speaker advised asking other teams during a migration about everything a crawler cannot reveal, such as sales workflows or legal requirements.

Speaker David Carrasco PamiesEvidence transcript

StageConsistent with docsD2-C903

A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.

Speaker David Carrasco PamiesEvidence transcript

Used byrequirement DEV-CAN-10

  • Extends D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
StageD2-C904

A community speaker's three post-migration checks per URL: the route (the old URL reaches its destination), the answer (the approved content is still there) and the lead (a demo request still reaches the CRM); automate them, but have people check critical journeys.

Speaker David Carrasco PamiesEvidence transcript

StageD2-C908

A community speaker advised telling users of a migrated brand that they are in the right place, through emails, a press release and a blog post, to keep the trust built with the old brand.

Speaker David Carrasco PamiesEvidence transcript

StageD2-C909

A community speaker said reputation travels with a migrated product: a good product takes its reviews and links along, and a product nobody likes takes its complaints along.

Speaker David Carrasco PamiesEvidence transcript

StageD2-C910

A community speaker summed up migration planning in three questions: what survives, who approves, and what test shows the migration actually works.

Speaker David Carrasco PamiesEvidence transcript

DocsSourceD2-C911

Google's site move guide says not to redirect many old URLs to one irrelevant destination such as the new site's home page, which can confuse users and might be treated as a soft 404, and to return a 404 or 410 for deleted or merged content that is not moved to the new site.

Publisher Google Search Central

Used byrequirement DEV-CAN-10

DocsSourceD2-C912

Google's site move guide says to submit a Change of Address in Search Console when moving from one domain or subdomain to another, to submit the new sitemap, and to change internal links on the new site from the old URLs to the new ones.

Publisher Google Search Central

Used byrequirement DEV-CAN-09glossary term Change of Address tool

AnalysisD2-C913

For URLs removed in a consolidation, return 404 or 410: Google's site move guide names those two codes, and Google's crawlers treat every 4xx code except 429 the same way, as content that does not exist, so the choice between them matters less than not redirecting to an unrelated page.

Author Ibrahim Anjro

AnalysisD2-C914

Google's crawl budget guide is written for sites with over a million pages changing weekly or over 10,000 pages changing daily, so for a consolidation of about 2,000 URLs the crawl-budget gain is likely minor; the benefit of pruning such a site more plausibly comes from one strong URL per intent and consolidated signals.

Author Ibrahim Anjro

AnalysisD2-C915

The 304-day median recovery a community speaker cited from an unnamed study and Google's site move guide (a few weeks or more for a medium-sized site until Google shows the new URLs, longer for larger sites) measure different things: indexing can switch within weeks while traffic recovery can take far longer, so plan and report on both.

Author Ibrahim Anjro

Things

13:30 · Finding the gold nuggets: structured data, media, and more! 11

SlideNot in docsD2-C441

Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
SlideNot in docsD2-C442

Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

Speaker Gary IllyesEvidence slide photo

  • Repeats D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
StageNot in docsD2-C443

Google's indexing includes a dedicated system, whose internal name Gary Illyes would not disclose, that extracts the parts of a page that are traditionally expensive to extract.

Speaker Gary IllyesEvidence transcript

DocsSourceD2-C445

Google's guide to how Search works says rendering happens during the crawl, and describes indexing as analysing a page's text, key tags and attributes such as title elements and alt attributes, images and videos, and deciding whether the page is a duplicate or the canonical.

Publisher Google Search Central

Things

Used byrequirement DEV-HTM-03

StageConsistent with docsD2-C446

The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such as JSON-LD), images and videos.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C323 Day 3: Most of Google's search features need nothing extra from the site owner; Google generates them from what it…
StageConsistent with docsD2-C447

For images, Google's feature extraction takes the img element with its src and other attributes, including inline images, and passes them on to Google's image indexing service.

Speaker Gary IllyesEvidence transcript

  • Repeated by D2-C524 Day 2: Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly…
StageConsistent with docsD2-C448

For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-01

  • Extended by D2-C930 Day 2: Google's media indexer processes the images and videos that feature extraction passes to it and attaches them…
StageNot in docsD2-C449

Gary Illyes said feature extraction, which extracts page structures into a form Google's systems can consume internally, is still expensive, though not the most expensive operation.

Speaker Gary IllyesEvidence transcript

DocsSourceD2-C451

Google's general structured data guidelines recommend placing the same structured data on all duplicate pages of the same content, not just on the canonical page.

“we recommend placing the same structured data on all page duplicates, not just on the canonical page”

Publisher Google Search Central

Used byrequirement DEV-SDA-10

13:35 · What is Structured Data and why we need it on the internet. 74

StageD2-C452

Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.

Speaker Ryan LeveringEvidence transcript

  • Extends D1-C011 Day 1: Ecosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such…
  • Extends D1-C158 Day 1: Google said that even the classic ten-blue-links layout was settled only after millions of experiments, as…
StageConfirmed by docsD2-C454

Structured data turns the loosely structured web into structured information that powers visual search features such as review stars and recipe filters (for example by preparation time).

Speaker Ryan LeveringEvidence transcript

Used byglossary term Rich results

  • Extended by D3-C337 Day 3: Review structured data lets a site specify how its users rated something; the review snippet shows an average…
SlideNot in docsD2-C458

Google gave four reasons why structured data is still valuable even though models can extract information from pages: precision, extra content, efficiency and focus.

Speaker Ryan LeveringEvidence slide photo, transcript

SlideNot in docsD2-C459

Structured data gives the high precision that complex schemas such as sale pricing need, with higher accuracy than large-scale extraction by large language models (LLMs).

“Structured data provides the high precision needed for complex schema (sale pricing), achieving higher accuracy than large-scale LLM extraction.”

Wording checked against the slide or recording

Speaker Ryan LeveringEvidence slide photo, transcript

StageNot in docsD2-C460

In the speaker's own tests, even the latest LLMs asked to generate schema.org markup for a page often invent properties that do not exist, get deeply nested schemas such as complex pricing models wrong, and duplicate content across several fields.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-11

DocsSourceD2-C462

Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.

Publisher Google Search Central

Used byrequirement DEV-SDA-11

  • Extends D1-C250 Day 1: In a community demo, a fix loop gave a large LLM (Claude Opus) the current markup, the errors Google's test…
  • Extended by D2-C947 Day 2: In a community speaker's case, an LLM asked to write alt text for a product image of a veterinary anxiety…
  • Extended by D2-C952 Day 2: A community speaker warned that articles still recommend rewriting alt text with AI, called AI a tool, and…
  • Extended by D3-C686 Day 3: Hallucinations can happen with any AI model, and with current training methods there is no way to get rid of…
SlideNot in docsD2-C463

Structured data often carries non-visible metadata that the page text lacks, such as full ISO dates or stable identifiers for user-generated content.

“It often contains non-visible metadata, such as full ISO dates or stable identifiers for UGC, that is not present in the page text.”

Wording checked against the slide or recording

Speaker Ryan LeveringEvidence slide photo, transcript

Used byrequirements DEV-SDA-04, DEV-SDA-05

DocsSourceD2-C467

Google's structured data documentation says not to mark up content that is not visible to readers of the page, and not to add structured data about information that users cannot see even if it is accurate.

“don't add structured data about information that is not visible to the user, even if the information is accurate”

Publisher Google Search Central

Used byrequirement DEV-SDA-02

AnalysisD2-C468

The 'non-visible metadata' argument is not a licence to mark up hidden content, since Google's structured data guidelines still require markup to describe what users can see; use markup for the machine-precise form of facts the page shows, such as a full ISO date with time zone for a visible event date or a homepage url for a named organisation.

Author Ibrahim Anjro

Used byrequirements DEV-SDA-02, DEV-SDA-04, DEV-SDA-05

SlideNot in docsD2-C469

Parsing structured data is significantly cheaper and more efficient than relying on LLMs for every complex extraction task.

Speaker Ryan LeveringEvidence slide photo, transcript

  • Repeated by D3-C405 Day 3: Web markup is an efficient and unambiguous way for sites to share product data with Google, Google Shopping…
StageNot in docsD2-C471

Rule-based parsing of markup is nearly free by comparison with AI models, so Google will always prefer extracting information from structured data over model-based extraction.

“So we're always going to prefer that particular approach.”

Speaker Ryan LeveringEvidence transcript

AnalysisD2-C472

Do not drop markup on the assumption that AI reads the page anyway: by Google's own account LLM extraction is not precise enough for prices and nested offers and too costly to run on every page, so explicit markup remains the dependable route for those facts.

Author Ibrahim Anjro

StageNot in docsD2-C474

Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-07

  • Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
StageConsistent with docsD2-C476

The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-AIF-02

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageNot in docsD2-C477

As far as the speaker knows, in most of Google's main AI uses a page's schema.org markup is not turned into text and put directly into the model's context; the data is first sorted out, checked for quality and indexed before it is passed on as grounding context.

Speaker Ryan LeveringEvidence transcript

  • Extends D1-C061 Day 1: Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in…
AnalysisD2-C479

There is no separate markup for AI features: by Google's account the same processed markup feeds classic results and AI answers, and raw schema.org is generally not passed into model context, so invest in the types Google documents for its features rather than in extra markup written for AI.

Author Ibrahim Anjro

Used byrequirement DEV-AIF-02

DocsSourceD2-C480

Google's guide to optimizing for generative AI features lists 'overfocusing on structured data' among the things site owners don't need to do: structured data is not required for generative AI search and no special schema.org markup is needed, though it remains worth using because it helps pages become eligible for rich results.

Publisher Google Search Central

Used byrequirement DEV-AIF-02

  • Extends D1-C061 Day 1: Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in…
SlideConfirmed by docsD2-C482

A Google case study shown on stage (published May 2018) reported that Rakuten Recipes got 2.7 times more traffic from search engines and a 1.5 times increase in session duration after implementing recipe structured data.

Speaker Ryan LeveringEvidence slide photo, transcript

StageD2-C483

The speaker called the Rakuten study old and said the world has changed a lot, but expects the link between structured data and more interactive, visually appealing results to hold for the foreseeable future.

Speaker Ryan LeveringEvidence transcript

StageConfirmed by docsD2-C484

Schema.org is a common vocabulary that several major search engines started together so that site owners can mark up pages and every consumer interprets the markup the same way; the speaker put its start 15 to 20 years ago (Google, Bing and Yahoo! announced it in June 2011).

Speaker Ryan LeveringEvidence transcript

Used byglossary term Schema.org

StageConsistent with docsD2-C489

Describing every semantic detail of a page in markup is probably not worth the effort; focus on the structured data that Google or other consumers actually use.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-03

  • Extended by D3-C378 Day 3: Google Shopping called rich data with light structure the AI sweet spot: a little structure that helps the AI…
SlideConfirmed by docsD2-C490

Google recommends using the Search gallery in its developer documentation to find the structured data features that suit a site; the gallery shows each feature and how Google uses the markup.

Speaker Ryan LeveringEvidence slide photo, transcript

Used byrequirement DEV-SDA-03

  • Repeated by D3-C329 Day 3: Google's structured data feature guide lists the kinds of structured data Google supports with a search…
StageConsistent with docsD2-C494

Google accepts three structured data syntaxes, JSON-LD, microdata and RDFa, which are all valid and are extracted at the very start into the same pipelines, so they are interpreted identically downstream.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-01glossary term JSON-LD

StageNot in docsD2-C496

Microdata has an advantage when payload size matters: embedded in the existing HTML, it avoids duplicating the page content in a separate JSON-LD block.

Speaker Ryan LeveringEvidence transcript

Things

Used byrequirement DEV-SDA-01

AnalysisD2-C497

Default to JSON-LD and switch to microdata only where page weight is critical, since Google interprets all three syntaxes identically and JSON-LD is the one people get wrong least often.

Author Ibrahim Anjro

Things

Used byrequirement DEV-SDA-01

SlideConfirmed by docsD2-C498

Google recommends testing and previewing structured data in the Rich Results Test, which shows the rich result features it detected and whether the markup is valid.

Speaker Ryan LeveringEvidence slide photo, transcript

Used byrequirement DEV-SDA-11glossary term Rich Results Test

StageConfirmed by docsD2-C500

Structured data that is not relevant to the page's content can be treated as abusive: Google's filters make it ineffective, and egregious cases can lead to a manual action.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-02

  • Extended by D3-C647 Day 3: Google may never use structured data from a site it does not trust: once it sees markup it does not trust, it…
StageNot in docsD2-C503

Several plug-ins emitting the same markup type is one of the most common structured data problems: the duplicates can make an event details page look like a list of events and change how Google interprets the page.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-06

AnalysisD2-C504

Audit CMS templates for duplicate markup by running one page per template through the Rich Results Test and checking whether an SEO plug-in and a theme or events plug-in emit the same type twice; 'more markup never hurts' covers relevant, non-duplicated markup only.

Author Ibrahim Anjro

Used byrequirement DEV-SDA-06

StageConfirmed by docsD2-C505

Google's shopping structured data launches of the previous year (2025) added support for merchant loyalty programs and shipping policies, letting merchants define a policy at organisation level and specify details at product level.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-09

  • Extended by D3-C383 Day 3: In product markup, an offer's shipping and return information can point through a JSON-LD identifier (@id) to…
StageConsistent with docsD2-C506

Google's structured data speaker said that in the months before the event Google added support for validity dates on sale prices in product structured data, so merchants no longer need to rush to remove a sale price when the sale ends for fear it shows wrongly in snippets.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-08

  • Extended by D3-C396 Day 3: Google works to make sure that what merchants can express in its shopping feeds can also be expressed in…
DocsSourceD2-C507

Google's documentation updates log calls the July 2026 sale price change a clarification: a new Sale duration section of the merchant listing guide explains validFrom with validThrough or priceValidUntil, aligned with Merchant Center's sale_price_effective_date attribute.

Publisher Google Search Central

Used byrequirement DEV-SDA-08

AnalysisD2-C508

For time-limited sales, give the sale price its validity dates in the product markup (Google's merchant listing documentation describes validFrom, validThrough and priceValidUntil) instead of editing markup by hand when the sale ends, and keep the dates aligned with the Merchant Center feed.

Author Ibrahim Anjro

Used byrequirement DEV-SDA-08

StageD2-C509

More shopping structured data news was left for a Day 3 talk by a Google colleague, Alex.

Speaker Ryan LeveringEvidence transcript

  • Extended by D3-C365 Day 3: Google added six Merchant Center feed attributes for AI shopping experiences: question and answer, documents…
StageNot in docsD2-C510

Schema.org, in which Google is a major participant, began publishing usage statistics for all its types and properties in 2026, showing in buckets how many domains use each one; the data is also in schema.org's GitHub repository.

Speaker Ryan LeveringEvidence transcript

StageNot in docsD2-C511

Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a chicken-and-egg problem that the schema.org usage statistics are meant to help break.

Speaker Ryan LeveringEvidence transcript

  • Extends D1-C120 Day 1: To get Google to build something, such as an addition to an API, report and request it publicly and in volume.
StageNot in docsD2-C512

Schema.org added support for RDF lists and sets, which give a way to express ordered values, because RDF triples are not ordered by nature.

Speaker Ryan LeveringEvidence transcript

SlideNot in docsD2-C514

Google announced server-side structured data validation as coming soon: it will publish downloadable validation rules in SHACL on each structured data feature guide, with other kinds of checks to follow.

Speaker Ryan LeveringEvidence slide photo, transcript

Used byrequirement DEV-SDA-12glossary term SHACL

SlideNot in docsD2-C515

Google's planned validation workflow has five steps: download the rules from the feature guide, generate the JSON or embedded microdata or RDFa, run the rules against the generated server-side markup as a first check, deploy and test in the Rich Results Test, and monitor ongoing performance in Search Console.

Speaker Ryan LeveringEvidence slide photo

Used byrequirement DEV-SDA-12

StageNot in docsD2-C516

The SHACL rules are meant to run inside a site's content generation, so markup is sanity-checked before it is published and does not silently regress later, a breakage site owners might otherwise discover only through a Search Console report.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-12glossary term SHACL

StageNot in docsD2-C517

The SHACL rules will not replace Search Console as the canonical place for structured data reports, because some checks use Google's internal libraries and cannot be expressed in SHACL, but they will catch problems such as a missing required field.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-12

SlideConsistent with docsD2-C518

Google's example SHACL shape for Event requires a name, treats a missing description only as a warning, and accepts the image either as an ImageObject or as a URL.

Speaker Ryan LeveringEvidence slide photo, transcript

AnalysisD2-C520

When the SHACL rules appear, wire them into the build or CMS publishing step as an automated test, so a template change that drops a required property fails before deployment instead of surfacing weeks later in Search Console.

Author Ibrahim Anjro

Used byrequirement DEV-SDA-12

StageConsistent with docsD2-C521

Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.

Speaker Ryan LeveringEvidence transcript

Used byglossary terms Rich results, Structured data

  • Extends D1-C122 Day 1: Check and focus on rich results for Google.
  • Extended by D3-C324 Day 3: Rich results differ from other search features because Google builds them from extra data that site owners…

13:50 · Using images to your advantage and Engaging Search users with videos 66

StageD2-C523

Gary Illyes said images and videos drive a large amount of traffic to publishers.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C116 Day 1: Images were stressed as important.
StageConsistent with docsD2-C524

Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-01

  • Repeats D2-C447 Day 2: For images, Google's feature extraction takes the img element with its src and other attributes, including…
StageConsistent with docsD2-C525

An image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-IMG-03

  • Extends D1-C116 Day 1: Images were stressed as important.
  • Extends D1-C117 Day 1: With one in six AI Mode searches being multimodal, original images with descriptive file names, alt text and…
  • Extended by D3-C319 Day 3: Image results shown among web results come from Google's image index and are roughly the same images that…
StageConsistent with docsD2-C526

Google supports the picture element only because a picture element must contain an img element, and that img element is what Google extracts and passes to its media indexer.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-01

StageConsistent with docsD2-C527

One of the most reliable ways to make sure Google finds an image is to include it in an img element in the HTML; Gary Illyes named image sitemaps as the other method.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirements DEV-IMG-01, DEV-IMG-04

StageConfirmed by docsD2-C528

Google does not extract CSS background images.

“we don't support CSS background extraction”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-01

DocsSourceD2-C529

Google's documented ways to keep a site's images out of search results are a robots.txt disallow rule (for example for Googlebot-Image) or a noindex X-Robots-Tag HTTP header, with the Removals tool for emergencies.

Publisher Google Search Central

Used byrequirements DEV-IDX-03, DEV-IMG-06glossary term X-Robots-Tag

AnalysisD2-C530

Hiding an image as a CSS background is a fragile way to keep it out of Google: it only stops extraction from that page, so the same image URL used in an img element elsewhere or listed in a sitemap can still be indexed; the documented robots.txt or noindex X-Robots-Tag methods are the reliable route.

Author Ibrahim Anjro

Used byrequirement DEV-IMG-06

StageNot in docsD2-C531

Of the attributes the HTML standard defines for the img element, Google uses three and ignores the rest; Gary Illyes named src and alt but not the third.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD2-C532

The src attribute is the most important img attribute, because without it Google does not know where the image bytes are and cannot index the image.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-01

StageNot in docsD2-C533

Gary Illyes ranked the alt attribute below src in importance, while saying alt attributes are still important.

Speaker Gary IllyesEvidence transcript

DocsSourceD2-C534

Google's image SEO guide calls alt text the most important attribute for providing more metadata about an image, and says Google uses it together with computer vision and the page content to understand the image.

“The most important attribute when it comes to providing more metadata for an image is the alt text”

Publisher Google Search Central

Used byrequirement DEV-IMG-02

StageConsistent with docsD2-C536

Google uses the words in an alt attribute to understand the image, may attach them to the image at serving time, and the image can rank for concepts the alt text describes.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-02

StageConsistent with docsD2-C537

The text around an image is critical: Google uses it as context to understand the image and to rank it, so an alt attribute alone is not enough.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-03

  • Repeats D1-C045 Day 1: Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour…
  • Extended by D2-C950 Day 2: Because not every generated alt text can be checked by hand, a community speaker's safer script adds the page…
AnalysisD2-C538

Images that should rank need an img element with a real src; hero or product images set as CSS backgrounds are invisible to Google Images and image features, so keep CSS backgrounds for decorative images you do not need indexed.

Author Ibrahim Anjro

Used byrequirement DEV-IMG-01

AnalysisD2-C539

Give important images a caption or explanatory sentence next to them rather than relying on alt text alone, since Gary Illyes ranked src above alt and called the surrounding text critical for ranking the image.

Author Ibrahim Anjro

Used byrequirement DEV-IMG-03

StageConfirmed by docsD2-C540

Besides img elements, image sitemaps tell Google about images: they are XML sitemaps that list, under a page's loc entry, the locations of the images on that page.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-IMG-04

StageNot in docsD2-C542

Gary Illyes said the AVIF image format currently has hiccups and Google may have problems ingesting it, although it should technically be supported.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-05

DocsSourceD2-C543

Google's image SEO guide lists AVIF among the supported image formats (with BMP, GIF, JPEG, PNG, WebP and SVG), and Google announced in August 2024 that AVIF files need nothing special to be indexed.

Publisher Google Search Central, Search Central blog (30 August 2024)

Used byrequirement DEV-IMG-05

StageConsistent with docsD2-C544

Use high-quality modern image formats such as WebP for a good balance of quality and compression, and keep image files small so people can enjoy them.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-05

AnalysisD2-C545

While the AVIF ingestion hiccup Gary Illyes mentioned lasts, serve AVIF only through source elements inside picture and keep a WebP or JPEG file in the img src, which is the URL Google extracts, even though Google's documentation lists AVIF as supported; then check in Google Images that key images appear.

Author Ibrahim Anjro

Used byrequirement DEV-IMG-05

StageNot in docsD2-C546

Gary Illyes cited a figure that more than 40% of Southeast Asian shoppers rely on videos to make purchase decisions.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C547

Gary Illyes said over 219 million people in Southeast Asia consume content on or through YouTube daily (the daily scope was heard in two independent recordings but is not verified).

Speaker Gary IllyesEvidence transcript

Things
StageNot in docsD2-C548

Gary Illyes said there are over 150 streaming apps in Southeast Asia.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD2-C549

Videos can appear in the main search results, in video-specific result tabs and in Discover (the tab names are unclear in the recording).

Speaker Gary IllyesEvidence transcript

Things
StageConfirmed by docsD2-C550

Google has video-specific features, such as key moments and previews, that help users interact with videos more easily.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-03

StageConsistent with docsD2-C551

Google strongly suggests describing a video's metadata with JSON-LD structured data, on top of placing the video in an HTML element on the page.

Speaker Gary IllyesEvidence transcript

Used byrequirements DEV-VID-01, DEV-VID-03

SlideConsistent with docsD2-C552

Google's slide listed seven key factors for video SEO success: high-quality video content, a dedicated watch page, compelling titles and descriptions, relevant thumbnails, video markup, fast-loading pages and sitemap inclusion.

Speaker Gary IllyesEvidence slide photo, transcript

Things

Used byrequirement DEV-VID-02

SlideConsistent with docsD2-C553

Google's video SEO slide said each video must be on its own dedicated HTML watch page.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-VID-02

DocsSourceD2-C554

Google's video SEO guide says a video must be embedded on an indexed watch page to be eligible for video features and recommends a dedicated watch page per video where it makes sense for the business; a non-watch page with the video can still appear as a text result or a Google Images result with a video badge.

Publisher Google Search Central

Used byrequirement DEV-VID-02

SlideConsistent with docsD2-C555

Google's video SEO slide warned that spending money on a generative AI subscription does not mean AI-generated videos will do well, because engaging, informative, well-produced content is foundational.

“just because you spent money on a genAI subscription, that does not mean you are going to do well with AI generated videos”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo

StageNot in docsD2-C917

Google knows from experiments that when a video's thumbnail is wrong, the share of viewers who drop out at the start of the video is extremely high.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-03

StageNot in docsD2-C919

Gary Illyes called the 'fast-loading pages' factor for video SEO a misnomer: what matters is that the video itself loads fast, because people no longer have the patience to wait for videos to load.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

StageNot in docsD2-C920

Serving videos from a content delivery network (CDN) that loads them faster than the site's own server is a win for video SEO, Gary Illyes said.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-VID-06

StageNot in docsD2-C922

Google recommends the MP4 container for videos because of its browser and device compatibility.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

StageNot in docsD2-C923

Encode videos with standard codecs, because some people will not be able to play videos that use unusual ones.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

StageNot in docsD2-C924

For a video to be discovered it must be embedded prominently, above the fold; Gary Illyes said a video placed below the fold is not going to be indexed.

“If it's not above the fold, then you basically lost the game.”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-01

StageConsistent with docsD2-C928

Gary Illyes suggested considering video hosting platforms such as YouTube or Vimeo, because they solved video search years ago.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-VID-06

StageNot in docsD2-C929

The more people talk about a site's videos, the more likely Google is to surface them in search results, so videos should be promoted.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C930

Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C448 Day 2: For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense…
StageConfirmed by docsD2-C932

If robots.txt disallows the location of an image or video file, Google does not index that image or video.

Speaker Gary IllyesEvidence transcript

Used byrequirements DEV-IMG-06, DEV-VID-06glossary term Googlebot-Image and Googlebot-Video

StageNot in docsD2-C933

Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

  • Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
StageConfirmed by docsD2-C934

The noimageindex robots meta tag tells Google not to index the images on the page.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IDX-09glossary term noimageindex

  • Repeats D2-C099 Day 2: The noimageindex rule tells Google not to index any of the images on the page, and John Mueller said he could…
StageConsistent with docsD2-C936

Setting the max-image-preview robots meta tag to large can make content perform surprisingly well in Discover, Gary Illyes said.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IDX-04

  • Repeats D2-C089 Day 2: max-image-preview:large matters mainly in Discover, where it allows a large image that draws people's…
  • Repeated by D3-C224 Day 3: Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels…
StageNot in docsD2-C938

Google serves AI-generated images and videos in search results when users are specifically looking for them.

“We will serve users AI-generated images and videos in search results if they are looking for them specifically.”

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C939

The diffusion models that generate images were built to generate images, not text, so they are typically poor at rendering text inside an image.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-08

  • Extended by D3-C688 Day 3: Generative models, including image diffusion models, make things up when they lack information or because of…
StageD2-C940

Gary Illyes showed a timeline image that Google made with an image generation model in 2025 and deliberately left unfixed as a counterexample: its text was garbled and its years jumped from 1994 to 1995 and back to 1994.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD2-C941

Sites that use AI-generated images or videos should make sure they work for users, check them for hallucinations and regenerate them where needed.

Speaker Gary IllyesEvidence transcript

Used byrequirements DEV-IMG-08, DEV-SPM-05

  • Extended by D3-C700 Day 3: Google's closing slide said to use AI responsibly because AI hallucinates, and, especially when creating…
DocsSourceD2-C942

Google's Video indexing report help says only videos on a watch page are eligible for indexing, and flags 'Cannot determine video position and size' when the player is not on the page at load, for example behind a click-to-play image, asking for the player to load at its real size and position without user interaction.

Publisher Google Search Console Help

Used byrequirement DEV-VID-01glossary term Video watch page

AnalysisD2-C944

Gary Illyes said a video below the fold is not indexed, but Google's video documentation only requires the player to be present at load, at a position and size Google can determine and not hidden behind other elements; to be safe, put the main video of a watch page in the first viewport and load the player without a click-to-play placeholder.

Author Ibrahim Anjro

AnalysisD2-C945

Do not set noimageindex on video watch pages: Gary Illyes said it also stops the video, through its thumbnail, and although the robots meta tag specification mentions only images, the Video indexing report treats a missing or blocked thumbnail as a reason a video is not indexed.

Author Ibrahim Anjro

14:00 · Lightning session F: Media 15

StageNot in docsD2-C946

Version 20 of the Screaming Frog SEO Spider, released in May 2024, was the first that could connect the crawler to ChatGPT through custom JavaScript, which let SEOs rewrite image alt text at scale.

Speaker not identifiedEvidence transcript

StageD2-C947

In a community speaker's case, an LLM asked to write alt text for a product image of a veterinary anxiety medicine for cats described only what it could see, a cat and a veterinarian, and dropped the stress, anxiety and medical-treatment intent of the page.

Speaker not identifiedEvidence transcript

  • Extends D2-C462 Day 2: Google's guidance on generative AI content warns that AI output can contain hallucinations and says…
StageNot in docsD2-C949

To test LLM-generated alt text before a rollout, a community speaker recommended running the crawler's custom JavaScript in Screaming Frog's List mode on a few chosen URLs and reviewing the generated alt text, instead of crawling the whole site in Spider mode.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-IMG-02

StageConsistent with docsD2-C950

Because not every generated alt text can be checked by hand, a community speaker's safer script adds the page title to the LLM's image description, so the alt text carries the page's intent and not only what the image shows.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-IMG-02

  • Extends D2-C537 Day 2: The text around an image is critical: Google uses it as context to understand the image and to rank it, so an…
StageD2-C951

A community speaker named the limits of the title-based alt-text script: it rewrites only the hero image, assumes the page title states the user's intent, had been tested in only one language or market, and runs on the free Gemini model.

Speaker not identifiedEvidence transcript

Things
StageD2-C952

A community speaker warned that articles still recommend rewriting alt text with AI, called AI a tool, and said that adopting such new AI implementations without applying existing SEO knowledge can damage ranking performance.

Speaker not identifiedEvidence transcript

  • Extends D2-C462 Day 2: Google's guidance on generative AI content warns that AI output can contain hallucinations and says…
StageD2-C953

A community talk described 'prompt journalism' as journalistic and content work, such as interviews and podcasts, done with a prompt-assisted workflow whose aim is to optimize material already recorded, which holds expertise, quotes and context, rather than to create more content.

Speaker Patrick DomanicoEvidence transcript

Used byglossary term Prompt journalism

StageD2-C954

The first step of a community speaker's prompt-assisted workflow is cleaning the interview transcript of fillers, a step that needs editorial and SEO judgment because a published page is read by many systems whose way of consuming it is unknown.

Speaker Patrick DomanicoEvidence transcript

StageD2-C956

In the mapping step of a prompt-assisted workflow, an AI system splits an interview transcript into chapters with timestamps, headings and entities; in a community speaker's example it correctly extracted the subsea cables 2Africa and Equiano from a passage on Internet infrastructure in Africa.

Speaker Patrick DomanicoEvidence transcript

StageD2-C958

The second step of the prompt-assisted workflow applies the newsroom's inverted pyramid (the five Ws and H first, then attribution and quotes, context and secondary text); because it is a fixed structure, it can be written as a prompt that generates the article.

Speaker Patrick DomanicoEvidence transcript

StageD2-C959

From one unified interview transcript, a community speaker's workflow produces a full article for the website and shorter angles for LinkedIn, Substack, X and YouTube.

Speaker Patrick DomanicoEvidence transcript

Things
StageConfirmed by docsD2-C960

A community speaker said publishers can also verify their social accounts in Search Console, which makes it easier to track the performance of content repurposed for social platforms.

Speaker Patrick DomanicoEvidence transcript

Used byrequirement DEV-MON-09

  • Extended by D3-C442 Day 3: Search Console launched platform properties, which bring social platforms into Search Console, about three or…
  • Extended by D3-C468 Day 3: Google's social and video performance guide says that if you already claimed your Search profile, all of its…

14:15 · Focusing on Internationalisation and Localisation 64

DocsSourceD2-C561

Google's documentation says Googlebot mostly crawls from US IP addresses and sends no Accept-Language header, so pages that change content or redirect by the visitor's perceived country or language may not have every version crawled, indexed or ranked; it recommends separate URLs annotated with hreflang.

Publisher Google Search Central

Used byrequirements DEV-INT-01, DEV-INT-02

SlideConfirmed by docsD2-C562

Language versions can also be language-plus-region variants: Google's example site had a generic English page (en), a UK English page (en-gb) and a Spanish page (es), each on its own URL.

Speaker GoogleEvidence slide photo, transcript

Used byrequirement DEV-INT-01

  • Extended by D2-C627 Day 2: A community speaker said language versions need adjusting for regional variants of one language: 90 is…
SlideConfirmed by docsD2-C565

Missing return links are a common hreflang mistake: if page X names page Y as a language version, page Y must link back to page X.

“If page X links to page Y, page Y must link back to page X.”

Wording checked against the slide or recording

Speaker GoogleEvidence slide photo, transcript

Things

Used byrequirement DEV-INT-03glossary term hreflang return link

DocsSourceD2-C571

Google's hreflang guide says a site that finds it hard to keep every language pair bidirectional may omit some languages on some pages, because Google still processes the pairs that point to each other; new language versions should at least link both ways with the original or dominant language.

Publisher Google Search Central

Things

Used byrequirement DEV-INT-03

AnalysisD2-C574

Audit hreflang per cluster rather than per page: check that every page lists itself and all its alternates, that every alternate links back, and that only one method (HTML head, HTTP header or sitemap) supplies the annotations.

Author Ibrahim Anjro

Used byrequirement DEV-INT-03

SlideConsistent with docsD2-C575

hreflang language codes must be ISO 639-1 codes: Google's list of common mistakes marked se, dk and cz (the country codes of Sweden, Denmark and Czechia) as wrong and sv, da and cs (Swedish, Danish, Czech) as right.

Speaker GoogleEvidence slide photo, transcript

Things

Used byrequirement DEV-INT-04

  • Extended by D2-C962 Day 2: Google said incorrect hreflang language or region codes are a mistake it finds very often, with the same…
AnalysisD2-C583

Validate hreflang values against the ISO 639-1 and ISO 3166-1 Alpha-2 lists in a crawler rule or build check rather than by eye: reserved codes such as UK and EU are simply ignored, but some plausible mistakes are valid codes for something else (se is Northern Sami, GE Georgia, LA Laos, SA Saudi Arabia) and silently point the annotation at the wrong audience.

Author Ibrahim Anjro

Things

Used byrequirement DEV-INT-04

StageConfirmed by docsD2-C585

Google determines a page's language for indexing from the page content, not from a language code in the URL.

Speaker GoogleEvidence transcript

Used byrequirement DEV-INT-07

  • Extended by D2-C651 Day 2: Google detects the language of each document and weights language in index selection so that the index is not…
StageConfirmed by docsD2-C587

Pages that mix several languages make it very hard for Google to decide which language a page is in, so each page should make its target language obvious and avoid mixing languages.

Speaker GoogleEvidence transcript

Used byrequirement DEV-INT-07

DocsSourceD2-C589

Google's hreflang guide says pages that translate only the template (navigation, footer) around main content in one language still count as duplicates, because localized versions are duplicates only if the main content stays untranslated; it recommends hreflang for such pages.

Publisher Google Search Central

Used byrequirement DEV-INT-07

StageConsistent with docsD2-C591

The ccTLD (country-code top-level domain) is one of the most important and strongest signals Google uses to decide which country a site targets; Google also considers other signals, which matter less.

Speaker GoogleEvidence transcript

Used byrequirement DEV-INT-09glossary term ccTLD

  • Extended by D2-C982 Day 2: At the point where one recording names the ccTLD as the most important country-targeting signal, a second…
StageNot in docsD2-C592

Server location is not really a reliable country-targeting signal nowadays, so Google does not use it much; the recording is unclear on the word 'server'.

Speaker GoogleEvidence transcript

AnalysisD2-C982

At the point where one recording names the ccTLD as the most important country-targeting signal, a second attendee recording heard 'IP address' instead; Google's multi-regional sites documentation names the ccTLD as a strong signal of a site's target country and server location only as a possible, not definitive one, so the ccTLD is the better-supported reading.

Author Ibrahim Anjro

  • Extends D2-C591 Day 2: The ccTLD (country-code top-level domain) is one of the most important and strongest signals Google uses to…
DocsSourceD2-C593

Google's documentation calls a ccTLD a strong signal that a site is meant for a certain country and still lists server location as a possible signal of a site's audience, though not a definitive one because sites use CDNs or are hosted abroad; it does not rank the signals.

Publisher Google Search Central

Things

Used byglossary term ccTLD

StageConsistent with docsD2-C594

For country versions of a site, a ccTLD, subdomains or subdirectories are all acceptable choices; the right one depends on the site's needs, goals and resources.

Speaker GoogleEvidence transcript

Used byrequirement DEV-INT-09

  • Extended by D2-C822 Day 2: Of the URL structures for country versions in a table on Google's slide (not photographed), the speaker…
DocsSourceD2-C595

Google's table of URL structures for country targeting gives pros and cons for a country-specific domain (example.de), a subdomain (de.example.com) and a subdirectory (example.com/de/) on a generic domain, and marks URL parameters (site.com?loc=de) as not recommended.

Publisher Google Search Central

Used byrequirements DEV-INT-01, DEV-INT-09

StageConsistent with docsD2-C822

Of the URL structures for country versions in a table on Google's slide (not photographed), the speaker called only one a wrong choice: the last option, which the speaker did not name but said they really do not recommend.

Speaker GoogleEvidence transcript

  • Extends D2-C594 Day 2: For country versions of a site, a ccTLD, subdomains or subdirectories are all acceptable choices; the right…
AnalysisD2-C598

For a business testing new countries, subdirectories on the existing domain are usually the cheapest start; a ccTLD's stronger country signal pays off only where domain availability, cost, local rules and long-term commitment justify it, and an established domain should not be moved just for the signal.

Author Ibrahim Anjro

Used byrequirement DEV-INT-09

StageNot in docsD2-C599

Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.

Speaker GoogleEvidence transcript

  • Extended by D2-C965 Day 2: Users of non-Latin-script languages do not always search in their own script: the same Persian query may be…
  • Extended by D3-C051 Day 3: Users expect content written the way they search: in some languages they search in Latin characters, in…
StageConfirmed by docsD2-C601

Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results, so what works for web results also works for the AI features.

Speaker GoogleEvidence transcript

Used byrequirement DEV-AIF-01story angle A-001

  • Repeats D1-C050 Day 1: Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
  • Repeats D1-C051 Day 1: Three reasons were given: generative AI features are built directly on the core ranking systems, query…
  • Extends D2-C116 Day 2: AI Overviews and AI Mode are built on top of Search results: they are a different experience of the same…
StageD2-C602

In classic results users can usually tell when content in their language is missing, poorly translated or from another country, but AI answers synthesize many sources and hide this; the speaker called it an invisible gap.

“an invisible gap”

Speaker GoogleEvidence transcript

StageConsistent with docsD2-C603

Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.

“AI is still language-dependent”

Speaker GoogleEvidence transcript

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
  • Extended by D2-C972 Day 2: In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English…
DocsSourceD2-C604

Google's Search Help says the language a query is typed in is an important factor in choosing the language of results, alongside interface and device language, location and site owners' annotations, and that Google may also show results in other languages when they are helpful or when there is not enough information in the query's language.

Publisher Google Search Help

StageD2-C605

In the speaker's own experience, the differences between languages are larger in AI answers than in classic search results.

Speaker GoogleEvidence transcript

AnalysisD2-C607

Check AI Overviews and AI Mode with native-language queries in each target market, not with translated English keywords, and compare with the country breakdown of Search Console's generative AI performance report; topics where competitors are cited and you are not point to missing or weak local content.

Author Ibrahim Anjro

Used byrequirement DEV-MON-07

  • Extends D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
StageNot in docsD2-C608

Whether machine-translated content is acceptable depends on the case and is the site owner's decision, after weighing three things machine translation can miss: translation quality, local conventions and cultural adaptation.

Speaker GoogleEvidence transcript

Used byrequirement DEV-INT-10

StageD2-C609

To show uneven machine-translation quality, the speaker showed Google Translate output of attendee questions from a visit to China with a Google colleague, Gary, some clear and some unintelligible, and a restaurant menu whose English listed an item called 'aggressive element'.

Speaker GoogleEvidence transcript

StageNot in docsD2-C610

Localisation should account for local conventions such as date formats, which differ between Europe, the UK, the US and other countries, and calendars (in Thailand the current year is 2569), or a date can point users to the wrong day.

Speaker GoogleEvidence transcript

Used byrequirements DEV-INT-10, DEV-SDA-05

DocsSourceD2-C611

Google's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.

Publisher Google Search Central

Used byrequirement DEV-INT-10glossary term Scaled content abuse

  • Extended by D3-C695 Day 3: The cheaper tokens become, the more AI slop is created, and Google counts AI slop as scaled content abuse.
  • Extended by D3-C697 Day 3: Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem…
AnalysisD2-C612

Treat machine translation as a first draft: have a native speaker review it and adapt dates, calendars and units before publishing, because unreviewed bulk translation that adds little value can also fall under Google's scaled content abuse policy.

Author Ibrahim Anjro

Used byrequirement DEV-INT-10

  • Extends D1-C048 Day 1: The answer is not a claim that Google detects AI text. It says ranking favours text that reads as natural to…
StageNot in docsD2-C613

According to the speaker, shoppers in Europe and the US pay a lot of attention to promotions and discounts when deciding to buy (no source for this was captured), one of the cultural factors localisation should take into account.

Speaker GoogleEvidence transcript

StageNot in docsD2-C614

According to consumer data shown on a slide in the talk (source not captured), US consumers judge product quality more by user feedback and reviews, while European shoppers seem to look more at brand reputation.

Speaker GoogleEvidence transcript

  • Extended by D3-C341 Day 3: Some countries rely a lot on social proof such as reviews, Google's speaker said, recalling a point from…
StageNot in docsD2-C615

According to consumer data shown on a slide in the talk (source not captured), German shoppers also rely heavily on expert recommendations and certifications when judging product quality.

Speaker GoogleEvidence transcript

StageNot in docsD2-C616

According to consumer data shown on a slide in the talk (source not captured), French shoppers particularly want to know where a product comes from.

Speaker GoogleEvidence transcript

AnalysisD2-C617

Localise the trust signals on product pages, not only the words: if the findings shown hold for your category, lead with reviews in the US, brand reputation elsewhere in Europe, expert certifications in Germany and origin information in France, and test the order per market.

Author Ibrahim Anjro

StageConsistent with docsD2-C962

Google said incorrect hreflang language or region codes are a mistake it finds very often, with the same wrong codes coming up again and again, especially in Europe, because site owners assume the codes are easy.

Speaker GoogleEvidence transcript

Things

Used byrequirement DEV-INT-04

  • Extends D2-C575 Day 2: hreflang language codes must be ISO 639-1 codes: Google's list of common mistakes marked se, dk and cz (the…

14:30 · Lightning session G: Internationalisation 46

StageD2-C618

In a community speaker's audit case, a premium multifunctional-furniture maker about to be acquired had multilingual sites with clean translations by native speakers and hreflang tags that were in fairly good shape.

Speaker Alizée BaudezEvidence transcript

Things
StageD2-C619

In the same furniture-maker audit, the Dutch site had only a fraction of the conversions of the brand's other language sites, although it showed the same catalogue, the same hero products and the same photography.

Speaker Alizée BaudezEvidence transcript

StageD2-C620

Checking a competitor during the furniture-maker audit, a community speaker found that a large furniture retailer's French and Dutch sites differ a little in the products they show and especially in their photography, with items apparently shown in much smaller rooms on the Dutch site.

Speaker Alizée BaudezEvidence transcript

StageD2-C621

A community speaker said her first guess, that Dutch homes are smaller than French ones, was wrong: Dutch and French homes have about the same floor area per person, and what differs is the type of house.

Speaker Alizée BaudezEvidence transcript

StageD2-C622

A community speaker said about 19% of homes in Europe are terraced houses on several floors, against 58% in the Netherlands; the matching Eurostat figures are 2019 shares of the EU and Dutch population living in semi-detached or terraced houses, not shares of homes.

Speaker Alizée BaudezEvidence transcript

StageD2-C623

In a community speaker's comparison, a typical Dutch home is narrow and spread over several levels, often with very steep, narrow staircases that make furniture harder to install upstairs, while a typical French home has a wide ground floor and usually wide stairs.

Speaker Alizée BaudezEvidence transcript

StageD2-C624

The lesson of the furniture-maker case was that the translation was correct but the target region had never been checked.

“the translation was right, but the region was never checked”

Speaker Alizée BaudezEvidence transcript

AnalysisD2-C625

When translations and hreflang are clean but one country version converts far worse, compare that market's product range, photography and delivery or installation constraints with local competitors before changing anything technical.

Author Ibrahim Anjro

Things
StageD2-C627

A community speaker said language versions need adjusting for regional variants of one language: 90 is 'quatre-vingt-dix' in France but 'nonante' in Belgium, and Swiss German differs from German, for example in whether the letter ß (Eszett) is used.

Speaker Alizée BaudezEvidence transcript

  • Extends D2-C562 Day 2: Language versions can also be language-plus-region variants: Google's example site had a generic English page…
StageD2-C630

Of language, region and authority, a community speaker called language easy to settle and region solvable with a little research, while authority (why local people should trust a brand) has to be earned and built over time.

Speaker Alizée BaudezEvidence transcript

StageConsistent with docsD2-C631

A community speaker stressed that an hreflang annotation for a country (her example was German for Austria, de-AT) does not by itself make a site established or trusted in that country; local authority has to be worked for.

Speaker Alizée BaudezEvidence transcript

Things
DocsSourceD2-C632

Google's documentation lists the signals it uses to decide which locale a page targets: ccTLDs, hreflang statements, server location, and other signals such as local addresses and phone numbers, local language and currency, links from other local sites and Business Profile signals.

Publisher Google Search Central

Things

Used byrequirement DEV-INT-08

StageD2-C634

A community speaker said the SEO industry has treated hreflang as the hard part of international SEO; hreflang is complicated, but the real hard part is the homework of understanding consumers in each market.

“the hard part is actually doing the homework to understand your consumers”

Speaker Alizée BaudezEvidence transcript

Things
StageD2-C639

The second question before international SEO work is what genuinely differs between the regions of each language, such as photography, content, vocabulary, the way products are presented, pricing, catalogues and payment options.

Speaker Alizée BaudezEvidence transcript

StageD2-C640

The third question is whether you can name three local sources that already treat the brand as present in a country; if you cannot, a community speaker advised not to expand there but to fix the markets you already serve.

“don't bother expanding to another country; just fix the ones you already have”

Speaker Alizée BaudezEvidence transcript

StageNot in docsD2-C644

Arabic and Persian contain letters that look identical to users but are different characters to software, with different Unicode code points; the example given was the letter ye, which has an Arabic and a Persian form.

Speaker a second community speakerEvidence transcript

StageConsistent with docsD2-C965

Users of non-Latin-script languages do not always search in their own script: the same Persian query may be typed in Persian script or in Latin letters, with the same intent and the same expected results.

Speaker a second community speakerEvidence transcript

Used byrequirement DEV-INT-11glossary term Transliterated queries

  • Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…
StageNot in docsD2-C967

The presenter of the non-Latin-script talk said that in competitive Persian, Turkish and Arabic searches, bought backlinks and paid editorial content still visibly influence rankings and are widespread (an observation; no data was shown); the presenter stressed this described the situation and was not a recommendation.

Speaker a second community speakerEvidence transcript

StageConfirmed by docsD2-C971

Google search features often launch in some languages or countries first and expand later (the presenter's example was site names, launched in several languages and then extended to all languages in 2023), so comparisons of performance across languages and countries should not assume a feature is live everywhere at once.

Speaker a second community speakerEvidence transcript

StageConsistent with docsD2-C972

In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English sources; the presenter explained that Persian content on a topic is often thinner, of lower quality or less relevant, so the systems retrieve from languages with better content, which the presenter called cross-lingual retrieval.

Speaker a second community speakerEvidence transcript

  • Extends D2-C603 Day 2: Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is…
StageConfirmed by docsD2-C977

The presenter of the non-Latin-script talk recalled that Google officially treats buying links for ranking as spam (the year of the announcement was not clear in the recording).

Speaker a second community speakerEvidence transcript

DocsSourceD2-C978

Google's October 2023 spam update post (4 October 2023) says the update improved coverage in many languages and spam types, cleaning up spam reported in Turkish, Vietnamese, Indonesian, Hindi, Chinese and other languages, particularly cloaking, hacked, auto-generated and scraped spam.

Publisher Search Central blog (4 October 2023)

AnalysisD2-C979

The October 2023 spam update cited in the non-Latin-script talk names neither Persian nor Arabic (only 'other languages') and lists cloaking, hacked, auto-generated and scraped spam, not link spam, so it does not show that the bought links the presenter saw working in Persian or Arabic search were addressed.

Author Ibrahim Anjro

DocsSourceD2-C981

Google's post on multilingual searches (8 September 2023) says that, because of typing difficulty on some keyboards, a person in India might search in Hindi using Latin rather than Devanagari characters and want and receive Hindi results written either way.

Publisher Search Central blog (8 September 2023)

Used byrequirement DEV-INT-11glossary term Transliterated queries

14:45 · Poster session H: Indexing 3

StageD2-C974

Google added a poster session to the Deep Dive because some topics need a longer, more one-to-one discussion than a seven-minute lightning talk allows.

Speaker Community speakersEvidence transcript

StageD2-C976

One Day 2 poster was presented by a Googler, about bringing Google's tools into WordPress sites.

Speaker Community speakersEvidence transcript

15:30 · Calculating (some) signals 33

StageConsistent with docsD2-C648

Among the many signals Google calculates during indexing, the ones singled out as having large effects on search results were country, language, freshness, SafeSearch and spam.

Speaker GoogleEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
StageConsistent with docsD2-C649

Country and language are among Google's most important signals and have been used since Google's early days.

“Both of them are extremely, extremely important”

Speaker GoogleEvidence transcript

  • Extended by D3-C079 Day 3: To order candidates at retrieval, Google uses signals collected during indexing, and the first two are…
  • Extended by D3-C082 Day 3: Country is the second retrieval signal: a user searching from Switzerland wants cheese from Switzerland, not…
StageConfirmed by docsD2-C654

In ranking, country and language signals help Google serve users the right content for their country and language.

Speaker GoogleEvidence transcript

  • Extended by D3-C006 Day 3: Google's first step in understanding almost any query is to detect its language, which tells Google roughly…
  • Extended by D3-C080 Day 3: At retrieval, Google tries to match results to the user's language wherever possible: someone searching in…
AnalysisD2-C655

A site serving a smaller country in a big language, such as British English or Swiss German, should make that country unmistakable (country-code domain or region-specific hreflang, local address, prices and currency) so its pages can benefit from index selection's country balancing instead of competing with the much larger US or German content pool.

Author Ibrahim Anjro

StageConfirmed by docsD2-C656

Freshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking event hits a city, such as possible closure of Barcelona's airport, users want really fresh results, not results from two weeks ago.

Speaker GoogleEvidence transcript

  • Extends D1-C045 Day 1: Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour…
  • Extended by D3-C487 Day 3: Google's ranking systems guide describes 'query deserves freshness' systems that show fresher content where…
StageConfirmed by docsD2-C660

Google uses multiple systems to protect users from unexpected explicit results: its algorithms detect whether a user is looking for explicit content and rank results accordingly, so the aim is not to never show explicit results but to avoid showing them to users who are not looking for them.

Speaker GoogleEvidence transcript

StageNot in docsD2-C662

Google calculates SafeSearch signals during indexing rather than in ranking, because ranking happens online with no time for such calculations, while indexing has the processing power.

“In ranking, everything has to happen online, and there's just not enough time to calculate those things.”

Speaker GoogleEvidence transcript

AnalysisD2-C665

Because explicit content on its own does not lower a page's chance of being indexed, the practical risk for a site that mixes explicit and general content is SafeSearch classification, which looks at the whole page and its links; keep explicit pages on a separate domain or subdomain, as Google advises.

Author Ibrahim Anjro

Used byrequirement DEV-IDX-12

StageConsistent with docsD2-C668

Google uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is highly accurate.

Speaker GoogleEvidence transcript

  • Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
  • Extended by D3-C205 Day 3: Google's quality talk said AI has fundamentally changed how Google builds spam updates, letting it evaluate…
  • Extended by D3-C206 Day 3: Google's quality talk said AI lets Google catch new types of spam and catch more of it.
StageNot in docsD2-C669

SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.

“built on top of, well, nowadays, Gemini, and fine-tuned to that specific purpose of finding spam”

Speaker GoogleEvidence transcript

Used byglossary term SpamBrain

  • Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageConfirmed by docsD2-C670

SpamBrain is central to Google's spam-fighting efforts and has been improved many times since its launch.

Speaker GoogleEvidence transcript

Things
  • Extended by D3-C287 Day 3: Google's spam updates page says its automated spam detection systems run constantly, and a notable…
AnalysisD2-C672

The SpamBrain launch year mentioned on stage, with hesitation, was 2022, which does not match Google's documented 2018; 2022 is the year of the improvements described in Google's 2022 webspam report, so cite 2018 as the launch year.

Author Ibrahim Anjro

Things
AnalysisD2-C675

The '5 times more spam sites' figure said on stage matches Google's 2022 webspam report, but the report compares 2022 with 2021 (and gives 200 times since launch), not SpamBrain with earlier algorithms; quote the documented comparison.

Author Ibrahim Anjro

Things
StageD2-C679

Combing Google's documentation for signals is not the best use of an SEO's time; creating content that users will like is a better one.

“If you are just focusing on actually creating the content that users will like, then that's probably a better use of your life.”

Speaker GoogleEvidence transcript

  • Repeats D1-C059 Day 1: The opening keynote closed with the advice to think about UEO, user engine optimisation, next to SEO and GEO…

15:40 · Deciding what goes in the index? 39

StageConsistent with docsD2-C680

Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.

“our index is immense. Like, immense. But it is a finite resource.”

Speaker GoogleEvidence transcript

  • Repeats D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeats D2-C344 Day 2: Google deduplicates pages because many sites have very many pages and Google's index does not have room for…
  • Repeated by D3-C256 Day 3: Google does not index every URL on the web; because it cannot index everything, it has to rank results better…
StageNot in docsD2-C681

Google gives two reasons for not indexing every URL it knows: most of them would not be useful to users, and including URLs in the index that users would never see would be an immense investment.

Speaker GoogleEvidence transcript

StageConsistent with docsD2-C682

Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.

Speaker GoogleEvidence transcript

Used byglossary term Index selection

  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
StageNot in docsD2-C685

Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

“the index selection system is going to be more forgiving when it sees a new URL from your site”

Speaker GoogleEvidence transcript

Used byrequirement DEV-URL-10

  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
AnalysisD2-C686

Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

Author Ibrahim Anjro

Used byrequirement DEV-URL-10

  • Extends D1-C095 Day 1: New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under…
StageNot in docsD2-C688

Index selection is the last step before documents enter Google's index.

Speaker GoogleEvidence transcript

Used byglossary term Index selection

  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
AnalysisD2-C692

The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.

Author Ibrahim Anjro

Used byrequirement DEV-INT-10

StageNot in docsD2-C694

News sites were given as the example of importance at work in index selection: they are generally very important on the web and their pages usually get indexed very fast ('indexed' is a likely but not certain reading of the recording).

Speaker GoogleEvidence transcript

StageConsistent with docsD2-C695

Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.

“focusing on the quality is the most reliable way to get stuff in the index”

Speaker GoogleEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
StageConsistent with docsD2-C696

Index selection applies the negative signals that immediately block indexing: noindex (the likely reading of one unclear word), expired unavailable_after dates, soft 404s, non-canonical duplicates, spam signals and other policies.

Speaker GoogleEvidence transcript

StageConsistent with docsD2-C697

Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).

Speaker GoogleEvidence transcript

  • Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
StageConsistent with docsD2-C698

Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.

Speaker GoogleEvidence transcript

Used byrequirement DEV-IDX-08glossary term unavailable_after

  • Extends D2-C100 Day 2: The unavailable_after rule lets a page drop out of search results after a set date and time, which suits…
AnalysisD2-C699

Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.

Author Ibrahim Anjro

Used byrequirement DEV-IDX-08

StageConsistent with docsD2-C700

Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

Speaker GoogleEvidence transcript

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D2-C334 Day 2: A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of…
StageConsistent with docsD2-C701

When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

Speaker GoogleEvidence transcript

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D2-C345 Day 2: Google's deduplication has three steps: identify and cluster duplicate web pages, pick representative URLs…
DocsSourceD2-C702

Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.

Publisher Google Search Central

AnalysisD2-C703

'Only canonicals end up in search results' as said on stage is a simplification: non-canonical duplicates are dropped from the index, but Google's documentation says an alternate from the same cluster can still be shown in some contexts, such as a mobile version to a mobile user.

Author Ibrahim Anjro

StageNot in docsD2-C706

'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

“The first one is kind of nastier.”

Speaker GoogleEvidence transcript

Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed

  • Extends D1-C097 Day 1: Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over…
  • Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
DocsSourceD2-C707

Google's Page indexing report help says a 'Discovered – currently not indexed' page was found but not crawled yet, typically because Google wanted to crawl it but expected the crawl to overload the site, so it rescheduled the crawl.

Publisher Google Search Console Help

Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed

AnalysisD2-C708

The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

Author Ibrahim Anjro

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
StageConsistent with docsD2-C709

Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.

Speaker GoogleEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
AnalysisD2-C710

Repeatedly submitting 'Discovered – currently not indexed' URLs does not change why they wait, because the status reflects a crawl-scheduling decision; raise the site's demonstrated quality instead, for example by improving or removing weak pages that are already indexed.

Author Ibrahim Anjro

StageNot in docsD2-C714

'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

“most of the time it is actually a quality issue”

Speaker GoogleEvidence transcript

Things

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
AnalysisD2-C716

Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

Author Ibrahim Anjro

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
StageConsistent with docsD2-C717

Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

Speaker GoogleEvidence transcript

Used byrequirement DEV-MON-03

  • Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…

16:15 · How does the index look like? 26

StageNot in docsD2-C721

Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.

“we don't have the full content of the page in our index”

Speaker GoogleEvidence transcript

  • Repeats D2-C318 Day 2: Google does not store the complete sentences or the full HTML of a page in the Search index, because large…
StageConsistent with docsD2-C722

Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.

Speaker GoogleEvidence transcript

  • Extends D1-C037 Day 1: For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as…
  • Repeated by D3-C076 Day 3: For retrieval, Google uses signals attached individually to each document in the index.
  • Extended by D3-C079 Day 3: To order candidates at retrieval, Google uses signals collected during indexing, and the first two are…
  • Extended by D3-C083 Day 3: Google called quality the most important of the signals used to order candidates at retrieval: a URL of high…
StageNot in docsD2-C724

The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.

“the snippet that you see was reconstructed from these tokens”

Speaker GoogleEvidence transcript

Things
  • Extended by D3-C316 Day 3: Google generates the parts of a text result, such as title link and snippet, from its understanding of the…
DocsSourceD2-C725

Google's snippet documentation says snippets are created automatically, primarily from the page content, to preview the part that best relates to the user's specific search, so one page can get different snippets for different searches; sometimes the meta description is used instead.

Publisher Google Search Central

Things

Used byrequirement DEV-HTM-03

StageConsistent with docsD2-C726

AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.

Speaker GoogleEvidence transcript

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
  • Extended by D3-C325 Day 3: AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to…
  • Repeated by D3-C702 Day 3: AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have…
StageConsistent with docsD2-C727

Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

Speaker GoogleEvidence transcript

  • Extends D1-C053 Day 1: Query fan-out means running several related searches at once to gather more results; a question about lawn…
  • Extends D1-C051 Day 1: Three reasons were given: generative AI features are built directly on the core ranking systems, query…
  • Extends D2-C073 Day 2: John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page…
  • Extends D1-C172 Day 1: Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further…
  • Extended by D3-C059 Day 3: Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding…
DocsSourceD2-C728

Google's robots meta tag specification says the nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI Mode, and max-snippet also limits how much of it may be used that way.

Publisher Google Search Central

Used byrequirement DEV-IDX-05glossary term max-snippet

DocsSourceD2-C729

Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several data sources, which for AI Mode include the Knowledge Graph and shopping data as well as web content (AI Mode launch post, March 2025).

Publisher Google Search Central, Google blog (5 March 2025)

Used byglossary term Query fan-out

  • Extends D1-C053 Day 1: Query fan-out means running several related searches at once to gather more results; a question about lawn…
  • Extended by D3-C061 Day 3: Google tries to make fan-out queries distinct from each other for better coverage, avoiding asking the same…
AnalysisD2-C730

There is no separate AI index to optimise for: when a page never shows up as a source in AI Overviews or AI Mode, first check that it is indexed, that no nosnippet rule blocks its snippet and that the site is not excluded in Search Console's generative AI setting, the eligibility conditions Google lists.

Author Ibrahim Anjro

Used byrequirement DEV-AIF-01

AnalysisD2-C731

Treat nosnippet, data-nosnippet and max-snippet as AI visibility settings too: Google lists them as the controls for content in AI features, and if AI answers are built from index snippets, as Google said on stage, a blocked or shortened snippet leaves AI Overviews and AI Mode less to use.

Author Ibrahim Anjro

StageConsistent with docsD2-C732

To find relevant pages, Google's serving system relies on posting lists, a long-established information retrieval structure taught in computer science courses, because simply asking for every page that contains a word would not work.

Speaker GoogleEvidence transcript

Used byglossary term Posting list

  • Extended by D2-C824 Day 2: Google said posting lists, which Google's serving system uses to find the pages that contain a query's words…
StageNot in docsD2-C824

Google said posting lists, which Google's serving system uses to find the pages that contain a query's words, are not new: they are at least 60 years old (as of 2026).

Speaker GoogleEvidence transcript

  • Extends D2-C732 Day 2: To find relevant pages, Google's serving system relies on posting lists, a long-established information…
DocsSourceD2-C734

Google's How Search Works site describes the Search index as like the index at the back of a book, with an entry for every word seen on every webpage Google indexes.

“It’s like the index in the back of a book - with an entry for every word seen on every webpage we index.”

Publisher Google Search (How Search Works)

AnalysisD2-C735

The speaker's 'most of the tokens' is more precise than the public explainer's 'an entry for every word': the explainer simplifies, and the speaker's self-correction suggests some tokens get no posting list, though the speaker did not say which.

Author Ibrahim Anjro

StageNot in docsD2-C736

In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.

Speaker GoogleEvidence transcript

Used byglossary term Posting list

  • Extended by D3-C075 Day 3: At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the…
StageNot in docsD2-C737

A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

Speaker GoogleEvidence transcript

  • Repeats D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageNot in docsD2-C738

At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.

Speaker GoogleEvidence transcript

  • Extended by D3-C075 Day 3: At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the…
StageConsistent with docsD2-C740

Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.

“just like you build a posting list, you can also build a vector space”

Speaker GoogleEvidence transcript

Used byglossary term Vector embeddings

  • Extended by D3-C077 Day 3: The first condition for retrieving a document is that the query's words, or its concepts in the case of…
  • Extended by D3-C311 Day 3: Google handles an image used as a search query much like a text query interpreted as an embedding: the image…
StageConsistent with docsD2-C741

In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the user's query decides which documents are returned.

Speaker GoogleEvidence transcript

Used byglossary term Vector embeddings

  • Extends D1-C129 Day 1: Google's guide says creating separate content for every variation of how people might search, including…
AnalysisD2-C742

Write for both retrieval routes without keyword stuffing: name the page's subject in the plain words people search with, because posting lists match the words on the page, and explain the topic fully, because embedding retrieval matches meaning, so every keyword variation is unnecessary.

Author Ibrahim Anjro

StageNot in docsD2-C744

Google said a vector space also holds embeddings for associations the web makes with a page, such as what is known about its author; most of them sit far from typical queries, and a query that names the association may move closer to them.

Speaker GoogleEvidence transcript

StageD2-C746

Google said, hedging with 'I think', that because retrieval is still based on content, the content mantra Google started about 25 years ago (as of 2026) still stands.

“this mantra that we started 25 years ago or whatever still stands, whether we like it or not”

Speaker GoogleEvidence transcript

16:20 · Google Trends 71

AnalysisD2-C753

By the author's arithmetic, more than five trillion Google searches a year means on average more than 13.7 billion searches a day, or roughly 160,000 a second.

Author Ibrahim Anjro

StageConsistent with docsD2-C756

Google said Google Trends data is almost real time, reaching up to a few minutes before the moment of viewing.

“we have data that goes all the way back from 2004 and up until three minutes ago”

Speaker Omri WeismanEvidence transcript

StageD2-C759

Google argued that people search authentically, for health conditions, career troubles or home repairs, so aggregated searches reflect what a country or the world cares about better than the curated picture on social media.

Speaker Omri WeismanEvidence transcript

StageConsistent with docsD2-C761

The Google Trends Explore page, which Google called the heart of Trends, shows search interest in a query or topic and how it changes over time.

Speaker Omri WeismanEvidence transcript

  • Extended by D3-C015 Day 3: Google said the difference between words and entities can be seen in Google Trends, where a term can be…
StageNot in docsD2-C763

The legacy Google Trends Explore page is still available but will be retired: Google is adding its features to the new Explore page until everyone can move to the new one (no date was given).

“the legacy Explore page, which is still available, but not for a long time”

Speaker Omri WeismanEvidence transcript

DocsSourceD2-C767

Google's August 2024 blog post said Trending now was available in 125 countries, with regional trends in 40 of them; the current Trends Help page says 100+ countries and regions.

Publisher Google blog (14 August 2024), Google Trends Help

StageNot in docsD2-C770

Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of Nano Banana (the image editing model in the Gemini app), as people searched for both Gemini and Nano Banana.

Speaker Omri WeismanEvidence transcript

  • Extends D1-C162 Day 1: Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the…
AnalysisD2-C772

On stage the Nano Banana launch was placed in September 2025, but Google announced it on 26 August 2025, so the September 2025 surge in Gemini search interest described in the talk came in the weeks after the launch.

Author Ibrahim Anjro

StageNot in docsD2-C777

An early-2012 worldwide spike in Google Trends search interest for 'ai' was caused by the global hit song 'Ai Se Eu Te Pego', whose title contains the Portuguese word 'ai', not by interest in artificial intelligence.

Speaker Omri WeismanEvidence transcript

StageConfirmed by docsD2-C784

Google Trends shows search interest, not search volume: each point on its 0 to 100 scale represents the share of searches for the term out of all searches.

“Search interest is not search volume.”

Speaker Omri WeismanEvidence transcript

Used byglossary term Search interest

StageNot in docsD2-C786

Google said its internal numbers show the actual volume of ski-related searches is going up, even though ski's share of all searches in Google Trends has declined since 2004.

“we can tell you from our own internal numbers that we're seeing on Google Trends: the actual volume is, in fact, going up.”

Speaker Omri WeismanEvidence transcript

StageConsistent with docsD2-C794

The 'Commonly searched queries' section of the new Google Trends Explore page, called related queries in the legacy Explore page, lists queries that people typed in the same search sessions as the entered term.

Speaker Omri WeismanEvidence transcript

StageConsistent with docsD2-C804

For ideation, Google recommended Google Trends' Trending now, where a real-world event shows up within about ten minutes as people hear the news and search for it.

Speaker Omri WeismanEvidence transcript

StageConfirmed by docsD2-C807

Google Trends' Suggest search terms picks the search terms or topics to compare from a plain-language request, which helps in a field you know little about, for example the most popular jazz singers in Japan.

Speaker Omri WeismanEvidence transcript

StageConsistent with docsD2-C818

Trends TV, at trends.google.com/tv, is a little-known Google Trends dashboard of what is trending on Google right now, meant for an office screen or as a screensaver.

Speaker Omri WeismanEvidence transcript

StageConsistent with docsD2-C819

Google has published a YouTube series of about eight episodes on using Google Trends for research, journalism, marketing and SEO.

Speaker Omri WeismanEvidence transcript

Day 3: Serving: Ranking, Search Console, and Performance 713

10:15 · Welcome to serving and ranking day! 1

SlideConsistent with docsD3-C001

Google's opening slide for serving day showed serving as the third stage after crawling and indexing, and drew Google's serving infrastructure as query understanding and retrieval leading into the index, then ranking and search features leading back to the user, for the example query 'Where to eat jamon'.

Speaker GoogleEvidence slide photo

  • Repeated by D3-C321 Day 3: Google's serving diagram shows the query passing through query understanding and retrieval to the index, then…

10:25 · Making sense of users' queries 95

StageD3-C002

Google's query understanding talk set out to show how the changes Google makes to users' queries relate to what a website can provide.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C003

A Google speaker said that a messy query, made of a hotel location copied from a website plus a question about Italian food, was figured out by Google's systems, which came up with answers for it.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C004

Google placed query understanding in the traditional part of Search, the part where answers are looked up in an index and then presented and ranked.

Speaker John MuellerEvidence transcript

  • Extended by D3-C304 Day 3: Which kinds of results Google shows for a query is decided by query understanding, which tries to predict the…
StageD3-C005

Google said it has no name for traditional, non-AI Search, only for its Search generative AI features, and that the two play together but are easier to understand separately.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C006

Google's first step in understanding almost any query is to detect its language, which tells Google roughly what content the user wants: a query in German suggests German content, a query in English English content.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-07glossary term Query understanding

  • Extends D2-C654 Day 2: In ranking, country and language signals help Google serve users the right content for their country and…
StageConsistent with docsD3-C007

Query language detection works poorly when someone searches only for a brand name, such as Facebook or Google, because the query does not show which language the user wants results in.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-03

  • Extends D1-C215 Day 1: Serving starts with interpreting the query, which includes cleaning it up, detecting its language and…
StageConsistent with docsD3-C008

For brand-only queries Google falls back on other information, such as the user's location and browser settings, to work out the language of the results.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-03

StageConsistent with docsD3-C009

Because a brand-only query does not reveal the user's language, Google sometimes struggles to show the right language version of a page for brand searches.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-03

StageNot in docsD3-C011

Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.

Speaker John MuellerEvidence transcript

  • Extends D2-C320 Day 2: Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings…
StageNot in docsD3-C012

After detecting the query language and separating the words, Google removes words it thinks matter little to the query, such as 'a' and 'of', known as stop words.

Speaker John MuellerEvidence transcript

Used byglossary terms Query understanding, Stop words

StageNot in docsD3-C013

Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

Speaker John MuellerEvidence transcript

  • Extends D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
  • Extends D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
StageNot in docsD3-C014

For some searches the stop words are important, and Google then tries to recognise the whole phrase, stop words included, as an entity; in indexing, such words are indexed together.

Speaker John MuellerEvidence transcript

Used byglossary term Stop words

StageConsistent with docsD3-C015

Google said the difference between words and entities can be seen in Google Trends, where a term can be searched as words or as an entity (Trends calls these a search term and a topic).

Speaker John MuellerEvidence transcript

  • Extends D2-C761 Day 2: The Google Trends Explore page, which Google called the heart of Trends, shows search interest in a query or…
StageNot in docsD3-C016

When a query names an entity, Google treats it as a request for that entity rather than as a collection of separate words.

Speaker John MuellerEvidence transcript

Used byglossary term Query understanding

StageNot in docsD3-C017

In ranking, Google can match a query's words or its entity, and, the speaker said with a 'probably', mixes both to some degree.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C019

In Google's example, the query word 'photograph' could be expanded to 'image', 'picture' or 'photo', but one German candidate had to be dropped because the similar German word means 'photographer', so expansions are language-specific.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C020

According to Google's ranking teams, the synonym system is one of the most important parts of how Google handles a query.

“they say that the synonym system is one of the most important parts”

Speaker John MuellerEvidence transcript

StageConfirmed by docsD3-C021

Synonyms matter because users often phrase a query differently from the way content is indexed; synonym expansion makes the words Google looks up findable in its index.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C022

For technical terms, Google's synonym swapping can return either a technical page or a simplified page for the same query.

Speaker John MuellerEvidence transcript

AnalysisD3-C024

Write each page in the words its audience uses and drop synonym lists added for search engines: Google adds synonyms at query time, and blocks of keyword variants read as keyword stuffing, which Google's spam policies prohibit.

Author Ibrahim Anjro

Used byrequirement DEV-INT-11

StageNot in docsD3-C026

Google suggested a test: a search that lists a term's synonyms joined with OR will probably return results very similar to the plain query, because Google adds the synonyms itself (part of the sentence is unclear in the recording).

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C027

Googlers write queries in brackets, for example [spicy cheese store near me], so when reporting a search problem to a Googler, putting the query in brackets shows that it is a query someone searched.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C029

Internally, a short query can turn into a much longer rewritten query, because Google adds entities, synonyms and other information before looking it up.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C030

In Google's rewrite example, [fried chicken place in Barcelona] keeps 'fried' and 'chicken' as two words, may add an entity for fried chicken, and replaces 'place' with alternatives such as 'area', 'location' or 'restaurant'.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C031

Some alternatives in a rewritten query make no sense, such as 'fried chicken area in Barcelona', which is harmless because few indexed pages match them.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C032

In a rewritten query, a place name such as Barcelona can stay a word or be swapped for an entity, possibly a more specific location.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C033

For a query containing 'near me', Google understands that the user wants results near their location, not pages containing the words 'near me', and rewrites the query accordingly.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C034

Some synonyms are contextual and depend on the rest of the query: 'GM' probably means General Motors in [GM car], general manager in [GM restaurants] and genetically modified in [GM barley].

Speaker John MuellerEvidence transcript

StageNot in docsD3-C036

Google's synonyms need not be synonyms linguistically: words people use interchangeably are treated as synonyms, because the aim is to find the right content in the index.

“we don't need to be technically accurate”

Speaker John MuellerEvidence transcript

Used byglossary term Synonyms and siblings

StageNot in docsD3-C037

When Google highlights a word in its results that looks like the wrong synonym, the reason is probably that many people use the two words interchangeably.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C038

Besides synonyms, Google detects 'siblings', words in the same category that are related but not interchangeable, such as Canon and Nikon.

Speaker John MuellerEvidence transcript

Used byglossary term Synonyms and siblings

StageNot in docsD3-C039

Because Canon and Nikon are siblings rather than synonyms, a search for [Canon camera] should not show Nikon cameras.

Speaker John MuellerEvidence transcript

Used byglossary term Synonyms and siblings

StageNot in docsD3-C040

Google learns synonyms and siblings from search behaviour: words people search with in the same way become synonyms, while frequent comparison queries mark words as not interchangeable (the end of the sentence is unclear in the recording).

Speaker John MuellerEvidence transcript

StageNot in docsD3-C043

To check whether Google treats two terms as the same, search for each: [Iberico ham] and [jamón ibérico] both brought up the same entity, so Google understands them as the same thing.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C044

When Google already matches a term's variants, there is no need to add them to pages artificially; mention another name only where visitors might not understand otherwise.

“I wouldn't artificially just stuff those variations in there.”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-11

StageNot in docsD3-C045

At the time of the talk, a search for [Spanish cured ham] brought up mainly a Wikipedia page, showing that Google does not automatically equate that phrase with Iberico ham.

Speaker John MuellerEvidence transcript

AnalysisD3-C047

Before writing, search for each variant of a key term: if the variants bring up the same entity or near-identical results, use the term your audience uses; if one variant returns unrelated results, write that variant on the page.

Author Ibrahim Anjro

StageConfirmed by docsD3-C048

Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a German 'ü' written as 'ü', as 'ue' or left out.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-11

  • Extends D2-C600 Day 2: Google usually understands a query word whether it is written with or without diacritics (accents).
StageNot in docsD3-C051

Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-11

  • Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…
StageD3-C054

Google advised double-checking any word you are unsure about by searching for it on Google and looking at what comes up.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C056

Google's generative AI features in Search build on the traditional ways of searching, so query understanding also flows into AI Overviews and AI Mode.

Speaker John MuellerEvidence transcript

  • Repeats D1-C051 Day 1: Three reasons were given: generative AI features are built directly on the core ranking systems, query…
SlideConsistent with docsD3-C057

Google's slide on how LLM features with grounding generally work showed a query going to both the search engine and an LLM, the search engine's results going to the LLM, the LLM generating fan-out queries that go back to the search engine, and the LLM returning answers with links.

“How LLM features with grounding generally work”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageConsistent with docsD3-C059

Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding how normal queries work explains fan-out queries too.

Speaker John MuellerEvidence transcript

  • Extends D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
  • Extends D1-C172 Day 1: Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further…
StageNot in docsD3-C060

Google said a short video showing how one query is expanded into several fan-out queries has been in its documentation for a while.

Speaker John MuellerEvidence transcript

StageConsistent with docsD3-C061

Google tries to make fan-out queries distinct from each other for better coverage, avoiding asking the same question several times, which would return the same answers.

Speaker John MuellerEvidence transcript

  • Extends D2-C729 Day 2: Google says query fan-out in AI Overviews and AI Mode issues related searches across subtopics and several…
  • Extends D1-C225 Day 1: Google said query fan-out is nothing new: it fires, for example, ten different searches in the background…
StageNot in docsD3-C063

Google said fan-out queries are not added to Search Console, because Google considers them part of its infrastructure.

“Fan-out queries are not added in Search Console, because they're basically a part of our infrastructure.”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-AIF-03glossary term Query fan-out

  • Extends D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
StageConsistent with docsD3-C065

Because every system runs fan-out differently, Google advised understanding that fan-out happens but not overfocusing on individual fan-out queries or on how to rank for them.

“I would not overfocus on which individual fan-out queries happen and how can I rank for those”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-AIF-03

  • Extends D1-C129 Day 1: Google's guide says creating separate content for every variation of how people might search, including…
  • Repeated by D3-C704 Day 3: Google warned that shiny new things, such as trying to work out fan-out queries, distract from the real…
SlideConsistent with docsD3-C067

Google's summary slide on query understanding told site owners not to worry about typos and plurals, which Google rewrites automatically.

“Don't worry about typos & plurals!”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

SlideNot in docsD3-C069

Google's summary slide on query understanding said Google's synonyms are not always language-based; on stage this was explained as words people use interchangeably counting as synonyms even when they are not synonyms linguistically.

“Google's synonyms aren't always language-based”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

SlideNot in docsD3-C070

Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.

Speaker John MuellerEvidence slide photo, transcript

  • Repeats D2-C320 Day 2: Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings…
SlideD3-C072

Google's summary slide on query understanding concluded that all this query expansion gives site owners many opportunities for their content to be found and shown.

“There are many opportunities to find and show your content.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

StageNot in docsD3-C075

At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.

Speaker Gary IllyesEvidence transcript

Used byglossary term Retrieval

  • Extends D2-C736 Day 2: In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked…
  • Extends D2-C738 Day 2: At retrieval, Google looks up the posting lists of the query words that are actually important rather than of…
StageConsistent with docsD3-C077

The first condition for retrieving a document is that the query's words, or its concepts in the case of vectors or embeddings, are in the document or related to it.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C740 Day 2: Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are…
StageNot in docsD3-C078

Because a query like [best fried chicken ever] can match millions of pages, Google already orders the candidates during retrieval, before ranking starts.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD3-C079

To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-INT-07glossary term Retrieval

  • Extends D2-C722 Day 2: Each document in Google's index has pretty much all the signals calculated for it attached, for example…
  • Extends D2-C649 Day 2: Country and language are among Google's most important signals and have been used since Google's early days.
StageConfirmed by docsD3-C080

At retrieval, Google tries to match results to the user's language wherever possible: someone searching in Spanish does not necessarily want results in Italian.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-INT-07

  • Extends D2-C654 Day 2: In ranking, country and language signals help Google serve users the right content for their country and…
StageConsistent with docsD3-C082

Country is the second retrieval signal: a user searching from Switzerland wants cheese from Switzerland, not from Germany, and a user in Spain is poorly served by results targeting a South American country.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-INT-07

  • Extends D2-C649 Day 2: Country and language are among Google's most important signals and have been used since Google's early days.
StageNot in docsD3-C083

Google called quality the most important of the signals used to order candidates at retrieval: a URL of high quality is more likely to be retrieved from the index for specific queries.

“if the quality of a URL is high, then it's more likely to be retrieved from the index for specific queries.”

Speaker Gary IllyesEvidence transcript

Used byglossary term Retrieval

  • Extends D2-C722 Day 2: Each document in Google's index has pretty much all the signals calculated for it attached, for example…
StageConsistent with docsD3-C084

The same quality signal that Google uses at retrieval is also used in ranking when results are served.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C131 Day 3: Google's quality talk said quality is one of the most important ranking signals, although Google uses…
DocsSourceD3-C087

Google's How Search Works page says understanding a query ranges from recognising and correcting spelling mistakes to a synonym system that finds relevant documents without the exact words searched, such as 'adjust laptop brightness' for 'change laptop brightness'.

Publisher Google Search (How Search Works)

DocsSourceD3-C088

Google's How Search Works page says the language models it builds to match a query's few words to the most useful content took over five years to develop and significantly improve results in over 30% of searches across languages.

Publisher Google Search (How Search Works)

DocsSourceD3-C089

Google's SEO Starter Guide advises anticipating the different words readers search with (some search for 'charcuterie', others for 'cheese board') but not worrying about every variation, because Google's language matching systems relate pages to queries without the exact terms.

Publisher Google Search Central

DocsSourceD3-C094

Google's Gemini API documentation on grounding with Google Search says the model automatically generates and runs one or more search queries when needed, and the response lists the search queries it executed.

Publisher Google AI for Developers (Gemini API docs)

AnalysisD3-C095

No Google documentation states that fan-out queries stay out of Search Console, but the docs fit the stage statement: the Performance report counts AI Overviews and AI Mode under the user's own queries, and the Generative AI reports have no query dimension.

Author Ibrahim Anjro

AnalysisD3-C096

The retrieval talk went further than Google's documentation, which says Search returns the highest-quality and most relevant results and lists quality among key ranking signals, but never describes quality as deciding which pages are retrieved from the index at all.

Author Ibrahim Anjro

10:45 · Lightning session K: Facets of quality 33

StageD3-C097

In Lightning session K (Facets of quality), a community speaker argued that SEO is for humans: write for people and the problems they have, rather than to please Googlebot.

“SEO is for humans”

Speaker Community speakersEvidence transcript

Things
StageNot in docsD3-C098

A community speaker recalled that around 2009 writers bent titles and phrases to include a keyword because their CMS counted keyword density, and said that hitting the ratio really worked at the time.

Speaker Community speakersEvidence transcript

DocsSourceD3-C100

Google's spam policies define keyword stuffing as filling a web page with keywords or numbers to manipulate rankings in Google Search, for example repeating the same words or phrases so often that it sounds unnatural.

“filling a web page with keywords or numbers in an attempt to manipulate rankings in Google Search results”

Publisher Google Search Central

DocsSourceD3-C101

Google's guide to people-first content lists writing to a particular word count, in the belief that Google prefers one, as a warning sign of search engine-first content, and says Google has no preferred word count.

“Google has a preferred word count? (No, we don't.)”

Publisher Google Search Central

StageD3-C102

Entering the Polish market late, against strong competition and with a small team, a community speaker's team could not be a 'department store' that covers everything, so it went really deep instead of going broad.

Speaker Community speakersEvidence transcript

DocsSourceD3-C103

Google's guide to people-first content lists producing lots of content on many different topics, in the hope that some of it performs well in search results, as a warning sign of search engine-first content.

Publisher Google Search Central

StageD3-C104

To choose what to cover in depth, a community speaker's team used Google Trends to find what people in their market were interested in and then gave them that content.

Speaker Community speakersEvidence transcript

  • Extends D2-C789 Day 2: Google presented three uses of Google Trends for content creation, marketing and SEO: keyword selection…
StageD3-C105

A community speaker reported that going deep instead of broad took the Polish edition of a large Spanish software-download website into the top three technology websites in Poland (period not stated).

Speaker Community speakersEvidence transcript

StageD3-C106

With a much smaller team, the Polish edition of a large Spanish software-download website reached double the organic traffic of the site's English edition, a community speaker reported (period and measure not stated).

Speaker Community speakersEvidence transcript

StageD3-C107

Because Googlebot is not a human, the speaker's team at the time thought more about Googlebot than about human readers, a community speaker said.

Speaker Community speakersEvidence transcript

Things
AnalysisD3-C109

The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.

Author Ibrahim Anjro

  • Extends D2-C318 Day 2: Google does not store the complete sentences or the full HTML of a page in the Search index, because large…
StageD3-C110

Because a bot that counts things was easy to work around, SEOs made the facade of their sites nice and fancy instead of focusing on what was inside, sometimes with no real shop behind the facade, a community speaker said.

Speaker Community speakersEvidence transcript

StageConsistent with docsD3-C116

Google's long-standing advice to write for people and give them what they want has become true in practice because Googlebot has become more and more human, a community speaker argued.

Speaker Community speakersEvidence transcript

Things
  • Repeats D1-C047 Day 1: Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so…
DocsSourceD3-C117

Google's guide to people-first content says its automated ranking systems are designed to prioritise helpful, reliable information created to benefit people, not content created to manipulate search engine rankings.

Publisher Google Search Central

StageD3-C119

In a community speaker's case study from the speaker's current company, a well-written, complex and complete article got 1,600,000 impressions, while an article written for a very specific audience got 46,000 impressions over the same period (data source and period length not stated).

Speaker Community speakersEvidence transcript

StageD3-C120

The niche article in a community speaker's case study was written for a very specific audience: a group of friends who wanted to play a game, had problems with a public server and looked for alternatives.

Speaker Community speakersEvidence transcript

StageD3-C121

Despite far fewer impressions (46,000 against 1,600,000), the niche article in a community speaker's case study got 28 times more clicks than the comprehensive article over the same period.

Speaker Community speakersEvidence transcript

StageD3-C122

In the same case study, visitors who landed on the niche article converted 10 times more than visitors of the comprehensive article, a community speaker reported.

Speaker Community speakersEvidence transcript

StageD3-C123

Conversions from content are real money that can be reported to stakeholders and shareholders, not an abstract KPI, a community speaker said.

“This is not some abstract KPI. This is real money that you can report to your stakeholders”

Speaker Community speakersEvidence transcript

StageD3-C124

The niche article in a community speaker's case study worked because the team stopped to analyse how people use the company's products and what the products bring them, then wrote about it.

Speaker Community speakersEvidence transcript

AnalysisD3-C126

The case-study figures were said on stage without a photographed slide, a data source or absolute click numbers, and together they imply a click-through rate nearly 1,000 times higher on the niche article, so use them as an illustration, not a benchmark.

Author Ibrahim Anjro

AnalysisD3-C127

For content planning, start briefs from a concrete problem a real group of users has with or around your product, such as a failing setup they need an alternative for, rather than from a broad keyword the page should rank for.

Author Ibrahim Anjro

StageD3-C128

A community speaker's exercise: step out of your shop, look at your site from across the street the way users do, and check what problems its pages solve.

“walk out of your shop. Stand on the other side of the street, look at the shop the way your users are looking at it”

Speaker Community speakersEvidence transcript

StageD3-C129

If you cannot say what problems your pages solve for users, or the answer is 'it depends', you are optimising for the search engine rather than for people, and that will not work for long, a community speaker warned.

Speaker Community speakersEvidence transcript

  • Repeats D1-C013 Day 1: Ecosystem principle 4, incentivise high-quality content: content made for Search will not be successful.

11:15 · How Google thinks about Quality 88

StageNot in docsD3-C130

Google's quality talk described quality as a ranking signal in its own right: a number calculated from many different parts (the word 'quality' is a repaired speech-to-text reading).

Speaker GoogleEvidence transcript

StageConsistent with docsD3-C131

Google's quality talk said quality is one of the most important ranking signals, although Google uses hundreds of ranking signals (the word 'quality' is a repaired speech-to-text reading).

Speaker GoogleEvidence transcript

  • Extends D3-C084 Day 3: The same quality signal that Google uses at retrieval is also used in ranking when results are served.
StageConfirmed by docsD3-C132

Google's quality talk defined content quality as the extent to which a human being put effort, originality, talent, skill and accuracy into creating the content.

“Quality is basically the extent to which a human being put effort, originality, talent, skill, and accuracy into creating content.”

Speaker GoogleEvidence transcript

Used byglossary term Content quality (effort, originality, talent or skill, accuracy)

SlideConfirmed by docsD3-C133

Google's slide said there is not one single ranking system and named spam detection systems, the reviews system, BERT, MUM, RankBrain, freshness systems, deduplication systems, crisis information systems and link analysis systems (PageRank).

“There’s not one single ranking system...”

Wording checked against the slide or recording

Speaker GoogleEvidence slide photo, transcript

AnalysisD3-C135

The quality talk's slide listed MUM among Google's ranking systems, while Google's ranking systems guide says MUM is not currently used for general ranking in Search; read the slide as a list of systems Google runs, not as proof that each one ranks every query.

Author Ibrahim Anjro

  • Extends D1-C130 Day 1: Google's ranking systems guide says MUM is not currently used for general ranking in Search, only for…
StageNot in docsD3-C136

Google's quality talk called PageRank the speaker's favourite ranking system, elegant in its time, but said Google does not really use it so much anymore.

“My favorite is probably PageRank, even though we don't really use them so much anymore.”

Speaker GoogleEvidence transcript

Things
AnalysisD3-C137

The remark that PageRank is not used so much anymore differs from Google's ranking systems guide, which says PageRank has evolved a lot and continues to be part of the core ranking systems; read it as PageRank weighing less among many signals, not as PageRank being switched off, so links still matter.

Author Ibrahim Anjro

Things
StageConsistent with docsD3-C142

Google's quality talk said ranking signals differ by result type: for web pages they include the text on the page, links and passages, while for news, probably, freshness, diversity and originality become more important.

Speaker GoogleEvidence transcript

  • Repeats D1-C045 Day 1: Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour…
AnalysisD3-C144

For content and development teams: plan improvements across content quality, internal links, page experience and freshness together; a project that pushes a single metric, such as keyword density or one link score, matches what Google said does not work.

Author Ibrahim Anjro

StageConsistent with docsD3-C145

Google's quality talk said users expect a richer set of results than ever, so Google brings different result types, such as web pages and videos, for different queries and intents.

Speaker GoogleEvidence transcript

StageConsistent with docsD3-C146

Google's quality talk said Google sometimes gives direct answers, and that for a query such as directions to the airport it should also show Maps features.

Speaker GoogleEvidence transcript

SlideConfirmed by docsD3-C149

Google's slide said that in 2023 Google ran 719,326 search quality tests, 124,942 side-by-side experiments and 16,871 live traffic experiments, and made 4,781 launches to Search.

Speaker GoogleEvidence 2 slide photos

  • Repeated by D3-C237 Day 3: Google's slide said Google made more than 4,700 launches to Search in 2023; the speaker rounded this to close…
AnalysisD3-C150

Two Day 3 slides headed 'In 2023' gave Google's testing figures in two versions: exact in the quality talk (719,326 search quality tests, 4,781 launches) and rounded in the updates talk (800,000+ tests, 4,700+ launches). Google's How Search Works page states the exact 719,326 and 4,781, so cite those with the year 2023; 800,000+ matches no figure on that page, although the page's three test counts (719,326 quality tests, 124,942 side-by-side and 16,871 live traffic experiments) add up to 861,139.

Author Ibrahim Anjro

DocsSourceD3-C155

Google's How Search Works page says a live traffic experiment first enables a feature for a small percentage of people, usually starting at 0.1%, and compares them with a control group on a very long list of metrics, such as clicks, number of queries, abandoned queries and time to click.

Publisher Google Search (How Search Works)

Used byglossary term Side-by-side and live traffic experiments

AnalysisD3-C157

Google made 4,781 launches to Search in 2023 while its status dashboard lists nine named ranking updates that year, so almost every change to Search goes unannounced; monitor rankings and traffic continuously, not only around announced core or spam updates.

Author Ibrahim Anjro

SlideConfirmed by docsD3-C158

Google's slide pointed to the Search Quality Rater Guidelines at goo.gle/raters-guidelines, with a QR code to the guidelines PDF, and the speaker said this is the exact document Google's raters use.

Speaker GoogleEvidence slide photo, transcript

SlideConfirmed by docsD3-C159

Google's slide said search quality raters are independent reviewers who help Google decide whether a change it is about to launch is good for users.

Speaker GoogleEvidence slide photo, transcript

Used byglossary term Search quality raters

SlideConfirmed by docsD3-C160

Google's slide said search quality raters cannot affect the rankings of individual sites.

“They can not affect individual site's rankings.”

Wording checked against the slide or recording

Speaker GoogleEvidence slide photo

Used byglossary term Search quality raters

  • Extended by D3-C690 Day 3: Quality raters cannot give a site a penalty or a manual action; their ratings are converted into labels that…
SlideConfirmed by docsD3-C162

The rater guidelines' table of contents shown on Google's slide covers page quality rating, YMYL topics, main and supplementary content, website and creator reputation, E-E-A-T and lowest quality pages.

Speaker GoogleEvidence slide photo

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConsistent with docsD3-C166

Google's quality talk showed five self-assessment questions (slide not photographed) that site owners can use to evaluate their own content, such as whether it is people-first or made mainly for search engines, and whether the author is an expert on the topic.

Speaker GoogleEvidence transcript

StageConfirmed by docsD3-C167

Google's quality talk said rater ratings are used to measure how effectively search engines deliver helpful content to people around the world.

Speaker GoogleEvidence transcript

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConsistent with docsD3-C168

Google's quality talk stressed that nothing in the Search Quality Rater Guidelines is itself a ranking factor; the guidelines explain how Google thinks about quality.

Speaker GoogleEvidence transcript, slide photo

Used byglossary term Search quality raters

StageConfirmed by docsD3-C170

Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.

Speaker GoogleEvidence transcript

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
  • Extends D2-C311 Day 2: Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks…
  • Extended by D3-C265 Day 3: Google sometimes couples a core update with an update to its Search Quality Rater Guidelines.
AnalysisD3-C173

The talk's originality test ('are you adding more value') matches Google's helpful-content question on content that draws on other sources, while the rater guidelines ask raters, where other websites have similar content, to consider whether the page is the original source; aim to be the original source, and add substantial value where you build on others.

Author Ibrahim Anjro

DocsSourceD3-C176

Google's helpful-content guidance lists effort, originality, talent or skill and accuracy as the main-content attributes search quality raters are trained to evaluate, and says using generative AI to produce large amounts of text without manual oversight or curation represents little to no effort.

Publisher Google Search Central

Used byrequirement DEV-SPM-02glossary term Content quality (effort, originality, talent or skill, accuracy)

StageD3-C177

Google's quality talk contrasted a commodity article, 'Top 10 things to consider when buying running shoes', with a non-commodity one, a wear pattern analysis of why certain running shoes collapse after 400 miles (slide not photographed).

Speaker GoogleEvidence transcript

StageConfirmed by docsD3-C179

Google's quality talk said non-commodity content offers a unique, experienced take, and advised writing content that readers will find very helpful and reliable.

Speaker GoogleEvidence transcript

Used byglossary term Commodity content

  • Extends D1-C055 Day 1: Myth: build content for every possible consumer need. Google's answer: prioritise unique perspectives…
DocsSourceD3-C180

Google's guide to optimizing for generative AI features contrasts commodity content, such as '7 Tips for First-Time Homebuyers', with non-commodity content that gives unique expert or experienced takes beyond common knowledge, such as a first-hand account of waiving a home inspection.

Publisher Google Search Central

Used byglossary term Commodity content

StageD3-C181

Google's quality talk said its article examples were illustrations, not a recipe: writing exactly that article is no guarantee that it will work.

Speaker GoogleEvidence transcript

StageD3-C183

Google's quality talk said AI is now everywhere and available to everyone, and that opinions differ on how the broad use of AI has affected search results.

Speaker GoogleEvidence transcript

StageD3-C184

Google's quality talk presented its AI-content examples as the speaker's personal opinion, not as the position of Google's search quality systems.

Speaker GoogleEvidence transcript

StageD3-C186

Google's quality talk asked why users would read AI-generated content when they could ask an LLM the same questions directly.

Speaker GoogleEvidence transcript

StageConfirmed by docsD3-C190

Google's quality talk said quality problems should be treated as quality issues, not as AI versus human content, because a lot of good AI-assisted or AI-written content exists.

“it's very important to think about these quality issues as quality issues and not AI versus human content.”

Speaker GoogleEvidence transcript

Used byrequirement DEV-SPM-05

  • Repeats D1-C210 Day 1: Google said it is not really trying to tell AI-written from human-written content, because it cares more…
  • Repeated by D3-C697 Day 3: Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem…
StageD3-C192

Google's quality talk said good content can be created for literally any vertical, answering site owners who say they do not know how to become experts or make their content stand out.

Speaker GoogleEvidence transcript

StageD3-C193

Google's quality talk showed, as an example of strong content, a running-shoe review site that cut the shoes apart, examined the upper and the foam density under a microscope, took high-quality images, built a test rig and created its own metrics to compare shoes (site not named, slide not photographed).

Speaker GoogleEvidence transcript

StageD3-C194

Google's quality talk said the speaker believes content in any vertical can go very far when its creators invest in their own knowledge, effort and expertise.

Speaker GoogleEvidence transcript

DocsSourceD3-C195

Google's helpful-content guidance says product reviews build trust when readers understand how many products were tested, what the test results were and how the tests were conducted, with evidence of the work such as photographs.

Publisher Google Search Central

AnalysisD3-C196

For content teams: turn reviews into non-commodity content by testing the products yourselves, publishing your own photos and measurements, explaining the test method and defining comparable metrics across products; a list of common buying tips is the commodity version.

Author Ibrahim Anjro

StageNot in docsD3-C198

Google's quality talk said a big portion of the new pages Google discovers every day is spam, adding that the speaker did not know of Google ever publishing the percentage.

Speaker GoogleEvidence transcript

StageConfirmed by docsD3-C201

Google's quality talk said Google had released four spam updates in 2026 by the time of the talk on 2 October, with one still rolling out.

Speaker GoogleEvidence transcript

DocsSourceD3-C202

Google's Search Status Dashboard says the September 2026 spam update began on 24 September 2026, applies globally to all languages and may take up to two weeks to roll out.

Publisher Google Search Status Dashboard (Ranking incident history), Google Search Status Dashboard (September 2026 spam update)

StageD3-C204

Google's quality talk said evaluating the quality of the web at scale is messy and hard, and that Google's work on it is never done.

Speaker GoogleEvidence transcript

StageConsistent with docsD3-C205

Google's quality talk said AI has fundamentally changed how Google builds spam updates, letting it evaluate many more candidates and drastically increasing its velocity, so it launches faster with more impact.

Speaker GoogleEvidence transcript

  • Extends D2-C668 Day 2: Google uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is…
StageConfirmed by docsD3-C207

Google's quality talk listed recent updates to Google's spam policies: back button hijacking, scaled content abuse and manipulating AI responses.

Speaker GoogleEvidence transcript

Used byrequirements DEV-SPM-01, DEV-SPM-02

  • Extended by D3-C284 Day 3: Reviewing the spam slide, the speaker said link spam would probably be replaced with scaled content abuse as…
StageConsistent with docsD3-C208

Google's quality talk said back button hijacking was added after many user complaints about pressing the back button and landing on an auto-generated page, or staying on the same site after pressing it twice.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SPM-01glossary term Back button hijacking

DocsSourceD3-C209

Google announced on 13 April 2026 that back button hijacking is an explicit violation of its malicious practices spam policy, with enforcement from 15 June 2026; pages that do it may get manual spam actions or automated demotions.

Publisher Search Central blog (13 April 2026), Google Search Central

Used byrequirement DEV-SPM-01glossary term Back button hijacking

StageConfirmed by docsD3-C212

Google's quality talk said Google clarified that traditional spam techniques aimed at manipulating AI responses also violate its spam policies, and that Google can take action against them.

Speaker GoogleEvidence transcript

  • Extends D1-C129 Day 1: Google's guide says creating separate content for every variation of how people might search, including…
AnalysisD3-C217

On 1 October 2026, during the event, Google updated its generative AI content guide with rater-guideline material to keep its documentation in sync with its developer-event presentations, and the helpful-content page, updated the same day, carries the effort, originality, talent or skill and accuracy definitions from the quality talk; cite those pages rather than the talk.

Author Ibrahim Anjro

11:45 · Uncovering Trustworthy Experiences on Discover 13

SlideConsistent with docsD3-C218

Google's Discover slide showed an eligibility gate: content with a policy violation goes to a policy compliance check and low-quality content to an E-E-A-T quality check, and both can end with the content filtered out, while safe and trusted content becomes eligible for Discover.

Speaker GoogleEvidence slide photo

SlideConsistent with docsD3-C221

Google's Discover slide said that for eligible content image quality decides click-through: a small thumbnail lowers the click-through rate and a large, high-resolution image raises it.

Speaker GoogleEvidence slide photo

Things

Used byrequirement DEV-IMG-07

SlideConfirmed by docsD3-C222

Google's Discover slide defined E-E-A-T: experience means first-hand, personal experience with the topic, expertise the necessary knowledge or skill, and authoritativeness the extent to which the creator or website is known as a go-to source for the topic.

Speaker GoogleEvidence slide photo

Used byglossary term E-E-A-T

SlideConfirmed by docsD3-C223

Google's Discover slide said trust comes first: untrustworthy pages have low E-E-A-T however experienced, expert or authoritative they seem.

“untrustworthy pages will have low E-E-A-T no matter how Experienced, Expert or Authoritative they may seem”

Wording checked against the slide or recording

Speaker GoogleEvidence slide photo

Used byglossary term E-E-A-T

SlideConfirmed by docsD3-C224

Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels and a 16x9 aspect ratio, enabled by the max-image-preview:large setting.

Speaker GoogleEvidence slide photo

Used byrequirements DEV-IDX-04, DEV-IMG-07

  • Extends D2-C089 Day 2: max-image-preview:large matters mainly in Discover, where it allows a large image that draws people's…
  • Repeats D2-C936 Day 2: Setting the max-image-preview robots meta tag to large can make content perform surprisingly well in…
DocsSourceD3-C227

Google's Discover documentation says to specify the large image with schema.org markup or the og:image meta tag, which can influence the thumbnail Discover chooses, and that a vertical image cropped to 16x9 must keep its important details.

Publisher Google Search Central

Used byrequirement DEV-IMG-07

SlideD3-C228

Google's closing Discover slide pointed to the Discover content policies in Google Search Help and to the Search Central documentation.

Speaker GoogleEvidence slide photo

Things
DocsSourceD3-C229

Google's Discover content policies ask news sources for clear dates and bylines, information about the authors, publication and publisher, and contact information, and say repeated or egregious violations can make a site ineligible for Discover.

Publisher Google Search Help

Things
AnalysisD3-C230

For Discover, set max-image-preview:large on article templates, give every article a relevant image at least 1,200 px wide in og:image or schema.org markup that is neither the site logo nor text-heavy, and show clear bylines and dates.

Author Ibrahim Anjro

Used byrequirement DEV-IMG-07

12:00 · What are quality updates 71

StageNot in docsD3-C234

The speaker opened the talk on why Search changes by noting that Google's logo has changed only a handful of times in about 30 years.

Speaker GoogleEvidence transcript

SlideNot in docsD3-C236

Google's slide said that in 2023 Google ran more than 800,000 search quality tests; the speaker added that a more recent figure might exist.

Speaker GoogleEvidence slide photo, transcript

SlideConfirmed by docsD3-C237

Google's slide said Google made more than 4,700 launches to Search in 2023; the speaker rounded this to close to 5,000.

Speaker GoogleEvidence slide photo, transcript

  • Repeats D3-C149 Day 3: Google's slide said that in 2023 Google ran 719,326 search quality tests, 124,942 side-by-side experiments…
SlideConsistent with docsD3-C239

Google's slide, headed 'In 2023, there were...', said 40 billion spammy pages are detected every day.

Speaker GoogleEvidence slide photo

  • Extends D3-C197 Day 3: Google's quality talk said Google discovers tens of billions of spam pages every day.
AnalysisD3-C241

Cite 40 billion spammy pages a day as Google's published figure, from its webspam report for 2020 and its How Search Works page: the slide's 'In 2023' heading does not make it a 2023 measurement, because the 2021 and 2022 reports give no daily count.

Author Ibrahim Anjro

SlideNot in docsD3-C243

Google's slide said that as content formats and types become more widespread, search users might start looking for them, and if enough people become interested Google might launch one or more Search features for those formats.

Speaker GoogleEvidence slide photo, transcript

SlideConsistent with docsD3-C245

Google's slide said Google is not perfect and spammers find loopholes to make low-effort content rank high in Search, so Google might launch targeted algorithms for a more level playing field in its search results.

Speaker GoogleEvidence slide photo, transcript

StageNot in docsD3-C246

Compared with the simple early-2000s Google results page, which showed a few expected sites, today's results page for the same query adds exploration features such as People Also Ask.

Speaker GoogleEvidence transcript

StageNot in docsD3-C247

Google added exploration features to its results pages because users' behaviour evolved: people wanted to explore more of the topic they were searching for.

Speaker GoogleEvidence transcript

StageD3-C248

The speaker called launches of new Search features 'feature updates', a personal label coined the year before, not an official Google term.

“I like to call that feature updates. It's not an official name.”

Speaker GoogleEvidence transcript

StageNot in docsD3-C249

The speaker said that without new Search features, search results would become obsolete and users would no longer find them useful.

Speaker GoogleEvidence transcript

StageNot in docsD3-C251

In the mid-1990s the web was small enough to list in one manually edited directory; by the time BackRub was being built, directories could no longer find information on the growing web, which is why Google was created.

Speaker GoogleEvidence transcript

Things
StageConsistent with docsD3-C256

Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.

Speaker GoogleEvidence transcript

  • Repeats D2-C680 Day 2: Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically…
StageNot in docsD3-C257

To show how much the web has grown, Google's deck compared result counts for the query 'durian': for every result in 2000 there were about 20,000 results in 2025.

Speaker GoogleEvidence transcript

StageConsistent with docsD3-C275

Google aims to keep more than 99% of search results free from spam and said it already achieves this, thanks to advances in AI.

Speaker GoogleEvidence transcript

  • Repeats D2-C676 Day 2: Google's testing shows that, thanks to SpamBrain, more than 99% of visits from Search are now spam-free.
AnalysisD3-C276

Google's webspam reports measure the 99% figure as visits from Search (2020 and 2022 reports) or searches (2021 report) that are spam-free, not as search results; cite it as more than 99% of visits from Search being spam-free.

Author Ibrahim Anjro

SlideConsistent with docsD3-C282

Google's slide defined hacked content as content placed on a site without permission through security vulnerabilities, said cleaning it up is hard and prevention is key, and linked to goo.gle/hack-tips.

Speaker GoogleEvidence slide photo

Used byrequirement DEV-SPM-04

StageNot in docsD3-C283

In the speaker's view, some of the egregious spam methods on the slide (cloaking, doorways, scraped content, link spam and hacked content) are not that common anymore or no longer matter much to Google.

“Some of these, I would say, are not that common anymore, or we don't care all that much about them.”

Speaker GoogleEvidence slide photo, transcript

StageNot in docsD3-C284

Reviewing the spam slide, the speaker said link spam would probably be replaced with scaled content abuse as the spam type worth talking about today.

Speaker GoogleEvidence slide photo, transcript

Used byrequirement DEV-SPM-02

  • Extends D3-C207 Day 3: Google's quality talk listed recent updates to Google's spam policies: back button hijacking, scaled content…
  • Extended by D3-C694 Day 3: Scaled content abuse is becoming a problem again: in the early 2000s pages were churned out with Perl or PHP…
DocsSourceD3-C287

Google's spam updates page says its automated spam detection systems run constantly, and a notable improvement to them, such as to the AI-based SpamBrain system, is called a spam update and listed with Google's ranking updates.

Publisher Google Search Central

Used byglossary term Spam update

  • Extends D2-C670 Day 2: SpamBrain is central to Google's spam-fighting efforts and has been improved many times since its launch.
StageConsistent with docsD3-C293

A ranking drop after a core update is not a penalty, so technically there is no recovery from a core update in the way there is after a spam update.

“technically, and from a purely technical perspective, there's no recovery, because you were not penalized”

Speaker GoogleEvidence slide photo, transcript

SlideConsistent with docsD3-C298

Google's 'Recoveries' slide advised sites affected by a core update to keep doing a great job, look at what competitors are doing better and learn from sites that are doing better.

Speaker GoogleEvidence slide photo

AnalysisD3-C301

For developers, Google's five listed spam types map to concrete checks: serve Googlebot the same content as users (cloaking), avoid near-identical location or keyword pages (doorways), add value to reused content (scraped content), and patch and monitor the CMS against injected pages and links (hacked content).

Author Ibrahim Anjro

Things

Used byrequirement DEV-SPM-03

13:25 · How Search results are born 50

StageNot in docsD3-C302

Google described its search results page as an auction system in which the individual results, of different types, all bid for a place on the page; the auction is a metaphor for result types competing for space, not a reference to ads.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C303

Text results are the most common and most prominent result type on Google's results page, whether they are generated by an LLM or not, or enhanced in some other way, Google's speaker said.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD3-C304

Which kinds of results Google shows for a query is decided by query understanding, which tries to predict the intent behind the query.

Speaker Gary IllyesEvidence transcript

  • Extends D3-C004 Day 3: Google placed query understanding in the traditional part of Search, the part where answers are looked up in…
AnalysisD3-C308

Google's 2007 launch post explained universal search by the growing number of separate search tools and traced it to a 2001 results-page mockup; the 2005-2006 image-seeking queries named on stage are a further motivation the post does not mention, and they predate the May 2007 launch.

Author Ibrahim Anjro

DocsSourceD3-C310

Google's How Search Works pages say Google uses aggregated and anonymized interaction data to assess whether search results are relevant to queries, turning that data into signals for its machine-learned systems.

Publisher Google Search (How Search Works)

StageNot in docsD3-C311

Google handles an image used as a search query much like a text query interpreted as an embedding: the image is broken down into vectors (embeddings) that are then searched for in the index.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C740 Day 2: Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are…
  • Repeats D1-C224 Day 1: For visual search, Google breaks an image down into vectors, sends the vectors to the index and returns…
StageNot in docsD3-C312

Google can still detect intent for image queries, even though they are searched as embeddings, the speaker said.

Speaker Gary IllyesEvidence transcript

StageConfirmed by docsD3-C316

Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-HTM-03

  • Extends D2-C724 Day 2: The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows…
StageNot in docsD3-C319

Image results shown among web results come from Google's image index and are roughly the same images that Google Images shows for the same query.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C525 Day 2: An image Google has extracted can appear almost anywhere Google shows results, including Discover, image…
SlideConsistent with docsD3-C321

Google's serving diagram shows the query passing through query understanding and retrieval to the index, then back through ranking and search features to the user, with the Search Features step highlighted (example query 'Where to eat orange').

Speaker Gary IllyesEvidence slide photo

  • Repeats D3-C001 Day 3: Google's opening slide for serving day showed serving as the third stage after crawling and indexing, and…
StageConsistent with docsD3-C323

Most of Google's search features need nothing extra from the site owner; Google generates them from what it extracted from the page during indexing.

Speaker Gary IllyesEvidence transcript

Used byglossary term Search features

  • Extends D2-C446 Day 2: The 'gold nuggets' that Google's feature extraction step pulls out of a page's HTML are structured data (such…
StageConfirmed by docsD3-C324

Rich results differ from other search features because Google builds them from extra data that site owners provide, usually structured data and usually in JSON-LD format.

“The magic behind rich results is usually structured data. Emphasis on "usually".”

Speaker Gary IllyesEvidence transcript

Used byglossary terms Rich results, Search features

  • Extends D2-C521 Day 2: Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its…
StageConfirmed by docsD3-C325

AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to function and work with the normal text results from Google's index.

“AI Mode and AI Overviews are not rich results.”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-AIF-02glossary terms Rich results, Search features

  • Extends D2-C726 Day 2: AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a…
StageConfirmed by docsD3-C329

Google's structured data feature guide lists the kinds of structured data Google supports with a search feature and what each can do to a site's search results.

Speaker Gary IllyesEvidence transcript

  • Repeats D2-C490 Day 2: Google recommends using the Search gallery in its developer documentation to find the structured data…
SlideConsistent with docsD3-C331

Google's rich results slide said Article structured data is not required to appear in Top Stories but is highly recommended for all articles.

“Not required for "Top Stories," but highly recommended for all articles.”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-SDA-13

SlideConfirmed by docsD3-C333

Google's Article rich results slide said structured data can be used to indicate that content is behind a paywall.

Speaker Gary IllyesEvidence slide photo

Used byrequirement DEV-SDA-13

SlideConfirmed by docsD3-C334

Google's Article slide showed the NewsArticle JSON-LD example from Google's Article documentation: a block in the page head with headline, three image URLs (1x1, 4x3 and 16x9), datePublished and dateModified with time zone offsets, and an author list of Person objects with name and url.

Speaker Gary IllyesEvidence slide photo

Things

Used byrequirement DEV-SDA-13

StageConfirmed by docsD3-C335

Google described the site name as an attribution feature rather than a rich result, and said site owners can influence which site name Google shows for their site in search results.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SDA-15glossary term Site name

StageConfirmed by docsD3-C337

Review structured data lets a site specify how its users rated something; the review snippet shows an average star rating and often the number of reviews of a product, service or piece of content.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SDA-14glossary term Review snippet

  • Extends D2-C454 Day 2: Structured data turns the loosely structured web into structured information that powers visual search…
StageNot in docsD3-C338

Google said users rely quite a bit on review stars, calling them a powerful signal of quality and trust for users and a way for sites to use social proof to increase clicks.

“Those little yellow stars are a powerful signal of quality and trust for our users.”

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C341

Some countries rely a lot on social proof such as reviews, Google's speaker said, recalling a point from Google's Day 2 internationalisation talk.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C614 Day 2: According to consumer data shown on a slide in the talk (source not captured), US consumers judge product…
StageConfirmed by docsD3-C343

Google Search displays product rich results from product structured data; it may also use the product data a merchant provides in Google Merchant Center, but Merchant Center is not required for them.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SHP-01glossary term Merchant Center

DocsSourceD3-C344

Google's Product structured data introduction says that providing both on-page structured data and a Merchant Center feed maximizes eligibility for shopping experiences and helps Google understand and verify the data; product snippets may take pricing from the feed when the markup lacks it.

Publisher Google Search Central

Used byrequirement DEV-SHP-01

StageConsistent with docsD3-C346

To appear in the Shopping tab of Google Search, a product must be submitted in Google Merchant Center; product structured data alone is not enough.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SHP-01glossary term Merchant Center

DocsSourceD3-C347

Google's Merchant Center help says products and details crawled from an online store can show up on Google even when they are not marked up with schema.org or added in Merchant Center, and lists the Shopping tab among the places Merchant Center free listings appear.

Publisher Google Merchant Center Help

AnalysisD3-C348

Google's free listings help says crawled store products can appear on Google without markup or Merchant Center but does not say whether that includes the Shopping tab; for a shop that wants Shopping tab visibility, the stage rule that Merchant Center is required is the safe working assumption.

Author Ibrahim Anjro

StageConfirmed by docsD3-C349

Product structured data can help Merchant Center in some cases, for example during data validation, but Merchant Center does not require it.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SHP-01

DocsSourceD3-C350

Merchant Center's automatic item updates use structured data and other product data found on landing pages to fix price, sale price, availability and condition mismatches in a merchant's product data; Google says they do not replace regular feed updates.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-01

13:40 · Shopping on Search: Beyond the blue links 60

StageConsistent with docsD3-C359

In 2026, a few months before the October event, Google added new Merchant Center feed attributes to help AI systems understand product nuances, prompted by feedback since 2025 from merchants worried whether AI experiences represented them correctly.

Speaker Alex JansenEvidence transcript

StageNot in docsD3-C363

Google decided against adding hundreds of thousands of highly specific feed attributes, such as heel height for shoes or lens details for cameras, and instead added a few flexible ones that leave merchants in control.

Speaker Alex JansenEvidence transcript

StageConfirmed by docsD3-C365

Google added six Merchant Center feed attributes for AI shopping experiences: question and answer, documents, related products, item group title and variant option for variants, and popularity rank.

Speaker Alex JansenEvidence transcript

Used byrequirement DEV-SHP-03glossary term Conversational attributes

  • Extends D2-C509 Day 2: More shopping structured data news was left for a Day 3 talk by a Google colleague, Alex.
AnalysisD3-C366

Google's January 2026 announcement spoke of 'dozens' of new Merchant Center attributes, while the Day 3 talk counted six; all six named on stage (question_and_answer, document_link, related_product, item_group_title, variant_option, popularity_rank) are in the Merchant Center product data specification as of 3 October 2026.

Author Ibrahim Anjro

DocsSourceD3-C367

Google's Merchant Center help marks the question_and_answer, document_link, related_product and popularity_rank attributes as primarily intended for conversational experiences such as AI Mode in Google Search.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-03glossary term Conversational attributes

DocsSourceD3-C369

Merchant Center's question_and_answer attribute takes up to 30 question-and-answer pairs per product, each question and answer up to 1,000 characters and 10,000 characters in total, and must not contain prices, shipping, dates or the company name.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-03

DocsSourceD3-C371

Merchant Center's document_link attribute takes up to five URLs of PDF documents about a product, such as manuals, user guides or assembly instructions, and the merchant must own the content licensing rights.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-03

DocsSourceD3-C375

Merchant Center's popularity_rank is a 0-100 value the merchant assigns to rank a product's popularity, based on its recent sales, against the rest of its own inventory; it does not reflect user ratings.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-03

AnalysisD3-C376

On stage the popularity rank was framed as how well a product sells on a marketplace, but Google's documentation defines it as the merchant's own 0-100 ranking against the rest of its inventory: a self-reported, relative value, not a sales figure across shops.

Author Ibrahim Anjro

AnalysisD3-C381

As of 3 October 2026 Google's Merchant Center help lists no schema.org property for question_and_answer, document_link, related_product, popularity_rank, variant_option or product_detail (only item_group_title maps, to ProductGroup.name), and Search Central does not document such markup: reading these attributes from schema.org is a stage statement, the feed is the documented route.

Author Ibrahim Anjro

Used byrequirement DEV-SHP-04

StageConfirmed by docsD3-C383

In product markup, an offer's shipping and return information can point through a JSON-LD identifier (@id) to shipping and return data defined elsewhere.

Speaker Alex JansenEvidence transcript

Things

Used byrequirement DEV-SDA-09

  • Extends D2-C505 Day 2: Google's shopping structured data launches of the previous year (2025) added support for merchant loyalty…
DocsSourceD3-C384

Google's merchant listing documentation recommends defining shipping and return policies once under Organization markup and referencing them from an Offer, even on another page, using only the @id keyword (through hasShippingService for shipping, hasMerchantReturnPolicy for returns).

Publisher Google Search Central

Used byrequirement DEV-SDA-09

StageConsistent with docsD3-C385

Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.

Speaker Alex JansenEvidence transcript

Used byrequirement DEV-SHP-02glossary term Storebot-Google

  • Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
DocsSourceD3-C388

Google's Merchant Center help says the StoreBot crawler goes through product detail, cart and checkout pages, can fill in checkout forms, and records price, shipping, availability, coupons and payment methods to verify the data merchants share in Merchant Center.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-02glossary term Storebot-Google

StageConfirmed by docsD3-C392

Google now emphasises two existing Merchant Center attributes for AI, product highlights (a short bulleted list) and product detail (detailed specifications), and updated its Help Center guidance to say they help AI systems.

Speaker Alex JansenEvidence transcript

Used byrequirement DEV-SHP-03glossary term Product highlights and product details

DocsSourceD3-C393

Merchant Center's product_highlight attribute takes 2 to 100 highlights of up to 150 characters each that describe only the product itself, without promotional text, keywords or search terms.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-03glossary term Product highlights and product details

StageConfirmed by docsD3-C394

Instead of adding feed attributes for individual specifications such as connectivity, memory, camera settings or box contents, Google asks merchants to send them as key-value pairs in product detail, because merchants know best what matters.

Speaker Alex JansenEvidence transcript

Used byrequirement DEV-SHP-03

DocsSourceD3-C395

Merchant Center's product_detail attribute takes up to 100 specifications, each a section name, attribute name and attribute value, and Google's help says clean key-value specifications improve product details on Google Shopping and AI-driven surfaces.

Publisher Google Merchant Center Help

Used byrequirement DEV-SHP-03glossary term Product highlights and product details

StageConsistent with docsD3-C396

Google works to make sure that what merchants can express in its shopping feeds can also be expressed in schema.org, adding vocabulary where schema.org lacks it, for example for product details.

Speaker Alex JansenEvidence transcript

  • Extends D2-C506 Day 2: Google's structured data speaker said that in the months before the event Google added support for validity…
StageConsistent with docsD3-C397

Google named the Universal Commerce Protocol (UCP), next to the Merchant Center feed, as an industry standard whose capabilities it also wants schema.org to cover.

Speaker Alex JansenEvidence transcript

Used byglossary term Universal Commerce Protocol (UCP)

DocsSourceD3-C398

Google launched the Universal Commerce Protocol (UCP) on 11 January 2026 as an open standard for agentic commerce across discovery, buying and post-purchase support, co-developed with retailers and platforms and compatible with A2A, AP2 and MCP.

Publisher Google blog (11 January 2026)

Used byglossary term Universal Commerce Protocol (UCP)

AnalysisD3-C400

Schema.org's release notes show version 30.1 (16 September 2026) added specification, isOftenBoughtWith and consumerNotice for Product, valueGroup for PropertyValue and itemPopularity for Offer as vocabulary for retail feed data; Google's Search Central docs did not yet document these properties on 3 October 2026.

Author Ibrahim Anjro

Used byrequirement DEV-SHP-04

StageNot in docsD3-C401

Rich, lightly structured information is critical to AI systems: product data with many levels of nesting is not needed, but some structure is very important.

“rich, lightly structured information is critical to AI systems.”

Speaker Alex JansenEvidence transcript

StageConsistent with docsD3-C405

Web markup is an efficient and unambiguous way for sites to share product data with Google, Google Shopping said, repeating the Day 2 structured data talk.

Speaker Alex JansenEvidence transcript

  • Repeats D2-C469 Day 2: Parsing structured data is significantly cheaper and more efficient than relying on LLMs for every complex…

13:55 · Inside Search Console: What’s New & How to Use It 57

StageConfirmed by docsD3-C413

Search Console's query groups put together queries that express a very similar user intent but are written differently, for example with more or fewer words or with spelling errors, so site owners can see the intent behind their traffic instead of every query string.

Speaker Ariel KroszynskiEvidence transcript

Used byglossary term Query groups

SlideConfirmed by docsD3-C414

To show how diverse queries for one intent are, Google's Search Console talk showed a slide titled 'How to spell Britney Spears?' with a long list of spelling variants, citing archive.google/jobs/britney.html.

“How to spell Britney Spears?”

Wording checked against the slide or recording

Speaker Ariel KroszynskiEvidence slide photo, transcript

AnalysisD3-C416

The Google archive page cited on the Britney Spears slide lists 593 spellings (the correct one plus 592 misspellings), each typed by at least two different users within a three-month period and corrected by Google's spelling system; it is an old jobs page (footer 2011), so the figures are historical.

Author Ibrahim Anjro

StageConsistent with docsD3-C417

Google said query groups are not built by matching query text alone: deciding which queries belong in which group goes beyond simple text filtering.

Speaker Ariel KroszynskiEvidence transcript

  • Extended by D3-C473 Day 3: Google's launch post for Search Console's branded queries filter (20 November 2025) says the branded versus…
StageNot in docsD3-C418

Google named granularity as a design challenge of query groups: putting many queries into one large group versus splitting them into smaller, more specific groups.

Speaker Ariel KroszynskiEvidence transcript

StageNot in docsD3-C420

Google acknowledged that query grouping is not perfect: some queries are not understood, and some are not grouped in a logical way.

Speaker Ariel KroszynskiEvidence transcript

StageConsistent with docsD3-C423

Drilling into a query group opens the Search Console Performance report with a regex filter that contains all the queries of the group, so each individual query can be analysed.

Speaker Ariel KroszynskiEvidence transcript

AnalysisD3-C426

The speaker placed the query groups launch, hedging, in December, but Google's announcement is dated 27 October 2025 with a gradual rollout over the following weeks, so many sites may first have seen the card around December 2025.

Author Ibrahim Anjro

StageConfirmed by docsD3-C429

Search Console's AI-powered configuration is a chat box in which users describe the analysis they want in natural language, and the request is transformed into the Performance report's configuration.

Speaker Ariel KroszynskiEvidence transcript

Used byglossary term AI-powered configuration

StageConfirmed by docsD3-C433

An example request for Search Console's AI-powered configuration was to show queries on phone searches containing the word sports in the last six months.

“Show me queries on phone searches that contain the word sports in the last 6 months.”

Speaker Ariel KroszynskiEvidence transcript

StageConsistent with docsD3-C434

In a test by Google's Search Console team, the vague request 'wine-related queries that lead to my site' made Search Console's AI-powered configuration build a regex filter of wine names, such as Merlot, joined by pipe characters, without the user typing the regex.

Speaker Ariel KroszynskiEvidence transcript

AnalysisD3-C435

Review the filters and regexes that Search Console's AI-powered configuration proposes before trusting the numbers: Google's launch post calls the feature experimental, warns that AI can misinterpret requests, and limits it to the Search results Performance report (not Discover or News) and to configuration, not sorting or exporting.

Author Ibrahim Anjro

  • Repeats D1-C243 Day 1: The first rule for using LLMs on search data, according to a community speaker, is never to let the model…
  • Extended by D3-C501 Day 3: Before feeding a branded versus non-branded split to an LLM, spot-check a sample of queries in each group…
StageConfirmed by docsD3-C436

Search Console has a tool with which every property owner can tell Google whether the site's content may be used in AI search results, because Google wants content owners to control their content.

Speaker Ariel KroszynskiEvidence transcript

  • Repeats D1-C029 Day 1: Search Console has a property setting called Search generative AI that gives direct control over AI Overviews…
StageConsistent with docsD3-C437

Search Console added reporting of the impressions a site's content gets in the AI surfaces of Google Search, found in the left navigation nested under Search results.

Speaker Ariel KroszynskiEvidence transcript

  • Repeats D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
StageConfirmed by docsD3-C438

The regular Search Console Performance report still shows all traffic including AI features, while the generative AI view shows only impressions from AI surfaces.

Speaker Ariel KroszynskiEvidence transcript

Used byrequirement DEV-MON-07glossary term Generative AI performance report

  • Extended by D3-C492 Day 3: Nik Vujic said GA4 setups that separate out traffic from LLMs do not capture visits from AI Overviews and AI…
  • Extended by D3-C498 Day 3: Google's AI features guide says traffic from sites appearing in AI features such as AI Overviews and AI Mode…
StageConfirmed by docsD3-C439

Search Console's generative AI report shows impressions broken down by pages, countries and devices.

Speaker Ariel KroszynskiEvidence transcript

Used byrequirement DEV-MON-07

  • Repeats D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
StageConsistent with docsD3-C440

Google called AI reporting in Search Console an evolving space and expects the generative AI report to get richer, with more information in the future.

Speaker Ariel KroszynskiEvidence transcript

  • Repeats D1-C196 Day 1: Google said Search Console's generative AI reporting launched alongside the generative AI control, starting…
  • Extends D1-C449 Day 1: A Google panelist said he expected more AI reporting to launch in Search Console, without a timeline or a…
  • Extended by D3-C495 Day 3: Nik Vujic said he hopes Search Console's generative AI report will get query data.
StageConfirmed by docsD3-C442

Search Console launched platform properties, which bring social platforms into Search Console, about three or four months before the October 2026 event.

Speaker Ariel KroszynskiEvidence transcript

  • Extends D2-C960 Day 2: A community speaker said publishers can also verify their social accounts in Search Console, which makes it…
StageConfirmed by docsD3-C443

Platform properties let a user verify ownership of an Instagram, TikTok, X or YouTube account in Search Console and see the traffic Google sends to that account.

Speaker Ariel KroszynskiEvidence transcript

Used byrequirement DEV-MON-09glossary term Platform properties

StageConsistent with docsD3-C446

A platform property is verified entirely from within Search Console: a pop-up from the platform, for example Instagram, asks the user to grant Google permission to identify the account and its handle, and completing it proves ownership without a token.

Speaker Ariel KroszynskiEvidence transcript

Used byrequirement DEV-MON-09glossary term Platform properties

StageConsistent with docsD3-C447

Platform properties have a newly designed landing page, similar to the Insights report, that shows overall clicks and the search types that lead to the account.

Speaker Ariel KroszynskiEvidence transcript

StageConsistent with docsD3-C448

The platform property landing page lists the account's top-performing content items, for example a channel's top YouTube videos, and the items trending up and trending down.

Speaker Ariel KroszynskiEvidence transcript

StageConfirmed by docsD3-C449

Platform properties also show the query groups that lead to the account's content and the countries that send it traffic.

Speaker Ariel KroszynskiEvidence transcript

StageConfirmed by docsD3-C450

Because platform properties bring many new users to Search Console, Google published documentation and strategy guidance on them, covering how to identify the Search audience, use the 24-hour real-time view, export data and compare platforms such as Instagram and YouTube.

Speaker Ariel KroszynskiEvidence transcript

StageConfirmed by docsD3-C451

In a YouTube platform property, long-form and short-form videos can be compared in the Performance report's comparison mode by URL pattern, /watch versus /shorts/; Instagram posts (/p/) and reels can be compared the same way.

Speaker Ariel KroszynskiEvidence transcript

Used byrequirement DEV-MON-09

StageNot in docsD3-C452

In the example YouTube report shown in Google's Search Console talk, long-form videos still got more traffic and more impressions from Google than Shorts; this was one example channel, not a general finding.

Speaker Ariel KroszynskiEvidence transcript

StageConfirmed by docsD3-C458

Multimodal searches reported in Search Console include searches with Google Lens, Circle to Search on Android, images uploaded to Google Search and Chrome's right-click search on an image.

Speaker Ariel KroszynskiEvidence transcript

Used byrequirement DEV-MON-07glossary term Multimodal search type

StageConsistent with docsD3-C459

To see multimodal traffic in the Search Console Performance report, select the multimodal search type instead of the default text-based one.

Speaker Ariel KroszynskiEvidence transcript

Used byrequirement DEV-MON-07glossary term Multimodal search type

StageConsistent with docsD3-C461

In a demo, asking Chrome about an image on a web page opened a side panel of visual matches, one of which was on Google's developers site and led on to more information about Search Console; Google described this as the full flow of the multimodal traffic now reported.

Speaker Ariel KroszynskiEvidence transcript

AnalysisD3-C462

Google's Performance report help says the queries dimension is not available for the multimodal search type, because these searches mostly use images, so judge image-led traffic by page, country and device, and make product and article images distinctive enough for Lens to match.

Author Ibrahim Anjro

Used byrequirement DEV-MON-07

DocsSourceD3-C464

Google's Search Console help says the experimental AI-powered configuration is visible only to a small percentage of users, allows each user 20 requests per day, keeps existing report filters when it adds new ones, and applies its suggested settings only after the user confirms them.

Publisher Google Search Console Help

Used byglossary term AI-powered configuration

DocsSourceD3-C465

Google's announcement of 24 September 2026 says multimodal search data appears both in the Performance report for Search results and in the generative AI performance report, can be exported, and began rolling out globally that day.

Publisher Search Central blog (24 September 2026)

Used byrequirement DEV-MON-07glossary terms Generative AI performance report, Multimodal search type

DocsSourceD3-C468

Google's social and video performance guide says that if you already claimed your Search profile, all of its verified accounts are added automatically as platform properties in Search Console.

Publisher Google Search Central

Used byrequirement DEV-MON-09glossary term Search profile

  • Extends D1-C025 Day 1: Publishers and creators can claim a dedicated Search profile with a name, profile photo, cover image, bio…
  • Extends D2-C960 Day 2: A community speaker said publishers can also verify their social accounts in Search Console, which makes it…

14:20 · Lightning session L: Understanding SERPs and your users 76

SlideConfirmed by docsD3-C472

A community slide (Workflow 1) said the first step is to apply the filters you need, even natively in Search Console, for example branded versus non-branded queries, with clicks and impressions.

“Apply the filters you need, even natively in GSC. For example branded vs non-branded, with clicks and impressions.”

Wording checked against the slide or recording

Speaker Nik VujicEvidence slide photo, transcript

DocsSourceD3-C473

Google's launch post for Search Console's branded queries filter (20 November 2025) says the branded versus non-branded split is made by an internal AI-assisted system, not by a regular expression, and that some queries may occasionally be misidentified.

Publisher Search Central blog (20 November 2025, updated 11 March 2026)

Used byglossary term Branded queries filter

  • Extends D3-C417 Day 3: Google said query groups are not built by matching query text alone: deciding which queries belong in which…
SlideConfirmed by docsD3-C475

A community slide said the Search Console bulk data export to BigQuery gives the most complete query dataset available, with anonymized queries excluded.

“The bulk export to BigQuery gives the most complete query dataset available - anonymized queries excluded.”

Wording checked against the slide or recording

Speaker Nik VujicEvidence slide photo

Used byrequirement DEV-MON-08glossary term Search Console bulk data export

AnalysisD3-C476

In the bulk export, 'anonymized queries excluded' means the query text, not the traffic: Google's table reference and query guidelines show those rows kept with an empty query field, often as the single most common 'query'. Tell an LLM to skip empty queries when ranking top queries, but keep them in totals.

Author Ibrahim Anjro

Used byrequirement DEV-MON-08

SlideD3-C478

Nik Vujic's slide said no agent is needed to start: pull the Search Console data, import it into any LLM and query it in conversation; his agency built its agent to have everything in one place.

“No agent needed to start: pull the data and import it into the LLM of your choice, and talk with it.”

Wording checked against the slide or recording

Speaker Nik VujicEvidence slide photo, transcript

  • Extends D1-C256 Day 1: An agency's automated monthly SEO report follows four simple steps: the data comes in, a script processes it…
DocsSourceD3-C480

A Search Central blog post on Search Console data limits (October 2022) says the interface exports at most 1,000 rows, while the Search Analytics API and the Looker Studio connector return up to 50,000 rows per day per site per search type.

Publisher Search Central blog (19 October 2022)

Used byrequirement DEV-MON-08

DocsSourceD3-C481

Google's blog post on Search Console data limits says anonymized queries are always left out of the report tables but counted in chart totals unless you filter by query, and are left out whenever a filter is applied, so filtered rows do not add up to the chart totals.

Publisher Search Central blog (19 October 2022)

Used byrequirement DEV-MON-08

AnalysisD3-C487

Google's ranking systems guide describes 'query deserves freshness' systems that show fresher content where it would be expected, which is narrower than a general preference for fresh content; refresh pages whose queries expect current information, and judge other refreshes by quality.

Author Ibrahim Anjro

  • Extends D2-C656 Day 2: Freshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking…
StageD3-C490

Nik Vujic advised starting every SEO experiment with a hypothesis, for example a bulk change to one page element across a group of pages, and using an LLM to report the group's Search Console data, for example click-through rate, for the test.

Speaker Nik VujicEvidence transcript

StageNot in docsD3-C493

Nik Vujic said Google Tag Manager events can be used to follow real traffic arriving from different LLMs.

Speaker Nik VujicEvidence transcript

StageD3-C495

Nik Vujic said he hopes Search Console's generative AI report will get query data.

Speaker Nik VujicEvidence transcript

  • Extends D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
  • Extends D3-C440 Day 3: Google called AI reporting in Search Console an evolving space and expects the generative AI report to get…
StageD3-C496

Nik Vujic said his agency uses the combined data to test whether good SEO is good GEO, and that the answer is 'it depends': for one client it found pages that performed better in generative AI features and LLMs than in standard search.

Speaker Nik VujicEvidence transcript

Things
AnalysisD3-C497

Pages that do better in AI answers than in classic results do not refute Google's line that optimising for people is optimising for generative AI Search, since the comparison mixed Google's AI features with third-party LLMs; measure each surface separately, per page, before deciding what to change.

Author Ibrahim Anjro

  • Extends D1-C058 Day 1: Optimising for people is optimising for generative AI Search.
DocsSourceD3-C498

Google's AI features guide says traffic from sites appearing in AI features such as AI Overviews and AI Mode is included in the overall search traffic in Search Console and reported in the Performance report within the Web search type.

Publisher Google Search Central

Used byrequirement DEV-MON-07

  • Extends D3-C438 Day 3: The regular Search Console Performance report still shows all traffic including AI features, while the…
DocsSourceD3-C499

Search Console's branded queries filter works only for top-level properties, not URL-path or subdomain properties, and only for sites with enough queries and impressions; Google made it available to all eligible sites on 11 March 2026.

Publisher Search Central blog (20 November 2025, updated 11 March 2026)

Used byglossary term Branded queries filter

AnalysisD3-C501

Before feeding a branded versus non-branded split to an LLM, spot-check a sample of queries in each group, since Google's AI-assisted classification can misidentify queries; on URL-path or subdomain properties, where the filter is unavailable, a regex query filter is the fallback.

Author Ibrahim Anjro

  • Extends D3-C435 Day 3: Review the filters and regexes that Search Console's AI-powered configuration proposes before trusting the…
AnalysisD3-C502

Looker Studio lifts Search Console's 1,000-row interface limit but keeps the connector's cap of 50,000 rows per day per site per search type; only the bulk export to BigQuery is free of the daily row limit, so large sites that need every non-anonymized query should use it.

Author Ibrahim Anjro

Used byrequirement DEV-MON-08

StageD3-C504

A community speaker argued that SEOs who panic when rankings rise but clicks do not are relying on a metric, rankings and visibility, that no longer stands up to scrutiny.

“we rely on a metric that no longer stands up to scrutiny”

Speaker not identifiedEvidence transcript

StageNot in docsD3-C505

A community speaker said most people in the room had probably seen clicks decline over the past couple of years because AI now takes part of the demand by summarising information from websites, a shift that hit informational queries first.

Speaker not identifiedEvidence transcript

DocsSourceD3-C506

In an August 2025 post, Google's Head of Search wrote that total organic click volume from Google Search to websites had been relatively stable year-over-year and that Google was sending slightly more quality clicks (clicks where users do not quickly click back) than a year earlier; the post gave no figures.

Publisher Google blog (6 August 2025)

DocsSourceD3-C507

Google's August 2025 post on AI in Search said traffic is shifting between sites, decreasing for some and increasing for others, and that for some questions where people want a quick answer they may be satisfied with the response and not click further.

Publisher Google blog (6 August 2025)

AnalysisD3-C508

The click decline described on stage is what many individual sites report, while Google's published position (August 2025) is that total organic clicks are relatively stable but shifting between sites, with fewer clicks for quick-answer questions; neither side gave figures, so measure the trend per site and per query type rather than assuming a market-wide fall.

Author Ibrahim Anjro

StageD3-C509

A community speaker said the loss of clicks on informational queries to AI summaries probably happened for a good reason, with benefits for the user journey (the word 'benefits' is an uncertain reading of the recording).

Speaker not identifiedEvidence transcript

StageNot in docsD3-C511

A community speaker reported, calling it worrying, that AI features have started to appear more often on results pages for commercial search terms too, in data the speaker tracks for particular brands.

Speaker not identifiedEvidence transcript

StageD3-C512

A community speaker said the problem SEOs face is not ranking but positioning SEO differently and measuring and showing its value to stakeholders while AI features take away clicks, some of which were probably not relevant anyway.

Speaker not identifiedEvidence transcript

StageD3-C513

Key stakeholders such as heads of department just want to see metrics and numbers, a community speaker said, which makes the way SEO reports its value the question to solve.

Speaker not identifiedEvidence transcript

StageD3-C514

A community speaker said digital marketing has been spoiled by about 20 years of plentiful attribution and that SEOs should relearn the mindset of marketers who worked with very little of it, while still using the performance data that remains.

Speaker not identifiedEvidence transcript

StageD3-C515

A community speaker advised SEOs to focus, as marketers did before detailed attribution, on downstream impact: the effect their work has on other channels.

Speaker not identifiedEvidence transcript

Used byglossary term Downstream impact

StageNot in docsD3-C516

A community speaker cited a recent Similarweb study as showing that being recommended in AI-driven services multiplies a brand's chance of getting traffic downstream; a Similarweb study reported in June 2026, probably the one meant, found brands recommended by ChatGPT 2.5 times more likely to get a site visit within 7 days (US desktop data, finance, travel and beauty).

Speaker not identifiedEvidence transcript

StageD3-C518

A community speaker concluded that rankings and clicks no longer equal real business outcomes.

“rankings and clicks don't really equal real outcomes anymore”

Speaker not identifiedEvidence transcript

  • Extends D1-C057 Day 1: Myth: the old metrics don't work in the AI era. Google's answer: measure success through metrics that matter…
StageD3-C521

A community speaker said the brand in the controlled test had a record month, the kind of real business value that business owners care about more than a dip in traffic.

Speaker not identifiedEvidence transcript

StageD3-C524

In the same controlled test the brand's top-of-page share in paid search dropped, so a slight dip in paid click-through rate would have been expected rather than the 36% rise, a community speaker said.

Speaker not identifiedEvidence transcript

StageD3-C525

A community speaker advised SEOs to work with paid search teams to understand changes such as a rise in paid click-through rate and what impact SEO work may have had, since those teams may not credit SEO for it.

Speaker not identifiedEvidence transcript

StageNot in docsD3-C526

A community speaker cited a study from about 10 years ago, repeated since, as finding that 82% of people clicked on a brand they already knew regardless of its position; the matching source is Red C's eye-tracking study of shopping-type searches, reported by Econsultancy in October 2018 (about eight years before the event).

Speaker not identifiedEvidence transcript

StageD3-C527

A community speaker suggested that people who meet a brand at several touchpoints, including in AI services, become familiar with it and are then more likely to choose it in a generic search, which can lift other metrics.

Speaker not identifiedEvidence transcript

StageD3-C528

A community speaker said brand familiarity should be built through AI-driven services and earned media as a whole, by improving content and visibility across services ('earned media' is an uncertain reading of the recording).

Speaker not identifiedEvidence transcript

StageD3-C529

A community speaker said the customer journey that starts in AI services still exists and only the direct attribution is missing, so SEOs should look at wider metrics they have not tracked before.

Speaker not identifiedEvidence transcript

StageD3-C531

When a brand bids on its own name, people who search the brand after an AI recommendation may arrive through paid search rather than organic, so organic reporting loses that visit, a community speaker noted.

Speaker not identifiedEvidence transcript

StageD3-C532

Clicks are down for a lot of businesses but rarely tell the whole story, a community speaker summarised, advising SEOs to ask the paid search team for Google Ads click-through rate reports.

Speaker not identifiedEvidence transcript

DocsSourceD3-C535

Beyond Search Console, Google's AI features guide suggests tracking conversions and time spent on the site in tools such as Google Analytics to understand the value of traffic from AI features.

Publisher Google Search Central

  • Extends D1-C062 Day 1: Google's guide for generative AI features recommends the Generative AI performance report in Search Console…
DocsSourceD3-C536

A May 2025 Search Central blog post advised site owners to look at the overall value of visits from Search rather than focusing too much on clicks, using indicators of conversion such as sales, sign-ups, a more engaged audience or information lookups about the business.

Publisher Search Central blog (21 May 2025)

  • Extends D1-C057 Day 1: Myth: the old metrics don't work in the AI era. Google's answer: measure success through metrics that matter…
StageD3-C537

For the brand in a community speaker's controlled test, organic sessions in GA4 roughly halved while direct sessions stayed fairly stable (an uncertain reading; a change of about 2% was mentioned, direction unclear in the recording).

Speaker not identifiedEvidence transcript

StageD3-C538

A community speaker advised comparing the relative performance of channels and choosing new metrics to report, rather than judging SEO on organic sessions alone.

Speaker not identifiedEvidence transcript

StageD3-C539

A fully controlled test of AI visibility was possible for one brand only because it ran no other marketing activity, and is not viable for most brands, a community speaker said.

Speaker not identifiedEvidence transcript

AnalysisD3-C542

Google's own line on Day 1 was to measure success by metrics that matter to the business, and its Generative AI performance report shows impressions but no clicks; together with the community talk this argues for a report that puts organic next to direct, branded paid search, conversions and revenue.

Author Ibrahim Anjro

  • Extends D1-C057 Day 1: Myth: the old metrics don't work in the AI era. Google's answer: measure success through metrics that matter…
  • Extends D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
AnalysisD3-C544

The two studies cited on stage are third-party and narrow (Similarweb, 2026: ChatGPT only, US desktop, three industries; Red C, 2018: eye-tracking on shopping-type searches), so quote them with that scope; Similarweb found that 55.9% of the resulting site traffic came from branded searches, the step after an AI recommendation in the speaker's customer journey.

Author Ibrahim Anjro

15:30 · Mastering the messy middle 53

SlideD3-C545

Google's talk 'Mastering the messy middle', given by a Google marketing research and insights manager, presented Google's research on how AI is changing the way consumers make purchase decisions.

Speaker Pablo PérezEvidence slide photo, transcript

StageConsistent with docsD3-C546

Google has used the 'messy middle' decision-making model since 2019; the messy middle is what happens between the trigger that starts a purchase journey and the purchase itself.

Speaker Pablo PérezEvidence transcript

Used byglossary term Messy middle

StageConfirmed by docsD3-C547

In the messy middle people are in one of two mindsets: exploring, which looks across the category, its price points, its brands and what is available, or evaluating, which works toward a position where a decision can be made.

Speaker Pablo PérezEvidence transcript

Used byglossary term Messy middle

DocsSourceD3-C549

Google's Think with Google report 'Marketing in the messy middle' places the messy middle between trigger and purchase and models it as two looping mindsets: exploration, an expansive mindset that adds brands, products and category information, and evaluation, a reductive mindset that narrows the options down to a decision.

Publisher Think with Google (Google and The Behavioural Architects; Decoding Decisions, part 2)

Used byglossary term Messy middle

DocsSourceD3-C550

Google's messy-middle report cites its original Decoding Decisions study, run with The Behavioural Architects in the UK in 2019 with 31,000 in-market online shoppers and published in 2020; the report says the findings were validated in more than 20 countries.

Publisher Think with Google (Google and The Behavioural Architects; Decoding Decisions, part 2)

StageConsistent with docsD3-C551

Google's researcher explained the looping in the messy middle as a tension between people's urge to collect more information and the mental effort of processing it, a subject studied by behavioural science.

Speaker Pablo PérezEvidence transcript

StageD3-C552

How AI is changing decision making is the number one question Google's consumer insights team receives, according to the speaker in October 2026.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C553

Google's research on how AI changes consumer purchase decisions, presented in October 2026, was based on 23,000 conversations with consumers (the talk did not say in which countries).

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C554

Because people do not always do what they say they will do, Google's research on AI and purchase decisions also analysed panels with 40,000 participants, together with research partners, besides the consumer conversations (the talk did not make clear whether 40,000 is one panel or the total).

Speaker Pablo PérezEvidence transcript

PressSourceD3-C555

The DMEXCO 2026 programme described a Google session, 'Beyond the Messy Middle: Mastering the consumer journey with AI' (24 September 2026), as based on tens of thousands of AI-moderated interviews and observed journeys across Europe, finding that consumers are being 'boosted' by AI.

Reported by DMEXCO 2026 event programme

AnalysisD3-C556

Keep the research numbers apart when quoting them: the 23,000 consumer conversations and 40,000-participant panels behind the 2026 AI findings are a newer research wave than the 2019 messy-middle study of 31,000 UK online shoppers.

Author Ibrahim Anjro

StageConsistent with docsD3-C558

Behavioural science distinguishes a fast, emotional way of deciding, used for frequent low-stakes decisions, from a slow, effortful way used for high-stakes decisions, which is tiring.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C559

Google's research found that consumers delegate the mental effort of a purchase decision to AI without surrendering the choice itself, which Google calls being 'boosted'.

“that doesn't mean they are surrendering their choices. What's happening is that consumers are feeling boosted.”

Speaker Pablo PérezEvidence transcript

Used byglossary term Boosted consumers

StageNot in docsD3-C560

According to Google's research, AI turns consumers who were confused and overwhelmed by information and options into consumers who feel empowered, like category experts, and able to make the best purchase for their needs.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C561

The first benefit consumers get from AI, according to Google's research, is ease: AI synthesises information quickly into an easy-to-digest form and makes choice overload manageable, so consumers can offload cognitive effort.

Speaker Pablo PérezEvidence transcript

Used byglossary term Boosted consumers

StageNot in docsD3-C563

The second benefit, assistance, is that AI reduces a complex category, such as smartphones or trip planning, to the few features that really matter (for a phone, the battery and the picture quality), so consumers feel like experts.

Speaker Pablo PérezEvidence transcript

Used byglossary term Boosted consumers

StageNot in docsD3-C564

Assistance also includes showing consumers what people in similar situations chose, because people are social and want that reassurance (the wording of this passage is partly uncertain in the recording).

Speaker Pablo PérezEvidence transcript

DocsSourceD3-C565

Google's messy-middle report names six behavioural principles at work in the messy middle: category heuristics, power of now, social proof, scarcity bias, authority bias and power of free; category heuristics are rules of thumb for a quick, satisfactory decision within a category, and social proof is the tendency to follow others' opinions and behaviour.

Publisher Think with Google (Google and The Behavioural Architects; Decoding Decisions, part 2)

StageNot in docsD3-C566

Google's research found that consumers who do not use AI pinball between exploration and evaluation, while AI users start a purchase journey mostly exploring and then move on to evaluating, a path that looks more like a funnel.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C571

Google's research found that people who use AI also visit more brand websites, because it is simply easier.

“the people that are using AI, they search more, and they also visit more websites of brands.”

Speaker Pablo PérezEvidence transcript

  • Extends D1-C005 Day 1: Google says AI Overviews help with new types of questions and lead users to visit a greater diversity of…
StageNot in docsD3-C573

AI users sometimes take even longer purchase journeys, but they perceive their journeys as shorter, a perception the data does not always support (part of this passage is uncertain in the recording).

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C574

For high-risk or high-involvement purchases, ease, assistance and suggestion are often not enough, and consumers add a reassurance step to reduce the risk of the decision.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C575

After exploring with AI, people often go back to Google Search and to the websites of the product or service providers to double-check before buying, for example trust in the company and cancellation policies.

Speaker Pablo PérezEvidence transcript

SlideConsistent with docsD3-C577

Google's slide defined the brand heuristic as a cognitive shortcut in which consumers use a brand's name and reputation as a proxy for product quality, reliability and value.

“A cognitive shortcut where consumers use a brand's name and reputation as a proxy for product quality, reliability, and value”

Wording checked against the slide or recording

Speaker Pablo PérezEvidence slide photo

Used byglossary term Brand heuristic

DocsSourceD3-C579

Google's messy-middle report says leading brands keep a meaningful share of consumer preference even against competitors with much better offers, and calls brand one of the strongest heuristics in consumer choice.

Publisher Think with Google (Google and The Behavioural Architects; Decoding Decisions, part 2)

StageNot in docsD3-C580

A new or less familiar brand should expect people to double-check it and should prepare assets that give them the extra reassurance they need at that last step.

Speaker Pablo PérezEvidence transcript

SlideNot in docsD3-C581

A Google slide showed nine Google Trends charts of trust-checking searches in several markets and languages (source line: Google Trends, January 2016 to March 2026, web search), all rising toward the end of the range: 'is [brand] legit?' (two panels), '[brand] è affidabile?', '[brand] jest bezpieczne?', 'Is [brand] betrouwbaar?', 'Ist [brand] seriös?', '[brand] est fiable?', '¿es [brand] falso?' and 'trustpilot'.

Speaker Pablo PérezEvidence slide photo

StageD3-C583

The speaker read the rise in 'is [brand] legit' searches as a signal that people go back to the brand's website and to Google Search to take an extra reassurance step.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C584

Consumers who use AI no longer feel frustrated by a lack of progress in a purchase decision; they feel they are moving through the funnel and feel like category experts.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C585

Consumers boosted by AI in their purchase decisions were still a small group in October 2026, but the group is growing as more people start using AI (the word 'boosted' is an uncertain reading at this point of the recording).

Speaker Pablo PérezEvidence transcript

SlideConsistent with docsD3-C586

Google's marketing research talk reached the same advice as Search: optimise for people to win in generative AI search.

“Optimise for people to win in generative AI search”

Wording checked against the slide or recording

Speaker Pablo PérezEvidence slide photo, transcript

  • Repeats D1-C058 Day 1: Optimising for people is optimising for generative AI Search.
StageConsistent with docsD3-C591

Because purchase journeys can be long and competitive, brands should be present where decisions are made: on Google Search and on their own brand website.

Speaker Pablo PérezEvidence transcript

StageD3-C592

A brand's own website matters in AI-assisted journeys because it is the one source where the brand speaks in its own voice and controls the message.

Speaker Pablo PérezEvidence transcript

AnalysisD3-C593

For high-involvement products and services, make the double-check easy on your own site: company details, reviews, cancellation and return policies, guarantees and contact options should be on crawlable, indexable pages linked from product and service pages.

Author Ibrahim Anjro

AnalysisD3-C595

The messy-middle findings come from Google's marketing research on consumer behaviour, not from Search documentation; they describe how people decide, not how Google ranks pages, and none of them is a ranking signal.

Author Ibrahim Anjro

AnalysisD3-C597

Content that helps AI-assisted shoppers names the few features that decide a category (for a phone, battery and camera), compares options on them and shows what similar buyers chose, which matches the ease, assistance and suggestion benefits Google described.

Author Ibrahim Anjro

15:45 · How long does it take to..? 80

StageD3-C598

Google presented its processing times for crawling, indexing and serving as an experiment: estimates from its internal logs, meant to be informative and not something to obsess about, with feedback from the audience invited.

“if you have an engineering mindset, then you can put a number to it.”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-MON-10

SlideD3-C599

Each of Google's processing-time charts marks a minimum, a typical value and an end point on a scale from seconds to one year, with a dotted line for corner cases, and carries the footnote that the times are estimations based on internal analysis.

“Times expressed on the diagram are estimations based on internal analysis.”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence 4 slide photos

StageNot in docsD3-C602

Google said it knows hundreds of trillions of URLs (as of October 2026).

Speaker Gary IllyesEvidence transcript

  • Extends D1-C201 Day 1: Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many…
StageConsistent with docsD3-C603

Google does not crawl all the URLs it knows, because crawling everything would generally waste resources and a much smaller crawl space is enough.

Speaker Gary IllyesEvidence transcript

SlideNot in docsD3-C604

Google estimated that refreshing (recrawling) a known URL takes about 30 days on average, with a minimum of seconds and an end point of weeks to never.

Speaker Gary IllyesEvidence slide photo, transcript

StageNot in docsD3-C606

For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).

Speaker Gary IllyesEvidence transcript

  • Extends D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
StageD3-C607

An audience member reported recrawl intervals on one very large, popular client site: the homepage about 5 times a day, first-level pages about every 1.5 days (uncertain reading), and pages clustered as soft 404s every 160 to 190 days, adding that other sites will differ.

From the audienceEvidence transcript

Things
StageD3-C608

An audience member said that on the large site they observed, recrawl frequency followed the site hierarchy, demand and how well each page is internally linked.

From the audienceEvidence transcript

AnalysisD3-C609

With a typical refresh of about 30 days and deep or soft-404-like pages recrawled only every few months, link important deep pages from strong hub pages and fix pages that look like errors; then check recrawl intervals by site depth in server logs or the Crawl Stats report.

Author Ibrahim Anjro

SlideNot in docsD3-C610

Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.

Speaker Gary IllyesEvidence slide photo, transcript

Things
  • Extends D1-C337 Day 1: An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a…
SlideNot in docsD3-C611

If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.

Speaker Gary IllyesEvidence slide photo, transcript

Things
  • Extends D1-C337 Day 1: An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a…
SlideConsistent with docsD3-C612

Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.

Speaker Gary IllyesEvidence slide photo, transcript

Things
  • Extends D1-C331 Day 1: Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically…
SlideConfirmed by docsD3-C613

Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirements DEV-MON-10, DEV-SRV-05

  • Extends D1-C085 Day 1: Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
StageConfirmed by docsD3-C614

Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C085 Day 1: Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
StageConfirmed by docsD3-C615

A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
StageConsistent with docsD3-C618

When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-03

  • Extends D1-C126 Day 1: Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down…
  • Extends D1-C354 Day 1: Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the…
  • Extends D1-C440 Day 1: If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its…
AnalysisD3-C621

Spoken and slide figures for capacity increases differ (one to three weeks, up to a month, versus 1-2 weeks typical and 1-3 weeks in recovery), but the lesson is the same: a burst of 5xx errors cuts crawling within hours and recovery takes weeks, so keep servers stable before launches and migrations.

Author Ibrahim Anjro

Things

Used byrequirement DEV-SRV-03

StageConsistent with docsD3-C623

When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
  • Extends D1-C377 Day 1: The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual…
  • Extends D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
StageConsistent with docsD3-C624

Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…
  • Repeats D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
StageConsistent with docsD3-C627

Content that JavaScript adds to a page is typically seen by Google's indexing system within a few hours, and at worst within weeks.

Speaker Gary IllyesEvidence transcript

Used byrequirements DEV-MON-10, DEV-REN-01

  • Extends D2-C172 Day 2: Google's JavaScript SEO basics guide says a page may wait in the render queue for a few seconds but that it…
StageNot in docsD3-C628

Gary Illyes said Google keeps saying it renders every URL on the internet and that he would say this is not true, though it is what he was told; he went on to say that Google's logs show the rendering queue cleared within weeks.

“we keep saying that we render every single URL on the internet. I would say that that's not true”

Speaker Gary IllyesEvidence slide photo, transcript

Things

Used byrequirement DEV-REN-01

  • Extends D2-C256 Day 2: Google renders nearly all of the web by replicating what a browser does, using a real browser's rendering…
SlideNot in docsD3-C635

Google defined indexing end to end as the time from a document entering indexing until its critical processes finish and it reaches the serving index, tokenized and ready to be served as a result.

Speaker Gary IllyesEvidence slide photo, transcript

StageD3-C637

Gary Illyes said the 1.5-hour indexing average is probably lowered by selection bias, because so many news sites push out articles.

Speaker Gary IllyesEvidence transcript

AnalysisD3-C638

By Google's own averages, getting found and crawled takes far longer than indexing (about 20 hours to discover a URL and 30 days to refresh one, against 1.5 hours to index), so for faster results work on discovery: internal links from often-crawled pages and accurate sitemaps.

Author Ibrahim Anjro

Things
StageConsistent with docsD3-C642

Google treats a site move as a complex canonicalization process in which every signal of the old site is recalculated and moved to the new one, and every indexing process has to run.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C351 Day 2: Google treats a site migration as deduplication across sites, in which the site owner says the old and the…
AnalysisD3-C645

Keep migration redirects in place for at least a year and judge a site move after one to three months, not days: Google's speaker said its slowest signal needs about a year to be recalculated (a partly uncertain passage of the recording), and Google's site move guide says to keep redirects generally at least one year.

Author Ibrahim Anjro

Used byrequirement DEV-CAN-09

StageConsistent with docsD3-C647

Google may never use structured data from a site it does not trust: once it sees markup it does not trust, it does not touch it.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SDA-02

  • Extends D2-C500 Day 2: Structured data that is not relevant to the page's content can be treated as abusive: Google's filters make…
StageNot in docsD3-C649

Google tries to index images from news sites much faster than usual, which lowers the average image indexing time (the reading 'news sites' is likely but not certain).

Speaker Gary IllyesEvidence transcript

SlideNot in docsD3-C655

Google estimated that updating the image shown with a text result takes 1-2 weeks on average, from a minimum of days up to several weeks to months.

Speaker Gary IllyesEvidence slide photo, transcript

StageNot in docsD3-C656

The image shown with a text result changes only after the image itself is indexed or reindexed and then goes through a further processing step involving embeddings.

Speaker Gary IllyesEvidence transcript

AnalysisD3-C658

After changing a title or meta description, wait at least a few days before judging the result in Search; if nothing changes after several weeks, check that the page was recrawled, since snippets and titles are only re-extracted from a reprocessed page.

Author Ibrahim Anjro

Things
SlideConsistent with docsD3-C662

Recovery after a core update typically takes 3 to 6 months once the site has put in work to regain traffic, and up to 6 months to a year, until the next core update.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extends D3-C294 Day 3: After a core update a site can win back some of its rankings by improving the site as a whole.
SlideConsistent with docsD3-C663

Google estimated that a spam update affects sites within 1-2 days of the rollout, 1-2 weeks for continuous processing (said on stage as 2 weeks on average) and months for batch refreshes.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extends D3-C286 Day 3: When spammers find loopholes that Google's systems miss, Google releases spam updates that change its systems…
DocsSourceD3-C667

Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.

Publisher Google Search Central, Google Search Console Help

Used byrequirement DEV-SRV-05

  • Extends D1-C528 Day 1: Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version…
AnalysisD3-C670

The canonicalization documentation update mentioned on stage appears in Google's documentation changelog on 10 July 2026 (re-evaluation time, page last updated 21 August 2026), earlier than 'about a month' before the event; its 'up to two weeks' sits at the short end of the one to three weeks said on stage.

Author Ibrahim Anjro

AnalysisD3-C676

The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.

Author Ibrahim Anjro

Used byrequirement DEV-REN-01

  • Extends D2-C131 Day 2: Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering…

16:00 · Wrapping all up: AI, Search, and making sense of everything. 36

StageD3-C678

Google said there is so much going on in Search that AI topics were covered across the event only where contextually important, and its closing message was to use AI but understand how it works so as to use it responsibly.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C679

Google described AI as an umbrella of many technologies working together, one of which is machine learning: systems that learn from large amounts of data to make informed decisions.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C680

Gary Illyes said machine learning is about 50 years old and that Google has been using it for 'probably 30 years', starting with statistical models that made the 'Did you mean' feature possible.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
  • Repeats D1-C209 Day 1: Google launched the 'Did you mean' feature around 2001-2002 using a statistical model, which Gary Illyes…
StageNot in docsD3-C682

Google described large language models as deep learning on internet-scale data sets, trained so that the model has an internal vector space in which concepts are mapped by context.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C683

Google distinguished predictive language models, whose primary job is to predict the next word or words in a text, from generative models that produce output from a prompt; next-word prediction is a guess, made with some accuracy, from the context given.

Speaker Gary IllyesEvidence transcript

StageConsistent with docsD3-C686

Hallucinations can happen with any AI model, and with current training methods there is no way to get rid of them.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SPM-05glossary term Hallucination

  • Extends D2-C462 Day 2: Google's guidance on generative AI content warns that AI output can contain hallucinations and says…
StageConsistent with docsD3-C688

Generative models, including image diffusion models, make things up when they lack information or because of issues in their training.

Speaker Gary IllyesEvidence transcript

Used byglossary term Hallucination

  • Extends D2-C939 Day 2: The diffusion models that generate images were built to generate images, not text, so they are typically poor…
StageConfirmed by docsD3-C690

Quality raters cannot give a site a penalty or a manual action; their ratings are converted into labels that Google uses to improve its algorithms.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SPM-05

  • Extends D3-C160 Day 3: Google's slide said search quality raters cannot affect the rankings of individual sites.
StageConsistent with docsD3-C691

Publishing AI output without checking it can indirectly harm a site's standing in Search, because deceptive content rated lowest by raters feeds the labels used to improve the algorithms.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SPM-05

StageConsistent with docsD3-C694

Scaled content abuse is becoming a problem again: in the early 2000s pages were churned out with Perl or PHP scripts, and now the same is done with LLMs.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SPM-02

  • Extends D3-C284 Day 3: Reviewing the spam slide, the speaker said link spam would probably be replaced with scaled content abuse as…
StageConsistent with docsD3-C695

The cheaper tokens become, the more AI slop is created, and Google counts AI slop as scaled content abuse.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SPM-02glossary term AI slop

  • Extends D2-C611 Day 2: Google's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings…
StageConsistent with docsD3-C697

Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem, but how it was used and what for.

“it's not the tool that was used that is the problem, but rather how the tool was used and what for.”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SPM-05

  • Extends D2-C611 Day 2: Google's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings…
  • Repeats D3-C190 Day 3: Google's quality talk said quality problems should be treated as quality issues, not as AI versus human…
StageD3-C699

Google encouraged brainstorming and even creating content and images with AI, while staying mindful of what AI can do, because it sometimes lies to its users and that hurts a business and its Google rankings.

Speaker Gary IllyesEvidence transcript

SlideConsistent with docsD3-C700

Google's closing slide said to use AI responsibly because AI hallucinates, and, especially when creating content briefs with AI, to make sure not to add to the sea of AI slop already flooding the internet.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-SPM-05

  • Extends D2-C941 Day 2: Sites that use AI-generated images or videos should make sure they work for users, check them for…
SlideConsistent with docsD3-C701

Google's closing slide said AI on Google is just SEO: AI features on Google Search use exactly the same processes as traditional results, so no new acronym is needed, as none was for mobile-first indexing or structured data.

“AI features on Google Search use exactly the same processes as traditional results.”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

Used bymyth M-001story angle A-001

  • Repeats D1-C050 Day 1: Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.
  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageConsistent with docsD3-C702

AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.

Speaker Gary IllyesEvidence transcript

Used bystory angle A-001

  • Repeats D2-C726 Day 2: AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a…
  • Repeats D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageD3-C704

Google warned that shiny new things, such as trying to work out fan-out queries, distract from the real mission of creating helpful content, since Search is just piping that connects users to information.

Speaker Gary IllyesEvidence transcript

  • Repeats D3-C065 Day 3: Because every system runs fan-out differently, Google advised understanding that fan-out happens but not…
StageNot in docsD3-C705

Search has changed every year since its start, for example the Florida update in 2003 (before core updates existed), Universal Search and the Knowledge Graph.

Speaker Gary IllyesEvidence transcript

StageD3-C706

Google said that each time Search changed someone declared SEO dead, but that new features just mean new opportunities.

Speaker Gary IllyesEvidence transcript

StageD3-C708

Google does not tell lightning-talk and poster speakers what to talk about; they bring their own ideas.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C227 Day 1: Google received close to 200 submissions for the community lightning talks of the Deep Dive, reviewed them…
StageD3-C709

Gary Illyes said he hoped the community lightning talks and posters confirmed Google's on-stage message, and judged from them that the vast majority of SEOs have a very good handle on what is happening in Search.

Speaker Gary IllyesEvidence transcript

DocsSourceD3-C710

Google's Search Quality Rater Guidelines (September 2025) list fake owner or content creator profiles, such as made-up author profiles with AI-generated images, as deception, and say pages using deception of any type should be rated Lowest.

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)

Used byrequirement DEV-SPM-05

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
DocsSourceD3-C711

Google's Search Quality Rater Guidelines say no single rating can directly change how a page appears in Google Search; ratings measure how well search works and give examples of helpful and unhelpful results to improve it.

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
DocsSourceD3-C713

Google updated its generative AI content guidance on 1 October 2026 with information from the Search Quality Rater Guidelines, to keep its documentation in sync with the presentations used at its developer events.

Publisher Google Search Central