Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Verification

Said at the event, not in the docs

Every slide and stage claim was checked against Google’s documentation. These are the ones the documentation does not cover: the most valuable claims, and the most fragile.

Slide and stage claims by verification grade
GradeDay 1Day 2Day 3
Confirmed by docs70153103
Consistent with docs147228183
Not in docs77217139
Nothing to verify149132102
Total443730527

How to quote these

A claim marked Not in docs was said or shown at the event, by Google or a community speaker, but is not in Google’s published documentation. Quote it as “said at Search Central Live Deep Dive 2026”, never as documented policy.

Where the event and the documentation say different things, the differences ledger puts both side by side and says what to follow.

Day 1: Crawling 77

Welcome and opening keynotes 12

SlideNot in docsD1-C004

Google says users like AI Overviews and search more because of them, with early feedback overwhelmingly positive and satisfaction highest among 18–24 year olds.

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D1-C184 Day 1: Google said AI Mode is growing almost exponentially and that people are searching more because of it, not…
  • Extended by D3-C569 Day 3: Google's research found that people who use AI run more searches than people who do not.
SlideNot in docsD1-C005

Google says AI Overviews help with new types of questions and lead users to visit a greater diversity of sites, with more total sites shown on the results page.

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D3-C571 Day 3: Google's research found that people who use AI also visit more brand websites, because it is simply easier.
SlideNot in docsD1-C007

Google says 94% of frequent LLM users are also frequent users of Google Search, as evidence that AI is growing search use rather than replacing it.

Speaker Lino CattaruzziEvidence slide photo, transcript

  • Extended by D1-C499 Day 1: The keynote added that people who also use a standalone LLM keep certain types of use anchored on Google (the…
SlideNot in docsD1-C008

Google says 71% of shoppers coming to Google Search agree they are open to trying new brands or products (the slide footnote cites a Google-commissioned Ipsos global consumer survey), and that people use Search to figure out what they want, not only to find what they already know.

Speaker Lino CattaruzziEvidence slide photo, transcript

StageNot in docsD1-C155

The keynote cited Google CEO Sundar Pichai as saying Google is the biggest contributor of clicks to the open web and wants to remain so.

Speaker Lino CattaruzziEvidence transcript

StageNot in docsD1-C158

Google said that even the classic ten-blue-links layout was settled only after millions of experiments, as part of using as much data as possible for product decisions.

Speaker Lino CattaruzziEvidence transcript

  • Extended by D2-C452 Day 2: Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
  • Extended by D3-C147 Day 3: Google's quality talk said that at any moment thousands of Search experiments are probably running.
StageNot in docsD1-C162

Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the number one app in the app store in 2025.

Speaker Lino CattaruzziEvidence transcript

Things
  • Extended by D2-C770 Day 2: Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of…
StageNot in docsD1-C166

Google said 18 to 24 year olds are more engaged than ever with the AI-powered features in Search: they requery and ask more, and more complex, questions, which leads to more searches.

Speaker Lino CattaruzziEvidence transcript

StageNot in docsD1-C167

Google said its searches are growing year on year in both commercial and non-commercial queries.

Speaker Lino CattaruzziEvidence transcript

StageNot in docsD1-C499

The keynote added that people who also use a standalone LLM keep certain types of use anchored on Google (the example given is unclear in both recordings).

Speaker Lino CattaruzziEvidence transcript

  • Extends D1-C007 Day 1: Google says 94% of frequent LLM users are also frequent users of Google Search, as evidence that AI is…

What's new in the world of Search 14

SlideNot in docsD1-C016

AI Mode queries have doubled every quarter since launch.

Speaker Gary IllyesEvidence slide photo

Things
  • Extended by D1-C184 Day 1: Google said AI Mode is growing almost exponentially and that people are searching more because of it, not…
SlideNot in docsD1-C017

The average AI Mode query is about 3 times the length of a traditional search (the opening keynote's slide and its speaker said two to three times).

Speaker Gary IllyesEvidence 2 slide photos, transcript

Things

Used byglossary term AI Mode

SlideNot in docsD1-C018

One in six AI Mode searches is multimodal, using voice or images.

Speaker Gary IllyesEvidence slide photo

Things
  • Extended by D3-C562 Day 3: Google's research also counted easy access as part of AI's ease benefit: multimodal input, such as taking…
SlideNot in docsD1-C020

Google describes four new ways people search in AI Mode: Explore (brainstorming queries, grown 30% faster than AI Mode queries overall in the US since launch), Learn (study guides, deep dives), Decide (searches beginning with 'which', up 40% faster than AI Mode queries overall in the past six months) and Do (planning queries, up 80% faster than AI Mode queries overall in the past six months).

Speaker Gary IllyesEvidence 2 slide photos, transcript

Things
SlideNot in docsD1-C021

Queries of five or more words are growing in volume 1.5 times faster than shorter queries. The slide footnote cites Google internal data on global English-language queries, for a comparison period that starts in November 2022 (the remaining dates are only partly legible).

Speaker Gary IllyesEvidence slide photo

  • Extended by D1-C185 Day 1: Google said a multi-part query, such as a white three-row SUV for two teenagers and a car-sick dog, cannot be…
SlideNot in docsD1-C022

Search Agents in AI Mode let users set standing requests for updates, such as being told when a product goes on sale, a team wins or a price drops below a level.

Speaker Gary IllyesEvidence slide photo, transcript

Things
  • Extended by D1-C192 Day 1: Google said Search agents in AI Mode had launched globally a day or two before 30 September 2026.
  • Extended by D1-C194 Day 1: A Search agent is set up by typing a natural-language request, which AI Mode interprets to create an alert…
SlideNot in docsD1-C023

Google listed eight ways AI Search helps users find and visit websites: preferred sources expansion, fresh perspectives and prominent links, subscription content, in-line quotes in AI Mode and AI Overviews, rich linking and embedded web, context on content sources, 'Highly Cited' labels and Search profiles.

Speaker Gary IllyesEvidence slide photo, transcript

SlideNot in docsD1-C024

Users can link their subscriptions so answers highlight paywalled content they already pay for; The Indian Express saw 34% more engagement among linked subscribers.

Speaker Gary IllyesEvidence slide photo, transcript

SlideNot in docsD1-C027

Sites that users mark as a preferred source are more visible in Top Stories and are labelled in AI Mode and AI Overviews.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extended by D3-C317 Day 3: Google can add elements to a text result, such as a 'highly cited' badge or a preferred-source badge.
StageNot in docsD1-C183

Google said Gen Z users are a very large part of Google Search's user base and search differently from older users, for example with images.

Speaker Gary IllyesEvidence transcript

StageNot in docsD1-C188

Google said it noticed young users uploading pictures to reverse image search and expecting results that answer something about the picture, and that Google Lens was its answer to that behaviour.

Speaker Gary IllyesEvidence transcript

Used byglossary term Google Lens

StageNot in docsD1-C192

Google said Search agents in AI Mode had launched globally a day or two before 30 September 2026.

Speaker Gary IllyesEvidence transcript

Things

Used byglossary term Search agents (information agents)

  • Extends D1-C022 Day 1: Search Agents in AI Mode let users set standing requests for updates, such as being told when a product goes…
StageNot in docsD1-C199

Gary Illyes said searching by image or by video, as Search offers it today, was not possible before transformers were invented.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C161 Day 1: Google said it created the Transformer architecture (the paper 'Attention Is All You Need', the T in ChatGPT)…

How Search works and where's AI? 10

SlideNot in docsD1-C041

Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Extended by D1-C209 Day 1: Google launched the 'Did you mean' feature around 2001-2002 using a statistical model, which Gary Illyes…
  • Extended by D2-C668 Day 2: Google uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is…
  • Extended by D2-C669 Day 2: SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically…
  • Extended by D3-C680 Day 3: Gary Illyes said machine learning is about 50 years old and that Google has been using it for 'probably 30…
SlideNot in docsD1-C044

MUM (Multitask Unified Model) understands information across text, images, audio and video, and processes information in more than 75 languages.

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo

  • Extended by D1-C218 Day 1: Google said MUM helps it understand the context of the words in a query: a search for hiking shoes that…
SlideNot in docsD1-C045

Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Repeated by D2-C537 Day 2: The text around an image is critical: Google uses it as context to understand the image and to rank it, so an…
  • Extended by D2-C656 Day 2: Freshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking…
  • Repeated by D3-C142 Day 3: Google's quality talk said ranking signals differ by result type: for web pages they include the text on the…
SlideNot in docsD1-C047

Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so they understand and promote natural content better.

“ML based ranking algorithms and signals are trained on content by humans for humans. They "understand" and promote natural content better.”

Wording checked against the slide or recording

Speaker Cherry Prommawin, Gary IllyesEvidence slide photo, transcript

  • Answers D1-C046 Day 1: A Google slide headed 'Your Question' raised whether Google can distinguish between AI-written and…
  • Extended by D1-C220 Day 1: Google said the more important documents in its index naturally tend to be documents created by humans.
  • Extended by D1-C221 Day 1: Google said the natural content its ranking algorithms understand and promote better includes content created…
  • Repeated by D3-C116 Day 3: Google's long-standing advice to write for people and give them what they want has become true in practice…
StageNot in docsD1-C201

Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many exist, and that some may never be discovered.

Speaker Cherry PrommawinEvidence transcript

  • Extended by D3-C602 Day 3: Google said it knows hundreds of trillions of URLs (as of October 2026).
StageNot in docsD1-C213

Google said its index, printed on paper, would reach the Moon and back twelve times.

Speaker Cherry PrommawinEvidence transcript

StageNot in docsD1-C220

Google said the more important documents in its index naturally tend to be documents created by humans.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C047 Day 1: Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so…
StageNot in docsD1-C221

Google said the natural content its ranking algorithms understand and promote better includes content created by humans or at least edited and reviewed by them.

“content that was created by humans, or at least edited and reviewed”

Speaker Gary IllyesEvidence transcript

  • Extends D1-C047 Day 1: Google's answer was that ML-based ranking systems are trained on content written by humans for humans, so…
StageNot in docsD1-C224

For visual search, Google breaks an image down into vectors, sends the vectors to the index and returns results based on them.

Speaker Gary IllyesEvidence transcript

  • Repeated by D3-C311 Day 3: Google handles an image used as a search query much like a text query interpreted as an embedding: the image…

Lightning session A: Automation and AI 6

StageNot in docsD1-C276

A community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which number is the price saves the agent from guessing.

Speaker Carlos OrtegaEvidence transcript

  • Extended by D1-C315 Day 1: Google's AI optimisation guide says structured data is not required for generative AI search and needs no…
StageNot in docsD1-C295

A community speaker showed a Spanish-language AI prompt about places to eat in Barcelona whose fan-out queries were all in English, so the content used to answer may be written by and for a different audience than the one asking.

Speaker Thiago PojdaEvidence transcript

How crawling works 6

StageNot in docsD1-C318

A URL states how a resource is requested (the protocol, HTTP or HTTPS), where (the host, meaning which computer on the network) and what (the path to the exact page or file).

Speaker Cherry PrommawinEvidence transcript

StageNot in docsD1-C323

AI agents are, technically, the same thing as crawlers: HTTP clients that accomplish something on behalf of a user or a service.

Speaker Gary IllyesEvidence transcript

StageNot in docsD1-C333

The scheduler hands the crawler an ordered list of URLs from the crawl queue, and the crawler works through the list from top to bottom.

Speaker Gary IllyesEvidence transcript

Used byglossary term Crawl scheduler

StageNot in docsD1-C335

Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

“the number of tokens is actually more important”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

Things
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageNot in docsD1-C338

Google described its crawling as a large-scale distributed swarm of simple HTTP clients, roughly what one would get by deploying many wget or curl libraries on cloud compute instances.

“a large-scale distributed swarm of simple HTTP clients”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence transcript

How crawling errors affect Search 3

StageNot in docsD1-C346

1xx informational status codes, which only say that a request was received and more data is coming, have no meaning of their own for crawling.

Speaker Cherry PrommawinEvidence transcript

StageNot in docsD1-C359

Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-SRV-01

  • Extends D1-C069 Day 1: DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that…
StageNot in docsD1-C509

Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-SRV-09

  • Extends D1-C366 Day 1: Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
  • Extends D1-C365 Day 1: Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication…

How Google interprets robots.txt 7

SlideNot in docsD1-C123

Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

“Unpredictability is never a good time.”

Wording checked against the slide or recording

Speaker GoogleEvidence slide photo, transcript

Used byrequirement DEV-SRV-06

  • Extended by D1-C526 Day 1: The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked…
StageNot in docsD1-C089

The robots.txt talk pointed to the robots.txt file of thebestfriedchickenever.com to show why comments in robots.txt are useful; the live file is mostly a large ASCII-art drawing written as # comment lines, followed by a single user-agent group (checked 2026-10-04).

Speaker GoogleEvidence notes, transcript

StageNot in docsD1-C525

Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-05

  • Extended by D1-C532 Day 1: A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts…
StageNot in docsD1-C526

The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-06

  • Extends D1-C087 Day 1: Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and…
  • Extends D1-C123 Day 1: Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's…

Lightning session B: Robots.txt 2

StageNot in docsD1-C534

Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.

Speaker Dave SmartEvidence transcript

Used byrequirement DEV-CAN-11glossary term Redirect chain

StageNot in docsD1-C536

Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.

Speaker Dave SmartEvidence transcript

Used byrequirement DEV-CAN-11

How Google thinks about crawl budget 2

SlideNot in docsD1-C094

If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

“If quality or popularity is unknown, use parent root's aggregate quality or popularity is used, then, that path's parent's, and so on.”

Wording checked against the slide or recording

Speaker Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-URL-10

  • Extended by D2-C369 Day 2: Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern…
  • Extended by D2-C685 Day 2: Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies…
StageNot in docsD1-C377

The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.

Speaker Cherry PrommawinEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
  • Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…

Lightning session C: Crawling 4

StageNot in docsD1-C402

Similar URLs usually come from programming errors or manually set links, but anyone can link to them from other sites, including malicious actors, so a site should be hardened against them.

Speaker Tobias SchwarzEvidence transcript

Used byrequirement DEV-URL-11

StageNot in docsD1-C407

A community speaker cited more than 900 million weekly active ChatGPT users and 2.5 billion monthly users of a Google AI feature (the recording is unclear which), adding that these are not comparable market-share figures.

Speaker Jovana AvramovicEvidence transcript

StageNot in docsD1-C414

A community speaker linked the serial-position effect in human memory (people best remember the first and last items of a list) to the 'lost in the middle' pattern that research has found in large language models.

Speaker Jovana AvramovicEvidence transcript

Used byglossary term Lost in the middle

Q&A 10

StageNot in docsD1-C425

Crawling for AI model training differs from search crawling because training mainly needs a very large number of tokens, and it matters little which pages they come from.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
StageNot in docsD1-C432

A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

Speaker not identifiedEvidence transcript

  • Extended by D1-C433 Day 1: The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices'…
StageNot in docsD1-C442

Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the site lies; many similar URLs being crawled may not matter if the site is nowhere near that limit.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-MON-04

  • Answers D1-C441 Day 1: An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl…
StageNot in docsD1-C446

Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
StageNot in docsD1-C467

To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Extended by D1-C543 Day 1: Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises…
StageNot in docsD1-C478

The first priority in Google's own site consolidation was to identify the popular URLs people care about and make sure the migration did not damage them.

Speaker not identifiedEvidence transcript

  • Answers D1-C476 Day 1: An audience member asked how to plan a site migration so that it does not leave large numbers of URLs not…
StageNot in docsD1-C487

A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Repeated by D1-C524 Day 1: Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.
StageNot in docsD1-C488

When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
StageNot in docsD1-C541

A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in Google's own site migration, because it was the only option available to the person doing it (the recording does not make fully clear whether the JavaScript performed the redirects or built the mapping).

Speaker not identifiedEvidence transcript

Used byrequirement DEV-CAN-01

  • Extends D1-C479 Day 1: Where content was duplicated across languages, Google's own site consolidation redirected two language…
  • Extended by D1-C542 Day 1: Google's redirects guide says Google Search follows JavaScript location redirects only after rendering, may…
StageNot in docsD1-C543

Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Extends D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…

Day 1, session not recorded 1

StageNot in docsD1-C122

Check and focus on rich results for Google.

Evidence notes

  • Extended by D2-C521 Day 2: Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its…

Day 2: Indexing 217

Welcome to indexing day! 6

SlideNot in docsD2-C002

Google answered on a Q&A slide that good HTML sitemaps are hard to make for large sites.

“It's hard to make good HTML sitemaps for large sites.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo

  • Answers D2-C001 Day 2: An audience member asked whether HTML sitemaps still matter for large sites, whether they improve Googlebot's…
SlideNot in docsD2-C010

Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it depends on how important that product page is to users.

“It depends on the importance of the PDPs to the users.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Things

Used byrequirement DEV-ERR-04

  • Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
  • Extended by D2-C841 Day 2: Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to…
SlideNot in docsD2-C020

Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

“Simply put, it's because of sites that are extremely important and like to disallow their most important pages, either accidentally or out of ignorance.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinEvidence slide photo, transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
  • Extended by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
StageNot in docsD2-C841

Google contrasted out-of-stock products a buyer will wait for, such as a rare watch battery that matters to them, with easily replaced products, such as a particular cheese, for which the buyer simply picks a similar product instead of waiting.

Speaker not identifiedEvidence transcript

  • Answers D2-C009 Day 2: An audience member asked whether a product detail page that is out of stock for two to three months should…
  • Extends D2-C010 Day 2: Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it…
StageNot in docsD2-C844

Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…

How is HTML interpreted 3

SlideNot in docsD2-C024

A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

Speaker Cherry PrommawinEvidence slide photo

Used byrequirement DEV-PRF-01

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C092 Day 1: Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status…
StageNot in docsD2-C046

The speaker said Google sometimes also extracts URLs that are typed out as plain text on a page without being hyperlinked; the remarks around this point were unclear in the recording.

Speaker Cherry PrommawinEvidence transcript

Controlling indexing 6

SlideNot in docsD2-C062

A slide named internal admin pages, temporary landing pages and thin content as typical use cases for the noindex rule.

Speaker John MuellerEvidence slide photo

Things
StageNot in docsD2-C066

John Mueller said AI crawlers do not really know what to do with nofollow links, because they look at the content rather than building a link graph; he did not say whether he meant Google's AI systems, other AI crawlers or both.

“AI crawlers don't really know what to do with a nofollow link, because they're looking at the content”

Speaker John MuellerEvidence transcript

Things
SlideNot in docsD2-C085

max-snippet:-1 removes the length limit and can produce a longer snippet than having no rule, because by default Google keeps snippets to a length it considers reasonable instead of quoting a page at length.

“max-snippet:-1 = No limit. (can be more than without)”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

StageNot in docsD2-C088

John Mueller said Google Search does not always show an image thumbnail for a result, and when it does, the thumbnail is usually limited to the size of the result entry.

Speaker John MuellerEvidence transcript

Lightning session D: Rendering and JavaScript 45

StageNot in docsD2-C113

Compared with other search engines and AI crawlers, Google said its effort to mimic what the user sees is essential to keeping its knowledge of the web up to date and comprehensive; it did not say what the others do.

Speaker Erin SparlingEvidence transcript

StageNot in docsD2-C115

Google described the mission of its rendering as simply executing JavaScript, while the implementation is complex, expensive and difficult.

“our mission is very simple these days: execute JavaScript”

Speaker Erin SparlingEvidence transcript

StageNot in docsD2-C118

To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.

“if it works for Search, it works for Gemini for training”

Speaker Erin SparlingEvidence transcript

Things

Used byrequirement DEV-IDX-11

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageNot in docsD2-C120

When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.

Speaker Erin SparlingEvidence transcript

Things

Used byrequirement DEV-PRF-02

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageNot in docsD2-C129

After the crawler has fetched a page, Google's rendering system executes it, checks that it loads properly and works out what it looks like at different sizes.

Speaker Erin SparlingEvidence transcript

Things
StageNot in docsD2-C138

Switching JavaScript off and on in the browser gives a quick first impression of which content on a page depends on rendering.

Speaker Sören BendigEvidence transcript

StageNot in docsD2-C149

A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal systems or APIs can be missing and unresolved placeholders can end up in the final snapshot, a community speaker said.

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-REN-06

  • Extended by D2-C259 Day 2: Erin Sparling named server-side or hybrid rendering, fallbacks and not leaving placeholders in the DOM, among…
  • Extended by D2-C263 Day 2: Content is missing from the rendered HTML either because the server does not serve it or because the content…
StageNot in docsD2-C153

To catch intermittent rendering problems, archive the full HTML and the resources of each page during a site audit and analyse them yourself, a community speaker advised.

Speaker Sören BendigEvidence transcript

Things

Used byrequirement DEV-MON-05

StageNot in docsD2-C155

Content Security Policy lowers the risk of cross-site scripting and clickjacking by defining trusted hosts; a resource from a host that is not trusted is not used.

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-REN-09

StageNot in docsD2-C162

A community speaker strongly advised SEOs to use Chrome DevTools more often to understand what happens when their pages render.

Speaker Sören BendigEvidence transcript

Things
  • Extended by D2-C860 Day 2: Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a…
SlideNot in docsD2-C171

A community speaker's slide put the usual wait in Google's render queue at seconds to a couple of minutes per page.

“Usually from seconds to a couple minutes.”

Wording checked against the slide or recording

Speaker Rebecca YuEvidence slide photo

StageNot in docsD2-C180

If content is missing from the rendered HTML, find the script responsible in Chrome DevTools: open the Network tab, filter by Fetch/XHR and reload the page.

Speaker Rebecca YuEvidence transcript

Things
StageNot in docsD2-C182

If the cause of missing JavaScript content is still unclear after checking the network requests, search the source code for a string related to the missing content.

Speaker Rebecca YuEvidence transcript

StageNot in docsD2-C219

Natalia Venditto named three typical failure modes by which an embedded app can break the application hosting it: collisions in the global JavaScript scope and module registry, CSS bleed, and fate sharing.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C220

CSS bleed, as Natalia Venditto described it, means styles leak between an embedded app and its host, so that suddenly everything looks like the embedded app or the embedded app takes on the look of the host.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C221

Fate sharing means a failure in an embedded app spreads to its host, for example an unhandled exception in the embedded app that ends up breaking the host application.

Speaker Natalia VendittoEvidence transcript

SlideNot in docsD2-C224

An iframe, the long-established way to embed an app, isolates both JavaScript and styles but is not part of the page: it is walled off from the host's DOM, navigation and layout, which Natalia Venditto said brings many problems with accessibility, layout and navigation.

“Walled off from DOM, navigation, layout”

Wording checked against the slide or recording

Speaker Natalia VendittoEvidence slide photo, transcript

SlideNot in docsD2-C228

Shadow DOM isolates styles and keeps the embedded content part of the page, but it does not isolate JavaScript: the embedded code still shares the host's JavaScript globals.

“Still shares JS globals”

Wording checked against the slide or recording

Speaker Natalia VendittoEvidence slide photo, transcript

Used byglossary term Shadow DOM

SlideNot in docsD2-C229

Rewriting everything into one application keeps the result part of the page, but JavaScript and style isolation then have to be done by hand.

Speaker Natalia VendittoEvidence slide photo, transcript

SlideNot in docsD2-C231

Reframing with Web Fragments was the only approach on the comparison slide ticked for all three criteria (JavaScript isolated, styles isolated, part of the page); its stated catch is that it relies on browser patches today.

“Needs patches today”

Wording checked against the slide or recording

Speaker Natalia VendittoEvidence slide photo, transcript

StageNot in docsD2-C232

Web Fragments is a lightweight library that puts a hidden iframe in the host page and uses it as an isolated sandbox to run the embedded app's JavaScript, not to render the app.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C233

Web Fragments fetches the embedded app's assets from a remote endpoint through middleware that acts as a gateway.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C234

Web Fragments reframes the embedded app's DOM by placing it inside a shadow DOM in the host page, while the app's JavaScript runs in the hidden iframe.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C236

With Web Fragments the page ends up as a single document in which the host application does not know the embedded app runs inside it and the embedded app does not know it lives in another document.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C237

Web Fragments monkey-patches browser APIs to containerize the browser, recreating inside it an architecture similar to Docker containers on the back end.

“we monkey patch the browser to containerize it”

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C238

The browser APIs Web Fragments patches include document, history and location; Natalia Venditto said they are virtualized rather than hijacked, so controls work exactly as in a normal application.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C239

Web Fragments retargets dispatchEvent calls to the fragment's shadow root and resolves DOM calls such as appendChild and element lookups against the main document.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C241

Natalia Venditto said Web Fragments suits AI-generated apps because a web fragment used as a custom element sandboxes the app's rendering, avoiding collisions and CSS bleed between the app and its host in either direction.

Speaker Natalia VendittoEvidence transcript

SlideNot in docsD2-C242

A slide set out Web Fragments in five steps: an LLM writes the app and adds a custom element such as <web-fragment fragment-id="some-id"> to the HTML, a standalone HTTP endpoint is set up, the library is imported and initialized with initializeWebFragments(), the fragment is registered in the gateway, and step 5 is a celebration emoji.

Speaker Natalia VendittoEvidence slide photo, transcript

StageNot in docsD2-C243

Natalia Venditto said the five Web Fragments steps are all that is needed to run a containerized application fully on the client side.

Speaker Natalia VendittoEvidence transcript

SlideNot in docsD2-C244

A web fragment is served from its own standalone HTTP endpoint, so the embedded app can be deployed anywhere, separately from the host page.

Speaker Natalia VendittoEvidence slide photo, transcript

StageNot in docsD2-C246

Natalia Venditto said Web Fragments is framework-agnostic and vendor-agnostic and works the same whatever JavaScript framework the embedded app uses.

Speaker Natalia VendittoEvidence transcript

StageNot in docsD2-C860

Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a modern crawler.

Speaker Sören BendigEvidence transcript

Used byrequirement DEV-MON-05

  • Extends D2-C162 Day 2: A community speaker strongly advised SEOs to use Chrome DevTools more often to understand what happens when…

What is Google friendly JavaScript 9

StageNot in docsD2-C263

Content is missing from the rendered HTML either because the server does not serve it or because the content has still not appeared after some period of time.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-REN-06

  • Extends D2-C149 Day 2: A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal…
StageNot in docsD2-C272

Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.

“it renders it in a very tall viewport. Specifically, around 10,000 pixels is what the viewport gets rendered as.”

Speaker Erin SparlingEvidence slide photo, transcript

Used byrequirements DEV-REN-03, DEV-URL-07glossary term Viewport expansion

  • Extends D2-C199 Day 2: An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders…
StageNot in docsD2-C289

Real URLs also work as deep links: Android, iOS and desktop operating systems accept full URLs as keys to specific content in an app, so clean URLs simplify the cross-platform user experience, not only indexing.

Speaker Erin SparlingEvidence transcript

Used byrequirement DEV-URL-03

StageNot in docsD2-C299

A headless content management system serves its content as an API; even WordPress, often seen as one monolithic application, can be used headless, with only its editor or only its renderer.

Speaker Erin SparlingEvidence transcript

Understanding what's on a page 16

StageNot in docsD2-C313

Words in the footer of a page get a lower weight, so text placed in the footer is unlikely to contribute much to ranking the page.

“if you put something in a footer, it's more likely that it's not going to contribute much to ranking”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-02

StageNot in docsD2-C316

Gary Illyes said not everything on a page can or should be important: if everything were placed in the main content, nothing would stand out as main content, which he called working as intended.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C318

Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

Speaker Gary IllyesEvidence slide photo, transcript

Used byglossary term Tokenization

  • Repeated by D2-C721 Day 2: Google's Search index does not hold the full content of pages; Google said storing full pages and pulling…
  • Extended by D3-C109 Day 3: The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens…
StageNot in docsD2-C320

Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
  • Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
StageNot in docsD2-C321

For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

Speaker Gary IllyesEvidence transcript

  • Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageNot in docsD2-C323

When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-HTM-07

  • Repeated by D2-C720 Day 2: Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the…
StageNot in docsD2-C325

Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

Speaker Gary IllyesEvidence slide photo, transcript

  • Contradicts D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
SlideNot in docsD2-C328

Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.

Speaker Gary IllyesEvidence 2 slide photos

StageNot in docsD2-C825

Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-03

  • Extends D2-C332 Day 2: Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller…
StageNot in docsD2-C335

Because error pages are worded in endless variations, of which 'page not found' is only the classic one, Google cannot detect soft 404s with simple error, word or keyword matching.

Speaker Gary IllyesEvidence transcript

Things
SlideNot in docsD2-C337

Mistakes in Google's own systems are a further cause of soft 404s, and Gary Illyes asked site owners to report such mistakes in Google's forums.

“BONUS: Mistakes in Google's systems (that you should notify us about)”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

Things
StageNot in docsD2-C338

Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

“This is basically an LLM thing, something like BERT, that is specifically trained to understand page structure”

Speaker Gary IllyesEvidence transcript

Things
  • Extends D1-C042 Day 1: BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at…
StageNot in docsD2-C339

For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-ERR-03

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…

Handling web duplication 10

StageNot in docsD2-C364

Google clusters duplicate pages by content in four ways: exact matches, near matches, structurally similar content, and soft 404s.

Speaker John MuellerEvidence transcript

Things
SlideNot in docsD2-C368

Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirements DEV-CAN-08, DEV-URL-09

SlideNot in docsD2-C369

Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

“Do we even need to crawl /buy/seo-service ?”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo

Used byrequirements DEV-CAN-08, DEV-URL-09

  • Extends D1-C128 Day 1: Google's canonicalization guide says some duplicate content on a site is normal and not a violation of its…
  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
SlideNot in docsD2-C370

City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-CAN-08

StageNot in docsD2-C377

AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-AIF-04

  • Extends D1-C436 Day 1: A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to…
StageNot in docsD2-C386

Google picks the canonical from a variety of criteria and uses some kind of machine learning to decide how much weight each criterion gets; the weighting changes from time to time.

“we use some kind of machine learning to understand how strong these criteria should be. And this changes from time to time.”

Speaker John MuellerEvidence transcript

StageNot in docsD2-C393

Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-URL-06

  • Contradicts D1-C115 Day 1: Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the…
SlideNot in docsD2-C397

Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.

“Don't block agents.”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

Used byrequirement DEV-AIF-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…

Lightning session E: Managing Duplicates and Site Moves 5

StageNot in docsD2-C426

A community speaker said that at best a search engine's heuristic would treat a canonical loop through redirects as self-referencing canonicals, and doubted that this is often done when the loop runs through a client-side redirect.

“But regarding the client-side redirect, I highly doubt that this is often done.”

Speaker Tobias SchwarzEvidence transcript

StageNot in docsD2-C429

A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.

Speaker Tobias SchwarzEvidence transcript

  • Contradicts D2-C348 Day 2: Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the…

Finding the gold nuggets: structured data, media, and more! 5

SlideNot in docsD2-C441

Google's indexing pipeline slide put a Processing stage between the crawler and the index and listed HTML parsing, rendering, deduplication and feature extraction above it, with feature extraction highlighted as the step of this talk.

Speaker Gary IllyesEvidence slide photo, transcript

  • Extends D1-C063 Day 1: The crawl pipeline runs from a crawl queue to a scheduler to the crawler, which fetches from the internet and…
  • Repeats D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
SlideNot in docsD2-C442

Google's indexing pipeline slide drew a feedback line from the Processing stage back to the crawl queue.

Speaker Gary IllyesEvidence slide photo

  • Repeats D2-C048 Day 2: Links extracted during processing are sent back to the crawl queue, where scheduling starts again for the…
StageNot in docsD2-C443

Google's indexing includes a dedicated system, whose internal name Gary Illyes would not disclose, that extracts the parts of a page that are traditionally expensive to extract.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C449

Gary Illyes said feature extraction, which extracts page structures into a form Google's systems can consume internally, is still expensive, though not the most expensive operation.

Speaker Gary IllyesEvidence transcript

What is Structured Data and why we need it on the internet. 24

SlideNot in docsD2-C458

Google gave four reasons why structured data is still valuable even though models can extract information from pages: precision, extra content, efficiency and focus.

Speaker Ryan LeveringEvidence slide photo, transcript

SlideNot in docsD2-C459

Structured data gives the high precision that complex schemas such as sale pricing need, with higher accuracy than large-scale extraction by large language models (LLMs).

“Structured data provides the high precision needed for complex schema (sale pricing), achieving higher accuracy than large-scale LLM extraction.”

Wording checked against the slide or recording

Speaker Ryan LeveringEvidence slide photo, transcript

StageNot in docsD2-C460

In the speaker's own tests, even the latest LLMs asked to generate schema.org markup for a page often invent properties that do not exist, get deeply nested schemas such as complex pricing models wrong, and duplicate content across several fields.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-11

SlideNot in docsD2-C463

Structured data often carries non-visible metadata that the page text lacks, such as full ISO dates or stable identifiers for user-generated content.

“It often contains non-visible metadata, such as full ISO dates or stable identifiers for UGC, that is not present in the page text.”

Wording checked against the slide or recording

Speaker Ryan LeveringEvidence slide photo, transcript

Used byrequirements DEV-SDA-04, DEV-SDA-05

SlideNot in docsD2-C469

Parsing structured data is significantly cheaper and more efficient than relying on LLMs for every complex extraction task.

Speaker Ryan LeveringEvidence slide photo, transcript

  • Repeated by D3-C405 Day 3: Web markup is an efficient and unambiguous way for sites to share product data with Google, Google Shopping…
StageNot in docsD2-C471

Rule-based parsing of markup is nearly free by comparison with AI models, so Google will always prefer extracting information from structured data over model-based extraction.

“So we're always going to prefer that particular approach.”

Speaker Ryan LeveringEvidence transcript

StageNot in docsD2-C474

Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-07

  • Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
StageNot in docsD2-C477

As far as the speaker knows, in most of Google's main AI uses a page's schema.org markup is not turned into text and put directly into the model's context; the data is first sorted out, checked for quality and indexed before it is passed on as grounding context.

Speaker Ryan LeveringEvidence transcript

  • Extends D1-C061 Day 1: Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in…
StageNot in docsD2-C496

Microdata has an advantage when payload size matters: embedded in the existing HTML, it avoids duplicating the page content in a separate JSON-LD block.

Speaker Ryan LeveringEvidence transcript

Things

Used byrequirement DEV-SDA-01

StageNot in docsD2-C503

Several plug-ins emitting the same markup type is one of the most common structured data problems: the duplicates can make an event details page look like a list of events and change how Google interprets the page.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-06

StageNot in docsD2-C510

Schema.org, in which Google is a major participant, began publishing usage statistics for all its types and properties in 2026, showing in buckets how many domains use each one; the data is also in schema.org's GitHub repository.

Speaker Ryan LeveringEvidence transcript

StageNot in docsD2-C511

Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a chicken-and-egg problem that the schema.org usage statistics are meant to help break.

Speaker Ryan LeveringEvidence transcript

  • Extends D1-C120 Day 1: To get Google to build something, such as an addition to an API, report and request it publicly and in volume.
StageNot in docsD2-C512

Schema.org added support for RDF lists and sets, which give a way to express ordered values, because RDF triples are not ordered by nature.

Speaker Ryan LeveringEvidence transcript

SlideNot in docsD2-C514

Google announced server-side structured data validation as coming soon: it will publish downloadable validation rules in SHACL on each structured data feature guide, with other kinds of checks to follow.

Speaker Ryan LeveringEvidence slide photo, transcript

Used byrequirement DEV-SDA-12glossary term SHACL

SlideNot in docsD2-C515

Google's planned validation workflow has five steps: download the rules from the feature guide, generate the JSON or embedded microdata or RDFa, run the rules against the generated server-side markup as a first check, deploy and test in the Rich Results Test, and monitor ongoing performance in Search Console.

Speaker Ryan LeveringEvidence slide photo

Used byrequirement DEV-SDA-12

StageNot in docsD2-C516

The SHACL rules are meant to run inside a site's content generation, so markup is sanity-checked before it is published and does not silently regress later, a breakage site owners might otherwise discover only through a Search Console report.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-12glossary term SHACL

StageNot in docsD2-C517

The SHACL rules will not replace Search Console as the canonical place for structured data reports, because some checks use Google's internal libraries and cannot be expressed in SHACL, but they will catch problems such as a missing required field.

Speaker Ryan LeveringEvidence transcript

Used byrequirement DEV-SDA-12

Using images to your advantage and Engaging Search users with videos 21

StageNot in docsD2-C531

Of the attributes the HTML standard defines for the img element, Google uses three and ignores the rest; Gary Illyes named src and alt but not the third.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C533

Gary Illyes ranked the alt attribute below src in importance, while saying alt attributes are still important.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C542

Gary Illyes said the AVIF image format currently has hiccups and Google may have problems ingesting it, although it should technically be supported.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-05

StageNot in docsD2-C546

Gary Illyes cited a figure that more than 40% of Southeast Asian shoppers rely on videos to make purchase decisions.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C547

Gary Illyes said over 219 million people in Southeast Asia consume content on or through YouTube daily (the daily scope was heard in two independent recordings but is not verified).

Speaker Gary IllyesEvidence transcript

Things
StageNot in docsD2-C548

Gary Illyes said there are over 150 streaming apps in Southeast Asia.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C917

Google knows from experiments that when a video's thumbnail is wrong, the share of viewers who drop out at the start of the video is extremely high.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-03

StageNot in docsD2-C919

Gary Illyes called the 'fast-loading pages' factor for video SEO a misnomer: what matters is that the video itself loads fast, because people no longer have the patience to wait for videos to load.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

StageNot in docsD2-C920

Serving videos from a content delivery network (CDN) that loads them faster than the site's own server is a win for video SEO, Gary Illyes said.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-VID-06

StageNot in docsD2-C922

Google recommends the MP4 container for videos because of its browser and device compatibility.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

StageNot in docsD2-C923

Encode videos with standard codecs, because some people will not be able to play videos that use unusual ones.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

StageNot in docsD2-C924

For a video to be discovered it must be embedded prominently, above the fold; Gary Illyes said a video placed below the fold is not going to be indexed.

“If it's not above the fold, then you basically lost the game.”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-01

StageNot in docsD2-C929

The more people talk about a site's videos, the more likely Google is to surface them in search results, so videos should be promoted.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C930

Google's media indexer processes the images and videos that feature extraction passes to it and attaches them to the URL of the page that hosts them.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C448 Day 2: For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense…
StageNot in docsD2-C933

Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-VID-06

  • Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
StageNot in docsD2-C938

Google serves AI-generated images and videos in search results when users are specifically looking for them.

“We will serve users AI-generated images and videos in search results if they are looking for them specifically.”

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C939

The diffusion models that generate images were built to generate images, not text, so they are typically poor at rendering text inside an image.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-IMG-08

  • Extended by D3-C688 Day 3: Generative models, including image diffusion models, make things up when they lack information or because of…

Lightning session F: Media 3

StageNot in docsD2-C946

Version 20 of the Screaming Frog SEO Spider, released in May 2024, was the first that could connect the crawler to ChatGPT through custom JavaScript, which let SEOs rewrite image alt text at scale.

Speaker not identifiedEvidence transcript

StageNot in docsD2-C949

To test LLM-generated alt text before a rollout, a community speaker recommended running the crawler's custom JavaScript in Screaming Frog's List mode on a few chosen URLs and reviewing the generated alt text, instead of crawling the whole site in Spider mode.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-IMG-02

Focusing on Internationalisation and Localisation 9

StageNot in docsD2-C592

Server location is not really a reliable country-targeting signal nowadays, so Google does not use it much; the recording is unclear on the word 'server'.

Speaker GoogleEvidence transcript

StageNot in docsD2-C599

Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.

Speaker GoogleEvidence transcript

  • Extended by D2-C965 Day 2: Users of non-Latin-script languages do not always search in their own script: the same Persian query may be…
  • Extended by D3-C051 Day 3: Users expect content written the way they search: in some languages they search in Latin characters, in…
StageNot in docsD2-C608

Whether machine-translated content is acceptable depends on the case and is the site owner's decision, after weighing three things machine translation can miss: translation quality, local conventions and cultural adaptation.

Speaker GoogleEvidence transcript

Used byrequirement DEV-INT-10

StageNot in docsD2-C610

Localisation should account for local conventions such as date formats, which differ between Europe, the UK, the US and other countries, and calendars (in Thailand the current year is 2569), or a date can point users to the wrong day.

Speaker GoogleEvidence transcript

Used byrequirements DEV-INT-10, DEV-SDA-05

StageNot in docsD2-C613

According to the speaker, shoppers in Europe and the US pay a lot of attention to promotions and discounts when deciding to buy (no source for this was captured), one of the cultural factors localisation should take into account.

Speaker GoogleEvidence transcript

StageNot in docsD2-C614

According to consumer data shown on a slide in the talk (source not captured), US consumers judge product quality more by user feedback and reviews, while European shoppers seem to look more at brand reputation.

Speaker GoogleEvidence transcript

  • Extended by D3-C341 Day 3: Some countries rely a lot on social proof such as reviews, Google's speaker said, recalling a point from…
StageNot in docsD2-C615

According to consumer data shown on a slide in the talk (source not captured), German shoppers also rely heavily on expert recommendations and certifications when judging product quality.

Speaker GoogleEvidence transcript

StageNot in docsD2-C616

According to consumer data shown on a slide in the talk (source not captured), French shoppers particularly want to know where a product comes from.

Speaker GoogleEvidence transcript

Lightning session G: Internationalisation 7

StageNot in docsD2-C644

Arabic and Persian contain letters that look identical to users but are different characters to software, with different Unicode code points; the example given was the letter ye, which has an Arabic and a Persian form.

Speaker a second community speakerEvidence transcript

StageNot in docsD2-C967

The presenter of the non-Latin-script talk said that in competitive Persian, Turkish and Arabic searches, bought backlinks and paid editorial content still visibly influence rankings and are widespread (an observation; no data was shown); the presenter stressed this described the situation and was not a recommendation.

Speaker a second community speakerEvidence transcript

Calculating (some) signals 9

StageNot in docsD2-C662

Google calculates SafeSearch signals during indexing rather than in ranking, because ranking happens online with no time for such calculations, while indexing has the processing power.

“In ranking, everything has to happen online, and there's just not enough time to calculate those things.”

Speaker GoogleEvidence transcript

StageNot in docsD2-C669

SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.

“built on top of, well, nowadays, Gemini, and fine-tuned to that specific purpose of finding spam”

Speaker GoogleEvidence transcript

Used byglossary term SpamBrain

  • Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…

Deciding what goes in the index? 14

StageNot in docsD2-C681

Google gives two reasons for not indexing every URL it knows: most of them would not be useful to users, and including URLs in the index that users would never see would be an immense investment.

Speaker GoogleEvidence transcript

StageNot in docsD2-C685

Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

“the index selection system is going to be more forgiving when it sees a new URL from your site”

Speaker GoogleEvidence transcript

Used byrequirement DEV-URL-10

  • Extends D1-C094 Day 1: If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is…
StageNot in docsD2-C688

Index selection is the last step before documents enter Google's index.

Speaker GoogleEvidence transcript

Used byglossary term Index selection

  • Extends D2-C026 Day 2: A Google pipeline slide placed processing between the crawler and the index and listed six processing steps…
  • Extends D1-C211 Day 1: Index selection runs after signals are collected and duplicates are dropped, and decides what goes into…
StageNot in docsD2-C694

News sites were given as the example of importance at work in index selection: they are generally very important on the web and their pages usually get indexed very fast ('indexed' is a likely but not certain reading of the recording).

Speaker GoogleEvidence transcript

StageNot in docsD2-C706

'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

“The first one is kind of nastier.”

Speaker GoogleEvidence transcript

Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed

  • Extends D1-C097 Day 1: Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over…
  • Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
StageNot in docsD2-C714

'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

“most of the time it is actually a quality issue”

Speaker GoogleEvidence transcript

Things

Used byrequirement DEV-MON-03

  • Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.

How does the index look like? 9

StageNot in docsD2-C721

Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.

“we don't have the full content of the page in our index”

Speaker GoogleEvidence transcript

  • Repeats D2-C318 Day 2: Google does not store the complete sentences or the full HTML of a page in the Search index, because large…
StageNot in docsD2-C724

The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.

“the snippet that you see was reconstructed from these tokens”

Speaker GoogleEvidence transcript

Things
  • Extended by D3-C316 Day 3: Google generates the parts of a text result, such as title link and snippet, from its understanding of the…
StageNot in docsD2-C824

Google said posting lists, which Google's serving system uses to find the pages that contain a query's words, are not new: they are at least 60 years old (as of 2026).

Speaker GoogleEvidence transcript

  • Extends D2-C732 Day 2: To find relevant pages, Google's serving system relies on posting lists, a long-established information…
StageNot in docsD2-C736

In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.

Speaker GoogleEvidence transcript

Used byglossary term Posting list

  • Extended by D3-C075 Day 3: At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the…
StageNot in docsD2-C737

A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

Speaker GoogleEvidence transcript

  • Repeats D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageNot in docsD2-C738

At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.

Speaker GoogleEvidence transcript

  • Extended by D3-C075 Day 3: At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the…
StageNot in docsD2-C744

Google said a vector space also holds embeddings for associations the web makes with a page, such as what is known about its author; most of them sit far from typical queries, and a query that names the association may move closer to them.

Speaker GoogleEvidence transcript

Google Trends 16

StageNot in docsD2-C763

The legacy Google Trends Explore page is still available but will be retired: Google is adding its features to the new Explore page until everyone can move to the new one (no date was given).

“the legacy Explore page, which is still available, but not for a long time”

Speaker Omri WeismanEvidence transcript

StageNot in docsD2-C770

Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of Nano Banana (the image editing model in the Gemini app), as people searched for both Gemini and Nano Banana.

Speaker Omri WeismanEvidence transcript

  • Extends D1-C162 Day 1: Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the…
StageNot in docsD2-C777

An early-2012 worldwide spike in Google Trends search interest for 'ai' was caused by the global hit song 'Ai Se Eu Te Pego', whose title contains the Portuguese word 'ai', not by interest in artificial intelligence.

Speaker Omri WeismanEvidence transcript

StageNot in docsD2-C786

Google said its internal numbers show the actual volume of ski-related searches is going up, even though ski's share of all searches in Google Trends has declined since 2004.

“we can tell you from our own internal numbers that we're seeing on Google Trends: the actual volume is, in fact, going up.”

Speaker Omri WeismanEvidence transcript

Day 3: Serving: Ranking, Search Console, and Performance 139

Making sense of users' queries 30

StageNot in docsD3-C011

Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.

Speaker John MuellerEvidence transcript

  • Extends D2-C320 Day 2: Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings…
StageNot in docsD3-C012

After detecting the query language and separating the words, Google removes words it thinks matter little to the query, such as 'a' and 'of', known as stop words.

Speaker John MuellerEvidence transcript

Used byglossary terms Query understanding, Stop words

StageNot in docsD3-C013

Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

Speaker John MuellerEvidence transcript

  • Extends D2-C321 Day 2: For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word…
  • Extends D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
StageNot in docsD3-C014

For some searches the stop words are important, and Google then tries to recognise the whole phrase, stop words included, as an entity; in indexing, such words are indexed together.

Speaker John MuellerEvidence transcript

Used byglossary term Stop words

StageNot in docsD3-C016

When a query names an entity, Google treats it as a request for that entity rather than as a collection of separate words.

Speaker John MuellerEvidence transcript

Used byglossary term Query understanding

StageNot in docsD3-C017

In ranking, Google can match a query's words or its entity, and, the speaker said with a 'probably', mixes both to some degree.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C019

In Google's example, the query word 'photograph' could be expanded to 'image', 'picture' or 'photo', but one German candidate had to be dropped because the similar German word means 'photographer', so expansions are language-specific.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C022

For technical terms, Google's synonym swapping can return either a technical page or a simplified page for the same query.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C026

Google suggested a test: a search that lists a term's synonyms joined with OR will probably return results very similar to the plain query, because Google adds the synonyms itself (part of the sentence is unclear in the recording).

Speaker John MuellerEvidence transcript

StageNot in docsD3-C029

Internally, a short query can turn into a much longer rewritten query, because Google adds entities, synonyms and other information before looking it up.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C030

In Google's rewrite example, [fried chicken place in Barcelona] keeps 'fried' and 'chicken' as two words, may add an entity for fried chicken, and replaces 'place' with alternatives such as 'area', 'location' or 'restaurant'.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C031

Some alternatives in a rewritten query make no sense, such as 'fried chicken area in Barcelona', which is harmless because few indexed pages match them.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C032

In a rewritten query, a place name such as Barcelona can stay a word or be swapped for an entity, possibly a more specific location.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C034

Some synonyms are contextual and depend on the rest of the query: 'GM' probably means General Motors in [GM car], general manager in [GM restaurants] and genetically modified in [GM barley].

Speaker John MuellerEvidence transcript

StageNot in docsD3-C036

Google's synonyms need not be synonyms linguistically: words people use interchangeably are treated as synonyms, because the aim is to find the right content in the index.

“we don't need to be technically accurate”

Speaker John MuellerEvidence transcript

Used byglossary term Synonyms and siblings

StageNot in docsD3-C037

When Google highlights a word in its results that looks like the wrong synonym, the reason is probably that many people use the two words interchangeably.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C038

Besides synonyms, Google detects 'siblings', words in the same category that are related but not interchangeable, such as Canon and Nikon.

Speaker John MuellerEvidence transcript

Used byglossary term Synonyms and siblings

StageNot in docsD3-C039

Because Canon and Nikon are siblings rather than synonyms, a search for [Canon camera] should not show Nikon cameras.

Speaker John MuellerEvidence transcript

Used byglossary term Synonyms and siblings

StageNot in docsD3-C040

Google learns synonyms and siblings from search behaviour: words people search with in the same way become synonyms, while frequent comparison queries mark words as not interchangeable (the end of the sentence is unclear in the recording).

Speaker John MuellerEvidence transcript

StageNot in docsD3-C043

To check whether Google treats two terms as the same, search for each: [Iberico ham] and [jamón ibérico] both brought up the same entity, so Google understands them as the same thing.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C045

At the time of the talk, a search for [Spanish cured ham] brought up mainly a Wikipedia page, showing that Google does not automatically equate that phrase with Iberico ham.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C051

Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-INT-11

  • Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…
StageNot in docsD3-C060

Google said a short video showing how one query is expanded into several fan-out queries has been in its documentation for a while.

Speaker John MuellerEvidence transcript

StageNot in docsD3-C063

Google said fan-out queries are not added to Search Console, because Google considers them part of its infrastructure.

“Fan-out queries are not added in Search Console, because they're basically a part of our infrastructure.”

Speaker John MuellerEvidence transcript

Used byrequirement DEV-AIF-03glossary term Query fan-out

  • Extends D1-C124 Day 1: Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to…
SlideNot in docsD3-C069

Google's summary slide on query understanding said Google's synonyms are not always language-based; on stage this was explained as words people use interchangeably counting as synonyms even when they are not synonyms linguistically.

“Google's synonyms aren't always language-based”

Wording checked against the slide or recording

Speaker John MuellerEvidence slide photo, transcript

SlideNot in docsD3-C070

Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.

Speaker John MuellerEvidence slide photo, transcript

  • Repeats D2-C320 Day 2: Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings…
StageNot in docsD3-C075

At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.

Speaker Gary IllyesEvidence transcript

Used byglossary term Retrieval

  • Extends D2-C736 Day 2: In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked…
  • Extends D2-C738 Day 2: At retrieval, Google looks up the posting lists of the query words that are actually important rather than of…
StageNot in docsD3-C078

Because a query like [best fried chicken ever] can match millions of pages, Google already orders the candidates during retrieval, before ranking starts.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C083

Google called quality the most important of the signals used to order candidates at retrieval: a URL of high quality is more likely to be retrieved from the index for specific queries.

“if the quality of a URL is high, then it's more likely to be retrieved from the index for specific queries.”

Speaker Gary IllyesEvidence transcript

Used byglossary term Retrieval

  • Extends D2-C722 Day 2: Each document in Google's index has pretty much all the signals calculated for it attached, for example…

Lightning session K: Facets of quality 2

StageNot in docsD3-C098

A community speaker recalled that around 2009 writers bent titles and phrases to include a keyword because their CMS counted keyword density, and said that hitting the ratio really worked at the time.

Speaker Community speakersEvidence transcript

How Google thinks about Quality 4

StageNot in docsD3-C130

Google's quality talk described quality as a ranking signal in its own right: a number calculated from many different parts (the word 'quality' is a repaired speech-to-text reading).

Speaker GoogleEvidence transcript

StageNot in docsD3-C136

Google's quality talk called PageRank the speaker's favourite ranking system, elegant in its time, but said Google does not really use it so much anymore.

“My favorite is probably PageRank, even though we don't really use them so much anymore.”

Speaker GoogleEvidence transcript

Things
StageNot in docsD3-C198

Google's quality talk said a big portion of the new pages Google discovers every day is spam, adding that the speaker did not know of Google ever publishing the percentage.

Speaker GoogleEvidence transcript

What are quality updates 14

StageNot in docsD3-C234

The speaker opened the talk on why Search changes by noting that Google's logo has changed only a handful of times in about 30 years.

Speaker GoogleEvidence transcript

SlideNot in docsD3-C236

Google's slide said that in 2023 Google ran more than 800,000 search quality tests; the speaker added that a more recent figure might exist.

Speaker GoogleEvidence slide photo, transcript

SlideNot in docsD3-C243

Google's slide said that as content formats and types become more widespread, search users might start looking for them, and if enough people become interested Google might launch one or more Search features for those formats.

Speaker GoogleEvidence slide photo, transcript

StageNot in docsD3-C246

Compared with the simple early-2000s Google results page, which showed a few expected sites, today's results page for the same query adds exploration features such as People Also Ask.

Speaker GoogleEvidence transcript

StageNot in docsD3-C247

Google added exploration features to its results pages because users' behaviour evolved: people wanted to explore more of the topic they were searching for.

Speaker GoogleEvidence transcript

StageNot in docsD3-C249

The speaker said that without new Search features, search results would become obsolete and users would no longer find them useful.

Speaker GoogleEvidence transcript

StageNot in docsD3-C251

In the mid-1990s the web was small enough to list in one manually edited directory; by the time BackRub was being built, directories could no longer find information on the growing web, which is why Google was created.

Speaker GoogleEvidence transcript

Things
StageNot in docsD3-C257

To show how much the web has grown, Google's deck compared result counts for the query 'durian': for every result in 2000 there were about 20,000 results in 2025.

Speaker GoogleEvidence transcript

StageNot in docsD3-C283

In the speaker's view, some of the egregious spam methods on the slide (cloaking, doorways, scraped content, link spam and hacked content) are not that common anymore or no longer matter much to Google.

“Some of these, I would say, are not that common anymore, or we don't care all that much about them.”

Speaker GoogleEvidence slide photo, transcript

StageNot in docsD3-C284

Reviewing the spam slide, the speaker said link spam would probably be replaced with scaled content abuse as the spam type worth talking about today.

Speaker GoogleEvidence slide photo, transcript

Used byrequirement DEV-SPM-02

  • Extends D3-C207 Day 3: Google's quality talk listed recent updates to Google's spam policies: back button hijacking, scaled content…
  • Extended by D3-C694 Day 3: Scaled content abuse is becoming a problem again: in the early 2000s pages were churned out with Perl or PHP…

How Search results are born 9

StageNot in docsD3-C302

Google described its search results page as an auction system in which the individual results, of different types, all bid for a place on the page; the auction is a metaphor for result types competing for space, not a reference to ads.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C303

Text results are the most common and most prominent result type on Google's results page, whether they are generated by an LLM or not, or enhanced in some other way, Google's speaker said.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C311

Google handles an image used as a search query much like a text query interpreted as an embedding: the image is broken down into vectors (embeddings) that are then searched for in the index.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C740 Day 2: Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are…
  • Repeats D1-C224 Day 1: For visual search, Google breaks an image down into vectors, sends the vectors to the index and returns…
StageNot in docsD3-C312

Google can still detect intent for image queries, even though they are searched as embeddings, the speaker said.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C319

Image results shown among web results come from Google's image index and are roughly the same images that Google Images shows for the same query.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C525 Day 2: An image Google has extracted can appear almost anywhere Google shows results, including Discover, image…
StageNot in docsD3-C338

Google said users rely quite a bit on review stars, calling them a powerful signal of quality and trust for users and a way for sites to use social proof to increase clicks.

“Those little yellow stars are a powerful signal of quality and trust for our users.”

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C341

Some countries rely a lot on social proof such as reviews, Google's speaker said, recalling a point from Google's Day 2 internationalisation talk.

Speaker Gary IllyesEvidence transcript

  • Extends D2-C614 Day 2: According to consumer data shown on a slide in the talk (source not captured), US consumers judge product…

Shopping on Search: Beyond the blue links 14

StageNot in docsD3-C363

Google decided against adding hundreds of thousands of highly specific feed attributes, such as heel height for shoes or lens details for cameras, and instead added a few flexible ones that leave merchants in control.

Speaker Alex JansenEvidence transcript

StageNot in docsD3-C401

Rich, lightly structured information is critical to AI systems: product data with many levels of nesting is not needed, but some structure is very important.

“rich, lightly structured information is critical to AI systems.”

Speaker Alex JansenEvidence transcript

Inside Search Console: What’s New & How to Use It 5

StageNot in docsD3-C418

Google named granularity as a design challenge of query groups: putting many queries into one large group versus splitting them into smaller, more specific groups.

Speaker Ariel KroszynskiEvidence transcript

StageNot in docsD3-C420

Google acknowledged that query grouping is not perfect: some queries are not understood, and some are not grouped in a logical way.

Speaker Ariel KroszynskiEvidence transcript

StageNot in docsD3-C452

In the example YouTube report shown in Google's Search Console talk, long-form videos still got more traffic and more impressions from Google than Shorts; this was one example channel, not a general finding.

Speaker Ariel KroszynskiEvidence transcript

Lightning session L: Understanding SERPs and your users 8

StageNot in docsD3-C493

Nik Vujic said Google Tag Manager events can be used to follow real traffic arriving from different LLMs.

Speaker Nik VujicEvidence transcript

StageNot in docsD3-C505

A community speaker said most people in the room had probably seen clicks decline over the past couple of years because AI now takes part of the demand by summarising information from websites, a shift that hit informational queries first.

Speaker not identifiedEvidence transcript

StageNot in docsD3-C511

A community speaker reported, calling it worrying, that AI features have started to appear more often on results pages for commercial search terms too, in data the speaker tracks for particular brands.

Speaker not identifiedEvidence transcript

StageNot in docsD3-C516

A community speaker cited a recent Similarweb study as showing that being recommended in AI-driven services multiplies a brand's chance of getting traffic downstream; a Similarweb study reported in June 2026, probably the one meant, found brands recommended by ChatGPT 2.5 times more likely to get a site visit within 7 days (US desktop data, finance, travel and beauty).

Speaker not identifiedEvidence transcript

StageNot in docsD3-C526

A community speaker cited a study from about 10 years ago, repeated since, as finding that 82% of people clicked on a brand they already knew regardless of its position; the matching source is Red C's eye-tracking study of shopping-type searches, reported by Econsultancy in October 2018 (about eight years before the event).

Speaker not identifiedEvidence transcript

Mastering the messy middle 21

StageNot in docsD3-C553

Google's research on how AI changes consumer purchase decisions, presented in October 2026, was based on 23,000 conversations with consumers (the talk did not say in which countries).

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C554

Because people do not always do what they say they will do, Google's research on AI and purchase decisions also analysed panels with 40,000 participants, together with research partners, besides the consumer conversations (the talk did not make clear whether 40,000 is one panel or the total).

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C559

Google's research found that consumers delegate the mental effort of a purchase decision to AI without surrendering the choice itself, which Google calls being 'boosted'.

“that doesn't mean they are surrendering their choices. What's happening is that consumers are feeling boosted.”

Speaker Pablo PérezEvidence transcript

Used byglossary term Boosted consumers

StageNot in docsD3-C560

According to Google's research, AI turns consumers who were confused and overwhelmed by information and options into consumers who feel empowered, like category experts, and able to make the best purchase for their needs.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C561

The first benefit consumers get from AI, according to Google's research, is ease: AI synthesises information quickly into an easy-to-digest form and makes choice overload manageable, so consumers can offload cognitive effort.

Speaker Pablo PérezEvidence transcript

Used byglossary term Boosted consumers

StageNot in docsD3-C563

The second benefit, assistance, is that AI reduces a complex category, such as smartphones or trip planning, to the few features that really matter (for a phone, the battery and the picture quality), so consumers feel like experts.

Speaker Pablo PérezEvidence transcript

Used byglossary term Boosted consumers

StageNot in docsD3-C564

Assistance also includes showing consumers what people in similar situations chose, because people are social and want that reassurance (the wording of this passage is partly uncertain in the recording).

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C566

Google's research found that consumers who do not use AI pinball between exploration and evaluation, while AI users start a purchase journey mostly exploring and then move on to evaluating, a path that looks more like a funnel.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C571

Google's research found that people who use AI also visit more brand websites, because it is simply easier.

“the people that are using AI, they search more, and they also visit more websites of brands.”

Speaker Pablo PérezEvidence transcript

  • Extends D1-C005 Day 1: Google says AI Overviews help with new types of questions and lead users to visit a greater diversity of…
StageNot in docsD3-C573

AI users sometimes take even longer purchase journeys, but they perceive their journeys as shorter, a perception the data does not always support (part of this passage is uncertain in the recording).

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C574

For high-risk or high-involvement purchases, ease, assistance and suggestion are often not enough, and consumers add a reassurance step to reduce the risk of the decision.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C575

After exploring with AI, people often go back to Google Search and to the websites of the product or service providers to double-check before buying, for example trust in the company and cancellation policies.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C580

A new or less familiar brand should expect people to double-check it and should prepare assets that give them the extra reassurance they need at that last step.

Speaker Pablo PérezEvidence transcript

SlideNot in docsD3-C581

A Google slide showed nine Google Trends charts of trust-checking searches in several markets and languages (source line: Google Trends, January 2016 to March 2026, web search), all rising toward the end of the range: 'is [brand] legit?' (two panels), '[brand] è affidabile?', '[brand] jest bezpieczne?', 'Is [brand] betrouwbaar?', 'Ist [brand] seriös?', '[brand] est fiable?', '¿es [brand] falso?' and 'trustpilot'.

Speaker Pablo PérezEvidence slide photo

StageNot in docsD3-C584

Consumers who use AI no longer feel frustrated by a lack of progress in a purchase decision; they feel they are moving through the funnel and feel like category experts.

Speaker Pablo PérezEvidence transcript

StageNot in docsD3-C585

Consumers boosted by AI in their purchase decisions were still a small group in October 2026, but the group is growing as more people start using AI (the word 'boosted' is an uncertain reading at this point of the recording).

Speaker Pablo PérezEvidence transcript

How long does it take to..? 26

StageNot in docsD3-C602

Google said it knows hundreds of trillions of URLs (as of October 2026).

Speaker Gary IllyesEvidence transcript

  • Extends D1-C201 Day 1: Google said there are trillions of URLs on the internet, or even more, that even Google cannot tell how many…
SlideNot in docsD3-C604

Google estimated that refreshing (recrawling) a known URL takes about 30 days on average, with a minimum of seconds and an end point of weeks to never.

Speaker Gary IllyesEvidence slide photo, transcript

StageNot in docsD3-C606

For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).

Speaker Gary IllyesEvidence transcript

  • Extends D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
SlideNot in docsD3-C610

Google estimated that processing a sitemap takes about 24 hours on average, with a minimum of minutes.

Speaker Gary IllyesEvidence slide photo, transcript

Things
  • Extends D1-C337 Day 1: An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a…
SlideNot in docsD3-C611

If a sitemap is useful to Google and the site is of high quality, Google tries to refetch the sitemap within 14 days at most.

Speaker Gary IllyesEvidence slide photo, transcript

Things
  • Extends D1-C337 Day 1: An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a…
StageNot in docsD3-C628

Gary Illyes said Google keeps saying it renders every URL on the internet and that he would say this is not true, though it is what he was told; he went on to say that Google's logs show the rendering queue cleared within weeks.

“we keep saying that we render every single URL on the internet. I would say that that's not true”

Speaker Gary IllyesEvidence slide photo, transcript

Things

Used byrequirement DEV-REN-01

  • Extends D2-C256 Day 2: Google renders nearly all of the web by replicating what a browser does, using a real browser's rendering…
SlideNot in docsD3-C635

Google defined indexing end to end as the time from a document entering indexing until its critical processes finish and it reaches the serving index, tokenized and ready to be served as a result.

Speaker Gary IllyesEvidence slide photo, transcript

StageNot in docsD3-C649

Google tries to index images from news sites much faster than usual, which lowers the average image indexing time (the reading 'news sites' is likely but not certain).

Speaker Gary IllyesEvidence transcript

SlideNot in docsD3-C655

Google estimated that updating the image shown with a text result takes 1-2 weeks on average, from a minimum of days up to several weeks to months.

Speaker Gary IllyesEvidence slide photo, transcript

StageNot in docsD3-C656

The image shown with a text result changes only after the image itself is indexed or reindexed and then goes through a further processing step involving embeddings.

Speaker Gary IllyesEvidence transcript

Wrapping all up: AI, Search, and making sense of everything. 6

StageNot in docsD3-C679

Google described AI as an umbrella of many technologies working together, one of which is machine learning: systems that learn from large amounts of data to make informed decisions.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C680

Gary Illyes said machine learning is about 50 years old and that Google has been using it for 'probably 30 years', starting with statistical models that made the 'Did you mean' feature possible.

Speaker Gary IllyesEvidence transcript

  • Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
  • Repeats D1-C209 Day 1: Google launched the 'Did you mean' feature around 2001-2002 using a statistical model, which Gary Illyes…
StageNot in docsD3-C682

Google described large language models as deep learning on internet-scale data sets, trained so that the model has an internal vector space in which concepts are mapped by context.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C683

Google distinguished predictive language models, whose primary job is to predict the next word or words in a text, from generative models that produce output from a prompt; next-word prediction is a guess, made with some accuracy, from the context given.

Speaker Gary IllyesEvidence transcript

StageNot in docsD3-C705

Search has changed every year since its start, for example the Florida update in 2003 (before core updates existed), Universal Search and the Knowledge Graph.

Speaker Gary IllyesEvidence transcript