Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · AI and Search

Gemini and Search

Gemini is not part of Search, but it uses Google's crawlers for data and grounds answers on the Search index; Google's crawler documentation defines that grounding as Search-index content given to the model at prompt time, controlled with Google-Extended. Day 2 contradicted part of Day 1's slide, which said Gemini shares tokenization with Search: Gary Illyes said Gemini's tokenization differs ('or mostly') and showed the AI-model tokenizer splitting 'robots.txt' and 'tl;dr' into pieces that the Search tokenizer keeps whole. Google also said Gemini training renders pages the same way Search does, that a Gemini user asking about a specific page triggers a live read that takes extra time, that Gemini in Chrome relies heavily on page screenshots and that SpamBrain is now built on Gemini (all said at the event; Google's docs say only, in general terms, that browser agents may analyse screenshots). Gary Illyes said Gemini's context window holds millions of tokens, and Erin Sparling urged making JavaScript content crawlable, renderable and fast, or server-side rendered, because AI systems increasingly ground answers on it. Author’s view: Google documents grounding only from the Search index, so whether a live page read follows Google-Extended or ignores robots.txt like a user-triggered fetcher is unclear. Day 3's closing talk explained the models behind it: large language models are deep learning on internet-scale data that map concepts by context in an internal vector space, and adding grounding or retrieval-augmented generation reduces hallucinations but cannot eliminate them (said at the event, not in Google's docs). A second recording of Day 1 added the keynote's framing: Google takes a full-stack approach from its own TPU chips to Gemini and its apps, and Gemini lets Search understand intent before fan-out adds further queries; an audio recording of the keynote added Google's claim that in 2025 it delivered a decade of innovation in 12 months, with its new Gemini model at the top of all the benchmarks. Gary Illyes said Google does not own third-party chatbots such as ChatGPT and has no insight into them, and that crawling for Gemini may care less about quality and more about the amount of content, because for LLMs the number of tokens matters more (said at the event). On Day 2 he put the context window at millions of tokens and then at perhaps 900,000 to a million. Author’s view: Google's long-context docs say Gemini models have context windows of one million or more tokens, so plan with about one million as the documented floor.

Based on D1-C039, D2-C121, D2-C325, D2-C328, D2-C327, D2-C118, D2-C120, D2-C465, D2-C669, D2-C331, D2-C332, D2-C124, D2-C134, D2-C123, D2-C122, D3-C682, D3-C687, D1-C163, D1-C172, D1-C226, D1-C335, D2-C869, D2-C872, D1-C496

28 claims · raised in 9 sessions · said or shown on Day 1 and Day 2 and Day 3

Open in Reef mapOpen in Graph

What to do

  • Decide Gemini training and grounding use with the Google-Extended token in robots.txt; rendering choices do not change it.
  • Server-render or speed up the main content of pages people ask assistants about, such as product, pricing, documentation and policy pages.
  • Do not cut content into small chunks for Gemini; its context window holds about a million tokens or more.

Day 1: Crawling 8

Shown on screen 1

SlideConsistent with docsD1-C039

Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

Speaker Cherry Prommawin, Gary IllyesIn Day 1, 11:45 · How Search works and where's AI?Evidence slide photo, transcript

  • Extended by D1-C226 Day 1: Google said it does not own third-party AI chatbots such as ChatGPT and has no insight into how they work or…
  • Extended by D1-C335 Day 1: Crawling for Gemini may be set to care less about quality and more about the amount of content, because for…
  • Extended by D2-C118 Day 2: To train Gemini models, Google renders every page just as it does for Search, so a page that renders…
  • Extended by D2-C120 Day 2: When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML…
  • Contradicted by D2-C325 Day 2: Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for…
  • Extended by D2-C669 Day 2: SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically…

Said on stage 7

StageNot in docsD1-C162

Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the number one app in the app store in 2025.

Speaker Lino CattaruzziIn Day 1, 11:00 · Welcome and opening keynotesEvidence transcript

Things
  • Extended by D2-C770 Day 2: Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of…
StageConsistent with docsD1-C172

Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.

Speaker Lino CattaruzziIn Day 1, 11:00 · Welcome and opening keynotesEvidence transcript

  • Extended by D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
  • Extended by D3-C059 Day 3: Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding…
StageConsistent with docsD1-C496

Google said that last year (2025) it delivered a decade of innovation in 12 months, with its new Gemini model at the top of all the benchmarks.

Speaker Lino CattaruzziIn Day 1, 11:00 · Welcome and opening keynotesEvidence transcript

Things
StageNot in docsD1-C335

Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

“the number of tokens is actually more important”

Wording checked against the slide or recording

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…

Day 2: Indexing 18

Shown on screen 2

SlideConsistent with docsD2-C326

In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byglossary term Tokenization

Said on stage 11

StageNot in docsD2-C118

To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.

“if it works for Search, it works for Gemini for training”

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Things

Used byrequirement DEV-IDX-11

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageNot in docsD2-C120

When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Things

Used byrequirement DEV-PRF-02

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageConsistent with docsD2-C124

Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Things

Used byrequirements DEV-PRF-02, DEV-REN-01

  • Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
StageConsistent with docsD2-C134

Google urged sites to make sure their JavaScript content can be crawled, rendered and indexed, calling this important today and also tomorrow, as AI systems increasingly ground answers to user requests.

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

  • Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
StageConsistent with docsD2-C327

Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byglossary term Tokenization

StageConsistent with docsD2-C331

Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Things

Used bymyth M-002

  • Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
  • Extended by D2-C869 Day 2: Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps…
StageConsistent with docsD2-C332

Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Things

Used byrequirement DEV-AIF-03myth M-002

  • Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
  • Extended by D2-C825 Day 2: Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window…
StageConsistent with docsD2-C869

Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.

“the context window is perhaps 900,000 or even closer to a million big”

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Things

Used byrequirement DEV-AIF-03

  • Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
  • Extends D2-C331 Day 2: Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
StageConsistent with docsD2-C465

Gemini in Chrome relies heavily on the screenshot it takes of a page.

Speaker Ryan LeveringIn Day 2, 13:35 · What is Structured Data and why we need it on the internet.Evidence transcript

Used byrequirement DEV-HTM-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
StageNot in docsD2-C669

SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.

“built on top of, well, nowadays, Gemini, and fine-tuned to that specific purpose of finding spam”

Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript

Used byglossary term SpamBrain

  • Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…

What Google's documentation says 1

DocsSourceD2-C121

Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

Publisher GoogleAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-IDX-11glossary term Grounding

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…

Analysis by the author 4

AnalysisD2-C123

Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.

Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Day 3: Serving: Ranking, Search Console, and Performance 2

Said on stage 2

Across days and sessions 17

  1. Stage D2-C325 Day 2 · Understanding what's on a page

    Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

    contradicts
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  2. Stage D1-C226 Day 1 · How Search works and where's AI?

    Google said it does not own third-party AI chatbots such as ChatGPT and has no insight into how they work or are built.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  3. Stage D1-C335 Day 1 · How crawling works

    Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  4. Stage D2-C118 Day 2 · Lightning session D: Rendering and JavaScript

    To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  5. Stage D2-C120 Day 2 · Lightning session D: Rendering and JavaScript

    When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  6. Docs D2-C121 Day 2 · Lightning session D: Rendering and JavaScript

    Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  7. Stage D2-C124 Day 2 · Lightning session D: Rendering and JavaScript

    Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.

    extends
    Slide D1-C056 Day 1 · How Search works and where's AI?

    Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.

  8. Stage D2-C134 Day 2 · Lightning session D: Rendering and JavaScript

    Google urged sites to make sure their JavaScript content can be crawled, rendered and indexed, calling this important today and also tomorrow, as AI systems increasingly ground answers to user requests.

    extends
    Slide D1-C056 Day 1 · How Search works and where's AI?

    Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.

  9. Stage D2-C332 Day 2 · Understanding what's on a page

    Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.

    extends
    Slide D1-C054 Day 1 · How Search works and where's AI?

    Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.

  10. Stage D2-C465 Day 2 · What is Structured Data and why we need it on the internet.

    Gemini in Chrome relies heavily on the screenshot it takes of a page.

    extends
    Docs D1-C131 Day 1 · session not recorded

    Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

  11. Stage D2-C669 Day 2 · Calculating (some) signals

    SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  12. Stage D2-C669 Day 2 · Calculating (some) signals

    SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.

    extends
    Slide D1-C041 Day 1 · How Search works and where's AI?

    Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.

  13. Stage D2-C727 Day 2 · How does the index look like?

    Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).

    extends
    Stage D1-C172 Day 1 · Welcome and opening keynotes

    Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.

  14. Stage D2-C770 Day 2 · Google Trends

    Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of Nano Banana (the image editing model in the Gemini app), as people searched for both Gemini and Nano Banana.

    extends
    Stage D1-C162 Day 1 · Welcome and opening keynotes

    Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the number one app in the app store in 2025.

  15. Stage D2-C825 Day 2 · Understanding what's on a page

    Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.

    extends
    Stage D2-C332 Day 2 · Understanding what's on a page

    Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.

  16. Stage D2-C869 Day 2 · Understanding what's on a page

    Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.

    extends
    Stage D2-C331 Day 2 · Understanding what's on a page

    Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.

  17. Stage D3-C059 Day 3 · Making sense of users' queries

    Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding how normal queries work explains fan-out queries too.

    extends
    Stage D1-C172 Day 1 · Welcome and opening keynotes

    Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.

Built on these claims 6

Developer requirements 5

Myths 1

Sources 8