Day 1: Crawling 8
Shown on screen 1
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
Speaker Cherry Prommawin, Gary IllyesIn Day 1, 11:45 · How Search works and where's AI?Evidence slide photo, transcript
- Extended by D1-C226 Day 1: Google said it does not own third-party AI chatbots such as ChatGPT and has no insight into how they work or…
- Extended by D1-C335 Day 1: Crawling for Gemini may be set to care less about quality and more about the amount of content, because for…
- Extended by D2-C118 Day 2: To train Gemini models, Google renders every page just as it does for Search, so a page that renders…
- Extended by D2-C120 Day 2: When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML…
- Contradicted by D2-C325 Day 2: Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for…
- Extended by D2-C669 Day 2: SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically…
Said on stage 7
Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the number one app in the app store in 2025.
Speaker Lino CattaruzziIn Day 1, 11:00 · Welcome and opening keynotesEvidence transcript
- Extended by D2-C770 Day 2: Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of…
Google described a full-stack approach to AI, from its own TPU chips for AI inference through research and frontier models such as Gemini to its apps.
Speaker Lino CattaruzziIn Day 1, 11:00 · Welcome and opening keynotesEvidence transcript
Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.
Speaker Lino CattaruzziIn Day 1, 11:00 · Welcome and opening keynotesEvidence transcript
- Extended by D2-C727 Day 2: Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's…
- Extended by D3-C059 Day 3: Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding…
Google said that last year (2025) it delivered a decade of innovation in 12 months, with its new Gemini model at the top of all the benchmarks.
Speaker Lino CattaruzziIn Day 1, 11:00 · Welcome and opening keynotesEvidence transcript
Google said it does not own third-party AI chatbots such as ChatGPT and has no insight into how they work or are built.
Speaker Gary IllyesIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
A community speaker named two signals that in their view strongly influence whether an AI selects a brand: the bias already in the model's memory before it searches, and the brand's presence in the search results the model is grounded on.
Speaker not identifiedIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.
“the number of tokens is actually more important”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
Day 2: Indexing 18
Shown on screen 2
In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byglossary term Tokenization
Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence 2 slide photos
Said on stage 11
To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.
“if it works for Search, it works for Gemini for training”
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
Used byrequirement DEV-IDX-11
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
Used byrequirement DEV-PRF-02
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
Used byrequirements DEV-PRF-02, DEV-REN-01
- Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
Google urged sites to make sure their JavaScript content can be crawled, rendered and indexed, calling this important today and also tomorrow, as AI systems increasingly ground answers to user requests.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
- Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
- Contradicts D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript
Used byglossary term Tokenization
Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
Used bymyth M-002
- Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
- Extended by D2-C869 Day 2: Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps…
Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
Used byrequirement DEV-AIF-03myth M-002
- Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
- Extended by D2-C825 Day 2: Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window…
Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.
“the context window is perhaps 900,000 or even closer to a million big”
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
Used byrequirement DEV-AIF-03
- Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
- Extends D2-C331 Day 2: Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Speaker Ryan LeveringIn Day 2, 13:35 · What is Structured Data and why we need it on the internet.Evidence transcript
Used byrequirement DEV-HTM-04
- Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.
“built on top of, well, nowadays, Gemini, and fine-tuned to that specific purpose of finding spam”
Speaker GoogleIn Day 2, 15:30 · Calculating (some) signalsEvidence transcript
Used byglossary term SpamBrain
- Extends D1-C041 Day 1: Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did…
- Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
What Google's documentation says 1
Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.
Publisher GoogleAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirement DEV-IDX-11glossary term Grounding
- Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
Analysis by the author 4
If Gemini training reuses the rendering done for Search, a site needs no separate rendering work for Gemini; whether its rendered content is used for training is decided with the Google-Extended token in robots.txt, not by rendering choices.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirement DEV-IDX-11
Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.
Author Ibrahim AnjroAnnotates Day 2, 11:30 · Understanding what's on a page
The talk gave Gemini's context window both as millions of tokens and as roughly 900,000 to a million; Google's long-context docs say Gemini models have context windows of 1 million or more tokens (about eight average novels per million), so plan with about one million tokens as the documented floor rather than several million.
Author Ibrahim AnjroAnnotates Day 2, 11:30 · Understanding what's on a page