Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
Thing · Product
Gemini
Google's family of AI models and the assistant app built on them; Gemini models also power AI features in Search and Google Trends.
- Claims
- 41
- In Google’s docs
- 7
- Said at the event
- 25
- Not in docs
- 10
- Kit items
- 8
Google’s documentation 7
Documented in
Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.
Google · Day 2 · Lightning session D: Rendering and JavaScript
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
Google · Day 2 · Lightning session D: Rendering and JavaScript
Google announced Nano Banana, a new image editing model from Google DeepMind in the Gemini app, on 26 August 2025.
Google blog (26 August 2025) · Day 2 · Google Trends
Google's Gemini API documentation on grounding with Google Search says the model automatically generates and runs one or more search queries when needed, and the response lists the search queries it executed.
Google AI for Developers (Gemini API docs) · Day 3 · Making sense of users' queries
On 11 January 2026 Google announced 'dozens of new data attributes' in Merchant Center for conversational commerce on AI Mode, Gemini and Business Agent, to roll out first with a small group of retailers.
Google blog (11 January 2026) · Day 3 · Shopping on Search: Beyond the blue links
Google's July 2025 blog post introduced Web Guide as a Search Labs experiment that uses a custom version of Gemini and a query fan-out technique to group web links by aspects of the query.
Google blog (24 July 2025) · Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Said at the event 25
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Not in docs 10
Google said that after the launch of Nano Banana, its image generation model, the Gemini app became the number one app in the app store in 2025.
Lino Cattaruzzi · Day 1 · Welcome and opening keynotes
Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.
Gary Illyes · Day 1 · How crawling works
To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.
Erin Sparling · Day 2 · Lightning session D: Rendering and JavaScript
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
Erin Sparling · Day 2 · Lightning session D: Rendering and JavaScript
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Gary Illyes · Day 2 · Understanding what's on a page
Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.
Gary Illyes · Day 2 · Understanding what's on a page
SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.
A Google Trends Explore chart presented on stage put worldwide search interest in Gemini above ChatGPT for the first time ever in September 2025, during a sharp surge for Gemini.
Omri Weisman · Day 2 · Google Trends
Google attributed the September 2025 surge in worldwide search interest for Gemini to the viral launch of Nano Banana (the image editing model in the Gemini app), as people searched for both Gemini and Nano Banana.
Omri Weisman · Day 2 · Google Trends
After the Nano Banana surge faded, worldwide baseline search interest in Gemini stayed much higher than before, which Google read as a sustained gain in brand recognition.
Omri Weisman · Day 2 · Google Trends
Consistent with docs 12
Google said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.
Lino Cattaruzzi · Day 1 · Welcome and opening keynotes
Google said that last year (2025) it delivered a decade of innovation in 12 months, with its new Gemini model at the top of all the benchmarks.
Lino Cattaruzzi · Day 1 · Welcome and opening keynotes
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
AI Overviews and AI Mode typically do not ground their answers by reading pages live, unlike Gemini when a user asks about a specific page, Google said.
Erin Sparling · Day 2 · Lightning session D: Rendering and JavaScript
Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
Gary Illyes · Day 2 · Understanding what's on a page
Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.
Gary Illyes · Day 2 · Understanding what's on a page
Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.
Gary Illyes · Day 2 · Understanding what's on a page
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Day 2 · What is Structured Data and why we need it on the internet.
Google Trends' Suggest search terms combines Gemini's knowledge of the world with Google Trends data.
Omri Weisman · Day 2 · Google Trends
Individual fan-out queries can be seen in some other places, such as Gemini, some search APIs and other LLM systems.
John Mueller · Day 3 · Making sense of users' queries
Query groups can change when the meaning of a query changes over time: a 2023 search for 'Gemini' may have meant the astrological sign, but after Google launched Gemini it means something else, so Search Console may now group it differently.
Ariel Kroszynski · Day 3 · Inside Search Console: What’s New & How to Use It
Confirmed by docs 1
Google described a full-stack approach to AI, from its own TPU chips for AI inference through research and frontier models such as Gemini to its apps.
Lino Cattaruzzi · Day 1 · Welcome and opening keynotes
Nothing to verify 2
A community speaker named the limits of the title-based alt-text script: it rewrites only the hero image, assumes the page title states the user's intent, had been tested in only one language or market, and runs on the free Gemini model.
Google's quality talk showed a surface-level example article (not photographed) that the speaker said Gemini could have generated, while stressing that humans have written such content for a long time too.
Press and analysis 9
The panel spoke of blocking Google's 'AI crawling or training' with Google-Extended, but Google documents Google-Extended as a usage token, not a crawler: disallowing it does not stop Googlebot fetching pages and only controls use for Gemini training and grounding.
Ibrahim Anjro · Day 1 · Q&A
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
Ibrahim Anjro · Day 2 · Controlling indexing
If Gemini training reuses the rendering done for Search, a site needs no separate rendering work for Gemini; whether its rendered content is used for training is decided with the Google-Extended token in robots.txt, not by rendering choices.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.
Ibrahim Anjro · Day 2 · Understanding what's on a page
Do not rewrite pages into short, self-contained chunks for AI systems; Google says Gemini reads context windows of millions of tokens, so structure content for readers, with clear headings and complete explanations.
Ibrahim Anjro · Day 2 · Understanding what's on a page
The talk gave Gemini's context window both as millions of tokens and as roughly 900,000 to a million; Google's long-context docs say Gemini models have context windows of 1 million or more tokens (about eight average novels per million), so plan with about one million tokens as the documented floor rather than several million.
Ibrahim Anjro · Day 2 · Understanding what's on a page
On stage the Nano Banana launch was placed in September 2025, but Google announced it on 26 August 2025, so the September 2025 surge in Gemini search interest described in the talk came in the weeks after the launch.
Ibrahim Anjro · Day 2 · Google Trends
Do not build pages for fan-out queries copied from Gemini or other LLM tools: they differ from system to system and Search Console does not report them. Cover a topic's real subquestions on solid pages instead.
Ibrahim Anjro · Day 3 · Making sense of users' queries
Built on these claims 8
Kit items about Gemini: their own words name it, or several of the claims they rest on do.
Developer requirements 3
Use the Google-Extended robots.txt token to control Gemini training and grounding
Rests on 8 claims, 6 of them naming Gemini; its own words name Gemini
Structure content for readers instead of splitting it into small chunks for AI
Rests on 9 claims, 5 of them naming Gemini; its own words name Gemini
Make JavaScript-generated content render quickly
Rests on 5 claims, 2 of them naming Gemini; its own words name Gemini
Myths 1
- DocumentedM-002
Myth: Cut your content into short, self-contained chunks so Google's AI can use it.
Rests on 4 claims, 2 of them naming Gemini; its own words name Gemini
Also inglossary terms Google-Extended, Grounding, SpamBrain, Web Guide
Connected things 21
Relations
- Affected by Google-Extended
3 claims, 3 documented
Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.
Most often named with it
Things named in the same claim, with the number of claims they share.
- Google-Extended 7
- robots.txt 6
- Grounding 5
- Query fan-out 4
- Search interest 4
- AI Mode 3
- Tokenization 3
- AI Overviews 2
- Search Console 2
- User-triggered fetchers 2
- Chrome 1
- Duplicate content 1
- Google Trends 1
- Googlebot 1
- Merchant Center 1
- nofollow 1
Topics that feature it
- Gemini and Search 19
- Google-Extended 8
- AI crawlers and agents on your site 7
- Chunking content for AI 6
- Query fan-out 5
- Reading Google Trends data 5
- AI systems already inside Search 3
- Content for people 2
- SEO vs GEO 2
- The generative AI opt-out in Search Console 2
- Tokenization: how text is stored 2
- AI-written vs human-written content 1
- Directives and crawl budget 1
- Google Trends tools and features 1
- How Google renders pages 1
- How Search is evolving 1
- Images 1
- Shopping data: feeds, markup and Storebot 1
- Spam detection and SpamBrain 1
- The pipeline: crawling, indexing, serving 1
- Using Google Trends for SEO and content 1
- What's new in Search Console 1
- robots.txt rules 1