The control does not affect AI training. Training of the models is limited with Google-Extended instead.
Google Search Console Help · Day 1 · What's new in the world of Search
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Thing · Crawler
A robots.txt control token, not a separate crawler, that decides whether crawled content may be used for Gemini models and grounding outside Search.
Glossary · Google-Extended
A robots.txt control token, not a separate crawler, that decides whether crawled content may train future Gemini models and ground Gemini apps and Vertex AI. It does not affect inclusion or ranking in Google Search. A site can disallow it for a directory or for the whole site.
Documented in
The control does not affect AI training. Training of the models is limited with Google-Extended instead.
Google Search Console Help · Day 1 · What's new in the world of Search
Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.
Google · Day 2 · Lightning session D: Rendering and JavaScript
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.
Ibrahim Anjro · Day 1 · How Google interprets robots.txt
The panel spoke of blocking Google's 'AI crawling or training' with Google-Extended, but Google documents Google-Extended as a usage token, not a crawler: disallowing it does not stop Googlebot fetching pages and only controls use for Gemini training and grounding.
Ibrahim Anjro · Day 1 · Q&A
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
Ibrahim Anjro · Day 2 · Controlling indexing
If Gemini training reuses the rendering done for Search, a site needs no separate rendering work for Gemini; whether its rendered content is used for training is decided with the Google-Extended token in robots.txt, not by rendering choices.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
Kit items about Google-Extended: their own words name it, or several of the claims they rest on do.
Use the Google-Extended robots.txt token to control Gemini training and grounding
Rests on 8 claims, 6 of them naming Google-Extended; its own words name Google-Extended
Rests on 13 claims, 3 of them naming Google-Extended; its own words name Google-Extended
Also inglossary terms Google-Extended, Grounding
Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.
Things named in the same claim, with the number of claims they share.