Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Thing · Crawler

Google-Extended

A robots.txt control token, not a separate crawler, that decides whether crawled content may be used for Gemini models and grounding outside Search.

Crawler

Open in Reef mapOpen in Graph

Narrative see the topic Google-Extended

Claims
12
In Google’s docs
3
Said at the event
4
Not in docs
1
Kit items
4

Glossary · Google-Extended

A robots.txt control token, not a separate crawler, that decides whether crawled content may train future Gemini models and ground Gemini apps and Vertex AI. It does not affect inclusion or ranking in Google Search. A site can disallow it for a directory or for the whole site.

Google’s documentation 3

Documented in

DocsSourceD1-C086

Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

Google · Day 1 · How Google interprets robots.txt

Said at the event 4

Slide and stage claims that name it, the ones Google’s documentation does not cover first.

Not in docs 1

StageNot in docsD1-C487

A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

Day 1 · Q&A

Consistent with docs 1

StageConsistent with docsD1-C522

Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

Day 1 · How Google interprets robots.txt

Confirmed by docs 1

Nothing to verify 1

SlideD1-C079

The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

Day 1 · How Google interprets robots.txt

Press and analysis 5

AnalysisD1-C523

The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

Ibrahim Anjro · Day 1 · How Google interprets robots.txt

AnalysisD1-C486

The panel spoke of blocking Google's 'AI crawling or training' with Google-Extended, but Google documents Google-Extended as a usage token, not a crawler: disallowing it does not stop Googlebot fetching pages and only controls use for Gemini training and grounding.

Ibrahim Anjro · Day 1 · Q&A

Built on these claims 4

Kit items about Google-Extended: their own words name it, or several of the claims they rest on do.

Developer requirements 2

Also inglossary terms Google-Extended, Grounding

Connected things 10

Relations

  • Affects Gemini
    3 claims, 3 documented
    • DocsSourceD1-C086

      Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

    • StageConsistent with docsD1-C522

      Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

    • DocsSourceD2-C121

      Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

Most often named with it

Things named in the same claim, with the number of claims they share.

Topics that feature it