Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Publisher controls

Google-Extended

Google-Extended is a robots.txt control token, not a crawler with its own user agent: it governs whether crawled content is used to train future Gemini models and to ground Gemini apps and Vertex AI, and it has no effect on inclusion or ranking in Google Search. Day 2 showed how it travels: in Google's example fetch record the applicable robots policy, shown as opt-out-google-extended, stays attached to the fetched page as it goes into processing, and Google's crawler documentation defines the grounding it controls as Search-index content given to the model at prompt time. Google also said Gemini training renders pages the same way Search does, provided the site allows training (said at the event, not in Google's docs). Author’s view: the live read of a page that a Gemini user asks about is not documented, so it is unclear whether Google-Extended applies to it or whether it behaves like a user-triggered fetcher, which generally ignores robots.txt. In the Day 1 Q&A, asked whether Google plans dedicated AI user agents, Google said a site can opt out of Google's AI training with the Google-Extended token. Author’s view: the panel spoke of blocking AI 'crawling or training', but Google documents Google-Extended as a usage token: disallowing it does not stop Googlebot fetching pages. An audio recording of Day 1's robots.txt talk added that a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token, which Google said it respects, and that robots.txt itself was designed only to control access, not how content is used; Google named Apple's Applebot-Extended as a similar token from another company. Author’s view: the speaker's belief that Google was the first to offer such an opt-out needs care, since OpenAI documented blocking its GPTBot crawler in August 2023, before Google announced Google-Extended on 28 September 2023; what Google-Extended added was a training control for pages fetched by Google's existing crawlers, separate from Search.

Based on D1-C086, D1-C032, D2-C025, D2-C121, D2-C118, D2-C123, D2-C122, D1-C485, D1-C486, D1-C522, D1-C513, D1-C524, D1-C523

13 claims · raised in 4 sessions · said or shown on Day 1 and Day 2

Open in Reef mapOpen in Graph

What to do

  • Set Google-Extended deliberately, knowing it governs Gemini training and grounding, not Search.
  • Use the Search Console generative AI setting, not Google-Extended, to leave AI Overviews and AI Mode.
  • Do not expect rendering choices to keep content out of Gemini training; only the token does.

Day 1: Crawling 7

Said on stage 3

StageConsistent with docsD1-C522

Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-IDX-11glossary term Google-Extended

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
  • Repeats D1-C485 Day 1: To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
StageD1-C481

An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that site owners can control AI access separately from Googlebot and analyse it in their logs.

From the audienceIn Day 1, 16:35 · Q&AEvidence transcript

Things
  • Answered by D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
  • Answered by D1-C483 Day 1: A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense…
  • Answered by D1-C484 Day 1: A Google panelist called blocking all AI crawlers while allowing search crawlers such as Googlebot, Bingbot…
  • Answered by D1-C485 Day 1: To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
  • Answered by D1-C487 Day 1: A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site…
  • Answered by D1-C488 Day 1: When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt…
  • Answered by D1-C489 Day 1: A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a…
StageConfirmed by docsD1-C485

To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirements DEV-AIF-05, DEV-IDX-11

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Repeated by D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…

What Google's documentation says 2

DocsSourceD1-C086

Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

“Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.”

Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-IDX-11glossary term Google-Extended

  • Extended by D1-C513 Day 1: The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in…
  • Extended by D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
  • Extended by D1-C523 Day 1: The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023…
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
  • Extended by D2-C121 Day 2: Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex…

Analysis by the author 2

AnalysisD1-C523

The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

Author Ibrahim AnjroAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-AIF-05

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…

Day 2: Indexing 6

Shown on screen 1

SlideConsistent with docsD2-C025

In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence slide photo

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Extends D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…

Said on stage 1

StageNot in docsD2-C118

To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.

“if it works for Search, it works for Gemini for training”

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Things

Used byrequirement DEV-IDX-11

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…

What Google's documentation says 1

DocsSourceD2-C121

Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

Publisher GoogleAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-IDX-11glossary term Grounding

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…

Analysis by the author 3

AnalysisD2-C067

nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing

  • Extends D1-C127 Day 1: Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed…
  • Extends D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
AnalysisD2-C123

Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.

Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Across days and sessions 11

  1. Stage D1-C513 Day 1 · How Google interprets robots.txt

    The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  2. Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  3. Analysis D1-C523 Day 1 · How Google interprets robots.txt

    The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  4. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  5. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  6. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

  7. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Docs D1-C127 Day 1 · How Google thinks about crawl budget

    Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

  8. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Stage D1-C482 Day 1 · Q&A

    Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

  9. Stage D2-C118 Day 2 · Lightning session D: Rendering and JavaScript

    To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  10. Docs D2-C121 Day 2 · Lightning session D: Rendering and JavaScript

    Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  11. Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

    repeats
    Stage D1-C485 Day 1 · Q&A

    To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.

Built on these claims 2

Developer requirements 2

Sources 4