Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 1 · Wednesday 30 September 2026 · 15:10

How Google interprets robots.txt

Speaker Google

TalkCoverageSlidesTranscript

Audio recording from the first minutes of the talk to its end, including a robots.txt quiz; the speaker is not named in the recording. The unknown-line slide stays with this talk: the recording reads out its example and value, although the photo's time stamp falls a few minutes after the hand-over to Lightning session B.

Shown on screen 3

SlideD1-C079

The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

Speaker GoogleEvidence slide photo, transcript

  • Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
  • Extended by D1-C521 Day 1: To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the…
  • Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
SlideConfirmed by docsD1-C087

Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.

Speaker GoogleEvidence slide photo, transcript

Things

Used byrequirement DEV-SRV-06

  • Extended by D1-C526 Day 1: The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked…
SlideNot in docsD1-C123

Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

“Unpredictability is never a good time.”

Wording checked against the slide or recording

Speaker GoogleEvidence slide photo, transcript

Used byrequirement DEV-SRV-06

  • Extended by D1-C526 Day 1: The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked…

Said on stage 22

StageNot in docsD1-C089

The robots.txt talk pointed to the robots.txt file of thebestfriedchickenever.com to show why comments in robots.txt are useful; the live file is mostly a large ASCII-art drawing written as # comment lines, followed by a single user-agent group (checked 2026-10-04).

Speaker GoogleEvidence notes, transcript

StageConfirmed by docsD1-C512

Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)

  • Extends D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
  • Repeated by D2-C050 Day 2: John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes…
StageConsistent with docsD1-C513

The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.

Speaker GoogleEvidence transcript

Used byglossary term RFC 9309 (Robots Exclusion Protocol)

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
StageConsistent with docsD1-C514

Robots.txt is not a security measure: the file sits at a predictable public address, so anyone can read the paths it disallows. Protect a secret folder with authentication, or do not put it on the internet.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-11glossary term robots.txt

StageConsistent with docsD1-C516

Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
StageConsistent with docsD1-C519

A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConsistent with docsD1-C520

In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.

Speaker GoogleEvidence transcript, slide photo

Used byrequirement DEV-SRV-05

  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
StageConsistent with docsD1-C521

To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.

Speaker GoogleEvidence transcript, slide photo

Used byrequirement DEV-SRV-05

  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extends D1-C082 Day 1: In robots.txt, * matches zero or more of any character and $ marks the end of the URL.
StageConsistent with docsD1-C522

Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

Speaker GoogleEvidence transcript

Used byrequirement DEV-IDX-11glossary term Google-Extended

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
  • Repeats D1-C485 Day 1: To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
StageNot in docsD1-C525

Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-05

  • Extended by D1-C532 Day 1: A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts…
StageNot in docsD1-C526

The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-06

  • Extends D1-C087 Day 1: Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and…
  • Extends D1-C123 Day 1: Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's…
StageConsistent with docsD1-C527

The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConfirmed by docsD1-C528

Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

Speaker GoogleEvidence transcript

Used byrequirements DEV-SRV-05, DEV-SRV-06glossary term robots.txt report

  • Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
  • Extended by D3-C667 Day 3: Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours…
StageConsistent with docsD1-C531

In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

Speaker GoogleEvidence transcript

Used byrequirement DEV-SRV-10

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…

What Google's documentation says 7

DocsSourceD1-C080

Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

Publisher Google

Used byrequirements DEV-SRV-05, DEV-SRV-10glossary term User-agent group

  • Extended by D1-C519 Day 1: A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several…
  • Extended by D1-C527 Day 1: The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another…
  • Extended by D1-C531 Day 1: In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of…
  • Extended by D1-C533 Day 1: Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that…
DocsSourceD1-C081

When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

“In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”

Publisher Google

Used byrequirement DEV-SRV-05

  • Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
  • Extended by D2-C208 Day 2: A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/…
DocsSourceD1-C084

Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

Publisher Google

Used byrequirement DEV-SRV-05glossary term robots.txt

  • Extended by D1-C516 Day 1: Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was…
  • Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
  • Extended by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
DocsSourceD1-C085

Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

Publisher Google

Used byrequirement DEV-SRV-05

  • Extended by D3-C613 Day 3: Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an…
  • Extended by D3-C614 Day 3: Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for…
DocsSourceD1-C086

Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

“Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.”

Publisher Google

Used byrequirement DEV-IDX-11glossary term Google-Extended

  • Extended by D1-C513 Day 1: The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in…
  • Extended by D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
  • Extended by D1-C523 Day 1: The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023…
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
  • Extended by D2-C121 Day 2: Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex…
DocsSourceD1-C140

Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

Publisher Google Search Console Help

Used byrequirement DEV-SRV-05

  • Extended by D1-C528 Day 1: Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version…
  • Extended by D1-C530 Day 1: Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
  • Extended by D3-C615 Day 3: A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

Analysis by the author 4

AnalysisD1-C083

Under the example file, Googlebot may crawl /, /politics/eu-vote and /sports/live/, and is blocked from /?utm_source=x, /index.html, /politics (no trailing slash), /live/ and /sports/live-score. An 'allow: /$' rule does not cover the homepage with tracking parameters.

Author Ibrahim Anjro

AnalysisD1-C088

Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files. Never place any non-standard directive between user-agent lines; put it after a group's rules and test the file with each search engine's tools, such as Bing's robots.txt tester and Google's open-source robots.txt parser.

Author Ibrahim Anjro

Used byrequirement DEV-SRV-06

AnalysisD1-C523

The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

Author Ibrahim Anjro

Used byrequirement DEV-AIF-05

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
AnalysisD1-C532

A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

Author Ibrahim Anjro

Used byrequirement DEV-SRV-05

  • Extends D1-C525 Day 1: Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule…
  1. Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

    extends
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

  2. Stage D1-C533 Day 1 · Lightning session B: Robots.txt

    Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  3. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Slide D1-C079 Day 1 · How Google interprets robots.txt

    The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

  4. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  5. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  6. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

  7. Docs D2-C121 Day 2 · Lightning session D: Rendering and JavaScript

    Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  8. Analysis D2-C208 Day 2 · Lightning session D: Rendering and JavaScript

    A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.

    extends
    Docs D1-C081 Day 1 · How Google interprets robots.txt

    When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

  9. Stage D2-C842 Day 2 · Welcome to indexing day!

    Google said listing the sitemap in robots.txt is fine, as many websites do.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  10. Slide D3-C613 Day 3 · How long does it take to..?

    Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.

    extends
    Docs D1-C085 Day 1 · How Google interprets robots.txt

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

  11. Stage D3-C614 Day 3 · How long does it take to..?

    Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

    extends
    Docs D1-C085 Day 1 · How Google interprets robots.txt

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

  12. Stage D3-C615 Day 3 · How long does it take to..?

    A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

    extends
    Docs D1-C140 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

  13. Docs D3-C667 Day 3 · How long does it take to..?

    Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.

    extends
    Stage D1-C528 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

  14. Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

    repeats
    Stage D1-C485 Day 1 · Q&A

    To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.

  15. Stage D1-C524 Day 1 · How Google interprets robots.txt

    Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.

    repeats
    Stage D1-C487 Day 1 · Q&A

    A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

  16. Stage D2-C050 Day 2 · Controlling indexing

    John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

    repeats
    Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.