Audio recording from the first minutes of the talk to its end, including a robots.txt quiz; the speaker is not named in the recording. The unknown-line slide stays with this talk: the recording reads out its example and value, although the photo's time stamp falls a few minutes after the hand-over to Lightning session B.
Said on stage 22
The robots.txt talk pointed to the robots.txt file of thebestfriedchickenever.com to show why comments in robots.txt are useful; the live file is mostly a large ASCII-art drawing written as # comment lines, followed by a single user-agent group (checked 2026-10-04).
Speaker GoogleEvidence notes, transcript
Google's robots.txt talk recalled that robots.txt began in 1994, when bots were crashing servers: Martijn Koster proposed a text file in the root of a site with rules for how automated clients may access it.
Speaker GoogleEvidence transcript
Google said it already supported robots.txt when it started crawling in 1996 as a Stanford PhD project, and has supported it ever since.
Speaker GoogleEvidence transcript
Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)
- Extends D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
- Repeated by D2-C050 Day 2: John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes…
The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.
Speaker GoogleEvidence transcript
Used byglossary term RFC 9309 (Robots Exclusion Protocol)
- Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
Robots.txt is not a security measure: the file sits at a predictable public address, so anyone can read the paths it disallows. Protect a secret folder with authentication, or do not put it on the internet.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-11glossary term robots.txt
The robots.txt file always sits in the root of the host; it cannot be placed anywhere else.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-04glossary term robots.txt
Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-05
- Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
Everything on a site is implicitly allowed, so an allow rule is only needed to re-open a specific path inside a disallowed one.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-05
Comments in robots.txt start with #, and a line without the # that a parser cannot read is ignored anyway, because the standard requires parsers to skip lines they cannot parse, so it acts like a comment.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)
A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-10glossary term User-agent group
- Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.
Speaker GoogleEvidence transcript, slide photo
Used byrequirement DEV-SRV-05
- Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
- Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.
Speaker GoogleEvidence transcript, slide photo
Used byrequirement DEV-SRV-05
- Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
- Extends D1-C082 Day 1: In robots.txt, * matches zero or more of any character and $ marks the end of the URL.
Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.
Speaker GoogleEvidence transcript
Used byrequirement DEV-IDX-11glossary term Google-Extended
- Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
- Repeats D1-C485 Day 1: To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
- Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.
Speaker GoogleEvidence transcript
Used byrequirement DEV-AIF-05
- Repeats D1-C487 Day 1: A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site…
Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-05
- Extended by D1-C532 Day 1: A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts…
The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-06
- Extends D1-C087 Day 1: Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and…
- Extends D1-C123 Day 1: Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's…
The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-10glossary term User-agent group
- Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.
Speaker GoogleEvidence transcript
Used byrequirements DEV-SRV-05, DEV-SRV-06glossary term robots.txt report
- Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
- Extended by D3-C667 Day 3: Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours…
Google said the robots.txt report helps catch CDNs that update a site's robots.txt without the owner's knowledge, and hosts that cloak the robots.txt file, which happens more often than people think.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-06
Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
Speaker GoogleEvidence transcript
Used byglossary term robots.txt report
- Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.
Speaker GoogleEvidence transcript
Used byrequirement DEV-SRV-10
- Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
What Google's documentation says 7
Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.
Publisher Google
Used byrequirements DEV-SRV-05, DEV-SRV-10glossary term User-agent group
- Extended by D1-C519 Day 1: A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several…
- Extended by D1-C527 Day 1: The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another…
- Extended by D1-C531 Day 1: In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of…
- Extended by D1-C533 Day 1: Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that…
When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.
“In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”
Publisher Google
Used byrequirement DEV-SRV-05
- Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
- Extended by D2-C208 Day 2: A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/…
In robots.txt, * matches zero or more of any character and $ marks the end of the URL.
Publisher Google
Used byrequirement DEV-SRV-05
- Extended by D1-C521 Day 1: To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the…
Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
Publisher Google
Used byrequirement DEV-SRV-05glossary term robots.txt
- Extended by D1-C516 Day 1: Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was…
- Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
- Extended by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
Publisher Google
Used byrequirement DEV-SRV-05
- Extended by D3-C613 Day 3: Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an…
- Extended by D3-C614 Day 3: Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for…
Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
“Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.”
Publisher Google
Used byrequirement DEV-IDX-11glossary term Google-Extended
- Extended by D1-C513 Day 1: The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in…
- Extended by D1-C522 Day 1: Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for…
- Extended by D1-C523 Day 1: The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023…
- Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
- Extended by D2-C121 Day 2: Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex…
Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.
Publisher Google Search Console Help
Used byrequirement DEV-SRV-05
- Extended by D1-C528 Day 1: Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version…
- Extended by D1-C530 Day 1: Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
- Extended by D3-C615 Day 3: A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.