Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

robots.txt group parsing differences

A user-agent line with its allow and disallow rules forms a group, and one group can name several crawlers; Google's Day 1 robots.txt talk advised giving a named crawler its own group only when it needs rules the * group should not grant to every crawler. Groups are not additive: a crawler that a named group matches follows only that group, so a Googlebot group that blocks /dogs/ leaves Googlebot free to crawl what the * group blocks, as Dave Smart showed in Lightning session B and Google's specification confirms. The talk's quiz made the same point: to let Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed, the answer was a Googlebot group with both the disallow and the allow, and Google's specification combines all groups that name the same user agent into one. Crawlers disagree on how to treat an unknown line between user-agent lines: Google said that while writing RFC 9309 it asked people whether the first crawler should inherit the rules that follow and opinions split roughly 50-50 (not in Google's docs). Googlebot ignores the line and merges the groups, as Google's robots.txt specification documents, while a Google slide showed Bing ending the group there instead, so the same file can block one crawler and not another. Author’s view: Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files, so put any non-standard line after a group's rules and test the file with each search engine's tools.

Based on D1-C087, D1-C123, D1-C080, D1-C519, D1-C533, D1-C531, D1-C527, D1-C526, D1-C088

9 claims · raised in 2 sessions · said or shown on Day 1

Open in Reef mapOpen in Graph

Things in this topic 5

Counts are claims that name the thing. All things

What to do

  • Never put non-standard lines such as Content-Signal between user-agent lines; put them after a group's rules.
  • When a crawler gets its own group, repeat in it every rule of the * group it should still follow; groups do not add up.
  • Give a named crawler its own group only when it needs rules the * group should not grant to every crawler.

Day 1: Crawling 9

Shown on screen 2

SlideConfirmed by docsD1-C087

Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence slide photo, transcript

Things

Used byrequirement DEV-SRV-06

  • Extended by D1-C526 Day 1: The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked…
SlideNot in docsD1-C123

Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

“Unpredictability is never a good time.”

Wording checked against the slide or recording

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence slide photo, transcript

Used byrequirement DEV-SRV-06

  • Extended by D1-C526 Day 1: The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked…

Said on stage 5

StageConsistent with docsD1-C519

A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageNot in docsD1-C526

The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-06

  • Extends D1-C087 Day 1: Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and…
  • Extends D1-C123 Day 1: Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's…
StageConsistent with docsD1-C527

The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConsistent with docsD1-C531

In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-10

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConfirmed by docsD1-C533

Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…

What Google's documentation says 1

DocsSourceD1-C080

Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirements DEV-SRV-05, DEV-SRV-10glossary term User-agent group

  • Extended by D1-C519 Day 1: A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several…
  • Extended by D1-C527 Day 1: The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another…
  • Extended by D1-C531 Day 1: In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of…
  • Extended by D1-C533 Day 1: Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that…

Analysis by the author 1

AnalysisD1-C088

Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files. Never place any non-standard directive between user-agent lines; put it after a group's rules and test the file with each search engine's tools, such as Bing's robots.txt tester and Google's open-source robots.txt parser.

Author Ibrahim AnjroAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-SRV-06

Across days and sessions 6

  1. Stage D1-C519 Day 1 · How Google interprets robots.txt

    A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  2. Stage D1-C526 Day 1 · How Google interprets robots.txt

    The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

    extends
    Slide D1-C087 Day 1 · How Google interprets robots.txt

    Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.

  3. Stage D1-C526 Day 1 · How Google interprets robots.txt

    The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

    extends
    Slide D1-C123 Day 1 · How Google interprets robots.txt

    Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

  4. Stage D1-C527 Day 1 · How Google interprets robots.txt

    The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  5. Stage D1-C531 Day 1 · How Google interprets robots.txt

    In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  6. Stage D1-C533 Day 1 · Lightning session B: Robots.txt

    Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

Built on these claims 3

Developer requirements 3

Sources 1