Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Working with Google

Web standards and how to take part

Robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored, but no RFC covers robots meta tags yet, and John Mueller invited attendees to help write one, saying standards groups are made up of ordinary people and need only passion and a willingness to join the discussions. Community speaker Natalia Venditto (Adobe) took up a similar remark, which two recordings have her crediting to Gary: the Web Fragments team wants to standardise the library's approach rather than maintain browser patches, and supports ShadowRealm, a JavaScript standard proposal that she said was at stage 2.7 at the time of the talk and would do much of the same isolation work (said at the event, not in Google's docs). Google's Erin Sparling said polyfills can also patch forward, letting a site build now for a feature expected in the future, and that web standards such as WebMCP, combined with schema-based interfaces, let a site expose tools that AI agents can operate (not in Google's docs). Schema.org is another open collaboration, and Google's structured data talk invited people who enjoy ontologies and data to join it. An audio recording of Day 1's robots.txt talk gave the standard's history: Martijn Koster proposed robots.txt in 1994 because bots were crashing servers, it was kept extremely simple so that anyone could implement and understand it, and its formal name is the Robots Exclusion Protocol; the standard requires parsers to skip lines they cannot read, and while writing RFC 9309 Google found opinions split roughly 50-50 on how to treat an unknown line between two user-agent lines (not in Google's docs). On Day 3 Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for. Day 1's lightning talks, captured by a second recording, presented WebMCP in more detail: a way for a site to declare the actions it offers to AI agents, such as booking an appointment, with declarative tools (mostly HTML annotations such as form fields) and imperative ones (booking, filtering, adding to a cart); Chrome's documentation calls it a proposed web standard that sites can test through an origin trial from Chrome 149. In the Day 1 Q&A a Google panelist mentioned crawler best practices being written that would exempt research, malware-scanning and privacy crawlers, apparently because they sometimes need to ignore robots.txt or probe URLs other crawlers would not touch. Author’s view: those best practices match the public IETF Internet-Draft 'Crawler best practices' (July 2025), a draft and not Google documentation.

Based on D2-C050, D2-C051, D2-C052, D2-C216, D2-C217, D2-C251, D2-C252, D2-C253, D2-C254, D2-C255, D2-C281, D2-C303, D2-C485, D3-C614, D1-C281, D1-C283, D1-C316, D1-C432, D1-C433, D1-C510, D1-C512, D1-C516, D1-C518, D1-C526

25 claims · raised in 8 sessions · said or shown on Day 1 and Day 2 and Day 3

Open in Reef mapOpen in Graph

What to do

  • If a standard you depend on is missing, such as a specification for robots meta tags, take part in writing it; standards groups are open to anyone who joins the discussions.
  • Join schema.org's public collaboration if you want a say in the structured data vocabulary.

Day 1: Crawling 11

Said on stage 9

StageConfirmed by docsD1-C281

A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Things

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extended by D2-C303 Day 2: Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents…
StageConfirmed by docsD1-C282

A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Things

Used byrequirement DEV-AIF-06

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extended by D1-C316 Day 1: Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through…
StageConfirmed by docsD1-C283

A community speaker said WebMCP has two kinds of tools: declarative ones, mostly HTML annotations such as the fields of a contact form, and imperative ones for other actions such as booking, filtering a catalogue, adding products to a cart, getting product specs or reordering.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Things

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
StageConfirmed by docsD1-C510

Google's robots.txt talk recalled that robots.txt began in 1994, when bots were crashing servers: Martijn Koster proposed a text file in the root of a site with rules for how automated clients may access it.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

StageConfirmed by docsD1-C512

Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)

  • Extends D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
  • Repeated by D2-C050 Day 2: John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes…
StageConsistent with docsD1-C516

Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
StageConfirmed by docsD1-C518

Comments in robots.txt start with #, and a line without the # that a parser cannot read is ignored anyway, because the standard requires parsers to skip lines they cannot parse, so it acts like a comment.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)

StageNot in docsD1-C526

The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-06

  • Extends D1-C087 Day 1: Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and…
  • Extends D1-C123 Day 1: Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's…
StageNot in docsD1-C432

A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

  • Extended by D1-C433 Day 1: The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices'…

What Google's documentation says 1

DocsSourceD1-C316

Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.

Publisher Chrome for DevelopersAnnotates Day 1, 13:10 · Lightning session A: Automation and AI

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extends D1-C282 Day 1: A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all…

Analysis by the author 1

AnalysisD1-C433

The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.

Author Ibrahim AnjroAnnotates Day 1, 16:35 · Q&A

  • Extends D1-C432 Day 1: A Google panelist said they were working on a set of crawler best practices and offering research…

Day 2: Indexing 13

Said on stage 13

StageConfirmed by docsD2-C050

John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript

  • Repeats D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
StageD2-C052

John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.

Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript

  • Repeated by D2-C217 Day 2: Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard…
StageNot in docsD2-C281

Polyfills work in two directions: they patch backwards where a browser lacks support, or patch forward to build now for a feature expected in the future, as Web Fragments does.

Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript

  • Extends D2-C254 Day 2: Natalia Venditto said the ShadowRealm proposal would do much of what Web Fragments does today with patches…
StageNot in docsD2-C303

Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.

Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript

Things
  • Extends D1-C281 Day 1: A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions…

Day 3: Serving: Ranking, Search Console, and Performance 1

Said on stage 1

StageConfirmed by docsD3-C614

Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C085 Day 1: Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

Across days and sessions 11

  1. Docs D1-C316 Day 1 · Lightning session A: Automation and AI

    Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.

    extends
    Stage D1-C282 Day 1 · Lightning session A: Automation and AI

    A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.

  2. Analysis D1-C433 Day 1 · Q&A

    The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.

    extends
    Stage D1-C432 Day 1 · Q&A

    A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

  3. Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

    extends
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

  4. Stage D1-C516 Day 1 · How Google interprets robots.txt

    Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  5. Stage D1-C526 Day 1 · How Google interprets robots.txt

    The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

    extends
    Slide D1-C087 Day 1 · How Google interprets robots.txt

    Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.

  6. Stage D1-C526 Day 1 · How Google interprets robots.txt

    The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

    extends
    Slide D1-C123 Day 1 · How Google interprets robots.txt

    Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

  7. Stage D2-C281 Day 2 · What is Google friendly JavaScript

    Polyfills work in two directions: they patch backwards where a browser lacks support, or patch forward to build now for a feature expected in the future, as Web Fragments does.

    extends
    Stage D2-C254 Day 2 · Lightning session D: Rendering and JavaScript

    Natalia Venditto said the ShadowRealm proposal would do much of what Web Fragments does today with patches: containerizing and isolating execution.

  8. Stage D2-C303 Day 2 · What is Google friendly JavaScript

    Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.

    extends
    Stage D1-C281 Day 1 · Lightning session A: Automation and AI

    A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.

  9. Stage D3-C614 Day 3 · How long does it take to..?

    Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

    extends
    Docs D1-C085 Day 1 · How Google interprets robots.txt

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

  10. Stage D2-C050 Day 2 · Controlling indexing

    John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

    repeats
    Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

  11. Stage D2-C217 Day 2 · Lightning session D: Rendering and JavaScript

    Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard, and that this is what the Web Fragments work is trying to do.

    repeats
    Stage D2-C052 Day 2 · Controlling indexing

    John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.

Built on these claims 3

Developer requirements 3

Sources 5