Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Topic · Working with Google
Web standards and how to take part
Robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored, but no RFC covers robots meta tags yet, and John Mueller invited attendees to help write one, saying standards groups are made up of ordinary people and need only passion and a willingness to join the discussions. Community speaker Natalia Venditto (Adobe) took up a similar remark, which two recordings have her crediting to Gary: the Web Fragments team wants to standardise the library's approach rather than maintain browser patches, and supports ShadowRealm, a JavaScript standard proposal that she said was at stage 2.7 at the time of the talk and would do much of the same isolation work (said at the event, not in Google's docs). Google's Erin Sparling said polyfills can also patch forward, letting a site build now for a feature expected in the future, and that web standards such as WebMCP, combined with schema-based interfaces, let a site expose tools that AI agents can operate (not in Google's docs). Schema.org is another open collaboration, and Google's structured data talk invited people who enjoy ontologies and data to join it. An audio recording of Day 1's robots.txt talk gave the standard's history: Martijn Koster proposed robots.txt in 1994 because bots were crashing servers, it was kept extremely simple so that anyone could implement and understand it, and its formal name is the Robots Exclusion Protocol; the standard requires parsers to skip lines they cannot read, and while writing RFC 9309 Google found opinions split roughly 50-50 on how to treat an unknown line between two user-agent lines (not in Google's docs). On Day 3 Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for. Day 1's lightning talks, captured by a second recording, presented WebMCP in more detail: a way for a site to declare the actions it offers to AI agents, such as booking an appointment, with declarative tools (mostly HTML annotations such as form fields) and imperative ones (booking, filtering, adding to a cart); Chrome's documentation calls it a proposed web standard that sites can test through an origin trial from Chrome 149. In the Day 1 Q&A a Google panelist mentioned crawler best practices being written that would exempt research, malware-scanning and privacy crawlers, apparently because they sometimes need to ignore robots.txt or probe URLs other crawlers would not touch. Author’s view: those best practices match the public IETF Internet-Draft 'Crawler best practices' (July 2025), a draft and not Google documentation.
If a standard you depend on is missing, such as a specification for robots meta tags, take part in writing it; standards groups are open to anyone who joins the discussions.
Join schema.org's public collaboration if you want a say in the structured data vocabulary.
A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.
A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.
A community speaker said WebMCP has two kinds of tools: declarative ones, mostly HTML annotations such as the fields of a contact form, and imperative ones for other actions such as booking, filtering a catalogue, adding products to a cart, getting product specs or reordering.
Google's robots.txt talk recalled that robots.txt began in 1994, when bots were crashing servers: Martijn Koster proposed a text file in the root of a site with rules for how automated clients may access it.
Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.
Comments in robots.txt start with #, and a line without the # that a parser cannot read is ignored anyway, because the standard requires parsers to skip lines they cannot parse, so it acts like a comment.
The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).
A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).
Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.
The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.
John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.
Natalia Venditto, a Principal Software Engineer at Adobe, gave a seven-minute community lightning talk on JavaScript and web standards, built around the Web Fragments library.
Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard, and that this is what the Web Fragments work is trying to do.
Natalia Venditto said the case for standardizing Web Fragments is that it already relies on many existing browser APIs, built that way so as not to rewrite much or reinvent the wheel.
Natalia Venditto said ShadowRealm, a JavaScript standard proposal, was at stage 2.7 at the time of the talk, which she glossed as meaning that only implementation is missing.
Polyfills work in two directions: they patch backwards where a browser lacks support, or patch forward to build now for a feature expected in the future, as Web Fragments does.
Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.
Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.
A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.
The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.
A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).
Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.
Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.
The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).
Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.
The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).
Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.
Polyfills work in two directions: they patch backwards where a browser lacks support, or patch forward to build now for a feature expected in the future, as Web Fragments does.
Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.
A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.
Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard, and that this is what the Web Fragments work is trying to do.
John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.