A community speaker framed the web as now visited by both humans and AI agents, so sites have to be prepared for both.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Topic · AI and Search
Google said its effort to render pages as users see them keeps its knowledge of the web current and that rendering JavaScript matters whether the client is a search crawler or an AI system, and a community speaker advised putting everything you want cited into the server-rendered HTML because some AI systems cannot render JavaScript yet. Google warned that AI agents browsing for users hit the same bot walls as scrapers and may buy elsewhere, and closed its duplication talk with 'Don't block agents'; John Mueller said AI crawlers do not really know what to do with nofollow links because they look at content (all said at the event, not in Google's docs). Google's documentation says user-triggered fetchers such as Google-Agent generally ignore robots.txt, and Google said a Gemini user asking about a specific page triggers a live read and that Gemini in Chrome relies heavily on screenshots (both said at the event; Google's docs describe neither specifically), extending Day 1's documented point that browser agents read screenshots, the DOM and the accessibility tree. Erin Sparling added that web standards such as WebMCP let a site expose tools that AI agents can operate (said at the event, not in Google's docs). Author’s view: nofollow is not an access control; at Google, Google-Extended governs Gemini training and grounding and the Search generative AI setting governs AI Overviews and AI Mode. Day 3 raised the bar for shopping data: Google said data quality matters more for AI and even more for agents acting for consumers ('garbage in, garbage out'), since an agent with wrong data might buy the wrong product (said at the event), and Google launched the Universal Commerce Protocol on 11 January 2026 as an open standard for agentic commerce, compatible with A2A, AP2 and MCP. The second recording of Day 1 added Google's view of crawlers in general. Gary Illyes said AI agents are technically the same thing as crawlers, HTTP clients acting for a user or a service; Google runs probably hundreds, if not thousands, of crawlers on one infrastructure (its Inside Googlebot post speaks of dozens of other clients, and both say only the larger ones are documented), Googlebot is the crawler for web search including Search's AI features, and the only Google-owned crawlers that ignore robots.txt are contractual ones that crawl sites by agreement (Google's docs call them special-case crawlers). Google said AI Mode sits on its existing crawling infrastructure, so it brings no extra Googlebot crawling, although sites should expect more crawling overall from other services, AI services included. In the Day 1 Q&A a Google panelist said many AI crawlers are less sophisticated than search crawlers and may simply work through a site in order, because model training mainly needs a very large number of tokens, while search crawling tracks which pages change (Gary Illyes likewise said crawling for Gemini may care more about the amount of content than its quality; said at the event). Mainstream crawlers from Google, other search engines and AI companies try to follow robots.txt, so correct rules are the control and crawlers that ignore them are a scraping problem; user-initiated fetchers, and AI agents acting on a user's request, generally do not check robots.txt, which the panelist argued makes business sense because an agent sent to buy something could not otherwise complete the purchase. Nearly all mainstream AI systems use their own user agents, so robots.txt can set a policy for each, but the panelist doubted per-crawler policies make practical sense yet, called blocking all AI crawlers while allowing search crawlers a personal decision and called a default-deny robots.txt a bad pattern. Gary Illyes added that Google now sees more 403 responses (described on stage as 'authentication required') and, more recently still, more 402 Payment Required responses (not in Google's docs), and treats both like a 404, dropping the pages from Search and its AI features. A community speaker described agent readiness in three layers: visual stability (CLS) for how agents see a page; schema, landmarks, a logical heading order and ARIA for how they understand it; and WebMCP, a proposed standard Chrome lets sites test through an origin trial, for how they act on it; the speaker advised against markdown copies of pages for agents, and another community speaker advised against blocking AI training bots in most cases. Author’s view: a logical heading order is an accessibility and agent concern rather than a Google Search requirement, structured data is not required for Google's generative AI search, and a Google user agent missing from the public lists is not proof of a fake request; verify with reverse DNS or Google's published IP ranges.
Things in this topic 28
Counts are claims that name the thing. All things
What to do
A community speaker framed the web as now visited by both humans and AI agents, so sites have to be prepared for both.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
A community speaker advised against publishing markdown copies of HTML pages for AI agents: the copy is a duplicate (which the speaker also called a possible source of cloaking, an uncertain word in the recordings), and the models are trained to read HTML, CSS and JavaScript.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-AIF-02
A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-AIF-05glossary term User-triggered fetchers
A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-HTM-04
A community speaker said layout shifts, such as a button or image popping in after the first load, annoy users and may confuse AI agents, and recommended watching Cumulative Layout Shift (CLS), the Core Web Vitals metric for visual stability.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byglossary term Cumulative Layout Shift (CLS)
A community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which number is the price saves the agent from guessing.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-HTM-04
A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-HTM-04
A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-HTM-04glossary term ARIA
A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-AIF-06glossary term WebMCP
A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-AIF-06
A community speaker said WebMCP has two kinds of tools: declarative ones, mostly HTML annotations such as the fields of a contact form, and imperative ones for other actions such as booking, filtering a catalogue, adding products to a cart, getting product specs or reordering.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-AIF-06glossary term WebMCP
A community speaker summed up agent readiness in three layers: a good CLS improves how agents see content, schema, landmarks and ARIA how they understand it, and WebMCP how they interact with it.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
A community speaker proposed an 'AI bot allow rate' for the know stage: the share of AI training bots a site allows (three of four is 75%), with 100% as the target.
Speaker not identifiedIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
A community speaker advised not blocking AI training bots in most cases, because a model that may not crawl a site finds it harder to represent the brand accurately in its parametric memory.
Speaker not identifiedIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
AI agents are, technically, the same thing as crawlers: HTTP clients that accomplish something on behalf of a user or a service.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Google runs probably hundreds, if not thousands, of crawlers on its crawler infrastructure; some of them are named and some are not.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Googlebot is the crawler Google uses for web search, including Search's AI features.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
The only Google-owned crawlers that do not obey robots.txt are contractual crawlers, which crawl a site whose owner has agreed that Google may crawl it however it likes.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byglossary term Contractual crawlers (special-case crawlers)
Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.
“the number of tokens is actually more important”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-09
Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
“402 will just mean 404 to us”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-09
Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-SRV-09
Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-AIF-05
An audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.
From the audienceIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
Google expects sites to see more crawling overall, because many other services, including AI services, now crawl the web besides Google.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
A community speaker said that more than half of internet traffic in the previous year (2025) was non-human; the source of the figure was not named.
Speaker Jovana AvramovicIn Day 1, 16:20 · Lightning session C: CrawlingEvidence transcript
An audience member asked, in a question submitted at registration, what the biggest difference is between crawling for search and crawling for AI models.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
A Google panelist said many AI crawlers are less sophisticated than search crawlers: in server logs, search crawlers tend to follow where a site changes and which pages are valuable, while AI crawlers may simply work through a site in order, which he put down to less crawling experience and different priorities.
“my feeling is a lot of the AI crawlers are still a bit stupid”
Wording checked against the slide or recording
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas many AI systems crawl a site with no understanding of it and take everything.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
Crawling for AI model training differs from search crawling because training mainly needs a very large number of tokens, and it matters little which pages they come from.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
An audience question asked how to make sure bots behave well beyond robots.txt.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
A Google panelist said crawlers that ignore robots.txt and cause a nuisance are better treated as a scraping problem than as a crawling problem.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
A Google panelist described research, malware-scanning and privacy crawlers as a separate category of crawler that is generally harmless and useful to the web.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05glossary term User-triggered fetchers
AI agents acting on a user's request fall into the same category as user-initiated fetchers: an agent told to look at a site reads the page without checking robots.txt.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05glossary term User-triggered fetchers
A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that site owners can control AI access separately from Googlebot and analyse it in their logs.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense yet, because nobody knows how these systems will develop.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
A Google panelist called blocking all AI crawlers while allowing search crawlers such as Googlebot, Bingbot and Applebot a personal, philosophical decision that every site owner can make.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.
“Personally, I think that's a bad pattern, because you don't know what you're blocking.”
Wording checked against the slide or recording
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.
Publisher Chrome for DevelopersAnnotates Day 1, 13:10 · Lightning session A: Automation and AI
Used byrequirement DEV-AIF-06glossary term WebMCP
Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.
“Googlebot is just a user of something that resembles a centralized crawling platform”
Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works
Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.
Publisher GoogleAnnotates Day 1, 14:05 · How crawling works
Used byglossary term Contractual crawlers (special-case crawlers)
Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirements DEV-MON-04, DEV-SRV-01
A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.
Author Ibrahim AnjroAnnotates Day 1, 13:10 · Lightning session A: Automation and AI
Google's AI optimisation guide says structured data is not required for generative AI search and needs no special schema.org markup, and on Day 2 Google said raw schema.org is generally not put into model context (D2-C477); use markup for rich-result eligibility and clear data, not as an AI-visibility lever.
Author Ibrahim AnjroAnnotates Day 1, 13:10 · Lightning session A: Automation and AI
On stage Google spoke of probably hundreds, if not thousands, of crawlers, while its Inside Googlebot post speaks of dozens of other clients; both agree that only the larger crawlers are documented, so a Google user agent missing from the public lists is not proof of a fake request, and reverse DNS or Google's published IP ranges are the test.
Author Ibrahim AnjroAnnotates Day 1, 14:05 · How crawling works
Do not answer verified Googlebot with 401, 402 or 403, for example from a login wall, a paywall or a pay-per-crawl setup: Google treats them like 404 and drops the pages, and Google's status code page says not to use 401 or 403 to limit crawling; to slow crawling temporarily, return 429 or 503.
Author Ibrahim AnjroAnnotates Day 1, 14:35 · How crawling errors affect Search
The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.
Author Ibrahim AnjroAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-AIF-05
Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.
“Don't block agents.”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript
Used byrequirement DEV-AIF-04
John Mueller said AI crawlers do not really know what to do with nofollow links, because they look at the content rather than building a link graph; he did not say whether he meant Google's AI systems, other AI crawlers or both.
“AI crawlers don't really know what to do with a nofollow link, because they're looking at the content”
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Compared with other search engines and AI crawlers, Google said its effort to mimic what the user sees is essential to keeping its knowledge of the web up to date and comprehensive; it did not say what the others do.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
Being able to render and read JavaScript matters whether the client fetching a page is a search crawler or an AI system, Google said.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
AI Overviews and AI Mode typically do not ground their answers by reading pages live, unlike Gemini when a user asks about a specific page, Google said.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
Used byrequirement DEV-PRF-02
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
Used byrequirement DEV-PRF-02
A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.
“everything you want cited, include it in the raw HTML, server-side rendered”
Speaker Sören BendigIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
Used byrequirement DEV-REN-01
Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.
Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript
When Search Console reports that content is not rendering, Erin Sparling suggested putting AI agents on the problem, but only after doing your own due diligence.
Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript
AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
Used byrequirement DEV-AIF-04
Inside Google, views on structured data split into two camps: one says it is useless because models can generate it or read the page directly, the other says it is the future of machines talking to each other through MCP servers and new standards; the speaker said the truth is in the middle.
Speaker Ryan LeveringIn Day 2, 13:35 · What is Structured Data and why we need it on the internet.Evidence transcript
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Speaker Ryan LeveringIn Day 2, 13:35 · What is Structured Data and why we need it on the internet.Evidence transcript
Used byrequirement DEV-HTM-04
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
Publisher GoogleAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirement DEV-IDX-11
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing
Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
For pages people are likely to ask an AI assistant about (product, pricing, documentation and policy pages), server-render the main content so a live, user-triggered read does not depend on client-side JavaScript finishing quickly.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirements DEV-PRF-02, DEV-REN-01
Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.
Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication
Used byrequirements DEV-AIF-04, DEV-SRV-02
Product data quality matters more than ever for AI ('garbage in, garbage out'), and even more for agents acting on behalf of consumers, because an agent with bad data cannot do the right thing.
“If the agent has bad data, it cannot do the right things.”
Speaker Alex JansenIn Day 3, 13:40 · Shopping on Search: Beyond the blue linksEvidence transcript
The bar for product data quality is higher than ever because agents and users need correct information: with wrong data an agent might buy the wrong product for a user.
Speaker Alex JansenIn Day 3, 13:40 · Shopping on Search: Beyond the blue linksEvidence transcript
Google launched the Universal Commerce Protocol (UCP) on 11 January 2026 as an open standard for agentic commerce across discovery, buying and post-purchase support, co-developed with retailers and platforms and compatible with A2A, AP2 and MCP.
Publisher Google blog (11 January 2026)Annotates Day 3, 13:40 · Shopping on Search: Beyond the blue links
Used byglossary term Universal Commerce Protocol (UCP)
A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.
A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.
A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.
Google's AI optimisation guide says structured data is not required for generative AI search and needs no special schema.org markup, and on Day 2 Google said raw schema.org is generally not put into model context (D2-C477); use markup for rich-result eligibility and clear data, not as an AI-visibility lever.
A community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which number is the price saves the agent from guessing.
Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.
A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.
Googlebot is the crawler Google uses for web search, including Search's AI features.
AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.
AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.
Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.
Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.
Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.
Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.
Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.
Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.
A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.
AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.
Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.
A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.
Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.
Rests on 13 claims, 12 of them in this topic
Use semantic, accessible HTML: real buttons, labelled forms and a logical heading order
Rests on 10 claims, 5 of them in this topic
Try WebMCP to declare the actions a site offers to AI agents
Rests on 4 claims, 4 of them in this topic
Do not block the AI agents you want as visitors
Rests on 3 claims, 3 of them in this topic
Make JavaScript-generated content render quickly
Rests on 5 claims, 3 of them in this topic
Do not answer crawlers with 401, 402 or 403 on pages that should be in Search
Rests on 5 claims, 3 of them in this topic
Server-render the main content and everything indexing depends on (SSR, static generation or hybrid)
Rests on 11 claims, 2 of them in this topic
Spend no effort on llms.txt, AI text files or AI-specific markup for Google
Rests on 15 claims, 1 of them in this topic
Use the Google-Extended robots.txt token to control Gemini training and grounding
Rests on 8 claims, 1 of them in this topic
Monitor Crawl stats and server logs for verified Googlebot traffic
Rests on 13 claims, 1 of them in this topic
Let verified Google crawlers through firewalls, CDNs and bot protection
Rests on 11 claims, 1 of them in this topic
Never serve bot challenges or error pages to crawlers with a 200 status
Rests on 7 claims, 1 of them in this topic