Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · AI and Search

AI crawlers and agents on your site

Google said its effort to render pages as users see them keeps its knowledge of the web current and that rendering JavaScript matters whether the client is a search crawler or an AI system, and a community speaker advised putting everything you want cited into the server-rendered HTML because some AI systems cannot render JavaScript yet. Google warned that AI agents browsing for users hit the same bot walls as scrapers and may buy elsewhere, and closed its duplication talk with 'Don't block agents'; John Mueller said AI crawlers do not really know what to do with nofollow links because they look at content (all said at the event, not in Google's docs). Google's documentation says user-triggered fetchers such as Google-Agent generally ignore robots.txt, and Google said a Gemini user asking about a specific page triggers a live read and that Gemini in Chrome relies heavily on screenshots (both said at the event; Google's docs describe neither specifically), extending Day 1's documented point that browser agents read screenshots, the DOM and the accessibility tree. Erin Sparling added that web standards such as WebMCP let a site expose tools that AI agents can operate (said at the event, not in Google's docs). Author’s view: nofollow is not an access control; at Google, Google-Extended governs Gemini training and grounding and the Search generative AI setting governs AI Overviews and AI Mode. Day 3 raised the bar for shopping data: Google said data quality matters more for AI and even more for agents acting for consumers ('garbage in, garbage out'), since an agent with wrong data might buy the wrong product (said at the event), and Google launched the Universal Commerce Protocol on 11 January 2026 as an open standard for agentic commerce, compatible with A2A, AP2 and MCP. The second recording of Day 1 added Google's view of crawlers in general. Gary Illyes said AI agents are technically the same thing as crawlers, HTTP clients acting for a user or a service; Google runs probably hundreds, if not thousands, of crawlers on one infrastructure (its Inside Googlebot post speaks of dozens of other clients, and both say only the larger ones are documented), Googlebot is the crawler for web search including Search's AI features, and the only Google-owned crawlers that ignore robots.txt are contractual ones that crawl sites by agreement (Google's docs call them special-case crawlers). Google said AI Mode sits on its existing crawling infrastructure, so it brings no extra Googlebot crawling, although sites should expect more crawling overall from other services, AI services included. In the Day 1 Q&A a Google panelist said many AI crawlers are less sophisticated than search crawlers and may simply work through a site in order, because model training mainly needs a very large number of tokens, while search crawling tracks which pages change (Gary Illyes likewise said crawling for Gemini may care more about the amount of content than its quality; said at the event). Mainstream crawlers from Google, other search engines and AI companies try to follow robots.txt, so correct rules are the control and crawlers that ignore them are a scraping problem; user-initiated fetchers, and AI agents acting on a user's request, generally do not check robots.txt, which the panelist argued makes business sense because an agent sent to buy something could not otherwise complete the purchase. Nearly all mainstream AI systems use their own user agents, so robots.txt can set a policy for each, but the panelist doubted per-crawler policies make practical sense yet, called blocking all AI crawlers while allowing search crawlers a personal decision and called a default-deny robots.txt a bad pattern. Gary Illyes added that Google now sees more 403 responses (described on stage as 'authentication required') and, more recently still, more 402 Payment Required responses (not in Google's docs), and treats both like a 404, dropping the pages from Search and its AI features. A community speaker described agent readiness in three layers: visual stability (CLS) for how agents see a page; schema, landmarks, a logical heading order and ARIA for how they understand it; and WebMCP, a proposed standard Chrome lets sites test through an origin trial, for how they act on it; the speaker advised against markdown copies of pages for agents, and another community speaker advised against blocking AI training bots in most cases. Author’s view: a logical heading order is an accessibility and agent concern rather than a Google Search requirement, structured data is not required for Google's generative AI search, and a Google user agent missing from the public lists is not proof of a fake request; verify with reverse DNS or Google's published IP ranges.

What to do

  • Check what your CDN or bot protection serves to verified Googlebot and to the agents you want to allow, and serve any challenge page with a 503, never as a 200 page.
  • Put the content you want cited in the server-rendered HTML, especially product, pricing, documentation and policy pages.
  • Control AI access with robots.txt rules for specific crawlers, not nofollow, remembering that user-triggered fetchers generally ignore robots.txt.
  • Use semantic HTML so agents that read the DOM, the accessibility tree or screenshots can use the page.
  • Do not answer verified Googlebot with 401, 402 or 403 from a login wall, paywall or pay-per-crawl setup; return 429 or 503 to slow crawling for a short time.
  • Before blocking an unknown crawler, look up its user agent to find the robots.txt token that controls it, and avoid a default-deny robots.txt that blocks crawlers you have not identified.
  • Verify a Google crawler by reverse DNS or Google's published IP ranges, not by whether its user agent is on a public list.
  • Serve one HTML page to people and agents rather than a markdown copy, with landmarks, a logical heading order and ARIA where native elements do not fit.

Day 1: Crawling 55

Said on stage 46

StageConsistent with docsD1-C272

A community speaker advised against publishing markdown copies of HTML pages for AI agents: the copy is a duplicate (which the speaker also called a possible source of cloaking, an uncertain word in the recordings), and the models are trained to read HTML, CSS and JavaScript.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-AIF-02

StageConfirmed by docsD1-C273

A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-AIF-05glossary term User-triggered fetchers

  • Repeated by D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…
  • Repeated by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
StageConfirmed by docsD1-C274

A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-HTM-04

  • Repeats D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
StageConsistent with docsD1-C275

A community speaker said layout shifts, such as a button or image popping in after the first load, annoy users and may confuse AI agents, and recommended watching Cumulative Layout Shift (CLS), the Core Web Vitals metric for visual stability.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byglossary term Cumulative Layout Shift (CLS)

StageConsistent with docsD1-C277

A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-HTM-04

  • Extended by D1-C314 Day 1: A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do…
StageConsistent with docsD1-C278

A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-HTM-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…
StageConfirmed by docsD1-C279

A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-HTM-04glossary term ARIA

  • Extends D1-C118 Day 1: Attendees were told to check ARIA and accessibility.
StageConfirmed by docsD1-C281

A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Things

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extended by D2-C303 Day 2: Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents…
StageConfirmed by docsD1-C282

A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Things

Used byrequirement DEV-AIF-06

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extended by D1-C316 Day 1: Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through…
StageConfirmed by docsD1-C283

A community speaker said WebMCP has two kinds of tools: declarative ones, mostly HTML annotations such as the fields of a contact form, and imperative ones for other actions such as booking, filtering a catalogue, adding products to a cart, getting product specs or reordering.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Things

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
StageConsistent with docsD1-C325

Googlebot is the crawler Google uses for web search, including Search's AI features.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things
  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageConsistent with docsD1-C330

The only Google-owned crawlers that do not obey robots.txt are contractual crawlers, which crawl a site whose owner has agreed that Google may crawl it however it likes.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Used byglossary term Contractual crawlers (special-case crawlers)

StageNot in docsD1-C335

Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

“the number of tokens is actually more important”

Wording checked against the slide or recording

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Things
  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageConsistent with docsD1-C365

Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Used byrequirement DEV-SRV-09

  • Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
StageConsistent with docsD1-C366

Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.

“402 will just mean 404 to us”

Wording checked against the slide or recording

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Used byrequirement DEV-SRV-09

  • Extended by D1-C509 Day 1: Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than…
StageNot in docsD1-C509

Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript

Things

Used byrequirement DEV-SRV-09

  • Extends D1-C366 Day 1: Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.
  • Extends D1-C365 Day 1: Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication…
StageD1-C387

An audience member asked, in questions submitted before the event, how often Googlebot should be expected to visit a site now that AI Mode has rolled out, and whether AI means more crawling.

From the audienceIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

  • Answered by D1-C388 Day 1: AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl…
  • Answered by D1-C389 Day 1: Google expects sites to see more crawling overall, because many other services, including AI services, now…
StageConsistent with docsD1-C388

AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript

Things
  • Answers D1-C387 Day 1: An audience member asked, in questions submitted before the event, how often Googlebot should be expected to…
  • Extends D1-C038 Day 1: AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on…
StageD1-C422

An audience member asked, in a question submitted at registration, what the biggest difference is between crawling for search and crawling for AI models.

From the audienceIn Day 1, 16:35 · Q&AEvidence transcript

  • Answered by D1-C423 Day 1: A Google panelist said many AI crawlers are less sophisticated than search crawlers: in server logs, search…
  • Answered by D1-C424 Day 1: Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas…
  • Answered by D1-C425 Day 1: Crawling for AI model training differs from search crawling because training mainly needs a very large number…
StageD1-C423

A Google panelist said many AI crawlers are less sophisticated than search crawlers: in server logs, search crawlers tend to follow where a site changes and which pages are valuable, while AI crawlers may simply work through a site in order, which he put down to less crawling experience and different priorities.

“my feeling is a lot of the AI crawlers are still a bit stupid”

Wording checked against the slide or recording

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

  • Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
StageConsistent with docsD1-C424

Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas many AI systems crawl a site with no understanding of it and take everything.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
StageNot in docsD1-C425

Crawling for AI model training differs from search crawling because training mainly needs a very large number of tokens, and it matters little which pages they come from.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
StageConfirmed by docsD1-C428

Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C427 Day 1: An audience question asked how to make sure bots behave well beyond robots.txt.
StageConfirmed by docsD1-C434

User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05glossary term User-triggered fetchers

  • Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
  • Extended by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
StageD1-C436

A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

  • Extended by D2-C377 Day 2: AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and…
StageD1-C481

An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that site owners can control AI access separately from Googlebot and analyse it in their logs.

From the audienceIn Day 1, 16:35 · Q&AEvidence transcript

Things
  • Answered by D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
  • Answered by D1-C483 Day 1: A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense…
  • Answered by D1-C484 Day 1: A Google panelist called blocking all AI crawlers while allowing search crawlers such as Googlebot, Bingbot…
  • Answered by D1-C485 Day 1: To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.
  • Answered by D1-C487 Day 1: A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site…
  • Answered by D1-C488 Day 1: When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt…
  • Answered by D1-C489 Day 1: A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a…
StageConsistent with docsD1-C482

Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
StageD1-C483

A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense yet, because nobody knows how these systems will develop.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
StageNot in docsD1-C487

A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Repeated by D1-C524 Day 1: Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.
StageNot in docsD1-C488

When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
StageD1-C489

A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.

“Personally, I think that's a bad pattern, because you don't know what you're blocking.”

Wording checked against the slide or recording

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…

What Google's documentation says 4

DocsSourceD1-C316

Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.

Publisher Chrome for DevelopersAnnotates Day 1, 13:10 · Lightning session A: Automation and AI

Used byrequirement DEV-AIF-06glossary term WebMCP

  • WebMCP Chrome for Developers · checked 3 October 2026
  • Extends D1-C282 Day 1: A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all…
DocsSourceD1-C342

Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.

“Googlebot is just a user of something that resembles a centralized crawling platform”

Publisher Search Central blog (31 March 2026)Annotates Day 1, 14:05 · How crawling works

DocsSourceD1-C343

Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.

Publisher GoogleAnnotates Day 1, 14:05 · How crawling works

Used byglossary term Contractual crawlers (special-case crawlers)

DocsSourceD1-C138

Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.

Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search

Used byrequirements DEV-MON-04, DEV-SRV-01

Analysis by the author 5

AnalysisD1-C314

A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.

Author Ibrahim AnjroAnnotates Day 1, 13:10 · Lightning session A: Automation and AI

  • Extends D1-C277 Day 1: A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section…
AnalysisD1-C315

Google's AI optimisation guide says structured data is not required for generative AI search and needs no special schema.org markup, and on Day 2 Google said raw schema.org is generally not put into model context (D2-C477); use markup for rich-result eligibility and clear data, not as an AI-visibility lever.

Author Ibrahim AnjroAnnotates Day 1, 13:10 · Lightning session A: Automation and AI

  • Extends D1-C276 Day 1: A community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which…
AnalysisD1-C523

The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

Author Ibrahim AnjroAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-AIF-05

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…

Day 2: Indexing 17

Shown on screen 1

SlideNot in docsD2-C397

Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.

“Don't block agents.”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-AIF-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…

Said on stage 11

StageNot in docsD2-C066

John Mueller said AI crawlers do not really know what to do with nofollow links, because they look at the content rather than building a link graph; he did not say whether he meant Google's AI systems, other AI crawlers or both.

“AI crawlers don't really know what to do with a nofollow link, because they're looking at the content”

Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript

Things
StageConsistent with docsD2-C117

AI Overviews and AI Mode typically do not ground their answers by reading pages live, unlike Gemini when a user asks about a specific page, Google said.

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Used byrequirement DEV-PRF-02

StageNot in docsD2-C120

When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.

Speaker Erin SparlingIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Things

Used byrequirement DEV-PRF-02

  • Extends D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageConsistent with docsD2-C139

A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.

“everything you want cited, include it in the raw HTML, server-side rendered”

Speaker Sören BendigIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

Used byrequirement DEV-REN-01

  • Extends D1-C056 Day 1: Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and…
StageNot in docsD2-C303

Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.

Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript

Things
  • Extends D1-C281 Day 1: A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions…
StageD2-C457

Inside Google, views on structured data split into two camps: one says it is useless because models can generate it or read the page directly, the other says it is the future of machines talking to each other through MCP servers and new standards; the speaker said the truth is in the middle.

Speaker Ryan LeveringIn Day 2, 13:35 · What is Structured Data and why we need it on the internet.Evidence transcript

StageConsistent with docsD2-C465

Gemini in Chrome relies heavily on the screenshot it takes of a page.

Speaker Ryan LeveringIn Day 2, 13:35 · What is Structured Data and why we need it on the internet.Evidence transcript

Used byrequirement DEV-HTM-04

  • Extends D1-C131 Day 1: Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM…

What Google's documentation says 1

DocsSourceD2-C122

Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

Publisher GoogleAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-IDX-11

  • Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
  • Extends D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…

Analysis by the author 4

AnalysisD2-C067

nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing

  • Extends D1-C127 Day 1: Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed…
  • Extends D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
AnalysisD2-C123

Google documents Gemini grounding only as content from the Search index at prompt time; the live read of a specific page at a user's request, described on stage, is not documented, so it is unclear whether it works like a user-triggered fetcher, which generally ignores robots.txt, or follows Google-Extended.

Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

AnalysisD2-C378

Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.

Author Ibrahim AnjroAnnotates Day 2, 11:55 · Handling web duplication

Used byrequirements DEV-AIF-04, DEV-SRV-02

Day 3: Serving: Ranking, Search Console, and Performance 3

Said on stage 2

What Google's documentation says 1

DocsSourceD3-C398

Google launched the Universal Commerce Protocol (UCP) on 11 January 2026 as an open standard for agentic commerce across discovery, buying and post-purchase support, co-developed with retailers and platforms and compatible with A2A, AP2 and MCP.

Publisher Google blog (11 January 2026)Annotates Day 3, 13:40 · Shopping on Search: Beyond the blue links

Used byglossary term Universal Commerce Protocol (UCP)

Across days and sessions 24

  1. Stage D1-C278 Day 1 · Lightning session A: Automation and AI

    A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.

    extends
    Docs D1-C131 Day 1 · session not recorded

    Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

  2. Stage D1-C279 Day 1 · Lightning session A: Automation and AI

    A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.

    extends
    Stage D1-C118 Day 1 · session not recorded

    Attendees were told to check ARIA and accessibility.

  3. Analysis D1-C314 Day 1 · Lightning session A: Automation and AI

    A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.

    extends
    Stage D1-C277 Day 1 · Lightning session A: Automation and AI

    A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.

  4. Analysis D1-C315 Day 1 · Lightning session A: Automation and AI

    Google's AI optimisation guide says structured data is not required for generative AI search and needs no special schema.org markup, and on Day 2 Google said raw schema.org is generally not put into model context (D2-C477); use markup for rich-result eligibility and clear data, not as an AI-visibility lever.

    extends
    Stage D1-C276 Day 1 · Lightning session A: Automation and AI

    A community speaker said schema markup helps AI agents interpret a page: on a product page, marking up which number is the price saves the agent from guessing.

  5. Docs D1-C316 Day 1 · Lightning session A: Automation and AI

    Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.

    extends
    Stage D1-C282 Day 1 · Lightning session A: Automation and AI

    A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.

  6. Stage D1-C325 Day 1 · How crawling works

    Googlebot is the crawler Google uses for web search, including Search's AI features.

    extends
    Slide D1-C038 Day 1 · How Search works and where's AI?

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

  7. Stage D1-C335 Day 1 · How crawling works

    Crawling for Gemini may be set to care less about quality and more about the amount of content, because for large language models the number of tokens matters more than quality.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  8. Stage D1-C388 Day 1 · How Google thinks about crawl budget

    AI Mode is a feature that sits on top of Google's existing crawling infrastructure, so Google does not crawl a site more because of AI Mode.

    extends
    Slide D1-C038 Day 1 · How Search works and where's AI?

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

  9. Stage D1-C509 Day 1 · How crawling errors affect Search

    Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

    extends
    Stage D1-C365 Day 1 · How crawling errors affect Search

    Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.

  10. Stage D1-C509 Day 1 · How crawling errors affect Search

    Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

    extends
    Stage D1-C366 Day 1 · How crawling errors affect Search

    Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.

  11. Analysis D1-C523 Day 1 · How Google interprets robots.txt

    The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  12. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Docs D1-C127 Day 1 · How Google thinks about crawl budget

    Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

  13. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Stage D1-C482 Day 1 · Q&A

    Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

  14. Stage D2-C120 Day 2 · Lightning session D: Rendering and JavaScript

    When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.

    extends
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  15. Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    extends
    Stage D1-C434 Day 1 · Q&A

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

  16. Stage D2-C139 Day 2 · Lightning session D: Rendering and JavaScript

    A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.

    extends
    Slide D1-C056 Day 1 · How Search works and where's AI?

    Myth: your website is no longer relevant. Google's answer: keep content crawlable, well structured, fast and easy to read, for readers and for AI tools.

  17. Stage D2-C303 Day 2 · What is Google friendly JavaScript

    Combining schema-based interfaces with web standards such as WebMCP lets a site expose tools to AI agents, which can then operate the interfaces, for example to test them.

    extends
    Stage D1-C281 Day 1 · Lightning session A: Automation and AI

    A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.

  18. Stage D2-C377 Day 2 · Handling web duplication

    AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

    extends
    Stage D1-C436 Day 1 · Q&A

    A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

  19. Slide D2-C397 Day 2 · Handling web duplication

    Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.

    extends
    Docs D1-C131 Day 1 · session not recorded

    Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

  20. Stage D2-C465 Day 2 · What is Structured Data and why we need it on the internet.

    Gemini in Chrome relies heavily on the screenshot it takes of a page.

    extends
    Docs D1-C131 Day 1 · session not recorded

    Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

  21. Stage D1-C274 Day 1 · Lightning session A: Automation and AI

    A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.

    repeats
    Docs D1-C131 Day 1 · session not recorded

    Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

  22. Stage D1-C434 Day 1 · Q&A

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

    repeats
    Stage D1-C273 Day 1 · Lightning session A: Automation and AI

    A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

  23. Stage D1-C524 Day 1 · How Google interprets robots.txt

    Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.

    repeats
    Stage D1-C487 Day 1 · Q&A

    A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

  24. Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    repeats
    Stage D1-C273 Day 1 · Lightning session A: Automation and AI

    A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

Built on these claims 12

Developer requirements 12

Sources 24