Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 1 · Wednesday 30 September 2026 · 16:35

Q&A

Speakers John Mueller, Gary Illyes, Search Relations

Q&ACoverageTranscript

Panel named in the moderator's invitation ('John and Gary'); an answer carries a name only where an on-stage cue names its speaker. Audience questioners are not named. Transcript from a second attendee recording; two audio recordings cover most of the panel as well (from the AI crawlers question to the llms.txt answers, and from the Search Console API answer to the end).

Said on stage 78

StageD1-C120

To get Google to build something, such as an addition to an API, report and request it publicly and in volume.

Speaker not identifiedEvidence notes, transcript

  • Extended by D2-C511 Day 2: Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a…

QQuestion and answers

StageD1-C422

An audience member asked, in a question submitted at registration, what the biggest difference is between crawling for search and crawling for AI models.

From the audienceEvidence transcript

StageD1-C423

A Google panelist said many AI crawlers are less sophisticated than search crawlers: in server logs, search crawlers tend to follow where a site changes and which pages are valuable, while AI crawlers may simply work through a site in order, which he put down to less crawling experience and different priorities.

“my feeling is a lot of the AI crawlers are still a bit stupid”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

StageNot in docsD1-C425

Crawling for AI model training differs from search crawling because training mainly needs a very large number of tokens, and it matters little which pages they come from.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

StageD1-C426

Asked whether optimising for AI search can hurt classic search or the reverse, a Google panelist said yes, pointing to questionable GEO advice published online that he declined to name.

Speaker not identifiedEvidence transcript

Things

QQuestion and answers

StageConfirmed by docsD1-C428

Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

StageNot in docsD1-C432

A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

Speaker not identifiedEvidence transcript

  • Extended by D1-C433 Day 1: The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices'…
StageConfirmed by docsD1-C434

User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05glossary term User-triggered fetchers

  • Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
  • Extended by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
StageD1-C436

A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

Speaker not identifiedEvidence transcript

  • Extended by D2-C377 Day 2: AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and…
StageD1-C437

A Google panelist said much of robots.txt is written with search engines in mind: a search engine should never add items to a cart or check out, but an agent acting for a user probably should be able to.

Speaker not identifiedEvidence transcript

QQuestion and answers

StageD1-C438

An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as sports sites, whose content changes constantly and is largely near-duplicate across seasons.

From the audienceEvidence transcript

StageConsistent with docsD1-C439

Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.

Speaker not identifiedEvidence transcript

  • Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
  • Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
  • Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
StageConfirmed by docsD1-C440

If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-PRF-01

  • Extended by D3-C618 Day 3: When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about…

QQuestion and answers

StageD1-C441

An audience member asked how to tell, from log files or by other methods, whether Google is spending crawl budget inefficiently on a site.

From the audienceEvidence transcript

StageNot in docsD1-C442

Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the site lies; many similar URLs being crawled may not matter if the site is nowhere near that limit.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-MON-04

StageConsistent with docsD1-C444

Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extends D1-C397 Day 1: A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary…
StageConsistent with docsD1-C445

One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-11

  • Repeats D1-C403 Day 1: A community speaker advised that a web application compute the expected URL for every request, for example…
StageNot in docsD1-C446

Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
StageConsistent with docsD1-C447

Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…

QQuestion and answers

StageD1-C449

A Google panelist said he expected more AI reporting to launch in Search Console, without a timeline or a promise, because many people at Google care about giving site owners the information to understand and adapt to what happens in Search.

Speaker not identifiedEvidence transcript

  • Extended by D3-C440 Day 3: Google called AI reporting in Search Console an evolving space and expects the generative AI report to get…
StageD1-C450

Google said reporting on AI features should not simply mirror classic search results, because users get more information before they click and interact with pages differently, so it first has to work out which data would be useful and actionable.

Speaker not identifiedEvidence transcript

QQuestion and answer

QQuestion and answers

StageD1-C453

An audience member asked whether llms.txt matters, given that Google had said it does not while PageSpeed Insights and Lighthouse advocate it and audit tools ask site owners to create the file.

From the audienceEvidence transcript

Things
StageConsistent with docsD1-C455

Gary Illyes said llms.txt matters to some people, which is why tools such as Lighthouse add checks for it, so both sides of the debate are right in their own context.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-02

  • llms.txt Chrome for Developers · checked 3 October 2026
StageD1-C456

A Google panelist said he knew of no plan for Google to use llms.txt, though it could happen, and that a possible future change is no reason to implement it now.

“Just because maybe something will change in the future doesn't mean you should take action now.”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Things
StageD1-C457

A Google panelist said most llms.txt files he had seen are generated automatically by a setting in SEO plugins, so if Google Search ever made llms.txt matter, a site could add one by ticking a checkbox.

Speaker not identifiedEvidence transcript

Things
StageD1-C458

A Google panelist objected that llms.txt is designed for supposedly intelligent systems that should be able to parse a website.

“it's designed for allegedly intelligent systems”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Things
StageD1-C459

A Google panelist compared llms.txt to the old meta keywords debate: an AI agent should not blindly trust what a site says about its own authority, just as a site calling itself the best car insurance site is no reason to stop looking at others.

Speaker not identifiedEvidence transcript

Things

QQuestion and answers

QQuestion and answers

StageD1-C465

An audience member asked how a large news site can tell whether crawl budget is limiting how fast new articles are discovered (within minutes), and which statistics in Search Console and the logs show this.

From the audienceEvidence transcript

StageConsistent with docsD1-C466

Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-MON-11

  • Repeats D1-C381 Day 1: Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few…
  • Extended by D3-C606 Day 3: For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading…
StageNot in docsD1-C467

To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Extended by D1-C543 Day 1: Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises…
StageConsistent with docsD1-C468

When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google crawls and recrawls a lot and judging whether they are useful.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-MON-11

StageConsistent with docsD1-C469

Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-URL-08

  • Extended by D1-C470 Day 1: Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless…

QQuestion and answers

StageD1-C471

An audience question noted that a cats.txt file received more crawler requests than llms.txt in the server logs and asked how important that is.

From the audienceEvidence transcript

Things
StageD1-C472

cats.txt was created by an SEO as a satirical take on llms.txt, to show that crawlers fetch whatever files they are given, so requests for llms.txt in server logs do not mean the file is important.

Speaker not identifiedEvidence transcript

Things

Used byglossary term cats.txt

QQuestion and answer

StageD1-C474

An audience member asked where the sweet spot is for a sitemap that misses nothing but does not overflow Search Console.

From the audienceEvidence transcript

StageConfirmed by docsD1-C475

There is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that Google can find each of them.

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-URL-05

QQuestion and answers

StageConsistent with docsD1-C477

In a migration about four or five years before the event, Google consolidated 12 or 14 of its blogs in different languages, a Help Center and its old developers site into one site.

Speaker not identifiedEvidence transcript

  • Extended by D1-C480 Day 1: The consolidation described on stage matches Google's November 2020 move from Google Webmasters to Google…
StageNot in docsD1-C478

The first priority in Google's own site consolidation was to identify the popular URLs people care about and make sure the migration did not damage them.

Speaker not identifiedEvidence transcript

StageConsistent with docsD1-C479

Where content was duplicated across languages, Google's own site consolidation redirected two language versions into one, giving both old URLs one target path to redirect to.

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-CAN-10

  • Extended by D1-C541 Day 1: A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in…
StageConsistent with docsD1-C540

Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-CAN-09

  • Extended by D2-C884 Day 2: The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a…

QQuestion and answers

StageConsistent with docsD1-C482

Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

  • Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
StageD1-C483

A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense yet, because nobody knows how these systems will develop.

Speaker not identifiedEvidence transcript

StageD1-C489

A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.

“Personally, I think that's a bad pattern, because you don't know what you're blocking.”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Used byrequirement DEV-AIF-05

QQuestion and answers

StageD1-C490

An audience member asked what happens to thousands of crawled and indexed paginated category and tag pages if they are replaced by a single page with a JavaScript load-more button, and whether the old pages should then redirect to the canonical first page.

From the audienceEvidence transcript

StageConfirmed by docsD1-C491

If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.

“Googlebot is not running around clicking on buttons”

Wording checked against the slide or recording

Speaker not identifiedEvidence transcript

Things

Used byrequirement DEV-URL-07

  • Extended by D2-C271 Day 2: Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on…
StageConsistent with docsD1-C114

A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

“the infinite scrolling or load more thing is kind of dangerous, depending on what you want to achieve”

Wording checked against the slide or recording

Speaker not identifiedEvidence notes, transcript

Things

Used byrequirements DEV-URL-06, DEV-URL-07

  • Extended by D2-C197 Day 2: Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll'…
  • Extended by D2-C271 Day 2: Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on…
StageConsistent with docsD1-C492

John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a rough view of whether discovery crawling is in a reasonable range, though measuring publish-to-first-crawl time is more accurate.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-MON-11

  • Extended by D1-C544 Day 1: John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering…

QQuestion and answers

StageConsistent with docsD1-C493

There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.

Speaker not identifiedEvidence transcript

Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)

  • Extended by D3-C617 Day 3: Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2…
StageConsistent with docsD1-C546

A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).

Speaker not identifiedEvidence transcript

  • Extends D1-C374 Day 1: Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can…
  • Extended by D1-C547 Day 1: Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the…
StageConsistent with docsD1-C538

A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.

Speaker not identifiedEvidence transcript

Used byglossary term Crawl demand

  • Extends D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
  • Extends D1-C334 Day 1: Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site…
  • Repeated by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
StageConsistent with docsD1-C539

A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

Speaker not identifiedEvidence transcript

Used byrequirement DEV-URL-08

  • Extends D1-C447 Day 1: Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its…
  • Extends D1-C446 Day 1: Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a…
StageNot in docsD1-C541

A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in Google's own site migration, because it was the only option available to the person doing it (the recording does not make fully clear whether the JavaScript performed the redirects or built the mapping).

Speaker not identifiedEvidence transcript

Used byrequirement DEV-CAN-01

  • Extends D1-C479 Day 1: Where content was duplicated across languages, Google's own site consolidation redirected two language…
  • Extended by D1-C542 Day 1: Google's redirects guide says Google Search follows JavaScript location redirects only after rendering, may…
StageNot in docsD1-C543

Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-MON-11

  • Extends D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
StageConsistent with docsD1-C544

John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering new pages, that is a lot of discovery crawling, which suggests that crawling of new pages is not the problem.

Speaker John MuellerEvidence transcript

Used byrequirement DEV-MON-11

  • Extends D1-C492 Day 1: John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a…

What Google's documentation says 1

DocsSourceD1-C460

Chrome's Lighthouse documentation lists an llms.txt audit among its agentic browsing audits: it flags a server error when llms.txt is fetched and marks the audit not applicable when the file is missing, because providing the file is optional for now.

Publisher Chrome for Developers

Used byrequirement DEV-AIF-02

  • llms.txt Chrome for Developers · checked 3 October 2026

Analysis by the author 6

AnalysisD1-C433

The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.

Author Ibrahim Anjro

  • Extends D1-C432 Day 1: A Google panelist said they were working on a set of crawler best practices and offering research…
AnalysisD1-C470

Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.

Author Ibrahim Anjro

  • Extends D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…
AnalysisD1-C480

The consolidation described on stage matches Google's November 2020 move from Google Webmasters to Google Search Central, which moved its main blog, 13 localized blogs and the Search help content of the Search Console Help Center to its developer site: about six years before the event, not the four or five said on stage.

Author Ibrahim Anjro

  • Extends D1-C477 Day 1: In a migration about four or five years before the event, Google consolidated 12 or 14 of its blogs in…
AnalysisD1-C542

Google's redirects guide says Google Search follows JavaScript location redirects only after rendering, may never see one if rendering fails, and should be used only when server-side or meta refresh redirects are impossible; its site move guide asks for server-side permanent redirects (301 or 308) where technically possible. The JavaScript in Google's own migration was a fallback ('my only option'), not a pattern to copy.

Author Ibrahim Anjro

Used byrequirement DEV-CAN-01

  • Extends D1-C541 Day 1: A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in…
AnalysisD1-C547

Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the total time a server spends holding connections open for Google, counting both the number of parallel connections and their duration. That fits the panel's advice to watch how many connections Googlebot opens (D1-C494): fewer connections is the documented form of a lower capacity limit.

Author Ibrahim Anjro

Used byrequirement DEV-PRF-01

  • Extends D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
  1. Stage D1-C439 Day 1 · Q&A

    Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.

    extends
    Slide D1-C093 Day 1 · How Google thinks about crawl budget

    Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.

  2. Stage D1-C444 Day 1 · Q&A

    Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

    extends
    Stage D1-C397 Day 1 · Lightning session C: Crawling

    A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.

  3. Stage D1-C538 Day 1 · Q&A

    A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.

    extends
    Stage D1-C334 Day 1 · How crawling works

    Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.

  4. Stage D1-C546 Day 1 · Q&A

    A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).

    extends
    Stage D1-C374 Day 1 · How Google thinks about crawl budget

    Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.

  5. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Stage D1-C482 Day 1 · Q&A

    Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

  6. Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    extends
    Stage D1-C434 Day 1 · Q&A

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

  7. Slide D2-C197 Day 2 · Lightning session D: Rendering and JavaScript

    Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.

    extends
    Stage D1-C114 Day 1 · Q&A

    A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

  8. Stage D2-C271 Day 2 · What is Google friendly JavaScript

    Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.

    extends
    Stage D1-C114 Day 1 · Q&A

    A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

  9. Stage D2-C271 Day 2 · What is Google friendly JavaScript

    Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.

    extends
    Stage D1-C491 Day 1 · Q&A

    If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.

  10. Stage D2-C377 Day 2 · Handling web duplication

    AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

    extends
    Stage D1-C436 Day 1 · Q&A

    A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

  11. Stage D2-C511 Day 2 · What is Structured Data and why we need it on the internet.

    Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a chicken-and-egg problem that the schema.org usage statistics are meant to help break.

    extends
    Stage D1-C120 Day 1 · Q&A

    To get Google to build something, such as an addition to an API, report and request it publicly and in volume.

  12. Stage D2-C884 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.

    extends
    Stage D1-C540 Day 1 · Q&A

    Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

  13. Stage D3-C440 Day 3 · Inside Search Console: What’s New & How to Use It

    Google called AI reporting in Search Console an evolving space and expects the generative AI report to get richer, with more information in the future.

    extends
    Stage D1-C449 Day 1 · Q&A

    A Google panelist said he expected more AI reporting to launch in Search Console, without a timeline or a promise, because many people at Google care about giving site owners the information to understand and adapt to what happens in Search.

  14. Stage D3-C606 Day 3 · How long does it take to..?

    For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).

    extends
    Stage D1-C466 Day 1 · Q&A

    Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

  15. Slide D3-C617 Day 3 · How long does it take to..?

    Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2 weeks, and 1-3 weeks in recovery.

    extends
    Stage D1-C493 Day 1 · Q&A

    There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.

  16. Stage D3-C618 Day 3 · How long does it take to..?

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    extends
    Stage D1-C440 Day 1 · Q&A

    If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.

  17. Stage D3-C623 Day 3 · How long does it take to..?

    When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.

    extends
    Stage D1-C439 Day 1 · Q&A

    Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.

  18. Stage D1-C434 Day 1 · Q&A

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

    repeats
    Stage D1-C273 Day 1 · Lightning session A: Automation and AI

    A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

  19. Stage D1-C445 Day 1 · Q&A

    One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

    repeats
    Stage D1-C403 Day 1 · Lightning session C: Crawling

    A community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.

  20. Stage D1-C466 Day 1 · Q&A

    Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

    repeats
    Stage D1-C381 Day 1 · How Google thinks about crawl budget

    Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few thousand URLs likely has no crawl budget problem, and a larger site does not necessarily have one either.

  21. Stage D1-C522 Day 1 · How Google interprets robots.txt

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

    repeats
    Stage D1-C485 Day 1 · Q&A

    To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.

  22. Stage D1-C524 Day 1 · How Google interprets robots.txt

    Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.

    repeats
    Stage D1-C487 Day 1 · Q&A

    A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

  23. Stage D3-C624 Day 3 · How long does it take to..?

    Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.

    repeats
    Stage D1-C538 Day 1 · Q&A

    A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.