Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Developer kit

Differences ledger

20 places where what was shown or said at Deep Dive Europe 2026 differs from Google’s documentation, and what to follow. Each row puts the event’s version next to the documented one, with the claims behind both.

Differences
20
Event claims
30
Docs claims and pages
51
Linked pairs
6
Say they are uncertain
35

The rule

Where the event and Google's documentation differ, follow the documentation and say that they differ (for example canonicals on paginated pages, max-snippet:-1, 'Discovered – currently not indexed', SpamBrain's launch year, PageRank's role in ranking, Google's 2023 testing figures, the source year of the 40 billion spam pages figure).

Media kit wording rule 13. A requirement rests on what Google documents: read the row before you build on what was said on stage.

How to read a row

  • What to follow is the author’s guidance, written from the claims below it.
  • At the event: what was shown on a slide or said on stage. Google’s documentation: the docs claims and the Google pages.
  • The author’s analysis explains the gap. It is the author’s view, never Google’s.
  • Cite the claim IDs, not the row.
At a glance
RowDifferenceWhat to follow
DIF-01How many crawlers Google runsFollow Google's Inside Googlebot post: dozens of other clients share Googlebot's crawling infrastructure and only the larger crawlers are documented. A Google user agent missing from the public lists is not proof of a fake request; verify with reverse DNS or Google's published IP ranges.
DIF-024xx responses, 429 and crawl budgetFollow Google's HTTP status code documentation: 4xx codes other than 429 do not slow crawling, while 429 slows it like a 5xx error. A 404 is still a fetch. Return 429 or 503 only for temporary overload; Google's crawl rate guide warns against doing so for longer than 1-2 days.
DIF-03Whether a robots.txt block frees crawl budget for other pagesFollow Google's crawl budget guide: use robots.txt only for sections you never want crawled, never to reallocate crawl budget for a while. Freed budget goes to other pages only when Google is already at the site's crawl capacity limit, so do not expect a block to speed up crawling elsewhere on a site that is not.
DIF-04A misspelt rule name in robots.txtGoogle's documentation is silent here: its robots.txt specification does not mention typos, so spell rule names correctly. On stage a misspelt rule name was said to make Google ignore the line, but Google's open-source parser accepts common misspellings of disallow and user-agent (not of allow) and other crawlers may be stricter: rely on neither.
DIF-05Why a URL is 'Discovered – currently not indexed'Follow the Page indexing report help first: rule out server overload (slow responses, 5xx errors, a falling crawl rate in the Crawl Stats report). Then treat the status as a demand problem and raise the quality of the pages already indexed; resubmitting the URLs does not change why they wait.
DIF-06Links Google cannot extractFollow Google's link best practices: only an a element with an href is a dependable link. The documentation says Google may still try to parse routerLink, href on a span, onclick-only links and javascript: URLs, where the slide said it cannot; either way, do not rely on them for discovery.
DIF-07A noindex page and its JavaScriptFollow the JavaScript SEO basics guide, which says only that Google may skip rendering a page served with noindex. The practical rule is the same either way: never serve a noindex that JavaScript is expected to remove.
DIF-08Whether every page gets renderedFollow the JavaScript SEO basics guide: every page with a 200 status is queued for rendering unless a robots rule blocks indexing. Queued is not the same as rendered, so put the content, links and meta tags that indexing needs in the HTML the server sends.
DIF-09How strong a signal rel=canonical isFollow Google's canonical guide: redirects and rel=canonical are strong signals, sitemap inclusion is a weak one. Use the strong signals as the main levers, keep sitemaps as support, and point all of them at the same URL.
DIF-10Canonicals on paginated pagesFollow Google's pagination guide: give every page of a series its own URL, a self-referencing canonical and a crawlable link to the next page. Pointing the canonical of later pages at page 1 is a deliberate trade-off, not the default: it folds them into page 1, so the items they list need links from elsewhere.
DIF-11Whether only canonicals are shown in resultsFollow Google's documentation: a result usually points to the canonical, but another page of the same duplicate cluster can be shown in some contexts, such as a mobile page to a user on a mobile device.
DIF-12Videos placed below the foldFollow the Video indexing report help: the video needs a watch page and a player that is there when the page loads, at a size and position Google can determine, with no click-to-play placeholder. The documentation does not say that a video below the fold is never indexed; putting the main video in the first viewport is the safe choice.
DIF-13What max-snippet:-1 doesFollow the robots meta tag specification: max-snippet:-1 removes the length limit and lets Google choose the snippet length it finds most effective. Do not promise longer snippets from it; use it to lift a lower limit that a template, plug-in or CDN sets.
DIF-14PageRank's and MUM's role in rankingFollow Google's ranking systems guide: PageRank has evolved a lot and is still part of the core ranking systems, and MUM is not used for general ranking, only for specific applications. Quote the stage remark that PageRank is not used so much anymore only as a remark, never as PageRank being switched off, and read the slide's list as systems Google runs, not as systems that rank every query.
DIF-15SpamBrain's launch year and its '5 times more spam sites'Cite 2018 as SpamBrain's launch year, from Google's 2021 webspam report; the 2022 heard on stage is the year of the improvements in the 2022 report. Quote '5 times more spam sites' as that report does, 2022 against 2021 (and 200 times more than at launch), not as SpamBrain against Google's earlier spam algorithms.
DIF-16The year of the 40 billion spammy pages a dayCite 40 billion spammy pages a day as Google's published figure, from its webspam report for 2020 (published in April 2021) and its How Search Works page. The slide's 'In 2023' heading does not make it a 2023 measurement.
DIF-17Google's 2023 testing figuresCite the figures on Google's How Search Works page, with the year 2023: 719,326 search quality tests, 124,942 side-by-side experiments, 16,871 live traffic experiments and 4,781 launches. Do not use the rounded 800,000+ tests, which match no figure on that page, or the spoken 'close to 5,000' launches.
DIF-18The size of Gemini's context windowFollow the Gemini API's long-context documentation: Gemini models have context windows of 1 million or more tokens, so plan with about one million tokens as the documented floor, not the several million said on stage. Do not quote the 900,000 heard on stage, which had no unit.
DIF-19What the popularity rank attribute measuresFollow Merchant Center's definition: popularity_rank is the merchant's own 0-100 rank of a product against the rest of its inventory, based on its recent sales. It is not a sales figure across shops or marketplaces.
DIF-20New Merchant Center attributes in schema.org markupSend the new attributes in the Merchant Center feed, the documented route, and keep the documented Product markup next to the feed, which Google says maximises eligibility. As of 3 October 2026 Google documents no schema.org property for the new attributes except item_group_title (ProductGroup.name); add markup for the others only once Google documents it.

The differences 20

Event and docs differDIF-01

How many crawlers Google runs

At the event

  • StageNot in docsD1-C324

    Google runs probably hundreds, if not thousands, of crawlers on its crawler infrastructure; some of them are named and some are not.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

Google’s documentation

  • DocsSourceD1-C342

    Google's Inside Googlebot post (March 2026) says Googlebot is today just one user of a centralized crawling platform, and that dozens of other clients, such as Google Shopping and AdSense, send their crawl requests through the same infrastructure under other crawler names, with only the larger ones documented.

    Search Central blog (31 March 2026)

  • DocsSourceD1-C138

    Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.

    Google

Google pages

The author’s analysis · 1 claim
  • AnalysisD1-C344

    On stage Google spoke of probably hundreds, if not thousands, of crawlers, while its Inside Googlebot post speaks of dozens of other clients; both agree that only the larger crawlers are documented, so a Google user agent missing from the public lists is not proof of a fake request, and reverse DNS or Google's published IP ranges are the test.

    Ibrahim Anjro (author)

Event and docs differDIF-02

4xx responses, 429 and crawl budget

At the event

  • StageConsistent with docsD1-C385

    4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

Google’s documentation

  • DocsSourceD1-C072

    4xx status codes other than 429 have no effect on crawl rate.

    Google

  • DocsSourceD1-C071

    5xx and 429 responses prompt Google's crawlers to slow down temporarily. Already indexed URLs are preserved in the index for a while but eventually dropped.

    Google

  • DocsSourceD3-C668

    Google's crawl rate guide says that when a significant number of URLs return 500, 503 or 429, Google reduces the site's crawl rate, which starts increasing again automatically once the errors drop; it warns against doing this for longer than 1-2 days.

    Google

Google pages

The author’s analysis · 2 claims
  • AnalysisD1-C386

    The stage point that 4xx responses do not affect crawl budget matches Google's documentation that 4xx codes have no effect on crawl rate (D1-C072), with one exception: 429 Too Many Requests counts as a server error and slows crawling like a 5xx (D1-C071, D1-C092).

    Ibrahim Anjro (author)

  • AnalysisD1-C419

    A 404 fetch is still a fetch: Google's 2017 crawl budget post says generally any URL Googlebot crawls counts towards a site's crawl budget, so the stage point that 4xx responses do not affect crawl budget is best read as 'they do not slow crawling, and a 404 tells Google to crawl that URL less over time'.

    Ibrahim Anjro (author)

Event and docs differDIF-03

Whether a robots.txt block frees crawl budget for other pages

At the event

  • StageConsistent with docsD1-C469

    Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

Google’s documentation

  • DocsSourceD1-C109

    Google advises against using noindex to save crawl budget and against using robots.txt to temporarily reallocate budget. Use robots.txt only for pages you never want crawled, and 404 or 410 for removed pages.

    Google

Google pages

The author’s analysis · 1 claim
  • AnalysisD1-C470

    Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.

    Ibrahim Anjro (author)

Event and docs differDIF-04

A misspelt rule name in robots.txt

At the event

  • StageNot in docsD1-C525

    Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

Google’s documentation

Google pages

The author’s analysis · 1 claim
  • AnalysisD1-C532

    A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

    Ibrahim Anjro (author)

Event and docs differDIF-05

Why a URL is 'Discovered – currently not indexed'

At the event

  • StageNot in docsD2-C706

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C707

    Google's Page indexing report help says a 'Discovered – currently not indexed' page was found but not crawled yet, typically because Google wanted to crawl it but expected the crawl to overload the site, so it rescheduled the crawl.

    Google Search Console Help

Google pages

The author’s analysis · 2 claims
  • AnalysisD2-C708

    The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

    Ibrahim Anjro (author)

  • AnalysisD2-C710

    Repeatedly submitting 'Discovered – currently not indexed' URLs does not change why they wait, because the status reflects a crawl-scheduling decision; raise the site's demonstrated quality instead, for example by improving or removing weak pages that are already indexed.

    Ibrahim Anjro (author)

Event and docs differDIF-06

Links Google cannot extract

At the event

  • SlideConsistent with docsD2-C041

    Google cannot extract a link from an href attribute placed on an element other than a, such as a span, because that is not a standard way to make a link.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C042

    A Google slide also listed as not extractable an a element with a routerLink attribute instead of an href, and javascript: URLs such as javascript:goTo('products') or javascript:window.location.href='/products'.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C043

    Google's link best practices say Google can generally crawl a link only if it is an a element with an href attribute, and list routerLink without href, href on a span, onclick-only a elements and javascript: URLs as not recommended, while noting that Google may still attempt to parse them.

    Google Search Central

Google pages

The author’s analysis · 1 claim
  • AnalysisD2-C044

    Google's link-extraction slide put routerLink, href on a span, onclick-only links and javascript: URLs under 'can not extract', which is stricter than Google's link documentation saying Google may still try to parse them; either way they are not dependable links for discovery.

    Ibrahim Anjro (author)

Event and docs differDIF-07

A noindex page and its JavaScript

At the event

  • StageConsistent with docsD2-C852

    When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript, so a script cannot switch the page back to indexable, John Mueller said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C108

    Removing a robots restriction such as noindex with JavaScript does not work, a slide said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C854

    Google's JavaScript SEO basics guide says that when Google encounters a noindex rule it may skip rendering and JavaScript execution, so using JavaScript to change or remove a noindex robots meta tag may not work as expected.

    Google Search Central

Google pages

The author’s analysis · 1 claim
  • AnalysisD2-C855

    On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex), while Google's guide says only that it may skip rendering; either way, a noindex in the served HTML must never be one that JavaScript is expected to lift.

    Ibrahim Anjro (author)

Event and docs differDIF-08

Whether every page gets rendered

At the event

  • StageNot in docsD2-C130

    Google said its pipeline needs to ensure that content indexable without JavaScript can pass through without rendering, while content that appears only through JavaScript and CSS takes a longer rendering pass.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD3-C628

    Gary Illyes said Google keeps saying it renders every URL on the internet and that he would say this is not true, though it is what he was told; he went on to say that Google's logs show the rendering queue cleared within weeks.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

Google’s documentation

  • DocsSourceD2-C131

    Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.

    Google Search Central

Google pages

The author’s analysis · 2 claims
  • AnalysisD2-C132

    The stage remark that content indexable without JavaScript can pass through without rendering does not mean such pages skip rendering, because Google's guide queues every 200 page for rendering; read it as: content already in the raw HTML does not depend on the slower rendering pass.

    Ibrahim Anjro (author)

  • AnalysisD3-C676

    The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.

    Ibrahim Anjro (author)

Event and docs differDIF-09

How strong a signal rel=canonical is

At the event

  • SlideConsistent with docsD2-C390

    Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C391

    Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.

    Google Search Central

Google pages

The author’s analysis · 1 claim
  • AnalysisD2-C392

    Google's duplication talk described rel=canonical as something that 'also helps a bit', while Google's canonical guide calls it a strong signal alongside redirects and calls sitemap inclusion weak; treat redirects and rel=canonical as the main levers and sitemaps as support.

    Ibrahim Anjro (author)

Event and docs differDIF-10

Canonicals on paginated pages

At the event

  • StageNot in docsD2-C393

    Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD1-C115

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

    Google Search Central

Google pages

The author’s analysis · 1 claim
  • AnalysisD2-C394

    Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.

    Ibrahim Anjro (author)

Event and docs differDIF-11

Whether only canonicals are shown in results

At the event

  • StageConsistent with docsD2-C701

    When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C702

    Google's documentation says a search result usually points to the canonical page, but the other pages in a duplicate cluster are alternate versions that may be served in different contexts, for example a mobile page for a user on a mobile device.

    Google Search Central

Google pages

The author’s analysis · 1 claim
  • AnalysisD2-C703

    'Only canonicals end up in search results' as said on stage is a simplification: non-canonical duplicates are dropped from the index, but Google's documentation says an alternate from the same cluster can still be shown in some contexts, such as a mobile version to a mobile user.

    Ibrahim Anjro (author)

Event and docs differDIF-12

Videos placed below the fold

At the event

  • StageNot in docsD2-C924

    For a video to be discovered it must be embedded prominently, above the fold; Gary Illyes said a video placed below the fold is not going to be indexed.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C942

    Google's Video indexing report help says only videos on a watch page are eligible for indexing, and flags 'Cannot determine video position and size' when the player is not on the page at load, for example behind a click-to-play image, asking for the player to load at its real size and position without user interaction.

    Google Search Console Help

Google pages

The author’s analysis · 1 claim
  • AnalysisD2-C944

    Gary Illyes said a video below the fold is not indexed, but Google's video documentation only requires the player to be present at load, at a position and size Google can determine and not hidden behind other elements; to be safe, put the main video of a watch page in the first viewport and load the player without a click-to-play placeholder.

    Ibrahim Anjro (author)

Event and docs differDIF-13

What max-snippet:-1 does

At the event

  • SlideNot in docsD2-C085

    max-snippet:-1 removes the length limit and can produce a longer snippet than having no rule, because by default Google keeps snippets to a length it considers reasonable instead of quoting a page at length.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C104

    To be as visible as possible in Google, John Mueller's closing slide recommended the robots rules max-image-preview:large and max-snippet:-1.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C086

    Google's robots meta tag specification says Google chooses the snippet length when no max-snippet rule is set, and with max-snippet:-1 chooses the length it believes most effective; it does not say that -1 produces longer snippets.

    Google Search Central

Google pages

The author’s analysis · 1 claim
  • AnalysisD2-C105

    Set max-snippet:-1 and max-image-preview:large on every indexable template unless licensing requires otherwise, and check that no CMS, plug-in or CDN setting adds lower snippet or image preview limits by default.

    Ibrahim Anjro (author)

Event and docs differDIF-14

PageRank's and MUM's role in ranking

At the event

  • StageNot in docsD3-C136

    Google's quality talk called PageRank the speaker's favourite ranking system, elegant in its time, but said Google does not really use it so much anymore.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C133

    Google's slide said there is not one single ranking system and named spam detection systems, the reviews system, BERT, MUM, RankBrain, freshness systems, deduplication systems, crisis information systems and link analysis systems (PageRank).

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

Google’s documentation

  • DocsSourceD3-C138

    Google's ranking systems guide says PageRank, one of its core ranking systems when Google first launched, has evolved a lot since then and continues to be part of its core ranking systems.

    Google Search Central

  • DocsSourceD1-C130

    Google's ranking systems guide says MUM is not currently used for general ranking in Search, only for specific applications such as COVID-19 vaccine searches and featured snippet callouts.

    Google Search Central

Google pages

The author’s analysis · 2 claims
  • AnalysisD3-C137

    The remark that PageRank is not used so much anymore differs from Google's ranking systems guide, which says PageRank has evolved a lot and continues to be part of the core ranking systems; read it as PageRank weighing less among many signals, not as PageRank being switched off, so links still matter.

    Ibrahim Anjro (author)

  • AnalysisD3-C135

    The quality talk's slide listed MUM among Google's ranking systems, while Google's ranking systems guide says MUM is not currently used for general ranking in Search; read the slide as a list of systems Google runs, not as proof that each one ranks every query.

    Ibrahim Anjro (author)

Event and docs differDIF-15

SpamBrain's launch year and its '5 times more spam sites'

At the event

  • StageConfirmed by docsD2-C670

    SpamBrain is central to Google's spam-fighting efforts and has been improved many times since its launch.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C673

    SpamBrain was said to detect 5 times more spam sites than the spam algorithms Google had launched before it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

  • DocsSourceD2-C671

    Google's 2021 webspam report says SpamBrain, its AI-based spam-prevention system, was launched in 2018 and has been continuously improved since.

    Search Central blog (21 April 2022)

  • DocsSourceD2-C674

    Google's 2022 webspam report says SpamBrain detected 5 times more spam sites in 2022 than in 2021, and 200 times more than when it first launched.

    Search Central blog (11 April 2023)

Google pages

The author’s analysis · 2 claims
  • AnalysisD2-C672

    The SpamBrain launch year mentioned on stage, with hesitation, was 2022, which does not match Google's documented 2018; 2022 is the year of the improvements described in Google's 2022 webspam report, so cite 2018 as the launch year.

    Ibrahim Anjro (author)

  • AnalysisD2-C675

    The '5 times more spam sites' figure said on stage matches Google's 2022 webspam report, but the report compares 2022 with 2021 (and gives 200 times since launch), not SpamBrain with earlier algorithms; quote the documented comparison.

    Ibrahim Anjro (author)

Things
Event and docs differDIF-16

The year of the 40 billion spammy pages a day

At the event

  • StageConfirmed by docsD3-C197

    Google's quality talk said Google discovers tens of billions of spam pages every day.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConsistent with docsD3-C239

    Google's slide, headed 'In 2023, there were...', said 40 billion spammy pages are detected every day.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

Google’s documentation

  • DocsSourceD3-C240

    Google's webspam report for 2020, published in April 2021, says Google discovers 40 billion spammy pages every day.

    Search Central blog (29 April 2021)

Google pages

The author’s analysis · 2 claims
  • AnalysisD3-C241

    Cite 40 billion spammy pages a day as Google's published figure, from its webspam report for 2020 and its How Search Works page: the slide's 'In 2023' heading does not make it a 2023 measurement, because the 2021 and 2022 reports give no daily count.

    Ibrahim Anjro (author)

  • AnalysisD3-C200

    The spoken 'tens of billions of spam pages' a day matches Google's How Search Works page, which says its systems find 40 billion spammy pages every day, and the later Day 3 slide '40B spammy pages detected every day' (under 'In 2023'); cite 40 billion a day.

    Ibrahim Anjro (author)

Event and docs differDIF-17

Google's 2023 testing figures

At the event

  • SlideNot in docsD3-C236

    Google's slide said that in 2023 Google ran more than 800,000 search quality tests; the speaker added that a more recent figure might exist.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C237

    Google's slide said Google made more than 4,700 launches to Search in 2023; the speaker rounded this to close to 5,000.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C149

    Google's slide said that in 2023 Google ran 719,326 search quality tests, 124,942 side-by-side experiments and 16,871 live traffic experiments, and made 4,781 launches to Search.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

Google’s documentation

  • DocsSourceD3-C238

    Google's page on how it tests Search says that in 2023 it ran 719,326 search quality tests, 124,942 side-by-side experiments and 16,871 live traffic experiments, and made 4,781 launches.

    Google Search (How Search Works)

Google pages

The author’s analysis · 1 claim
  • AnalysisD3-C150

    Two Day 3 slides headed 'In 2023' gave Google's testing figures in two versions: exact in the quality talk (719,326 search quality tests, 4,781 launches) and rounded in the updates talk (800,000+ tests, 4,700+ launches). Google's How Search Works page states the exact 719,326 and 4,781, so cite those with the year 2023; 800,000+ matches no figure on that page, although the page's three test counts (719,326 quality tests, 124,942 side-by-side and 16,871 live traffic experiments) add up to 861,139.

    Ibrahim Anjro (author)

Event and docs differDIF-18

The size of Gemini's context window

At the event

  • StageConsistent with docsD2-C331

    Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C869

    Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Google’s documentation

Google pages

  • Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
The author’s analysis · 1 claim
  • AnalysisD2-C872

    The talk gave Gemini's context window both as millions of tokens and as roughly 900,000 to a million; Google's long-context docs say Gemini models have context windows of 1 million or more tokens (about eight average novels per million), so plan with about one million tokens as the documented floor rather than several million.

    Ibrahim Anjro (author)

Things
Event and docs differDIF-19

What the popularity rank attribute measures

At the event

  • StageConsistent with docsD3-C374

    Google added a popularity rank feed attribute because consumers want to know how well a product sells on a marketplace.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

Google’s documentation

  • DocsSourceD3-C375

    Merchant Center's popularity_rank is a 0-100 value the merchant assigns to rank a product's popularity, based on its recent sales, against the rest of its own inventory; it does not reflect user ratings.

    Google Merchant Center Help

Google pages

The author’s analysis · 1 claim
  • AnalysisD3-C376

    On stage the popularity rank was framed as how well a product sells on a marketplace, but Google's documentation defines it as the merchant's own 0-100 ranking against the rest of its inventory: a self-reported, relative value, not a sales figure across shops.

    Ibrahim Anjro (author)

Event and docs differDIF-20

New Merchant Center attributes in schema.org markup

At the event

  • StageNot in docsD3-C380

    Most of the new conversational feed attributes were already available in schema.org, so Google ties them back to structured data and can use them from product markup as well as from feeds.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

Google’s documentation

  • DocsSourceD3-C344

    Google's Product structured data introduction says that providing both on-page structured data and a Merchant Center feed maximizes eligibility for shopping experiences and helps Google understand and verify the data; product snippets may take pricing from the feed when the markup lacks it.

    Google Search Central

Google pages

The author’s analysis · 1 claim
  • AnalysisD3-C381

    As of 3 October 2026 Google's Merchant Center help lists no schema.org property for question_and_answer, document_link, related_product, popularity_rank, variant_option or product_detail (only item_group_title maps, to ProductGroup.name), and Search Central does not document such markup: reading these attributes from schema.org is a stage statement, the feed is the documented route.

    Ibrahim Anjro (author)

Claims linked as contradicting or updating 6

Set with a relation on the later claim. A pair that no row covers is usually two views at the event (Google and a community speaker, or two talks) or a later page of Google’s documentation, not a difference between the event and the documentation. DIF-10 covers D2-C393 contradicts D1-C115.

  1. Docs D1-C125 Day 1 · What's new in the world of Search

    Since 16 September 2026, US publishers and creators qualify for a Search profile with 10,000 followers across YouTube, Instagram, X or TikTok, and media organisations can claim and manage profiles for all their sub-brands from one login.

    updates
    Docs D1-C026 Day 1 · What's new in the world of Search

    Search profiles launched on 4 June 2026, in the US first, for creators and publishers with a sizable following on at least one major social or video platform. They appear in knowledge panels and Discover.

  2. Stage D1-C296 Day 1 · Lightning session A: Automation and AI

    A community speaker argued for adopting the GEO label as the industry's chance to leave behind the bad reputation SEO built, unlike Google's view earlier the same day that the new name is not needed.

    contradicts
    Stage D1-C049 Day 1 · What's new in the world of Search

    Gary Illyes argued that GEO is a label invented to create a new field and is not needed. Understanding how SEO works is enough.

  3. Stage D1-C296 Day 1 · Lightning session A: Automation and AI

    A community speaker argued for adopting the GEO label as the industry's chance to leave behind the bad reputation SEO built, unlike Google's view earlier the same day that the new name is not needed.

    contradicts
    Slide D1-C050 Day 1 · How Search works and where's AI?

    Google's answer to 'SEO is dead, long live GEO' is not to worry about the name: good SEO is good GEO and AEO.

  4. Stage D2-C325 Day 2 · Understanding what's on a page

    Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

    contradicts
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  5. Stage D2-C393 Day 2 · Handling web duplication

    Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

    contradicts
    Docs D1-C115 Day 1 · session not recorded

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

  6. Stage D2-C429 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker said pages carry different link equity, and a canonical leader that is not the strongest page in its group is technically valid but most likely suboptimal for ranking, so the strongest page should be the leader.

    contradicts
    Slide D2-C348 Day 2 · Handling web duplication

    Google forwards the signals attached to every URL in a duplicate cluster, such as links, to the representative URL, so that nothing is lost by showing only one URL.

Claims that say they are uncertain 35

Their text states the uncertainty: an unclear recording, an uncertain word, a best reading. Keep the hedge in your wording or leave the point out (media kit wording rule 10).

Day 1 11

StageNot in docsD1-C499

The keynote added that people who also use a standalone LLM keep certain types of use anchored on Google (the example given is unclear in both recordings).

Lino Cattaruzzi · Day 1 · Welcome and opening keynotes

StageConsistent with docsD1-C272

A community speaker advised against publishing markdown copies of HTML pages for AI agents: the copy is a duplicate (which the speaker also called a possible source of cloaking, an uncertain word in the recordings), and the models are trained to read HTML, CSS and JavaScript.

Carlos Ortega · Day 1 · Lightning session A: Automation and AI

StageConsistent with docsD1-C501

A community demo's script used both input modes of Google's Rich Results Test: the URL mode for public pages, which Google fetches itself, and the code mode, into which the script pasted the page's HTML, for private pages and pages Google cannot fetch (best reading of a largely unintelligible recording; the second kind of page was heard as 'dead', possibly 'dev').

Day 1 · Lightning session A: Automation and AI

StageConsistent with docsD1-C502

In a community demo, the structured-data problems Google's Rich Results Test reported on the test page included an empty name, a breadcrumb problem and a price of zero, with errors shown in pink and warnings in orange (best readings of a largely unintelligible recording; each item is heard in only one of the two recordings).

Day 1 · Lightning session A: Automation and AI

StageNot in docsD1-C526

The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

Day 1 · How Google interprets robots.txt

StageNot in docsD1-C407

A community speaker cited more than 900 million weekly active ChatGPT users and 2.5 billion monthly users of a Google AI feature (the recording is unclear which), adding that these are not comparable market-share figures.

Jovana Avramovic · Day 1 · Lightning session C: Crawling

StageNot in docsD1-C432

A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

Day 1 · Q&A

StageConsistent with docsD1-C539

A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

Day 1 · Q&A

StageNot in docsD1-C543

Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

Gary Illyes · Day 1 · Q&A

StageConsistent with docsD1-C546

A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example from 10 to 5, which the panelist said means Googlebot then makes up to five requests per second (the recording adds 'per connection', which is unclear).

Day 1 · Q&A

Day 2 13

StageNot in docsD2-C046

The speaker said Google sometimes also extracts URLs that are typed out as plain text on a page without being hyperlinked; the remarks around this point were unclear in the recording.

Day 2 · How is HTML interpreted

StageD2-C146

On a retail brand's page, the rendered version promised a bigger discount for a newsletter sign-up than the non-rendered version that was served, a mismatch that can hurt customer satisfaction; the two recordings disagree on the figure the rendered page promised.

Sören Bendig · Day 2 · Lightning session D: Rendering and JavaScript

StageConsistent with docsD2-C379

rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

Day 2 · Handling web duplication

StageNot in docsD2-C691

Under-served languages were presented as an opportunity: where Google's index holds a lot of spam in a language such as Basque, a site that starts publishing in that language can very likely replace that spam with its own content and rank for those keywords (one clause of the reasoning was inaudible).

Day 2 · Deciding what goes in the index?

Day 3 11

StageNot in docsD3-C026

Google suggested a test: a search that lists a term's synonyms joined with OR will probably return results very similar to the plain query, because Google adds the synonyms itself (part of the sentence is unclear in the recording).

John Mueller · Day 3 · Making sense of users' queries

StageNot in docsD3-C040

Google learns synonyms and siblings from search behaviour: words people search with in the same way become synonyms, while frequent comparison queries mark words as not interchangeable (the end of the sentence is unclear in the recording).

John Mueller · Day 3 · Making sense of users' queries

StageNot in docsD3-C564

Assistance also includes showing consumers what people in similar situations chose, because people are social and want that reassurance (the wording of this passage is partly uncertain in the recording).

Day 3 · Mastering the messy middle

StageNot in docsD3-C573

AI users sometimes take even longer purchase journeys, but they perceive their journeys as shorter, a perception the data does not always support (part of this passage is uncertain in the recording).

Day 3 · Mastering the messy middle

StageNot in docsD3-C585

Consumers boosted by AI in their purchase decisions were still a small group in October 2026, but the group is growing as more people start using AI (the word 'boosted' is an uncertain reading at this point of the recording).

Day 3 · Mastering the messy middle

StageD3-C607

An audience member reported recrawl intervals on one very large, popular client site: the homepage about 5 times a day, first-level pages about every 1.5 days (uncertain reading), and pages clustered as soft 404s every 160 to 190 days, adding that other sites will differ.

From the audience · Day 3 · How long does it take to..?

StageConsistent with docsD3-C644

Gary Illyes said a site move can take up to about a year in the worst case, because Google's slowest signal is recalculated only about once a year (some words of this passage are uncertain readings).

Gary Illyes · Day 3 · How long does it take to..?

AnalysisD3-C645

Keep migration redirects in place for at least a year and judge a site move after one to three months, not days: Google's speaker said its slowest signal needs about a year to be recalculated (a partly uncertain passage of the recording), and Google's site move guide says to keep redirects generally at least one year.

Ibrahim Anjro · Day 3 · How long does it take to..?

Take it with you

Looking after many sites? This kit can run as an agent on every release.