Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Developer kit

For developers

138 requirements for sites that Google can crawl, render, index and show well, drawn from Deep Dive Europe 2026 and checked against Google’s documentation. Each one says why it matters, how to build it, how to test it, and what it rests on.

Requirements
138
Must
37
Should
67
May
15
Avoid
19

How to read a requirement

  • Must Required: the site is not ready to launch without it.
  • Should Recommended: skip it only for a reason you can name.
  • May Optional: worth doing where it fits.
  • Avoid Must not: the site is not ready to launch while it does this.
  • Documented A cited Google page states or supports it.
  • Said at Search Central Live Said or shown at the event, not in Google’s documentation. Present it as said at Search Central Live, never as documented policy.

Use it in your tracker

Import the CSV into Jira, Linear, GitHub Issues or a spreadsheet: one ticket per requirement, with the test as the acceptance criterion and the level as the priority. Each requirement cites the claims it rests on: follow an ID to see exactly what was said and how it stands against Google’s documentation.

Requirements by area 138

Server, status codes and crawling 11

How the server answers Google's crawlers decides whether anything else on the site matters. These requirements cover reachability through firewalls and CDNs, status codes for outages and overload, robots.txt on every host, HTTP versions and caching headers.

MustDocumentedDEV-SRV-01

Let verified Google crawlers through firewalls, CDNs and bot protection

Why

Google does not know that a firewall or bot-protection rule is blocking it; it only sees failing requests. Timeouts, connection resets and DNS errors are treated like 5xx errors: crawling slows down at once, and indexed URLs that stay unreachable drop out of the index within days.

How

Allowlist Google's crawlers in WAF, CDN and bot-management rules by the IP ranges Google publishes for its common and special-case crawlers (common-crawlers.json, special-crawlers.json), or by reverse DNS to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com hosts; never by user-agent string alone (it is trivial to fake), and never by a whole googleusercontent.com host, where Google Cloud customers' machines resolve too. Do not rate-limit or challenge verified Googlebot with CAPTCHAs or JavaScript puzzles, which it cannot solve. Make sure DNS resolves reliably for every host that serves pages, scripts, styles or APIs. Review the rules after every CDN or security-vendor change, including rules written for HTTP/3 traffic or AI scrapers. When Search Console reports DNS or network errors, first ask the hosting provider, CDN or DNS provider what changed and check for recently added firewall rules: said at Search Central Live, most network errors happen between the origin and Google's data centres, where neither Google nor the site owner can see them.

Test

Search Console Crawl stats report: host status shows no robots.txt fetch, DNS resolution or server connectivity failures, and no Search Console alert for DNS errors, which can get a site removed from Google Search very aggressively and, with it, from every feature that depends on Search. URL Inspection live test succeeds for one URL per template. WAF/CDN logs show no 403, 429 or challenge responses to IPs that verify as Googlebot.

Code · Verify that a request really comes from Googlebot

Before a firewall, CDN or bot-protection rule blocks or challenges a "Googlebot" request, verify it: the user-agent string alone can be faked. In WAF rules, prefer matching the IP against the ranges Google publishes as common-crawlers.json and special-crawlers.json on its page on verifying its crawlers; for log analysis, a reverse DNS lookup followed by a forward lookup works too. Never allowlist a whole googleusercontent.com host: Google Cloud customer VMs resolve there too, and only *.gae.googleusercontent.com belongs to Google's user-triggered fetchers, which are not Googlebot.

Shell
# 1. Reverse DNS: Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com
#    or geo-crawl-*.geo.googlebot.com; special-case crawlers to rate-limited-proxy-*.google.com.
#    A host such as 81.59.117.34.bc.googleusercontent.com is a Google Cloud customer, not Googlebot.
host 66.249.66.1
# -> 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.

# 2. Forward DNS: that host name must resolve back to the same IP
host crawl-66-249-66-1.googlebot.com
# -> crawl-66-249-66-1.googlebot.com has address 66.249.66.1

The published IP ranges are easier to keep in sync with WAF and CDN allowlists than DNS lookups:

Shell
# Common crawlers (Googlebot and others) and special-case crawlers; refresh the lists regularly
curl -s https://developers.google.com/static/crawling/ipranges/common-crawlers.json
curl -s https://developers.google.com/static/crawling/ipranges/special-crawlers.json
Evidence · 11 claims · 4 Google pages
  • StageConsistent with docsD1-C069

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C126

    Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.

    Google

  • StageConsistent with docsD2-C374

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C375

    Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD1-C138

    Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.

    Google

  • StageConsistent with docsD1-C358

    DNS errors get a site removed from Google Search very aggressively, and Search Console alerts site owners to them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C359

    Most network errors happen somewhere between the site's origin server and Google's data centers, where they are invisible to both Google and the site owner; the way TCP networks work leaves no way to see what went wrong there.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C360

    When Search Console reports network errors, the first step Google recommends is to ask the hosting provider, or the CDN, whether they changed anything.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C361

    Network timeouts are usually caused close to the site, very often by a firewall, a CDN, the hosting provider or the DNS provider, and Google has no visibility into them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C362

    Network errors and timeouts, like DNS errors, can get a site removed from Google Search and, with it, from every feature that depends on Search.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C363

    For DNS and network errors caused by a firewall or CDN, Google advises checking whether new firewall rules were set recently and otherwise asking in the CDN's forum, as Google itself cannot see or help with these errors.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

AvoidDocumentedDEV-SRV-02

Never serve bot challenges or error pages to crawlers with a 200 status

Why

Google finds a bot-challenge page very hard to recognise as an error. Served with 200 on many URLs, the identical page can make Google cluster those URLs as duplicates and drop them, and recovery is slow because Google has little reason to recrawl duplicates. A 200 error page is also a soft 404.

How

If a crawler must get a challenge or block page, send it with 503 (temporary), as Google's CDN guidance recommends for bot-verification interstitials, or 429 for rate limiting; never 200. Most CDN and bot-management products let you set the status code of challenge responses. Better still, do not challenge verified Google crawlers at all (DEV-SRV-01). A CDN that throttles crawlers by injecting 429 or 503 responses on the network path shows up as a rise in HTTP errors in crawl reports.

Test

Request a few URLs with a Googlebot user agent from an unverified IP: any challenge comes back as 503 or 429, never 200. The Page indexing report shows no sudden rise of "Duplicate" or "Soft 404" on unrelated URLs.

Evidence · 7 claims · 2 Google pages
  • DocsSourceD2-C376

    Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

    Search Central blog (24 December 2024)

  • StageConsistent with docsD2-C374

    A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C375

    Google's systems find it very hard to recognise a bot-challenge page as an error page, and Googlebot cannot solve the puzzles, such as selecting all the cats, that those pages ask for.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD1-C078

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

    Ibrahim Anjro (author)

  • AnalysisD2-C378

    Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.

    Ibrahim Anjro (author)

  • StageConsistent with docsD1-C367

    CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C364

    A rise in HTTP errors in crawl reports can come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path between Google's crawler and the site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-SRV-03

Answer planned maintenance and short outages with 503, for a day or two at most

Why

5xx and 429 responses make Google slow down temporarily and keep already indexed URLs for a while, but URLs that keep failing are eventually dropped. A 200 "we'll be back soon" page becomes a soft 404 or replaces your content in the index, and a 404 removes the URLs. Google's crawl rate guide says crawling picks up again automatically once the errors drop. Said at Search Central Live (October 2026): Google lowers a site's crawl capacity within about four hours of a burst of 500 errors, while increases take one to three weeks, so a botched maintenance window can cost weeks of crawling.

How

During maintenance return 503 for every page (Retry-After is optional standard HTTP), keep robots.txt answering normally, and limit the 503 period to one or two days. For a longer pause keep the site online with reduced functionality, for example a disabled cart and an information banner, and update Product markup to the real availability. Use 503 or 429 only for genuine outages or overload, never as a permanent throttle.

Test

In a maintenance drill, curl -I on any page returns 503 while curl -I on /robots.txt returns 200. Afterwards, the 5xx share in Crawl stats falls back to zero.

Code · Planned maintenance with 503 (robots.txt stays available)

During a short outage every page answers 503 Service Unavailable, never a 200 "we'll be back" page and never a 404. Keep it to a day or two, and keep serving robots.txt normally: a 5xx on robots.txt stops crawling of the whole site.

HTTP
HTTP/1.1 503 Service Unavailable
Content-Type: text/html; charset=utf-8
Retry-After: 3600
Cache-Control: no-store
nginx
server {
    listen 443 ssl;
    server_name www.example.com;
    ssl_certificate     /etc/ssl/www.example.com.crt;
    ssl_certificate_key /etc/ssl/www.example.com.key;

    error_page 503 /maintenance.html;

    location = /maintenance.html {
        root /var/www/static;
        internal;
        # add_header here replaces server-level add_header lines: repeat HSTS or CSP if you set them
        add_header Retry-After 3600 always;
        add_header Cache-Control "no-store" always;
    }

    # robots.txt keeps answering 200 during maintenance
    location = /robots.txt {
        root /var/www/static;
    }

    location / {
        return 503;
    }
}
Evidence · 11 claims · 4 Google pages
  • DocsSourceD1-C071

    5xx and 429 responses prompt Google's crawlers to slow down temporarily. Already indexed URLs are preserved in the index for a while but eventually dropped.

    Google

  • DocsSourceD1-C126

    Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.

    Google

  • DocsSourceD2-C012

    Google's guide to temporarily pausing an online business recommends that a shop expecting to sell again within weeks or months stays online with limited functionality, such as a disabled cart, and updates its Product structured data to show current availability.

    Google Search Central

  • AnalysisD1-C078

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

    Ibrahim Anjro (author)

  • AnalysisD2-C342

    Return real error status codes for error states, 404 or 410 for missing content and 503 for outages such as a failed database connection, also in single-page apps; a 200 page whose main content is only an error message is treated as a soft 404 even when header and navigation look normal.

    Ibrahim Anjro (author)

  • StageConsistent with docsD3-C618

    When a site starts serving 500 errors, Google lowers the crawl capacity allocated to the site within about four hours on average.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C619

    Increases in crawl capacity take longer than decreases, within one to three weeks, because Google first needs to know that the higher demand will last.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C668

    Google's crawl rate guide says that when a significant number of URLs return 500, 503 or 429, Google reduces the site's crawl rate, which starts increasing again automatically once the errors drop; it warns against doing this for longer than 1-2 days.

    Google

  • AnalysisD3-C621

    Spoken and slide figures for capacity increases differ (one to three weeks, up to a month, versus 1-2 weeks typical and 1-3 weeks in recovery), but the lesson is the same: a burst of 5xx errors cuts crawling within hours and recovery takes weeks, so keep servers stable before launches and migrations.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD1-C353

    429 Too Many Requests is the exception among 4xx codes: instead of meaning there is nothing at the URL, it tells Google's crawler to slow down, and Google slows its crawling.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C354

    Google slows crawling when a site returns 5xx errors, because a 5xx usually means the server, and often the whole site, cannot serve requests, and Google does not want to break the site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-SRV-04

Keep robots.txt reachable on every host: 200 with rules, or 404 if there are none

Why

If robots.txt answers with a 5xx error, Google stops crawling the site for the first 12 hours and then relies on a cached copy for up to 30 days. robots.txt rules apply per host, so the file on www.example.com does nothing for api.example.com or a CDN host name, and the file is read only at the root of a host: a robots.txt in a subdirectory is never used.

How

Serve the file at the root path /robots.txt on every host name that serves pages or rendering resources: www, the bare domain if it answers, API, static and CDN hosts, every country subdomain. Answer 200 with the rules, or 404 if the host has no rules (Google then crawls everything). Exclude robots.txt from maintenance modes, WAF challenges and authentication, and never let an API or CDN host fall back to a blanket Disallow or an error page.

Test

curl -I https://<each-host>/robots.txt returns 200 or 404, never 5xx, 401/403 or a redirect to a login page. Search Console's robots.txt report (Settings) shows each host's file as fetched.

Code · robots.txt with crawl controls and a rendering carve-out

Disallow only URLs that should never be crawled (internal search, cart and checkout actions, filter parameters), keep every script, style and API path that pages need for rendering crawlable, and list the sitemap. Google reads only user-agent, allow, disallow and sitemap; each host (www, api, cdn) needs its own file at its root. Robots.txt is public, so never list secret paths in it.

Text
# https://www.example.com/robots.txt
User-agent: *
# Internal search results and cart/checkout actions (adapt to your URL patterns)
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
# (Shops in Google Shopping: give Storebot-Google its own group that leaves cart and checkout open.
#  A named group replaces this * group for that crawler, so repeat in it every rule that should still apply.)
# Filter and sort parameters of faceted navigation, as the first parameter (?) or a later one (&);
# /*?*size= would also block ?pagesize= and /*?*color= would block ?bgcolor=
Disallow: /*?color=
Disallow: /*&color=
Disallow: /*?size=
Disallow: /*&size=
Disallow: /*?sort=
Disallow: /*&sort=
# API: blocked in general, but the endpoints pages render from stay crawlable
# (the longer, more specific Allow rule wins)
Disallow: /api/
Allow: /api/products/
Allow: /api/reviews/
# Never disallow /static/, /assets/ or other JS and CSS folders

# Optional: keep content out of Gemini training and grounding (no effect on Google Search)
User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
Evidence · 4 claims · 2 Google pages
  • DocsSourceD1-C074

    If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.

    Google

  • SlideConfirmed by docsD2-C209

    robots.txt rules apply per host, so the rules in example.com/robots.txt do not apply at all to an API served from api.example.com.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C210

    An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD1-C515

    The robots.txt file always sits in the root of the host; it cannot be placed anywhere else.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-SRV-05

Rely only on the fields Google supports (user-agent, allow, disallow, sitemap) and keep robots.txt under 500 KiB

Why

Robots.txt is an IETF standard, RFC 9309, that basically all major crawlers follow, and it was kept simple on purpose. Google supports only user-agent, allow, disallow and sitemap in robots.txt; it ignores other lines, such as crawl-delay or Content-Signal, so they save no crawl budget with Google but can stay for other crawlers as long as they never sit between user-agent lines (DEV-SRV-06). Google reads only the first 500 KiB of the file and generally caches it for up to 24 hours, so a change can take a day to apply. Google refreshes its cached copy every 24 hours (its service level objective, as RFC 9309 asks), and the Request a recrawl option in Search Console's robots.txt report refreshes it sooner.

How

Everything is allowed by default, so use allow only to re-open a path inside a disallowed one. Cover URL patterns with the * and $ wildcards instead of listing URLs: * matches any run of characters (Disallow: /*/live/ blocks /science/live/ and /sports/live/, and Allow: /science/live/ re-opens one of them), and $ ends the match (Disallow: / with Allow: /$ leaves only the home page open to that group). The longest matching rule wins and, when rules conflict, Google uses the least restrictive one. A crawler follows only the most specific group that names it, so give each named crawler a complete group (DEV-SRV-10). Start comments with # and spell field names exactly: said at Search Central Live, robots.txt is forgiving, Google skips a line it cannot parse and still uses the rest of the file, so a misspelt rule fails silently (Google's open-source parser happens to accept some misspellings of disallow and user-agent, but not of allow, and other crawlers may be stricter). Add the Sitemap line with an absolute URL. To slow Google down in an emergency, return 503 or 429 briefly instead of inventing directives. Plan robots.txt changes a day ahead; for an urgent change, request a recrawl in the robots.txt report. Listing the sitemap in robots.txt is fine, but it makes the sitemap public; submit it only in Search Console if that matters.

Test

Check the file's fetch status, warnings and errors in Search Console's robots.txt report and test it with an RFC 9309 parser such as Google's open-source robots.txt library (github.com/google/robotstxt); check the size is well under 500 KiB; spot-check important URLs (homepage with tracking parameters, JS bundles, API endpoints, filter URLs) against the rules.

Code · robots.txt with crawl controls and a rendering carve-out

Disallow only URLs that should never be crawled (internal search, cart and checkout actions, filter parameters), keep every script, style and API path that pages need for rendering crawlable, and list the sitemap. Google reads only user-agent, allow, disallow and sitemap; each host (www, api, cdn) needs its own file at its root. Robots.txt is public, so never list secret paths in it.

Text
# https://www.example.com/robots.txt
User-agent: *
# Internal search results and cart/checkout actions (adapt to your URL patterns)
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
# (Shops in Google Shopping: give Storebot-Google its own group that leaves cart and checkout open.
#  A named group replaces this * group for that crawler, so repeat in it every rule that should still apply.)
# Filter and sort parameters of faceted navigation, as the first parameter (?) or a later one (&);
# /*?*size= would also block ?pagesize= and /*?*color= would block ?bgcolor=
Disallow: /*?color=
Disallow: /*&color=
Disallow: /*?size=
Disallow: /*&size=
Disallow: /*?sort=
Disallow: /*&sort=
# API: blocked in general, but the endpoints pages render from stay crawlable
# (the longer, more specific Allow rule wins)
Disallow: /api/
Allow: /api/products/
Allow: /api/reviews/
# Never disallow /static/, /assets/ or other JS and CSS folders

# Optional: keep content out of Gemini training and grounding (no effect on Google Search)
User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
Evidence · 20 claims · 4 Google pages
  • DocsSourceD1-C084

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

    Google

  • DocsSourceD1-C085

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

    Google

  • DocsSourceD1-C080

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

    Google

  • DocsSourceD1-C081

    When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

    Google

  • DocsSourceD1-C082

    In robots.txt, * matches zero or more of any character and $ marks the end of the URL.

    Google

  • SlideConfirmed by docsD1-C108

    The non-standard crawl-delay rule is not processed by Googlebot, so it does not save any crawl budget.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C140

    Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

    Google Search Console Help

  • SlideConfirmed by docsD3-C613

    Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C614

    Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C615

    A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C667

    Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.

    Google Search Central, Google Search Console Help

  • StageConfirmed by docsD1-C512

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C516

    Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C517

    Everything on a site is implicitly allowed, so an allow rule is only needed to re-open a specific path inside a disallowed one.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C518

    Comments in robots.txt start with #, and a line without the # that a parser cannot read is ignored anyway, because the standard requires parsers to skip lines they cannot parse, so it acts like a comment.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C520

    In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C521

    To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C525

    Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C532

    A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD1-C528

    Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

AvoidDocumentedDEV-SRV-06

Do not put Content-Signal or other non-standard lines between user-agent lines

Why

Googlebot ignores an unknown line placed between two user-agent lines and merges both user agents into one group, so the rules that follow apply to both. A Google slide at Search Central Live showed Bing closing the group at that line instead, so the same file blocks one crawler and not the other. Google's robots.txt talk called this the real trouble spot of the format: while RFC 9309 was being written, opinions on whether the first crawler should inherit the rules that follow split roughly 50-50. Managed robots.txt files from some CDNs add such lines, and said at Search Central Live, CDNs sometimes update a site's robots.txt without the owner's knowledge.

How

Keep each group as an uninterrupted run of user-agent lines followed by its allow and disallow rules. If a CDN or plug-in adds Content-Signal or other non-standard lines, make sure they sit after a group's rules, never between user-agent lines, and review the generated file after every CDN setting change. Check the version history in Search Console's robots.txt report for changes nobody on the team made.

Test

Fetch the live file (curl https://www.example.com/robots.txt) and check that every run of user-agent lines is uninterrupted; test it in Bing Webmaster Tools' robots.txt tester and with Google's open-source robots.txt parser (github.com/google/robotstxt). The robots.txt report's latest fetched version matches the file the team deployed.

Evidence · 6 claims · 2 Google pages
  • SlideConfirmed by docsD1-C087

    Googlebot ignores an unknown line, such as a Content-signal line, placed between two user-agent lines and joins both user agents into one group, so the rules that follow apply to both; in the slide's example both are blocked.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideNot in docsD1-C123

    Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C088

    Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files. Never place any non-standard directive between user-agent lines; put it after a group's rules and test the file with each search engine's tools, such as Bing's robots.txt tester and Google's open-source robots.txt parser.

    Ibrahim Anjro (author)

  • StageNot in docsD1-C526

    The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C529

    Google said the robots.txt report helps catch CDNs that update a site's robots.txt without the owner's knowledge, and hosts that cloak the robots.txt file, which happens more often than people think.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C528

    Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-SRV-07

Answer crawlers over HTTP/1.1 (HTTP/2 optional); never run a crawlable host on HTTP/3 only

Why

Google's crawlers support HTTP/1.1, their default, and HTTP/2, and use whichever crawls best; Googlebot does not support HTTP/3 today. HTTP/2 can save server resources but gives no ranking benefit, and a site can opt out of HTTP/2 crawling by answering such requests with 421.

How

Make every crawlable host answer over HTTP/1.1. HTTP/2 may save server resources and is worth enabling on the origin and the CDN; to opt out of HTTP/2 crawling, return 421 to Google's HTTP/2 requests. Treat HTTP/3 as an extra for browsers. Check that CDN, firewall and load-balancer rules written for HTTP/3 (QUIC) traffic do not break HTTP/1.1 or HTTP/2 responses. Support gzip or Brotli compression, which Google's crawlers accept.

Test

curl -I --http1.1 https://www.example.com/ returns 200, and curl -I --http2 https://www.example.com/ returns the same status, canonical, robots and caching headers (Date and similar headers may differ), unless the site deliberately answers 421 to opt out of HTTP/2 crawling; Crawl stats shows normal response times.

Evidence · 4 claims · 1 Google page
  • DocsSourceD1-C076

    Google's crawlers support HTTP/1.1 and HTTP/2, use whichever gives the best crawling performance, and may switch between sessions. HTTP/2 can save server resources but gives no ranking benefit.

    Google

  • StageConsistent with docsD1-C075

    Googlebot does not support HTTP/3 today.

    a speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C077

    This does not mean avoiding HTTP/3, which benefits browsers. The rule is never to run a host that only answers over HTTP/3, and to check that CDN or firewall rules written for HTTP/3 traffic do not break HTTP/1.1 and HTTP/2.

    Ibrahim Anjro (author)

  • DocsSourceD1-C139

    Google's crawler overview says HTTP/1.1 is the default protocol of Google's crawlers and that a site can opt out of crawling over HTTP/2 by answering Google's HTTP/2 requests with a 421 status code.

    Google

ShouldDocumentedDEV-SRV-08

Support conditional requests with ETag or Last-Modified and answer 304 when nothing changed

Why

A 304 Not Modified response tells Google to reuse its copy, so unchanged pages cost almost nothing to recrawl. Google lists HTTP cache control among the ways to manage crawl budget, and its crawling team has said it prefers ETag.

How

Send a strong ETag (a hash of the content) or Last-Modified on HTML and on the JS, CSS and API responses used for rendering, and answer If-None-Match or If-Modified-Since with 304 and no body when nothing changed. Derive the ETag from the content only, not from timestamps, nonces or session data, and keep it identical across CDN nodes.

Test

curl -sI https://www.example.com/page shows an ETag; repeating the request with -H 'If-None-Match: "<etag>"' returns 304. The Crawl stats response breakdown shows 304 responses.

Code · ETag and 304 Not Modified for crawlers

Send a validator (ETag, which Google's crawling team prefers, or Last-Modified) and answer a matching conditional request with 304 and no body, so unchanged pages cost almost nothing to recrawl. Express does this for you when it sends the body, but its ETag hashes the whole response: if the HTML embeds a CSP nonce, a CSRF token or a timestamp, compute the ETag from the content data instead (for example a hash of the product record and the template version) and set it with res.set('ETag', ...).

HTTP
GET /chairs/oak-dining-chair HTTP/1.1
Host: www.example.com
If-None-Match: "a3f9c2e1"

HTTP/1.1 304 Not Modified
ETag: "a3f9c2e1"
Cache-Control: no-cache
JavaScript
// Express: strong ETags, and an automatic 304 when If-None-Match matches
const express = require('express');
const app = express();
app.set('etag', 'strong');

app.get('/chairs/:slug', async (req, res) => {
  const html = await renderChairPage(req.params.slug); // your server-side renderer
  res.set('Cache-Control', 'no-cache'); // may be stored, but must be revalidated
  res.type('html').send(html); // res.send() compares the ETag and answers 304 when fresh
});

// Pages with a per-request nonce or CSRF token: an ETag from the data, not from the body
const crypto = require('crypto');
const TEMPLATE_VERSION = '2026-10-02';
app.get('/products/:slug', async (req, res) => {
  const product = await loadProduct(req.params.slug); // your data access
  const etag = '"' + crypto.createHash('sha256').update(JSON.stringify(product) + TEMPLATE_VERSION).digest('hex').slice(0, 16) + '"';
  res.set('ETag', etag);
  res.set('Cache-Control', 'no-cache');
  if (req.fresh) return res.status(304).end(); // If-None-Match matches the data ETag
  res.type('html').send(renderProductPage(product, res.locals.cspNonce));
});
Evidence · 5 claims · 2 Google pages
  • DocsSourceD1-C104

    HTTP caching for crawlers means supporting conditional requests (ETag with If-None-Match, or Last-Modified with If-Modified-Since) and answering 304 Not Modified when nothing changed. Google's crawling team has said it prefers ETag.

    Google, Search Central blog (9 December 2024)

  • SlideConsistent with docsD1-C103

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C321

    Gary Illyes pointed to the Last-Modified and ETag headers as the fields of an HTTP response that matter for caching.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C350

    A 304 Not Modified response tells Google's crawler that the content has not changed since its last visit.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C384

    Every fetch that returns a 2xx success response consumes crawl budget.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

AvoidDocumentedDEV-SRV-09

Do not answer crawlers with 401, 402 or 403 on pages that should be in Search

Why

Google does not index URLs that return a 4xx status (except 429) and removes already indexed ones. Said at Search Central Live by Gary Illyes: Google has recently seen more 403 responses to its crawlers (described on stage as "authentication required"), treats them as client errors like 404 or 410 and drops those pages from Search and its AI features; more recently still it has seen an uptick in 402 Payment Required responses, which it treats as a 404 because Googlebot cannot pay. A 403 tells the crawler it may not see the content, so there is nothing to index.

How

Check that login walls, paywalls, geo-blocks and bot-management rules never return 401, 402 or 403 to verified Google crawlers on public pages. For paywalled articles serve the content with the paywalled-content markup instead of an error (DEV-SDA-13). Keep 401 and 403 for pages that must stay out of Search, such as account areas.

Test

Crawl stats and server logs show no 401, 402 or 403 responses to verified Googlebot on indexable templates, and URL Inspection's live test returns 200 for one URL per template.

Evidence · 5 claims · 3 Google pages
  • StageConsistent with docsD1-C365

    Google has recently seen more 403 responses to its crawlers (described on stage as 'authentication required'); it treats them as client errors, technically equivalent to 404 or 410, and drops those pages from Search and its AI features.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C366

    Google treats a 402 Payment Required response as a 404 Not Found, because Googlebot cannot pay for content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C352

    A 403 response tells Google's crawler it has no permission to see the requested content, so there is nothing for Google to index.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C351

    404 Not Found and 410 Gone both tell Google there is nothing at the URL, so the URL is not indexable.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C509

    Google has seen an uptick in 402 Payment Required responses to its crawlers, an even more recent trend than the rise in 403 responses.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-SRV-10

Give a crawler its own robots.txt group only when needed, and make that group complete

Why

Robots.txt groups are not additive. A crawler obeys only the most specific group that names it and ignores the * group, so a Googlebot group that blocks only /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks (Dave Smart's example at Search Central Live). One group can name several crawlers, and Google combines all groups that name the same user agent into one. In the robots.txt quiz of Google's talk, a Googlebot group with only an allow rule, or an allow rule in the * group, was the wrong answer for opening /staging/preview/ to Googlebot while keeping the rest of /staging/ closed.

How

Put the rules every crawler shares in the * group. Add a named group (Googlebot, Storebot-Google, an AI crawler token) only when that crawler needs different rules, and copy into it every rule from the * group that should still apply to it. To give several crawlers the same rules, list their user-agent lines on top of one group. Generate the file from one template so that a change to the shared rules updates every named group too.

Test

For each named group, run Google's open-source robots.txt parser (github.com/google/robotstxt) with that crawler's token against a list of sample URLs covering every path the * group disallows: each result is the one intended for that crawler. For the quiz case, /staging/preview/ is allowed and /staging/other/ disallowed for Googlebot.

Code · Complete robots.txt groups for named crawlers

A crawler obeys only the most specific group that names it and ignores the * group, so groups are not additive: a named group must repeat every shared rule that should still apply to that crawler. One group can name several crawlers, and Google combines all groups that name the same user agent into one.

Text
# WRONG: the googlebot group replaces the * group instead of adding to it,
# so Googlebot may crawl /goats/ and /cows/ and only /dogs/ is closed to it
User-agent: *
Disallow: /goats/
Disallow: /cows/

User-agent: googlebot
Disallow: /dogs/

Right: the shared rules sit in the * group, and the named group repeats them before adding its own. To open /staging/preview/ to Googlebot while the rest of /staging/ stays closed to it, the Googlebot group needs both the disallow and the allow; an allow in the * group, or a Googlebot group with only the allow, does not do it.

Text
# Shared rules for every crawler without a group of its own
User-agent: *
Disallow: /goats/
Disallow: /cows/
Disallow: /staging/

# Googlebot: the shared rules again, plus its own
User-agent: googlebot
Disallow: /goats/
Disallow: /cows/
Disallow: /dogs/
Disallow: /staging/
Allow: /staging/preview/

# Several crawlers can share one group: stack their user-agent lines with nothing in between
User-agent: bingbot
User-agent: applebot
Disallow: /goats/
Disallow: /cows/
Disallow: /staging/

Sitemap: https://www.example.com/sitemap.xml
Evidence · 5 claims · 2 Google pages
  • StageConfirmed by docsD1-C533

    Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C531

    In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C527

    The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C519

    A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C080

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

    Google

AvoidDocumentedDEV-SRV-11

Never use robots.txt to hide private paths; protect them with authentication

Why

Robots.txt is not a security measure: it sits at a predictable public address, so anyone can read the paths it disallows, and it only asks well-behaved crawlers not to fetch them. Google's guide says it is not a mechanism for keeping a page out of Google: a disallowed URL can still be indexed if other sites link to it, and Google points to password protection for private files.

How

Put admin areas, internal tools, exports and staging behind authentication or an IP allowlist (DEV-CAN-07), or keep them off the internet; do not list their paths in robots.txt. To keep a public page out of Search, use noindex and leave it crawlable (DEV-IDX-01).

Test

The live robots.txt names no admin, backup, export or internal paths, and curl -I on each private path returns 401 or 403 without credentials.

Evidence · 1 claim · 1 Google page
  • StageConsistent with docsD1-C514

    Robots.txt is not a security measure: the file sits at a predictable public address, so anyone can read the paths it disallows. Protect a secret folder with authentication, or do not put it on the internet.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

Google discovers pages through the links it extracts from HTML and sends them back to the crawl queue; links also tell it how a site is structured and feed ranking. Every page that should rank needs a real URL and a real link to it.

MustDocumentedDEV-URL-01

Make every navigational link an <a> element whose href holds a real URL

Why

Google extracts links from a elements with an href holding an absolute or relative URL, once from the server HTML and again from the rendered page, and uses them to discover pages, understand the site's structure and rank. Googlebot does not click, so a link that only works through JavaScript leaves its target without a link Google can reliably extract.

How

Audit every component that outputs links: main and sub-navigation, breadcrumbs, pagination, filters, product tiles, footers and language selectors. Each needs <a href="/real/path">, and framework link components must render that href. Client-side routing may intercept the click as long as the href stays (DEV-URL-03). Do not rely on plain-text URLs or on sitemaps alone for discovery.

Test

In URL Inspection (rendered HTML) and in a crawler with JavaScript off and on, every navigation item appears as <a href="...">; a crawl from the homepage reaches all key templates through links; search the rendered HTML for onclick=, javascript:, routerLink and href="#".

Code · Crawlable links versus links Google cannot rely on

Google extracts links from <a> elements with an href that holds a real URL, both from the server HTML and from the rendered page. Click handlers, javascript: URLs, routerLink without href, href on other elements and # routes are not dependable.

HTML
<!-- Crawlable: <a> with a real relative or absolute URL -->
<a href="/chairs/">Chairs</a>
<a href="https://www.example.com/chairs/oak-dining-chair">Oak dining chair</a>

<!-- Crawlable and still client-side routed: keep the href, intercept the click in JavaScript -->
<a href="/chairs/" data-link>Chairs</a>

<!-- Not dependable: do not use for anything that should be discovered -->
<a onclick="goTo('/chairs/')">Chairs</a>
<a routerLink="/chairs/">Chairs</a>
<a href="javascript:goTo('chairs')">Chairs</a>
<a href="#/chairs">Chairs</a>
<span data-href="/chairs/" onclick="location.href=this.dataset.href">Chairs</span>
<button type="button" onclick="location.href='/chairs/'">Chairs</button>
Evidence · 6 claims · 2 Google pages
  • SlideConfirmed by docsD2-C039

    Google can extract links written as an a element with an href attribute that holds an absolute or a relative URL, which the speaker called the good old normal way.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C043

    Google's link best practices say Google can generally crawl a link only if it is an a element with an href attribute, and list routerLink without href, href on a span, onclick-only a elements and javascript: URLs as not recommended, while noting that Google may still attempt to parse them.

    Google Search Central

  • StageConsistent with docsD2-C038

    Google uses the links it extracts for three purposes: discovering new pages, determining a site's structure, and ranking.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C169

    Google's JavaScript SEO basics guide says Googlebot extracts links twice, from the HTML response before rendering and again from the rendered HTML, so links injected with JavaScript can be found if they use crawlable <a href> markup.

    Google Search Central

  • StageConfirmed by docsD2-C287

    A link that is an <a> element but does not point to a real URL gives Google something to look at, but Google will not know where the link goes.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C036

    Links and anchors are among the things Google extracts from a page's HTML, and the slide card for them simply read 'We like links.'

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

AvoidDocumentedDEV-URL-02

Do not navigate with onclick handlers, javascript: URLs, routerLink without href or # pseudo-links

Why

A Google slide put onclick-only links, javascript: URLs, routerLink without href and href on non-a elements under "cannot extract"; Google's link documentation calls them not recommended and says Google may still try to parse them. Either way they are not dependable for discovery, and they are common in single-page apps and faceted navigation.

How

Replace each pattern with <a href="/path">. Keep buttons for actions that do not change the URL (add to cart, open a dialog). Use URL fragments only where you deliberately do not want Google to follow, such as filter combinations (DEV-URL-08).

Test

Search templates and the rendered HTML (DevTools Elements panel, URL Inspection) for onclick navigation, javascript:, routerLink without href, href on span or div, and href="#/"; a JavaScript-rendering crawler finds no orphaned templates.

Code · Crawlable links versus links Google cannot rely on

Google extracts links from <a> elements with an href that holds a real URL, both from the server HTML and from the rendered page. Click handlers, javascript: URLs, routerLink without href, href on other elements and # routes are not dependable.

HTML
<!-- Crawlable: <a> with a real relative or absolute URL -->
<a href="/chairs/">Chairs</a>
<a href="https://www.example.com/chairs/oak-dining-chair">Oak dining chair</a>

<!-- Crawlable and still client-side routed: keep the href, intercept the click in JavaScript -->
<a href="/chairs/" data-link>Chairs</a>

<!-- Not dependable: do not use for anything that should be discovered -->
<a onclick="goTo('/chairs/')">Chairs</a>
<a routerLink="/chairs/">Chairs</a>
<a href="javascript:goTo('chairs')">Chairs</a>
<a href="#/chairs">Chairs</a>
<span data-href="/chairs/" onclick="location.href=this.dataset.href">Chairs</span>
<button type="button" onclick="location.href='/chairs/'">Chairs</button>
Evidence · 6 claims · 3 Google pages
  • SlideConsistent with docsD2-C040

    Google cannot extract a link from an a element that only has an onclick handler, because Googlebot does not click, so the JavaScript is never triggered.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C041

    Google cannot extract a link from an href attribute placed on an element other than a, such as a span, because that is not a standard way to make a link.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C042

    A Google slide also listed as not extractable an a element with a routerLink attribute instead of an href, and javascript: URLs such as javascript:goTo('products') or javascript:window.location.href='/products'.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C043

    Google's link best practices say Google can generally crawl a link only if it is an a element with an href attribute, and list routerLink without href, href on a span, onclick-only a elements and javascript: URLs as not recommended, while noting that Google may still attempt to parse them.

    Google Search Central

  • SlideConfirmed by docsD2-C184

    Google follows links in <a href> elements; a link that only runs an onclick handler, or a hash pseudo-link such as href=#/products, may be invisible to Google.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C186

    Non-crawlable link markup, such as onclick links and hash pseudo-links, is common in single-page web apps, and sites that use faceted navigation should check their links for it.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-URL-03

Route single-page apps with the History API and real paths, not # fragments

Why

A fragment (the part of a URL after #) exists only in the browser and is often ignored by crawlers, so hash-routed views are not separate pages for Google. Google recommends the History API for clean URLs, and moving from hash routes to real paths keeps the same client-side behaviour.

How

Run the router in history mode. Render links as <a href="/path">, intercept the click, call history.pushState and handle popstate. Configure the server or hosting rewrites so a direct request to every valid path returns 200 with the page, ideally server-rendered, and unknown paths return 404 (DEV-ERR-02). Real paths also work as deep links on Android, iOS and desktop.

Test

Open a deep URL directly in a fresh browser session and with curl: it returns 200 and the right content. The rendered HTML in URL Inspection contains no href="#/..." routes.

Code · Single-page app routing with the History API instead of #/ routes

Every view gets a real path in a real <a href>; JavaScript intercepts the click, updates the address with pushState and renders the view, so users skip full reloads while Google can follow the links. The server must answer a direct request for each path with the full page (ideally server-rendered) and unknown paths with 404. The not-found view is a route of its own, so a shell served with 404 renders it instead of redirecting again.

HTML
<nav>
  <a href="/" data-link>Home</a>
  <a href="/chairs/" data-link>Chairs</a>
  <a href="/tables/" data-link>Tables</a>
</nav>
<main id="app"></main>

<script>
  const routes = {
    '/': () => '<h1>Example Shop</h1>',
    '/chairs/': () => '<h1>Chairs</h1>',
    '/tables/': () => '<h1>Tables</h1>',
    '/not-found': () => '<h1>Page not found</h1>', // the server answers this URL with 404
  };

  function renderRoute(path) {
    const view = routes[path];
    if (!view) {
      // Unknown route on a 200 URL: go to the URL the server answers with 404 (once; never loop)
      if (path !== '/not-found') window.location.replace('/not-found');
      return;
    }
    document.getElementById('app').innerHTML = view();
  }

  document.addEventListener('click', (event) => {
    const link = event.target.closest('a[data-link]');
    if (!link || link.origin !== location.origin) return;
    if (event.button !== 0 || event.metaKey || event.ctrlKey || event.shiftKey || event.altKey) return;
    event.preventDefault();
    history.pushState({}, '', link.pathname + link.search);
    renderRoute(link.pathname);
  });

  window.addEventListener('popstate', () => renderRoute(location.pathname));
  renderRoute(location.pathname);
</script>
Evidence · 6 claims · 2 Google pages
  • SlideConfirmed by docsD2-C185

    The crawlable pattern for single-page app navigation is a real URL in the href, such as <a href=/products>, combined with History API routing (window.history.pushState) instead of hash routes.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C284

    URL fragments (#) are often ignored by crawlers: a fragment exists only in the browser, so Google cannot request it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C285

    Google recommends the History API to give single-page apps clean URLs instead of fragment-based routes.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C288

    With the History API, a single-page app can use real links and attach event listeners that intercept the click, rewrite the URL with pushState and load the new content, so Google can follow the links while users avoid full page reloads.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C286

    Moving a single-page app from fragment-based routes to real paths with the History API keeps the same client-side behaviour and was described as not free but relatively straightforward.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C289

    Real URLs also work as deep links: Android, iOS and desktop operating systems accept full URLs as keys to specific content in an app, so clean URLs simplify the cross-platform user experience, not only indexing.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-URL-04

Link every indexable page from the normal navigation or a hub page

Why

Discovery works through links: the homepage links to sections, which link to further pages, and Google infers a page's importance from its internal links (how many point to it and how many steps it is from the homepage). A page reachable only through a sitemap is found slowly and attracts little crawl demand. For large sites, Google pointed to category hub pages rather than HTML sitemaps.

How

Link menus to category pages, category pages to sub-categories and sub-categories to every product or article, all with <a href>. Keep important pages within a few clicks of the homepage. Where not every item can be linked, list the rest in the XML sitemap or a Merchant Center feed as a supplement, not a replacement. Never let a page be reachable only through a canonical tag (DEV-CAN-04).

Test

Crawl from the homepage following only <a href> links and compare the result with the sitemap: sitemap URLs missing from the crawl are orphans. Check the click depth of key templates.

Evidence · 8 claims · 2 Google pages
  • SlideConfirmed by docsD1-C040

    URL discovery works through links: a homepage links to section pages, which link to further pages.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD2-C005

    Google's ecommerce documentation recommends linking from menus to category pages, from category pages to sub-category pages and from sub-category pages to all product pages; where not every product can be linked, it recommends a sitemap or a Merchant Center feed.

    Google Search Central

  • DocsSourceD2-C006

    Google's ecommerce documentation says Google can infer a page's relative importance within a site from its internal links, such as how many links point to the page and how many links Google must follow to reach it.

    Google Search Central

  • SlideConsistent with docsD2-C820

    Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their important pages, instead of on HTML sitemaps.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideD2-C003

    Google's Q&A slide on HTML sitemaps added that category hub pages linking to a site's important pages have the extra benefit that they may also be useful to users.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD1-C067

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C838

    Answering whether HTML sitemaps still matter for large sites, Google said a site that wants one can make one, but that the effort is better spent on hub pages such as category pages.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD1-C202

    Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

ShouldDocumentedDEV-URL-05

Publish XML sitemaps listing only canonical, indexable URLs with accurate lastmod

Why

A sitemap listed in robots.txt can be picked up by any crawler; one that is not listed has to be submitted, for example in Search Console. Sitemap inclusion is a weak canonical signal, so a sitemap full of redirects, parameters or noindexed URLs sends mixed signals.

How

Generate sitemaps from the same source as the canonical tags: only URLs that return 200, are indexable and are their own canonical. Change lastmod only for significant updates (main content, structured data or links), never for footer, copyright or timestamp changes; Google uses lastmod only when it is consistently accurate. Split files at 50,000 URLs or 50 MB uncompressed under a sitemap index, and add hreflang (DEV-INT-05), image (DEV-IMG-04) and video (DEV-VID-04) entries where used. Reference the index in robots.txt and submit it in Search Console.

Test

Crawl the sitemap URLs: every one returns 200, has a self-referencing canonical and no noindex. The Sitemaps report shows each file as read, and the Page indexing report filtered by sitemap shows how many listed URLs are indexed.

Code · XML sitemap with canonical URLs and real lastmod dates

List only canonical, indexable URLs that answer 200 (no redirects, no noindex, no parameters you canonicalise away), with a lastmod that changes only for significant updates (the main content, structured data or links), never for footer, copyright or timestamp changes; Google uses lastmod only when it is consistently accurate. Split large sites with a sitemap index (at most 50,000 URLs or 50 MB uncompressed per file), reference it in robots.txt and submit it in Search Console.

XML
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://www.example.com/sitemaps/products-1.xml</loc>
    <lastmod>2026-09-30T06:00:00+00:00</lastmod>
  </sitemap>
</sitemapindex>
XML
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://www.example.com/chairs/oak-dining-chair</loc>
    <lastmod>2026-09-28T10:15:00+00:00</lastmod>
  </url>
  <url>
    <loc>https://www.example.com/chairs/beech-stool</loc>
    <lastmod>2026-09-12T08:00:00+00:00</lastmod>
  </url>
</urlset>
Evidence · 11 claims · 4 Google pages
  • SlideConfirmed by docsD2-C017

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C018

    Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in Search Console.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C391

    Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.

    Google Search Central

  • SlideConsistent with docsD2-C390

    Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD1-C067

    A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.

    Ibrahim Anjro (author)

  • DocsSourceD2-C833

    Google's sitemap guide says Google uses lastmod only when it is consistently and verifiably accurate, and counts a change to the main content, the structured data or the links of a page as significant, but not a changed copyright date.

    Google Search Central

  • StageConfirmed by docsD2-C842

    Google said listing the sitemap in robots.txt is fine, as many websites do.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C843

    Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD1-C475

    There is no sweet spot to find for a sitemap: technically it should list every URL you want indexed, so that Google can find each of them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C337

    An XML sitemap, a format from around 2005, gives Google, other search engines and potentially AI systems a list of the URLs a site wants crawled; Google called it nothing fancy but said it still uses sitemaps.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C336

    Google finds what to crawl mainly by extracting URLs from previously crawled pages, and additionally from sitemaps.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-URL-06

Give each page of a paginated series its own URL, its own canonical and an <a href> to the next page

Why

Google's crawlers do not click "next" buttons, and rel=next and rel=prev are no longer used. Pointing the canonical of page 2 and later at page 1 is incorrect because the pages are not duplicates, and items listed only on later pages may never be indexed. A remark at Search Central Live that a canonical to page 1 can sometimes make sense was hedged; the documented rule stands.

How

Use ?page=n (or /page/n/) URLs, with page 1 at the clean listing URL. Each page carries a self-referencing canonical and <a href> links to the next page and back to page 1, ideally also to nearby pages; titles and descriptions may stay the same across the series, so a page number in the title is optional. Never put page numbers in # fragments. Keep sort and filter variants out of the crawl (DEV-URL-08).

Test

For pages 1 to 3 of a listing, curl the HTML: the canonical is self-referencing and the next-page <a href> is present without JavaScript. URL Inspection shows later pages as indexable.

Code · Paginated listing with crawlable page links

Each page of a series has its own URL (?page=n), its own self-referencing canonical and plain <a href> links to the next page and back to the first page; the pages may share one title and description, so a page number in the title is optional. Do not canonicalise page 2 and later to page 1, and do not put page numbers in # fragments; Google no longer uses rel=next/rel=prev.

HTML
<!-- https://www.example.com/chairs/?page=2 -->
<head>
  <title>Chairs, page 2 | Example Shop</title> <!-- the page number is optional -->
  <link rel="canonical" href="https://www.example.com/chairs/?page=2">
</head>
<body>
  <main>
    <h1>Chairs</h1>
    <ul>
      <li><a href="/chairs/oak-dining-chair">Oak dining chair</a></li>
      <li><a href="/chairs/beech-stool">Beech stool</a></li>
    </ul>
    <nav aria-label="Pagination">
      <a href="/chairs/">1</a>
      <a href="/chairs/?page=2" aria-current="page">2</a>
      <a href="/chairs/?page=3">3</a>
      <a href="/chairs/?page=3">Next page</a>
    </nav>
  </main>
</body>
Evidence · 6 claims · 2 Google pages
  • DocsSourceD1-C115

    Google's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.

    Google Search Central

  • DocsSourceD2-C395

    Google's 2013 post on rel=canonical mistakes says pointing rel=canonical from page 2 or later of a paginated series to page 1 is incorrect because the pages are not duplicates, and that it would result in the content on later pages not being indexed at all.

    Search Central blog (8 April 2013)

  • AnalysisD2-C394

    Google's pagination guide still says not to use page 1 as the canonical of a paginated series, so keep self-referencing canonicals on paginated pages unless you deliberately want later pages folded into page 1 and the items they list are linked from elsewhere.

    Ibrahim Anjro (author)

  • StageNot in docsD2-C393

    Google's speaker said that pointing rel=canonical from the pages of a paginated set to the first page can sometimes make sense depending on the goal, for example to make a category page more visible, but that it affects canonicalization and deduplication.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD1-C114

    A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD2-C832

    Google's pagination guide says the pages of a paginated sequence may share the same title and description, and suggests linking every page of the sequence back to the first page.

    Google Search Central

MustDocumentedDEV-URL-07

Back infinite scroll and load-more buttons with paginated URLs

Why

Google renders a page in a very tall viewport instead of scrolling, which is not efficient and can miss content; a figure of around 10,000 pixels was said at Search Central Live but is not documented. Content that appears only after scrolling beyond that, or after clicking "load more", stays invisible unless each chunk also has its own URL. Said at Search Central Live: a Google panelist called pagination one of the trickiest things in web development and a switch to infinite scroll or a load-more button risky, depending on what you want to achieve, because Googlebot does not click buttons; without crawlable links to them, Google does not see the further pages at all.

How

Build load-more and infinite scroll on top of the paginated series (DEV-URL-06): each chunk has a persistent, unique URL that also works on direct load, chunks link to each other with <a href>, and the History API updates the address when a new chunk becomes the main visible element. Trigger loading with an IntersectionObserver, never a scroll listener (DEV-REN-03).

Test

With JavaScript disabled, follow the "more" link through every page of a listing. In URL Inspection the rendered HTML of page 1 contains a crawlable link to page 2, and server logs show Googlebot fetching the paginated URLs.

Code · Lazy loading that works without scrolling

Googlebot does not scroll; it renders with a very tall viewport. Use native loading="lazy" for images and iframes, and load list chunks with an IntersectionObserver on a sentinel element, never with a scroll event listener. The plain <a href> links to the paginated URLs stay in the page whatever the script does, so the rendered HTML always links to the next page.

HTML
<!-- Images and iframes: native lazy loading, real src in the HTML -->
<img src="/img/oak-chair-800.webp" alt="Oak dining chair with a woven paper-cord seat" width="800" height="600" loading="lazy">

<!-- Lists: page 1 of /chairs/, a sentinel for the observer, and pagination links the script never removes -->
<ul id="products">
  <li><a href="/chairs/oak-dining-chair">Oak dining chair</a></li>
</ul>
<div id="load-more-sentinel" style="height: 1px"></div>
<nav aria-label="Pagination">
  <a href="/chairs/" aria-current="page">1</a>
  <a href="/chairs/?page=2">2</a>
  <a href="/chairs/?page=3">3</a>
  <a href="/chairs/?page=2">Next page</a>
</nav>

<!-- Avoid: window.addEventListener('scroll', loadMoreProducts) never runs for Googlebot -->

<script>
  const list = document.getElementById('products');
  const sentinel = document.getElementById('load-more-sentinel');
  let next = '/chairs/?page=2'; // written by the server: the next page's URL, empty on the last page

  const observer = new IntersectionObserver(async ([entry]) => {
    if (!entry.isIntersecting || !next) return;
    observer.unobserve(sentinel);
    const url = new URL(next, location.href);
    const res = await fetch('/fragments' + url.pathname + url.search); // returns the <li> items of that page
    if (!res.ok) { observer.disconnect(); return; } // keep the links; never insert an error page
    list.insertAdjacentHTML('beforeend', await res.text());
    history.replaceState({}, '', url.pathname + url.search); // the URL follows the chunk once it is in place
    next = res.headers.get('X-Next-Page') || ''; // e.g. "/chairs/?page=3", absent on the last page
    if (next) observer.observe(sentinel); else observer.disconnect();
  });

  if (next) observer.observe(sentinel);
</script>
Evidence · 6 claims · 3 Google pages
  • DocsSourceD2-C276

    To make infinite scroll indexable, Google's lazy-loading guide says to support paginated loading: give each chunk its own persistent, unique URL, link sequentially to those URLs, and update the displayed URL with the History API when a new chunk becomes the main visible element.

    Google Search Central

  • DocsSourceD2-C273

    Google's March 2023 SEO office hours say Google sees infinite-scroll content through viewport expansion, rendering a page like a very long phone, which is not particularly efficient and can miss infinite content.

    Google Search Central (SEO office hours transcript, March 2023)

  • StageConfirmed by docsD2-C271

    Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C272

    Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD1-C114

    A Google panelist called pagination one of the trickiest things in web development and said switching to infinite scroll or a load-more button is risky, depending on what you want to achieve, because Googlebot does not click buttons.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C491

    If paginated pages are replaced by a load-more button or infinite scroll without crawlable links to them, Google will not see the further pages at all.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-URL-08

Keep faceted navigation out of the crawl with robots.txt rules or URL fragments

Why

Facets create near-infinite URL spaces (the same filters in another order, conflicting or excessive combinations, facets on paginated or search pages), one of the things named as burning crawl budget. Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt or use URL fragments, which Google generally does not crawl; rel=canonical and nofollow are weaker, slower options.

How

Decide per facet whether it deserves an indexable landing page, for example a main category or brand with real search demand, and give those a clean URL, unique content and <a href> links. Put every other filter, sort and view parameter behind a robots.txt Disallow pattern or in the # fragment, and do not link filter combinations with crawlable hrefs. Also disallow internal search results and action URLs (cart, wishlist, compare). Audit CMS plug-ins for generated URL spaces: Google's crawling team said many crawling complaints it receives turn out to come from a plug-in, such as a calendar plug-in that adds calendar parameters to the URLs of every post, page, category and tag page, and Google crawls such URLs until its systems learn from large samples that they are useless. A Google panelist added that useless plug-in parameter URLs can also be handled at the web server, for example with a rewrite rule or a custom module, so that crawl budget goes to the URLs that matter.

Test

A crawl that respects robots.txt finds a bounded number of listing URLs; server logs and Crawl stats show Googlebot not spending fetches on parameter URLs; Google's open-source robots.txt parser (github.com/google/robotstxt), run with the Googlebot token, returns disallowed for sample filter URLs and allowed for item pages, and URL Inspection shows item pages as Crawl allowed.

Code · robots.txt with crawl controls and a rendering carve-out

Disallow only URLs that should never be crawled (internal search, cart and checkout actions, filter parameters), keep every script, style and API path that pages need for rendering crawlable, and list the sitemap. Google reads only user-agent, allow, disallow and sitemap; each host (www, api, cdn) needs its own file at its root. Robots.txt is public, so never list secret paths in it.

Text
# https://www.example.com/robots.txt
User-agent: *
# Internal search results and cart/checkout actions (adapt to your URL patterns)
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
# (Shops in Google Shopping: give Storebot-Google its own group that leaves cart and checkout open.
#  A named group replaces this * group for that crawler, so repeat in it every rule that should still apply.)
# Filter and sort parameters of faceted navigation, as the first parameter (?) or a later one (&);
# /*?*size= would also block ?pagesize= and /*?*color= would block ?bgcolor=
Disallow: /*?color=
Disallow: /*&color=
Disallow: /*?size=
Disallow: /*&size=
Disallow: /*?sort=
Disallow: /*&sort=
# API: blocked in general, but the endpoints pages render from stay crawlable
# (the longer, more specific Allow rule wins)
Disallow: /api/
Allow: /api/products/
Allow: /api/reviews/
# Never disallow /static/, /assets/ or other JS and CSS folders

# Optional: keep content out of Gemini training and grounding (no effect on Google Search)
User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
Evidence · 12 claims · 2 Google pages
  • DocsSourceD1-C101

    Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.

    Google

  • SlideConsistent with docsD1-C100

    Six faceted-URL problems were shown: the same facets in a different order, irrelevant or conflicting facet combinations, excessive facet selection, facets on paginated series, facets combined with search queries, and optional facets with default values.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideConsistent with docsD1-C099

    Three things burn crawl budget: server errors unrelated to server load, useless pages and resources, and infinite URL spaces.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideConsistent with docsD1-C103

    Four ways to manage crawl budget: use HTTP cache control, have good site navigation, restrict crawlers' access to faceted navigation and action URLs, and improve or remove useless content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD2-C188

    Hash-fragment links are a problem only where Google should follow them: product, category and language links need a real URL in an <a href>, while fragments can deliberately keep filter combinations out of the crawl, as Google's faceted navigation guide allows.

    Ibrahim Anjro (author)

  • StageConsistent with docsD1-C469

    Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C430

    Google follows robots.txt partly in its own interest: crawling an infinite URL space such as a calendar that robots.txt blocks would waste Google's crawling time as well as the site's resources.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C446

    Many crawling complaints that reach Google's crawling team turn out to be caused by a plugin, for example a WordPress calendar plugin that adds calendar parameters to the URLs of every post, page, category and tag page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C447

    Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C378

    Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C444

    Log files are worth checking on sites with filter parameters or complicated URLs, because crawlers can easily wander off into URLs that make no sense for the site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C539

    A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example with a rule or a custom module in Apache, which saves crawl budget so that Google may pick up the URLs that matter instead (the exact mechanism is not clear in the recording).

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

ShouldDocumentedDEV-URL-09

If facet URLs must be crawlable, use & separators, a fixed filter order and 404 for empty results

Why

Google's faceted navigation guide asks crawlable facet URLs to use the standard & separator, keep filters in a consistent order and return 404 when a combination has no results. Said at Search Central Live: many similar-looking URLs leading to the same or empty content can make Google fold a whole URL pattern into one canonical and stop crawling new URLs that fit it.

How

Normalise the parameter order on the server (redirect or canonicalise ?size=m&color=red and ?color=red&size=m to one order). Use key=value pairs joined with &, not custom separators. Return 404, not an empty 200 page, for combinations with no items. Make indexable facet pages differ in their main content (heading, introduction, items), not only in the title.

Test

Request one combination in two parameter orders: one redirects or canonicalises to the other. curl -I on an empty combination returns 404. The Page indexing report shows no growth of duplicate reasons on facet URLs.

Evidence · 5 claims · 2 Google pages
  • DocsSourceD1-C102

    If faceted URLs must be crawled, use the standard & separator, keep filters in a consistent order, and return a 404 when a filter combination has no results.

    Google

  • SlideConsistent with docsD2-C371

    To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C368

    Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C369

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C372

    Make location and variant pages differ in their main content (local stock, staff, prices, addresses) and return 404 for empty or invalid combinations; otherwise every URL that fits the pattern can be folded into one canonical.

    Ibrahim Anjro (author)

ShouldSaid at Search Central LiveDEV-URL-10

Group URLs into meaningful section paths and launch new content under sections Google already indexes well

Why

Said at Search Central Live: when Google does not yet know a URL's quality or popularity, it uses the aggregate of the URL's parent path, then that path's parent, and index selection is more forgiving with new URLs from a site or a section that already satisfies users well. Neither rule is in Google's documentation. Author's view: URL paths are a lever for crawl demand and index selection, not only for tidiness.

How

Plan the URL structure so each template lives under a section path that reflects the site's structure (/chairs/oak-dining-chair rather than /p?id=123 or a flat root). Launch new content types under sections Google already crawls and indexes well rather than under weak or brand-new ones, and improve or remove weak sections before adding to them.

Test

The URL plan maps every template to a section path, and after launch each new section is reviewed on its own (a URL-prefix property or a sitemap per section) in the Page indexing report and in Crawl stats.

Evidence · 4 claims
  • SlideNot in docsD1-C094

    If the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C095

    New content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.

    Ibrahim Anjro (author)

  • StageNot in docsD2-C685

    Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C686

    Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-URL-11

Compute the expected URL for every request and redirect variants to it

Why

Google treats URLs that differ as different URLs and crawls them all, even when they show the same content. A community speaker showed five kinds of similar URLs on well-known brand sites: another protocol or host, other capitalisation, extra slashes, a space encoded as %20 in one URL and + in another, and parameters in another order. Such URLs signal duplicate content, broken links and unnecessary redirects, and anyone can link to them from other sites, so the application should be hardened against them.

How

In the application, build the expected URL for each request from the route, the page type and ID (reverse routing) and the parameters the page actually uses, in a fixed order. When the requested URL differs, answer one 301 to the expected URL, or an error page when it maps to nothing. Generate internal links with the same function so they are never variants.

Test

Requests for the uppercase, double-slash, reordered-parameter and unused-parameter variants of a sample URL each return one 301 to the expected URL; a crawl that normalises URLs (lowercase scheme and host, default port removed, dot segments resolved, fragment dropped) finds no two crawled URLs that collapse to the same normalised URL.

Code · Redirect URL variants to the one URL the application expects

Build the expected URL from the route and the parameters the page really uses, in a fixed order, and answer any other spelling with one 301. Generate internal links with the same function so the site never links to a variant. The example is Express middleware; lowercasing the path assumes the site's routes are all lowercase.

JavaScript
const ORIGIN = 'https://www.example.com';
// Parameters each route actually uses, in the order they must appear.
const PARAMS = { '/chairs': ['colour', 'page'], '/search': ['q'] };

function expectedUrl(pathname, searchParams) {
  let path = pathname.toLowerCase().replace(/\/{2,}/g, '/');
  if (path.length > 1 && path.endsWith('/')) path = path.slice(0, -1);
  const kept = new URLSearchParams();
  for (const name of PARAMS[path] || []) {
    const value = searchParams.get(name);
    if (value) kept.set(name, value); // unused and empty parameters are dropped
  }
  const query = kept.toString();      // spaces always come out as +
  return path + (query ? '?' + query : '');
}

app.set('trust proxy', true);         // so req.protocol is right behind a CDN or load balancer
app.use((req, res, next) => {
  const requested = new URL(ORIGIN + req.originalUrl); // not new URL(path, base): '//x' would parse as a host
  const expected = expectedUrl(requested.pathname, requested.searchParams);
  const wrongOrigin = req.protocol !== 'https' || req.hostname !== 'www.example.com';
  if (wrongOrigin || expected !== requested.pathname + requested.search) {
    return res.redirect(301, ORIGIN + expected); // one hop, straight to the final URL
  }
  next();
});
Evidence · 8 claims · 3 Google pages
  • StageConsistent with docsD1-C397

    A community speaker said similar URLs are an indicator of duplicate content, broken links, unnecessary redirects and other technical issues.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C399

    A community speaker showed five kinds of similar URLs on well-known brand sites: a different protocol or host (HTTP vs HTTPS, www vs non-www), different capitalisation, a different number of delimiters such as slashes, a space encoded as %20 in one URL and + in another, and parameters in a different order.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C400

    To find similar URLs in a URL list, a community speaker first applies basic normalisation that keeps the meaning: lowercase the scheme and host, remove default ports such as 80 for HTTP, resolve relative path segments and usually drop the fragment, which matters only to clients.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C402

    Similar URLs usually come from programming errors or manually set links, but anyone can link to them from other sites, including malicious actors, so a site should be hardened against them.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C403

    A community speaker advised that a web application compute the expected URL for every request, for example with reverse routing from the page type and ID, and redirect or return an error page when the requested URL differs.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C404

    For query parameters, a community speaker advised checking that every parameter in a request is actually used and in the expected order, and otherwise redirecting to the expected URL with only the used parameters in the correct order.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C445

    One fix for parameter variants of a URL is to work out the normalised URL and redirect the variants to it: the redirects cost crawl budget at first, but leave the site with a clean slate.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C378

    Google treats URLs that differ as different URLs and crawls all of them even when they lead to the same content, so infinite URL spaces such as calendars or many versions of a page burn crawl budget.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

Rendering and JavaScript 10

Google renders pages with a headless Chromium in a separate, queued system and indexes what is in the DOM when rendering ends. Content and links that depend on clicks, scrolling, blocked resources, permission prompts or slow APIs can be missing from what Google indexes.

ShouldDocumentedDEV-REN-01

Server-render the main content and everything indexing depends on (SSR, static generation or hybrid)

Why

Every 200 page waits in a render queue, for seconds or longer, and rendering can fail or time out; content already in the server HTML does not depend on that slower pass. Google named server-side rendering as the fix for slow live reads of a page, and a community speaker stressed that AI systems that cannot render JavaScript see only the raw HTML. Said at Search Central Live (October 2026): content that JavaScript adds is usually seen by indexing within a few hours but at worst within weeks, and a Google speaker doubted that every URL is actually rendered; Google's JavaScript guide says every 200 page is queued for rendering, and queued is not the same as rendered.

How

Choose the rendering mode per template before launch: server-render or pre-render every template that must rank (product, category, article, landing pages), including the title, meta description, canonical, robots meta tag, hreflang, main text, prices and internal links. Hydrate on the client for interactivity only. With a headless CMS, render on the server or at build time rather than only in the browser.

Test

curl -s the page and find the main heading, price and key links in the raw HTML; compare with the rendered HTML in URL Inspection; with JavaScript disabled in DevTools the page still shows its main content.

Code · Page template with a clearly delimited main content area

Everything indexing depends on (title, description, canonical, robots rules, main text, links) is in the server HTML. Header, navigation and footer are separated from one <main> element that holds the title, headings, opening text and media, because Google weighs words by where they appear and treats the main content as the most important part.

HTML
<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <title>Oak dining chair with woven seat | Example Shop</title>
  <meta name="description" content="Solid oak dining chair with a hand-woven paper-cord seat. Seat height 45 cm, delivered assembled.">
  <link rel="canonical" href="https://www.example.com/chairs/oak-dining-chair">
  <meta name="robots" content="max-snippet:-1, max-image-preview:large">
</head>
<body>
  <header>
    <a href="/">Example Shop</a>
    <nav aria-label="Main">
      <a href="/chairs/">Chairs</a>
      <a href="/tables/">Tables</a>
    </nav>
  </header>

  <main>
    <article>
      <h1>Oak dining chair with woven seat</h1>
      <p>A solid oak dining chair with a hand-woven paper-cord seat, made for everyday family meals.</p>
      <figure>
        <img src="/img/oak-chair-1200.webp" alt="Oak dining chair with a woven paper-cord seat, seen from the front" width="1200" height="900">
        <figcaption>Natural oak finish, seat height 45 cm.</figcaption>
      </figure>
      <h2>Dimensions and materials</h2>
      <p>Width 46 cm, depth 52 cm, height 80 cm. <strong>Solid European oak</strong>, paper-cord seat.</p>
    </article>
  </main>

  <footer>
    <nav aria-label="Footer">
      <a href="/delivery/">Delivery</a>
      <a href="/returns/">Returns</a>
      <a href="/contact/">Contact</a>
    </nav>
  </footer>
</body>
</html>
Evidence · 11 claims · 2 Google pages
  • DocsSourceD2-C131

    Google's JavaScript SEO guide says Googlebot sends every page with a 200 HTTP status code to the rendering queue, whether or not it contains JavaScript, unless a robots meta tag or header tells Google not to index it, and Google uses the rendered HTML to index the page.

    Google Search Central

  • DocsSourceD2-C172

    Google's JavaScript SEO basics guide says a page may wait in the render queue for a few seconds but that it can take longer, and it gives no upper limit.

    Google Search Central

  • StageConsistent with docsD2-C124

    Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C139

    A community speaker strongly advised putting everything you want cited into the raw, server-side rendered HTML, especially for AI systems that cannot render JavaScript yet.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C133

    Put everything indexing depends on (title, meta description, canonical, robots meta tag, main text and links) in the raw HTML so it is available without rendering, let JavaScript add enhancements only, and check the rendered HTML in URL Inspection for the rest.

    Ibrahim Anjro (author)

  • AnalysisD2-C125

    For pages people are likely to ask an AI assistant about (product, pricing, documentation and policy pages), server-render the main content so a live, user-triggered read does not depend on client-side JavaScript finishing quickly.

    Ibrahim Anjro (author)

  • StageConsistent with docsD3-C627

    Content that JavaScript adds to a page is typically seen by Google's indexing system within a few hours, and at worst within weeks.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C628

    Gary Illyes said Google keeps saying it renders every URL on the internet and that he would say this is not true, though it is what he was told; he went on to say that Google's logs show the rendering queue cleared within weeks.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C676

    The stage doubt that every URL gets rendered sits beside Google's JavaScript guide, which says every page with a 200 status is queued for rendering unless a robots rule blocks indexing: queued is not the same as rendered, so do not rely on rendering for critical content.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C247

    Natalia Venditto recommended that web fragments be server-rendered.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C249

    Natalia Venditto said a server-rendered app reframed into the host page is fully indexable.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-REN-02

Have all indexable content in the DOM after load, without clicks, typing or scrolling

Why

Googlebot does not click, type or scroll. Content that loads only after a user action is not in the DOM when Google renders the page, so it cannot be indexed, while content that is in the DOM but hidden (tabs, accordions, read-more) can be. Google's slide called content missing from the rendered HTML the most common JavaScript indexing issue.

How

Ship tab, accordion and read-more content in the initial HTML or DOM and hide it with CSS or the hidden attribute; never fetch it on click. Product specifications, reviews, FAQs and descriptions belong in the page at load. Keep interaction-only loading for content you do not need indexed.

Test

In URL Inspection (rendered HTML) or the Rich Results Test, find text from every tab and accordion panel. In the DevTools Network panel, no content request waits for a click.

Code · Tabs and accordions whose content is in the DOM from the start

Googlebot does not click, so content fetched only when a tab is clicked never exists for it. Ship every panel in the HTML and only hide the inactive ones with the hidden attribute or CSS: hidden content in the DOM can be indexed, absent content cannot.

HTML
<div class="tabs">
  <div role="tablist" aria-label="Product details">
    <button type="button" role="tab" id="tab-specs" aria-controls="panel-specs" aria-selected="true">Specifications</button>
    <button type="button" role="tab" id="tab-care" aria-controls="panel-care" aria-selected="false">Care</button>
  </div>
  <section role="tabpanel" id="panel-specs" aria-labelledby="tab-specs">
    <p>Seat height 45 cm, solid oak frame, paper-cord seat.</p>
  </section>
  <section role="tabpanel" id="panel-care" aria-labelledby="tab-care" hidden>
    <p>Wipe with a damp cloth; re-oil the frame once a year.</p>
  </section>
</div>

<!-- Avoid: tab.onclick = () => fetch('/api/specs') ... content that exists only after a click -->

<script>
  const tabs = document.querySelectorAll('[role="tab"]');
  tabs.forEach((tab) => {
    tab.addEventListener('click', () => {
      tabs.forEach((other) => {
        const selected = other === tab;
        other.setAttribute('aria-selected', String(selected));
        document.getElementById(other.getAttribute('aria-controls')).hidden = !selected;
      });
    });
  });
</script>
Evidence · 8 claims · 3 Google pages
  • SlideConfirmed by docsD2-C195

    Content that loads only after a user action such as a click or a scroll is not in the DOM while Google renders the page, so Google cannot index it.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C268

    Content that loads only when a user clicks an element is not supported in the way Google renders pages for indexing.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C201

    Tab or accordion content fetched from an API only when a user clicks the tab, as in tab.onclick = () => fetch('/api/specs'), does not exist for Google until someone clicks.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C196

    For rendering, what matters is whether content is present in the DOM, not whether it is visible on screen: hidden content can be indexed, absent content cannot.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C202

    Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden attribute; Google indexes such hidden content.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C262

    Content that is not in the final DOM after rendering cannot be seen by Google, so it cannot be indexed.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C261

    Google's slide called content missing from the rendered HTML the most common JavaScript issue for indexing.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C866

    Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-REN-03

Lazy-load with native lazy loading or an IntersectionObserver, never with scroll events

Why

Googlebot does not scroll, so a scroll event listener never fires and the content it would load never appears. Google renders with a very tall viewport instead (said at Search Central Live to be around 10,000 pixels), so content loaded when it enters the viewport does load during rendering.

How

Use loading="lazy" on img and iframe elements, with a real src plus width and height, and an IntersectionObserver for other components. Never make indexable content depend on scroll position, scroll depth or a timer. Treat the tall viewport as a safety margin, not a target, and give long lists paginated URLs (DEV-URL-07).

Test

The rendered HTML in URL Inspection contains below-the-fold images and list items. A code search finds no addEventListener('scroll') or onscroll handler that loads content.

Code · Lazy loading that works without scrolling

Googlebot does not scroll; it renders with a very tall viewport. Use native loading="lazy" for images and iframes, and load list chunks with an IntersectionObserver on a sentinel element, never with a scroll event listener. The plain <a href> links to the paginated URLs stay in the page whatever the script does, so the rendered HTML always links to the next page.

HTML
<!-- Images and iframes: native lazy loading, real src in the HTML -->
<img src="/img/oak-chair-800.webp" alt="Oak dining chair with a woven paper-cord seat" width="800" height="600" loading="lazy">

<!-- Lists: page 1 of /chairs/, a sentinel for the observer, and pagination links the script never removes -->
<ul id="products">
  <li><a href="/chairs/oak-dining-chair">Oak dining chair</a></li>
</ul>
<div id="load-more-sentinel" style="height: 1px"></div>
<nav aria-label="Pagination">
  <a href="/chairs/" aria-current="page">1</a>
  <a href="/chairs/?page=2">2</a>
  <a href="/chairs/?page=3">3</a>
  <a href="/chairs/?page=2">Next page</a>
</nav>

<!-- Avoid: window.addEventListener('scroll', loadMoreProducts) never runs for Googlebot -->

<script>
  const list = document.getElementById('products');
  const sentinel = document.getElementById('load-more-sentinel');
  let next = '/chairs/?page=2'; // written by the server: the next page's URL, empty on the last page

  const observer = new IntersectionObserver(async ([entry]) => {
    if (!entry.isIntersecting || !next) return;
    observer.unobserve(sentinel);
    const url = new URL(next, location.href);
    const res = await fetch('/fragments' + url.pathname + url.search); // returns the <li> items of that page
    if (!res.ok) { observer.disconnect(); return; } // keep the links; never insert an error page
    list.insertAdjacentHTML('beforeend', await res.text());
    history.replaceState({}, '', url.pathname + url.search); // the URL follows the chunk once it is in place
    next = res.headers.get('X-Next-Page') || ''; // e.g. "/chairs/?page=3", absent on the last page
    if (next) observer.observe(sentinel); else observer.disconnect();
  });

  if (next) observer.observe(sentinel);
</script>
Evidence · 6 claims · 2 Google pages
  • SlideConfirmed by docsD2-C198

    Lazy loading should be triggered by an Intersection Observer that watches a sentinel element instead of by a scroll event listener.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C197

    Lazy loading triggered by a scroll event listener, such as window.addEventListener('scroll', loadMoreProducts), never runs for Googlebot because Googlebot does not scroll.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C271

    Google does not scroll a page when it renders it for indexing, so content that infinite scroll loads on scroll-depth triggers works for users but is never loaded for Google.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C274

    Content loaded as elements enter the viewport, for example with an Intersection Observer, does load when Google renders a page, because Google's rendering viewport is very tall.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C199

    An Intersection Observer fires when the observed element enters the viewport, and the viewport Google renders with is tall, so content lazy-loaded this way can load during rendering.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C272

    Instead of scrolling, Google renders a page in a very tall viewport, around 10,000 pixels high.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

AvoidDocumentedDEV-REN-04

Do not block JavaScript, CSS or API endpoints that pages need for rendering in robots.txt

Why

Google's renderer fetches scripts, styles and API requests through Googlebot and obeys robots.txt: if it cannot fetch a resource, it cannot run it. Blocking a script folder or a generic /api/ folder leaves blank modules or an empty page, which can be treated as a soft 404 although users see a full page.

How

Allow all JS, CSS and font paths. If an API folder must stay disallowed, carve out the endpoints rendering needs (Disallow: /api/ plus Allow: /api/products/), since the longer Allow rule wins. Check the robots.txt of every host the page loads data from, and re-test a rendered page after every robots.txt change touching script or API paths.

Test

URL Inspection live test, page resources: nothing needed for content is "Blocked by robots.txt". The DevTools Network panel (Fetch/XHR) lists each API host, and each host's robots.txt allows those paths.

Code · robots.txt with crawl controls and a rendering carve-out

Disallow only URLs that should never be crawled (internal search, cart and checkout actions, filter parameters), keep every script, style and API path that pages need for rendering crawlable, and list the sitemap. Google reads only user-agent, allow, disallow and sitemap; each host (www, api, cdn) needs its own file at its root. Robots.txt is public, so never list secret paths in it.

Text
# https://www.example.com/robots.txt
User-agent: *
# Internal search results and cart/checkout actions (adapt to your URL patterns)
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
# (Shops in Google Shopping: give Storebot-Google its own group that leaves cart and checkout open.
#  A named group replaces this * group for that crawler, so repeat in it every rule that should still apply.)
# Filter and sort parameters of faceted navigation, as the first parameter (?) or a later one (&);
# /*?*size= would also block ?pagesize= and /*?*color= would block ?bgcolor=
Disallow: /*?color=
Disallow: /*&color=
Disallow: /*?size=
Disallow: /*&size=
Disallow: /*?sort=
Disallow: /*&sort=
# API: blocked in general, but the endpoints pages render from stay crawlable
# (the longer, more specific Allow rule wins)
Disallow: /api/
Allow: /api/products/
Allow: /api/reviews/
# Never disallow /static/, /assets/ or other JS and CSS folders

# Optional: keep content out of Gemini training and grounding (no effect on Google Search)
User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
Evidence · 7 claims · 4 Google pages
  • SlideConfirmed by docsD2-C204

    Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C206

    Google says its Web Rendering Service fetches the resources a page references through Googlebot, including JavaScript, CSS and XHR requests to APIs, but not images or videos.

    Search Central blog (3 December 2024), Search Central blog (31 March 2026)

  • DocsSourceD2-C297

    Google's introduction to robots.txt says unimportant image, script or style files may be blocked, but resources whose absence makes a page harder for Google to understand should not be blocked.

    Google Search Central

  • SlideConsistent with docsD2-C207

    If an API folder must stay blocked in robots.txt, allow the endpoints that rendering needs, for example Disallow: /api/ together with Allow: /api/products/.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C265

    Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the crawling of critical JavaScript files or API endpoints.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C296

    If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C213

    A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

AvoidDocumentedDEV-REN-05

Do not add, change or remove robots meta tags with JavaScript

Why

Robots rules added with JavaScript take longer to be seen, and removing a noindex with JavaScript does not work: Google may skip rendering a page that arrives with noindex, so it stays out of the index. Google's documentation recommends avoiding JavaScript to inject or change meta tags.

How

Decide indexability on the server and send the final robots meta tag or X-Robots-Tag header in the first response. Never ship noindex in the HTML shell and remove it after hydration, and never toggle data-nosnippet with JavaScript. The one documented exception is adding noindex with JavaScript to an error view in a client-rendered app (DEV-ERR-02).

Test

The robots meta tag in curl output and in URL Inspection's rendered HTML is identical on every indexable template. A search of the client bundle finds no code that writes meta[name="robots"] except error views.

Evidence · 9 claims · 3 Google pages
  • SlideConsistent with docsD2-C108

    Removing a robots restriction such as noindex with JavaScript does not work, a slide said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C106

    A slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C107

    Google's meta tags documentation recommends avoiding JavaScript to inject or change meta tags whenever possible and testing the implementation thoroughly when it must be used.

    Google Search Central

  • DocsSourceD2-C257

    Google's JavaScript SEO basics says every page with a 200 status code is queued for rendering, whether or not it uses JavaScript, unless a robots meta tag or header says not to index it; for non-200 pages such as 404 error pages, rendering might be skipped.

    Google Search Central

  • AnalysisD2-C109

    Ship robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may skip rendering a page that arrives with noindex, so the page can stay out of the index even if a script removes the tag later.

    Ibrahim Anjro (author)

  • DocsSourceD2-C079

    Google's robots meta tag specification says data-nosnippet may be extracted both before and after rendering, so the attribute should not be added to or removed from existing elements with JavaScript.

    Google Search Central

  • StageConsistent with docsD2-C852

    When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript, so a script cannot switch the page back to indexable, John Mueller said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C853

    John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding them with JavaScript only where that is not possible, as in a JavaScript web app.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C854

    Google's JavaScript SEO basics guide says that when Google encounters a noindex rule it may skip rendering and JavaScript execution, so using JavaScript to change or remove a noindex robots meta tag may not work as expected.

    Google Search Central

ShouldDocumentedDEV-REN-06

Make rendering robust: no slow API dependencies and no unresolved placeholders

Why

Said at Search Central Live by a community speaker: a renderer does not wait forever before taking its snapshot, so content from slow internal systems can be missing and template placeholders can end up in indexed titles and snippets; because timing varies, the broken page changes from crawl to crawl. Server-side or hybrid rendering, fallbacks and never leaving placeholders in the DOM were named as safeguards.

How

Render critical content on the server from data available at request time, give slow services timeouts with sensible fallbacks, and never write a template variable, undefined or null into the DOM, title or meta description (render nothing or a neutral default). Keep JavaScript bundles and API waterfalls small so the main content appears quickly. Never make content depend on cookies, localStorage, sessionStorage or a consent state from an earlier page, because Google's renderer clears them between page loads, and give JS and CSS bundles content-hashed file names (main.2bb85551.js), because the renderer may ignore caching headers and use stale files.

Test

Render the same URLs several times with a JavaScript-rendering crawler and search titles, descriptions and body text for {{, undefined, null and NaN; investigate any URL whose rendered output differs between runs.

Evidence · 8 claims · 2 Google pages
  • StageConsistent with docsD2-C259

    Erin Sparling named server-side or hybrid rendering, fallbacks and not leaving placeholders in the DOM, among other measures, as ways to guard against failed JavaScript rendering.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C149

    A search engine's renderer does not wait forever before taking its snapshot, so content from slow internal systems or APIs can be missing and unresolved placeholders can end up in the final snapshot, a community speaker said.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C263

    Content is missing from the rendered HTML either because the server does not serve it or because the content has still not appeared after some period of time.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C175

    If a page keeps loading content indefinitely, not all of that content will be indexed, because Google's rendering does not go on forever (this passage of the recording is partly unclear).

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C151

    Unresolved placeholders caused by rendering timeouts are hard to catch because the problem moves around: it is not always the same page that is broken.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C154

    Add a check for unresolved template placeholders (such as {{brand}}, undefined or null) in rendered titles, meta descriptions and visible text to recurring crawls, and render the same URLs more than once, because such failures are intermittent.

    Ibrahim Anjro (author)

  • DocsSourceD2-C828

    Google's guide to fixing Search-related JavaScript problems says its Web Rendering Service keeps no state across page loads: local storage, session storage and HTTP cookies are cleared.

    Google Search Central

  • DocsSourceD2-C829

    Google's guide to fixing Search-related JavaScript problems says the Web Rendering Service may ignore caching headers and so use outdated JavaScript or CSS, and recommends content fingerprinting, which puts a hash of the content in the file name, as in main.2bb85551.js.

    Google Search Central

ShouldSaid at Search Central LiveDEV-REN-07

Keep prices, offer counts and other commercial data identical in raw and rendered HTML

Why

Said at Search Central Live by a community speaker: offer counts, prices and promotions sometimes differ between the server HTML and the rendered page, for example a different newsletter discount in each, so bots and users can get different versions. Google's documentation does not cover this case specifically.

How

Generate both versions from the same data source with the same filters, and treat any difference in prices, availability, offer counts, links or meta tags between raw and rendered HTML as a bug.

Test

On every release, diff the raw HTML (curl) and the rendered HTML (headless Chrome or URL Inspection) of each key template for prices, offer counts, links and meta tags.

Evidence · 3 claims
  • StageNot in docsD2-C147

    Offer counts, discounts and other data should always be the same in the raw HTML and the rendered page, a community speaker advised.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideD2-C137

    The first rendering blind spot a community speaker showed was content changes: text, images and recommendations that differ between the raw HTML and the rendered page, illustrated by a shop page whose rendered version had a whole extra section.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C148

    Diff the raw and the rendered HTML of each key template for prices, offer counts, links and meta tags, and treat any difference in commercial data as a bug, because bots and users may get different versions.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-REN-08

Feature-detect permission-based APIs and WebGL, and render the content without them

Why

Browser APIs that need user permission, such as location and payments, are not supported when Google renders a page, and WebGL effects are not applied. Google recommends feature detection with polyfills or differential serving, and prerendering on the server for content that depends on WebGL.

How

Render indexable content first; ask for permissions only after a user action and keep a complete default (for example the full store list) if the request is declined. Wrap WebGL and other optional APIs in feature detection with a plain HTML fallback. Add polyfills for missing APIs the content needs, knowing that some features cannot be polyfilled.

Test

In URL Inspection the screenshot and rendered HTML show the main content, and the JavaScript console messages of the live test show no uncaught errors from missing APIs.

Code · Feature detection with a fallback for unsupported or permission-based APIs

Google's renderer does not grant permission prompts (location, payments, camera) and does not run WebGL well. Render the indexable content first and treat such features as optional enhancements behind feature detection.

HTML
<main>
  <h1>Our stores</h1>
  <ul id="stores">
    <li>Barcelona, Carrer Example 1</li>
    <li>Madrid, Calle Example 2</li>
  </ul>
  <button type="button" id="near-me" hidden>Sort by distance</button>
  <canvas class="hero-effect" width="1200" height="400"></canvas>
</main>

<script>
  // Location: optional, only after a user click; the full list is already in the page
  if ('geolocation' in navigator) {
    const button = document.getElementById('near-me');
    button.hidden = false;
    button.addEventListener('click', () => {
      navigator.geolocation.getCurrentPosition(
        (position) => sortStoresByDistance(position.coords),
        () => {} // declined or unavailable: keep the default order
      );
    });
  }

  function sortStoresByDistance(coords) {
    console.log('Sort stores near', coords.latitude, coords.longitude);
  }

  // WebGL: decoration only; without it the server-rendered text and images stay as they are
  const canvas = document.querySelector('canvas.hero-effect');
  const gl = canvas.getContext('webgl');
  if (gl) {
    gl.clearColor(0.1, 0.3, 0.5, 1.0);
    gl.clear(gl.COLOR_BUFFER_BIT);
  } else {
    canvas.remove();
  }
</script>
Evidence · 5 claims · 2 Google pages
  • StageConfirmed by docsD2-C269

    Browser APIs that need user permission, such as payments and location, are not supported when Google renders a page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C277

    WebGL does not work well in Google's rendering: a WebGL shader that makes a page look like it is underwater will not be applied when Google renders the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C278

    For effects that need WebGL, which Googlebot does not support, Google's guide to fixing JavaScript problems suggests skipping the effect or prerendering it with server-side rendering so that the content is accessible to Googlebot.

    Google Search Central

  • DocsSourceD2-C280

    Google's JavaScript SEO basics recommends differential serving and polyfills when feature detection finds a missing browser API, and warns that some browser features cannot be polyfilled.

    Google Search Central

  • AnalysisD2-C282

    Wrap permission-based features (location, payments, camera) and WebGL effects in feature detection with a fallback, so the main content renders when the API is missing or the permission is declined; never make indexable content wait for a permission prompt.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-REN-09

Monitor JavaScript console errors and Content Security Policy violations on key templates

Why

Console messages reveal problems that remove content, such as uncaught exceptions and Content Security Policy violations that block a script or a video. An unhandled exception in an embedded third-party or AI-generated app can break the host page as well.

How

Collect console errors in automated browser tests and in a rendering crawl of one URL per template. List in the CSP every host the page needs (scripts, APIs, media, fonts) and test CSP changes in staging, for example with Content-Security-Policy-Report-Only first. Isolate third-party widgets so their errors cannot stop the main content from rendering.

Test

URL Inspection live test, JavaScript console messages: no errors on key templates. Browser tests fail on new console errors or securitypolicyviolation events.

Evidence · 4 claims · 2 Google pages
  • StageConsistent with docsD2-C144

    Browser console messages reveal further problems on a page, Content Security Policy violations among them; few teams analyse console messages at scale, but they should, a community speaker said.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C155

    Content Security Policy lowers the risk of cross-site scripting and clickjacking by defining trusted hosts; a resource from a host that is not trusted is not used.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C222

    Pages that embed third-party or AI-generated apps can share their fate, since an unhandled exception in the embedded code can break the host page; check the rendered HTML and JavaScript console output of such pages in Google's testing tools to make sure the host's own content still renders.

    Ibrahim Anjro (author)

  • AnalysisD2-C158

    When something is missing from a rendered page, check both robots.txt and the Content Security Policy: the URL Inspection live test shows the page resources, the JavaScript console output and a screenshot of the rendered page.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-REN-10

Keep indexable content out of iframes, and check web components in the rendered HTML

Why

Google generally tries to index iframe content as part of the embedding page, but that is not guaranteed because both are normal pages of their own. Google flattens shadow DOM and light DOM when it renders web components, and content that does not appear in the rendered HTML cannot be indexed.

How

Put the main content in the page's own DOM and use iframes for third-party widgets you do not need indexed. For web components and micro-frontend libraries that patch browser APIs, verify the result in Google's tools rather than trusting a library's promise of indexability.

Test

The rendered HTML in URL Inspection or the Rich Results Test shows the component's and embedded content's text as part of the page.

Evidence · 4 claims · 4 Google pages
  • DocsSourceD2-C227

    Google says its systems generally try to index the content of a page embedded with an iframe as part of the page that embeds it, but this is not guaranteed because both pages are also normal HTML pages on their own.

    Google Search Central (SEO office hours transcript, December 2023), Search Central blog (21 January 2022)

  • DocsSourceD2-C235

    Google's documentation says Google supports web components and flattens shadow DOM and light DOM content when it renders a page, and that content not visible in the rendered HTML cannot be indexed.

    Google Search Central

  • AnalysisD2-C250

    Because Google flattens shadow DOM when it renders, content that Web Fragments places in a shadow root can in principle be indexed with the host page, but only if it appears in the rendered HTML; check that in URL Inspection rather than relying on a library's promise of indexability.

    Ibrahim Anjro (author)

  • AnalysisD2-C240

    Because Web Fragments monkey-patches core browser APIs such as document, history and location, test pages that use it in Google's rendering (the rendered HTML in Search Console's URL Inspection tool or the Rich Results Test), not only in a normal desktop browser.

    Ibrahim Anjro (author)

Indexing and snippet controls 13

robots.txt controls crawling; the robots meta tag and the X-Robots-Tag header control indexing and how a page appears, including how much of it AI Overviews and AI Mode may use. Search Console and the Google-Extended token add controls for generative AI.

MustDocumentedDEV-IDX-01

Use noindex to keep a page out of Search, and leave that URL crawlable

Why

robots.txt does not control indexing: a disallowed URL can still appear in results without its content when other pages link to it, and Google never sees a noindex on a page it may not fetch. Google said it does not treat a disallow as noindex because important sites sometimes disallow their most important pages by mistake, and robots meta rules only work if robots.txt lets Google fetch the page.

How

For admin, thin, temporary or internal-search pages that may be crawled but should not be indexed, send <meta name="robots" content="noindex"> or X-Robots-Tag: noindex in the server response and do not disallow the URL. Use robots.txt only for URLs that should never be fetched. Use name="googlebot" only for Google-specific rules. To de-index URLs that are already disallowed, add the noindex first (or in the same release), then remove the disallow so Google can see it.

Test

For each noindexed URL: URL Inspection shows Crawl allowed: Yes, curl shows the noindex, and URL Inspection and the Page indexing report list it as excluded by the noindex tag.

Code · Robots meta tags for common page types

Put robots rules in the <head> of the HTML the server sends, one variant per page type below; never add, change or remove them with JavaScript. A noindex only works if the URL is not disallowed in robots.txt, because Google has to fetch the page to see it.

HTML
<!-- Indexable templates: allow full-length snippets and large image and video previews -->
<meta name="robots" content="max-snippet:-1, max-image-preview:large, max-video-preview:-1">

<!-- Keep a page out of Search (internal admin, thin or temporary pages) -->
<meta name="robots" content="noindex">

<!-- The same rule for Google only -->
<meta name="googlebot" content="noindex">

<!-- Time-limited page (event, offer, job ad): drop it from results after the end date -->
<meta name="robots" content="unavailable_after: 2026-12-31T23:59:59+01:00">
Evidence · 12 claims · 3 Google pages
  • DocsSourceD2-C022

    Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.

    Google Search Central

  • DocsSourceD2-C069

    Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.

    Google Search Central

  • SlideConfirmed by docsD2-C021

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C020

    Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C061

    The noindex robots rule tells Google not to show the page in search results.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD1-C110

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

    Ibrahim Anjro (author)

  • AnalysisD2-C023

    Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD2-C850

    John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C851

    When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C846

    Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C844

    Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C845

    Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no figure given).

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

AvoidDocumentedDEV-IDX-02

Do not use noindex, nofollow or crawl-delay to save crawl budget

Why

noindex consumes crawl budget because Google must fetch the page to see it; nofollow is a hint, and the linked URLs may still be crawled through other links or sitemaps; crawl-delay is not processed at all. Google advises against noindex for crawl budget and against using robots.txt to reallocate it temporarily.

How

To stop Google fetching URLs on your own site, disallow them in robots.txt; return 404 or 410 for removed pages; consolidate duplicates with redirects or canonicals. Keep crawl rules stable instead of switching them to steer Googlebot around.

Test

Server logs show few or no Googlebot fetches of parameter and action URLs; robots.txt has no crawl-delay meant for Google; Crawl stats shows fewer fetches of low-value URLs after fixes.

Evidence · 6 claims · 3 Google pages
  • SlideConfirmed by docsD1-C106

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C109

    Google advises against using noindex to save crawl budget and against using robots.txt to temporarily reallocate budget. Use robots.txt only for pages you never want crawled, and 404 or 410 for removed pages.

    Google

  • DocsSourceD1-C127

    Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

    Google Search Central, Search Central blog (10 September 2019)

  • SlideConfirmed by docsD1-C107

    The nofollow rule can still consume crawl budget: Google does not crawl through the nofollow link itself, but it still crawls the linked page when it finds that page through other links.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideConfirmed by docsD1-C108

    The non-standard crawl-delay rule is not processed by Googlebot, so it does not save any crawl budget.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideConsistent with docsD1-C105

    URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

ShouldDocumentedDEV-IDX-03

Send robots rules for PDFs, images and other non-HTML files in the X-Robots-Tag header

Why

Files without an HTML head can carry robots rules only in the HTTP response. X-Robots-Tag accepts the same rules as the meta tag and, like the tag, is seen only when the URL is crawled.

How

Add X-Robots-Tag in the web server or CDN for the file types or folders that should stay out of Search (internal PDFs, exports, private images), and do not also disallow those URLs in robots.txt. Use the user-agent form (X-Robots-Tag: googlebot: noindex) only for crawler-specific rules.

Test

curl -I on a sample file shows the X-Robots-Tag header; URL Inspection of the file reports it excluded by noindex.

Code · X-Robots-Tag header for PDFs, images and other non-HTML files

Files that cannot carry a meta tag get their robots rules as an HTTP response header. The same rules as the meta tag apply (noindex, nosnippet, max-snippet...), and the URL must stay crawlable for Google to see the header.

HTTP
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex
nginx
# Inside the server block: office documents out of the index.
# An add_header in a location replaces every add_header set at server level for those responses:
# repeat your HSTS and CSP lines in each such location (or use the headers-more module).
location ~* \.(pdf|docx?|xlsx?)$ {
    add_header X-Robots-Tag "noindex" always;
    add_header Strict-Transport-Security "max-age=31536000" always;  # repeated from the server block
}

# Images that must not appear in Google Images (they still display on your pages)
location ^~ /internal-images/ {
    add_header X-Robots-Tag "noindex" always;
    add_header Strict-Transport-Security "max-age=31536000" always;  # repeated from the server block
}
Apache
<FilesMatch "\.(pdf|docx?|xlsx?)$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>
Evidence · 3 claims · 2 Google pages
  • DocsSourceD2-C069

    Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.

    Google Search Central

  • DocsSourceD2-C529

    Google's documented ways to keep a site's images out of search results are a robots.txt disallow rule (for example for Googlebot-Image) or a noindex X-Robots-Tag HTTP header, with the Removals tool for emergencies.

    Google Search Central

  • DocsSourceD2-C095

    Google's Search documentation gives notranslate as a rule for a robots or googlebot meta tag or an X-Robots-Tag header, and says the rule opts a page out of all translation features in Google Search.

    Google Search Central

ShouldDocumentedDEV-IDX-04

Allow full snippets and large previews with max-snippet:-1 and max-image-preview:large

Why

Robots meta rules can make a page more visible than having none. max-image-preview:large allows the largest image previews, which matters mainly in Discover, and max-snippet:-1 removes the length limit so Google picks the length it finds most effective. Google's closing slide on robots meta tags recommended exactly these two rules for maximum visibility.

How

Add <meta name="robots" content="max-snippet:-1, max-image-preview:large"> (optionally with max-video-preview:-1) to every indexable template unless licensing requires limits, and check that no CMS, SEO plug-in or CDN feature adds lower limits by default. Provide large, high-quality images, at least 1,200 px wide for Discover.

Test

curl the HTML of each template and check the robots meta content; the Discover performance report in Search Console shows impressions for article templates.

Code · Robots meta tags for common page types

Put robots rules in the <head> of the HTML the server sends, one variant per page type below; never add, change or remove them with JavaScript. A noindex only works if the URL is not disallowed in robots.txt, because Google has to fetch the page to see it.

HTML
<!-- Indexable templates: allow full-length snippets and large image and video previews -->
<meta name="robots" content="max-snippet:-1, max-image-preview:large, max-video-preview:-1">

<!-- Keep a page out of Search (internal admin, thin or temporary pages) -->
<meta name="robots" content="noindex">

<!-- The same rule for Google only -->
<meta name="googlebot" content="noindex">

<!-- Time-limited page (event, offer, job ad): drop it from results after the end date -->
<meta name="robots" content="unavailable_after: 2026-12-31T23:59:59+01:00">
Evidence · 8 claims · 2 Google pages
  • SlideConfirmed by docsD2-C087

    The max-image-preview rule sets the maximum size of a page's image previews: none shows no preview, standard a default-sized one and large the largest possible, for example in Discover.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C082

    The max-snippet:[number] rule limits a page's search result snippet to that number of characters (100 characters, for example), and max-snippet:0 shows no snippet, the same as nosnippet.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C086

    Google's robots meta tag specification says Google chooses the snippet length when no max-snippet rule is set, and with max-snippet:-1 chooses the length it believes most effective; it does not say that -1 produces longer snippets.

    Google Search Central

  • SlideConsistent with docsD2-C104

    To be as visible as possible in Google, John Mueller's closing slide recommended the robots rules max-image-preview:large and max-snippet:-1.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C053

    John Mueller opened with a true-or-false quiz slide asking whether robots meta tags can make a page more visible in Search results than having none, and later answered that the statement is true.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C105

    Set max-snippet:-1 and max-image-preview:large on every indexable template unless licensing requires otherwise, and check that no CMS, plug-in or CDN setting adds lower snippet or image preview limits by default.

    Ibrahim Anjro (author)

  • SlideConfirmed by docsD3-C224

    Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels and a 16x9 aspect ratio, enabled by the max-image-preview:large setting.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD2-C936

    Setting the max-image-preview robots meta tag to large can make content perform surprisingly well in Discover, Gary Illyes said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

AvoidDocumentedDEV-IDX-05

Do not add nosnippet or a short max-snippet unless you accept losing snippets and AI feature use

Why

nosnippet removes the text snippet and video preview and also stops the page's content being used as a direct input for AI Overviews and AI Mode; max-snippet limits that use too. A page must be eligible for a snippet to be a supporting link in AI features, and Google said its AI answers are built from snippets. Said at Search Central Live: a value too short for a useful snippet may lead to no snippet at all.

How

Keep nosnippet and max-snippet limits off by default. To hide one detail, such as a phone number or a member price, use data-nosnippet on that element instead (DEV-IDX-07). If licensing requires limits, set them per template deliberately and record the trade-off.

Test

Search the codebase, CMS settings and CDN transforms for nosnippet and max-snippet; curl key templates to confirm none is present unless intended.

Evidence · 6 claims · 3 Google pages
  • DocsSourceD2-C728

    Google's robots meta tag specification says the nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI Mode, and max-snippet also limits how much of it may be used that way.

    Google Search Central

  • SlideConfirmed by docsD2-C072

    The nosnippet rule also prevents a page's content from being used as a direct input for AI Overviews and AI Mode in Search.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C074

    Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements.

    Google Search Central

  • SlideConfirmed by docsD2-C070

    The nosnippet rule stops Google from showing a text snippet or video preview for a page in search results, while the page's title is still shown.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C073

    John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C084

    A max-snippet value too short for a useful snippet may lead Google to show no snippet at all, John Mueller said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-IDX-06

Prefer rel attributes on individual links to a page-level nofollow

Why

The page-level nofollow rule tells search engines not to pass signals through any link on the page, a very broad rule that, as Google put it at Search Central Live, makes the page stand on its own, while an explicit robots all rule does nothing. Google recommends rel attributes on individual links so a site chooses which links to qualify; its documentation does not forbid a page-level nofollow.

How

Remove page-level nofollow and redundant all or index, follow values from template and plug-in robots meta tags; where a template uses none, replace it with noindex if the page must stay out of Search (none also means nofollow), never simply delete it. Mark individual links instead: rel="sponsored" for paid links, rel="ugc" for user-generated content and rel="nofollow" for other links you do not vouch for. nofollow is not an access control; use robots.txt to keep crawlers out (DEV-IDX-02).

Test

A crawl lists no page with a page-level nofollow (none included) that lacks a recorded reason, and every page that must stay out of Search still carries noindex; outbound links in comments, ads and affiliate modules carry the right rel values.

Evidence · 6 claims · 3 Google pages
  • StageConsistent with docsD2-C063

    The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page, which John Mueller called a weird and very broad rule that makes the page stand on its own.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C064

    John Mueller recommends rel=nofollow on individual links instead of the page-level nofollow robots rule, so a site can choose which links are useful and which are not.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C059

    The robots meta rule all is the default value and does nothing: it places no restrictions on indexing the page or following its links, the same as having no robots meta tag, and it is not an instruction that search engines must index the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD1-C127

    Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

    Google Search Central, Search Central blog (10 September 2019)

  • AnalysisD2-C065

    Audit robots meta tags set by templates and plug-ins: an explicit all rule does nothing and can go, while a page-level nofollow strips link signals from every link on the page and is better replaced by qualifying only the specific links that need it.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD2-C068

    The robots rule none is equivalent to noindex plus nofollow, so a page that carries it will not show up in Search, provided Google can see the tag.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MayDocumentedDEV-IDX-07

Use data-nosnippet on a span, div or section to keep one detail out of snippets

Why

data-nosnippet keeps a passage such as a phone number or a member price out of snippets while the page can still be found for it. It works on only a handful of elements, may be read before and after rendering, and broken HTML can widen it: an unclosed element marked data-nosnippet can hide the rest of the page from snippets.

How

Wrap the passage in <span data-nosnippet>, <div data-nosnippet> or <section data-nosnippet> in the server HTML, never add or remove the attribute with JavaScript, and validate the HTML of templates that use it so every element is closed.

Test

An HTML validator (W3C Nu or html-validate in CI) reports no unclosed elements on templates that use data-nosnippet; after recrawl, the page's snippet avoids the marked text.

Code · data-nosnippet for one detail instead of the whole snippet

data-nosnippet keeps a specific passage out of search snippets (and so out of the snippet-based inputs of AI features) while the page can still rank for it. It works only on span, div and section, must be in the server HTML (not toggled by JavaScript) and needs valid, closed markup: an unclosed element can hide the rest of the page.

HTML
<p>
  Call our workshop on <span data-nosnippet>+1 555-0100</span> for a repair quote.
</p>

<section data-nosnippet>
  <h2>Member prices</h2>
  <p>Log in to see your personal discount.</p>
</section>
Evidence · 6 claims · 1 Google page
  • StageConfirmed by docsD2-C077

    The data-nosnippet attribute can be used on only a handful of HTML elements.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C078

    Broken HTML can widen data-nosnippet: if a div marked data-nosnippet is not closed properly, the rest of the page can be blocked from the snippet as well.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C079

    Google's robots meta tag specification says data-nosnippet may be extracted both before and after rendering, so the attribute should not be added to or removed from existing elements with JavaScript.

    Google Search Central

  • StageConsistent with docsD2-C076

    John Mueller said data-nosnippet is rarely needed but lets a site keep a specific piece of text, such as a business phone number, out of the snippet, so the page can still be found for it while searchers have to visit the page to see it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C080

    John Mueller said valid HTML is technically not a ranking factor but does matter for controls such as data-nosnippet.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C081

    To hide one detail, such as a phone number, a price or a direct answer, instead of the whole snippet, wrap it in data-nosnippet on a span, div or section in the server HTML and validate the HTML so that an unclosed element cannot hide the rest of the page from snippets.

    Ibrahim Anjro (author)

MayDocumentedDEV-IDX-08

Use unavailable_after on pages with a known end date

Why

The unavailable_after rule drops a page from results after a set date and time, and Googlebot considerably reduces how often it crawls the URL after that date. Index selection was described as checking the rule and dropping the page once the date is reached.

How

Add <meta name="robots" content="unavailable_after: 2026-12-31T23:59:59+01:00"> (ISO 8601, RFC 822 or RFC 850 dates) to event pages, time-limited offers and job ads when they are published. After the end date, either keep the page as an archive without the rule or return 404 or 410 if it is gone.

Test

curl shows the rule with a valid date; after the date, URL Inspection reports the page as not indexed.

Code · Robots meta tags for common page types

Put robots rules in the <head> of the HTML the server sends, one variant per page type below; never add, change or remove them with JavaScript. A noindex only works if the URL is not disallowed in robots.txt, because Google has to fetch the page to see it.

HTML
<!-- Indexable templates: allow full-length snippets and large image and video previews -->
<meta name="robots" content="max-snippet:-1, max-image-preview:large, max-video-preview:-1">

<!-- Keep a page out of Search (internal admin, thin or temporary pages) -->
<meta name="robots" content="noindex">

<!-- The same rule for Google only -->
<meta name="googlebot" content="noindex">

<!-- Time-limited page (event, offer, job ad): drop it from results after the end date -->
<meta name="robots" content="unavailable_after: 2026-12-31T23:59:59+01:00">
Evidence · 4 claims · 1 Google page
  • StageConfirmed by docsD2-C100

    The unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C101

    Google's robots meta tag specification says Googlebot considerably decreases the crawl rate of a URL after the date and time set in its unavailable_after rule.

    Google Search Central

  • StageConsistent with docsD2-C698

    Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C699

    Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-IDX-09

Leave notranslate and noimageindex out of templates unless translation or image indexing must be prevented

Why

notranslate, as a robots or googlebot rule or in X-Robots-Tag, opts a page out of all translation features in Google Search, and the meta tag named google with notranslate also turns off Chrome's automatic translation for visitors who cannot read the language. noimageindex stops Google indexing every image on the page; said at Search Central Live, it also affects the page's videos, because Google has to index a video's thumbnail, which is an image. Both are often inherited from templates without a reason.

How

Remove notranslate and noimageindex from global templates. Where a specific passage must not be translated, the standard HTML translate="no" attribute on that element is narrower than a page-wide rule. To keep specific images out of Search, use DEV-IMG-06.

Test

A crawl lists every robots, googlebot or google meta tag and X-Robots-Tag header containing notranslate or noimageindex, and each one has a documented reason.

Evidence · 8 claims · 4 Google pages
  • StageConfirmed by docsD2-C093

    The notranslate robots rule tells Google not to offer a translation of the page in search results.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C095

    Google's Search documentation gives notranslate as a rule for a robots or googlebot meta tag or an X-Robots-Tag header, and says the rule opts a page out of all translation features in Google Search.

    Google Search Central

  • StageConsistent with docsD2-C094

    John Mueller said notranslate can also be set in a meta tag named google, and that in that form it also turns off Chrome's automatic translation of the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C096

    A 2008 Search Central blog post says a meta tag named google with the value notranslate stops Google Translate from translating any of a page's content.

    Search Central blog (14 October 2008)

  • StageConfirmed by docsD2-C099

    The noimageindex rule tells Google not to index any of the images on the page, and John Mueller said he could not see why a site would want that.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C098

    International sites should remove notranslate from their templates unless translation must be prevented: the robots or googlebot form switches off Google Search's translation features, and John Mueller said the form named google also blocks Chrome's translation for visitors who do not read the language.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD2-C934

    The noimageindex robots meta tag tells Google not to index the images on the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C935

    Gary Illyes said noimageindex also affects videos on the page, because Google has to index a video's thumbnail, which is an image.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MayDocumentedDEV-IDX-10

Use the Search generative AI setting in Search Console only to take a property out of AI Overviews and AI Mode

Why

This property-level setting (include, exclude or inherit from parent) controls AI Overviews, AI Mode and generative AI features in Discover without affecting web rankings. Excluding removes both links to the site and the use of its content for grounding, so the site gets no traffic or impressions from those features; include is the default and changes generally apply within one to two days.

How

Leave most properties on include or inherit. If a business decision requires exclusion, set it on a child property such as a folder first and measure the effect. The setting lives in Search Console, not in code, so record it in the launch checklist. Use nosnippet only if you also want out of regular snippets.

Test

Search Console settings show the intended state for each property; the Generative AI performance report shows impressions for included properties.

Evidence · 6 claims · 1 Google page
  • SlideConfirmed by docsD1-C029

    Search Console has a property setting called Search generative AI that gives direct control over AI Overviews and AI Mode without affecting web rankings. Its states are inherit from parent, include and exclude.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C030

    The Search generative AI control covers AI Overviews, AI Mode and generative AI features in Discover. Excluding removes both links to the site and use of its content to ground answers.

    Google Search Console Help

  • DocsSourceD1-C031

    Google says the control is not used as a ranking or inclusion signal for regular results.

    Google Search Console Help

  • DocsSourceD1-C033

    Changes to the control generally apply within 1–2 days, though some content takes longer. It works at property level only, and a child property inherits from its closest parent that changed the setting.

    Google Search Console Help

  • SlideConfirmed by docsD2-C103

    Setting the Search generative AI control in Search Console to exclude keeps a site's links and content out of Search generative AI features such as AI Overviews and AI Mode, so the site gets no traffic or impressions from them; include is the default.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD1-C034

    Lead-generation, local-service and e-commerce sites should normally stay included, because AI answers cite and link sources. Sites selling paywalled or licensed content should test exclusion on a child property first. Use nosnippet only if you also want out of regular snippets.

    Ibrahim Anjro (author)

MayDocumentedDEV-IDX-11

Use the Google-Extended robots.txt token to control Gemini training and grounding

Why

Google-Extended is a robots.txt control token, not a crawler: it decides whether crawled content may be used to train future Gemini models and to ground Gemini Apps and Vertex AI, and it does not affect inclusion or ranking in Google Search. The Search generative AI setting does not affect training, and user-triggered fetchers, which fetch a page because a user asked, generally ignore robots.txt.

How

To opt out, add a User-agent: Google-Extended group with Disallow: / for the whole site, or Disallow lines for specific directories, to robots.txt; never block Googlebot for this purpose. Google's robots.txt talk confirmed that Google respects that policy for Gemini model training. Said at Search Central Live: rendering for Gemini training reuses the rendering done for Search. Author's view: so no separate rendering work is needed for Gemini.

Test

Google's open-source robots.txt parser (github.com/google/robotstxt), run with the Google-Extended token, returns disallowed for the intended paths, and run with the Googlebot token returns allowed for the same URLs; URL Inspection shows them as Crawl allowed.

Code · robots.txt with crawl controls and a rendering carve-out

Disallow only URLs that should never be crawled (internal search, cart and checkout actions, filter parameters), keep every script, style and API path that pages need for rendering crawlable, and list the sitemap. Google reads only user-agent, allow, disallow and sitemap; each host (www, api, cdn) needs its own file at its root. Robots.txt is public, so never list secret paths in it.

Text
# https://www.example.com/robots.txt
User-agent: *
# Internal search results and cart/checkout actions (adapt to your URL patterns)
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
# (Shops in Google Shopping: give Storebot-Google its own group that leaves cart and checkout open.
#  A named group replaces this * group for that crawler, so repeat in it every rule that should still apply.)
# Filter and sort parameters of faceted navigation, as the first parameter (?) or a later one (&);
# /*?*size= would also block ?pagesize= and /*?*color= would block ?bgcolor=
Disallow: /*?color=
Disallow: /*&color=
Disallow: /*?size=
Disallow: /*&size=
Disallow: /*?sort=
Disallow: /*&sort=
# API: blocked in general, but the endpoints pages render from stay crawlable
# (the longer, more specific Allow rule wins)
Disallow: /api/
Allow: /api/products/
Allow: /api/reviews/
# Never disallow /static/, /assets/ or other JS and CSS folders

# Optional: keep content out of Gemini training and grounding (no effect on Google Search)
User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
Evidence · 8 claims · 3 Google pages
  • DocsSourceD1-C086

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

    Google

  • DocsSourceD1-C032

    The control does not affect AI training. Training of the models is limited with Google-Extended instead.

    Google Search Console Help

  • DocsSourceD2-C121

    Google's crawler documentation defines grounding in Gemini Apps and in Grounding with Google Search on Vertex AI as providing content from the Google Search index to the model at prompt time, and sites manage whether their content is used for it with the Google-Extended robots.txt token.

    Google

  • DocsSourceD2-C122

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    Google

  • StageNot in docsD2-C118

    To train Gemini models, Google renders every page just as it does for Search, so a page that renders correctly for Search also works for Gemini training, provided the site allows its content to be used for training.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C119

    If Gemini training reuses the rendering done for Search, a site needs no separate rendering work for Gemini; whether its rendered content is used for training is decided with the Google-Extended token in robots.txt, not by rendering choices.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD1-C485

    To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C522

    Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

ShouldDocumentedDEV-IDX-12

On sites with explicit content, keep it on its own subdomain or domain and label it with rating markup

Why

Google's guidelines for sites with explicit content recommend grouping explicit pages on a separate domain or subdomain; otherwise Google's systems may judge the whole site explicit and filter all of it when SafeSearch is on. Explicit pages are marked with rating markup set to adult, as a meta tag or an HTTP response header. Author's view: SafeSearch signals are calculated at indexing from the whole page and its links, so a fix counts only after Google recrawls and reprocesses the pages.

How

Only for sites that host explicit material: serve it from a dedicated subdomain or domain, add <meta name="rating" content="adult"> (or the same value as an HTTP response header) to every explicit page, and keep explicit images, links and teasers off general templates. After moving or relabelling pages, request recrawling of the most important URLs.

Test

A crawl finds every explicit URL on the dedicated host with the rating markup, and no general template links to or embeds explicit content.

Evidence · 4 claims · 1 Google page
  • DocsSourceD2-C664

    Google's guidelines for sites with explicit content recommend grouping explicit pages on a separate domain or subdomain; otherwise Google's systems may judge the whole site explicit and filter all of it when SafeSearch is on.

    Google Search Central

  • DocsSourceD2-C837

    Google's guidelines for sites with explicit content say to mark each explicit page with rating markup set to adult, either as a meta tag named rating or as an HTTP response header.

    Google Search Central

  • AnalysisD2-C665

    Because explicit content on its own does not lower a page's chance of being indexed, the practical risk for a site that mixes explicit and general content is SafeSearch classification, which looks at the whole page and its links; keep explicit pages on a separate domain or subdomain, as Google advises.

    Ibrahim Anjro (author)

  • AnalysisD2-C677

    Language, country, SafeSearch and spam signals, and some freshness signals, are calculated when a page is indexed, so a fix such as correcting a page's language or removing content that triggers SafeSearch only counts once Google recrawls and reprocesses the page; request recrawling of the most important URLs after the fix.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-IDX-13

For an urgent removal, use the Removals tool in Search Console and also return 404 or noindex

Why

Google's Removals tool help says a temporary removal usually takes up to a day and lasts only about six months. Said at Search Central Live: removals take about two hours on average because Google pushes them to its serving system about every two hours, while a 404 or noindex usually drops a page from the serving index within one to three weeks after Google processes it. Only the permanent signal keeps the page out once the temporary removal expires.

How

For leaked, outdated or legally sensitive URLs, file a temporary removal in Search Console at once and, at the same time, return 404 or 410, add noindex, or put the content behind a login. Never rely on the Removals tool alone, and keep the URL crawlable so Google can see the 404 or noindex (DEV-IDX-01).

Test

The Removals report shows the request processed within a day; curl -I on the URL returns 404 or 410, or the page carries noindex; months later the URL is still absent from Search.

Evidence · 5 claims · 1 Google page
  • SlideConsistent with docsD3-C651

    Google estimated that a removal requested by the site owner in Search Console takes effect in about 2 hours on average, from minutes up to 24 hours.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C673

    Google's Removals tool help says a temporary removal request usually takes up to a day to process and lasts only about six months, so the content must also be removed permanently, for example with a 404 or 410, a noindex or a password.

    Google Search Console Help

  • StageConsistent with docsD3-C639

    After Google processes a page that now returns a 404 or a noindex, it usually removes the page from its serving index within one to three weeks, sometimes much sooner.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C652

    Search Console removals average about 2 hours because Google pushes removals out to its serving system about every 2 hours.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C653

    To hide a URL from Google Search urgently, use the Removals tool in Search Console (about 2 hours) and also add a noindex or a 404, which takes one to three weeks to drop the page from the index.

    Ibrahim Anjro (author)

Canonicals, redirects and duplicates 11

Google clusters duplicate pages, picks one representative URL (the canonical) and forwards signals to it. Redirects, rel=canonical, internal links and sitemaps tell it which URL you prefer: when they agree Google follows them, and when they conflict it decides on its own.

MustDocumentedDEV-CAN-01

Use permanent server-side redirects (301 or 308) for moved URLs, straight to the final URL

Why

Google trusts redirects very much when it clusters duplicates, and whether a redirect is permanent or temporary decides which URL becomes canonical: with a temporary redirect Google keeps showing the source URL. Google's crawlers follow up to 10 redirect hops, and long chains hurt crawling. Google follows JavaScript redirects only after rendering and may never see one if rendering fails; said at Search Central Live, even Google's own site consolidation used JavaScript only because it was the only option available to the person doing it, a fallback rather than a pattern to copy.

How

Answer every moved, merged or renamed URL with one 301 or 308 to the final URL, and collapse old chains in the redirect map into single hops. Update internal links, canonicals, hreflang and sitemaps to the new URL. Use 302 or 307 only for genuinely temporary moves, and avoid meta refresh and JavaScript redirects for permanent ones.

Test

curl -sIL on an old URL shows one 301 or 308 hop ending in 200; a crawler's redirect-chain report is empty; after recrawl, URL Inspection of the new URL shows it as the Google-selected canonical.

Code · Permanent redirects to one host, one protocol, one hop

Send http:// and the other host name to the preferred HTTPS host with a single 301, and send moved pages straight to their final URL (301 or 308). Keep migration redirects in place long term; avoid chains, meta refresh and JavaScript redirects for permanent moves.

nginx
# http:// on both host names -> https://www. in one hop
server {
    listen 80;
    server_name example.com www.example.com;
    return 301 https://www.example.com$request_uri;
}

# https:// on the bare domain -> https://www.
server {
    listen 443 ssl;
    server_name example.com;
    ssl_certificate     /etc/ssl/example.com.crt;
    ssl_certificate_key /etc/ssl/example.com.key;
    return 301 https://www.example.com$request_uri;
}

server {
    listen 443 ssl;
    http2 on;  # nginx 1.25.1+; older versions use "listen 443 ssl http2;"
    server_name www.example.com;
    ssl_certificate     /etc/ssl/www.example.com.crt;
    ssl_certificate_key /etc/ssl/www.example.com.key;

    # Moved pages: permanent and straight to the final URL
    location = /old-chairs/ {
        return 301 https://www.example.com/chairs/;
    }
    location ^~ /shop/chairs/ {
        rewrite ^/shop/chairs/(.*)$ https://www.example.com/chairs/$1 permanent;
    }

    root /var/www/example;
}
Evidence · 10 claims · 3 Google pages
  • DocsSourceD2-C013

    Google's redirect documentation says that with a temporary redirect, such as a 302, Google Search shows the source page in search results and does not use the redirect as a signal that the target should be canonical.

    Google Search Central

  • StageConsistent with docsD2-C361

    Google trusts redirects very much for clustering, because a redirect is a clear sign that there is one version of the content; Google keeps track of both URLs but stores only one copy of the content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C362

    Whether a redirect is permanent or temporary matters only for choosing the canonical, not for clustering the URLs together.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C363

    Use permanent redirects (301 or 308) for moves you want reflected in Search: a temporary redirect still groups the URLs, but it changes which URL Google is likely to pick as the canonical.

    Ibrahim Anjro (author)

  • DocsSourceD2-C357

    Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.

    Google Search Central

  • SlideConsistent with docsD2-C387

    Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD1-C349

    Google follows 3xx redirects, permanent (301, 308) and temporary (302, 307), to the new location, and whether content gets indexed depends on what the redirect target returns.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C370

    Google's crawlers follow up to 10 redirect hops by default (some products' crawlers have other limits); content served by the redirecting URL is ignored and the final target's content is processed instead.

    Google

  • StageNot in docsD1-C541

    A Google panelist said, as a fun fact, that JavaScript was used for the language-consolidation redirects in Google's own site migration, because it was the only option available to the person doing it (the recording does not make fully clear whether the JavaScript performed the redirects or built the mapping).

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C542

    Google's redirects guide says Google Search follows JavaScript location redirects only after rendering, may never see one if rendering fails, and should be used only when server-side or meta refresh redirects are impossible; its site move guide asks for server-side permanent redirects (301 or 308) where technically possible. The JavaScript in Google's own migration was a fallback ('my only option'), not a pattern to copy.

    Ibrahim Anjro (author)

MustDocumentedDEV-CAN-02

Serve one host over HTTPS with a valid certificate and redirect every other origin to it

Why

The www and non-www versions of a page are exact duplicates that Google clusters, keeping one. Whether a page loads properly counts in canonical selection: a broken certificate, failing JavaScript or a page that cannot load counts against a URL, and Google's slide also listed meta refresh and security.

How

Pick one origin (for example https://www.example.com), redirect every other combination of protocol and host name with a single 301, serve a valid TLS certificate on every host name that answers, including the redirecting ones, and use only the chosen origin in internal links, canonicals, hreflang and sitemaps. Enable HSTS once HTTPS works everywhere.

Test

curl -I on http://example.com, http://www.example.com and https://example.com each returns one 301 to the https://www. URL; a TLS check (openssl s_client or an online scanner) shows a valid certificate on every host name.

Code · Permanent redirects to one host, one protocol, one hop

Send http:// and the other host name to the preferred HTTPS host with a single 301, and send moved pages straight to their final URL (301 or 308). Keep migration redirects in place long term; avoid chains, meta refresh and JavaScript redirects for permanent moves.

nginx
# http:// on both host names -> https://www. in one hop
server {
    listen 80;
    server_name example.com www.example.com;
    return 301 https://www.example.com$request_uri;
}

# https:// on the bare domain -> https://www.
server {
    listen 443 ssl;
    server_name example.com;
    ssl_certificate     /etc/ssl/example.com.crt;
    ssl_certificate_key /etc/ssl/example.com.key;
    return 301 https://www.example.com$request_uri;
}

server {
    listen 443 ssl;
    http2 on;  # nginx 1.25.1+; older versions use "listen 443 ssl http2;"
    server_name www.example.com;
    ssl_certificate     /etc/ssl/www.example.com.crt;
    ssl_certificate_key /etc/ssl/www.example.com.key;

    # Moved pages: permanent and straight to the final URL
    location = /old-chairs/ {
        return 301 https://www.example.com/chairs/;
    }
    location ^~ /shop/chairs/ {
        rewrite ^/shop/chairs/(.*)$ https://www.example.com/chairs/$1 permanent;
    }

    root /var/www/example;
}
Evidence · 4 claims · 3 Google pages
  • StageConsistent with docsD2-C365

    Exact-match duplicates, such as the www and non-www versions of the same page, are clustered and Google keeps only one of them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C389

    Whether a page can load is really important for canonical selection: a broken certificate, failing JavaScript or a page that cannot be loaded counts against a URL, and the slide also listed meta refresh and security.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C387

    Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C396

    Google's closing suggestions on duplication: use redirects for site migrations, use HTTP result codes, don't block agents, check your rel=canonical links, use hreflang links to help Google localize, report weird canonicals in the forums, make secure pages that work, and keep canonical signals clear.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-CAN-03

Put exactly one absolute rel=canonical in the server-rendered head of every indexable page

Why

Google extracts rel=canonical from the head and uses it in deduplication and canonical selection, but the tag is very often wrong, so Google treats it as a strong hint rather than an order. With more than one canonical on a page, Google will likely ignore all of them.

How

Give every indexable page a self-referencing canonical, and point duplicates (tracking parameters, print views, sort orders) at the preferred URL. Use absolute URLs on the chosen origin. Output the tag from one layer only, making sure the CMS, an SEO plug-in and the front-end framework do not each add one, and put it in the server-sent head, not only in JavaScript-rendered HTML, with no invalid element (such as an img or iframe) before it, because Google stops reading the head at the first invalid element. For PDFs use an HTTP Link header.

Test

A crawl finds no indexable page with zero or several canonicals, or with a relative or placeholder value; raw and rendered HTML carry the same canonical; URL Inspection shows the user-declared and Google-selected canonical as equal on key templates.

Code · One rel=canonical per page, pointing straight at the final URL

Exactly one absolute canonical in the server-sent <head>, self-referencing on the preferred URL and identical on its duplicates (tracking parameters, sort orders, print views). The target must answer 200, be indexable, not redirect and be in the same language. Non-HTML files can declare it in an HTTP Link header.

HTML
<!-- On https://www.example.com/chairs/oak-dining-chair
     and on https://www.example.com/chairs/oak-dining-chair?utm_source=newsletter -->
<head>
  <link rel="canonical" href="https://www.example.com/chairs/oak-dining-chair">
</head>
HTTP
HTTP/1.1 200 OK
Content-Type: application/pdf
Link: <https://www.example.com/guides/chair-care.pdf>; rel="canonical"
Evidence · 8 claims · 3 Google pages
  • SlideConfirmed by docsD2-C032

    The rel=canonical link element is placed in the head section of the HTML and tells Google that one page is the representative of another.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C031

    Google extracts the rel=canonical link, through which site owners state their preferred canonical URL, and uses it in deduplication and in canonical selection.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C418

    Google's 2013 Search Central blog post on rel=canonical mistakes says to specify no more than one rel=canonical per page: when a page has more than one, Google will likely ignore all of them, and any benefit of a legitimate canonical is lost.

    Search Central blog (8 April 2013)

  • StageConsistent with docsD2-C417

    Multiple canonical declarations on one page conflict and are invalid, so the search engine applies its own heuristic and picks the canonical for the site owner; a page should declare only one canonical.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C379

    rel=canonical is probably the most common way site owners signal duplicates, but it is very often wrong, for example a tag whose value reads 'canonical target' instead of a real URL (the example is partly unclear in the recording), so Google can only sometimes trust it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C035

    Put rel=canonical and hreflang link elements in the head of the HTML the server sends, not only in JavaScript-rendered HTML, so Google can read them during HTML parsing without depending on rendering.

    Ibrahim Anjro (author)

  • DocsSourceD2-C831

    Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.

    Google Search Central

  • StageConfirmed by docsD2-C873

    A community speaker said that, under the canonical link specification, an improperly declared canonical tag can also be ignored completely by the application that processes it, not only replaced by that application's own heuristic.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-CAN-04

Point every canonical directly at a final URL that returns 200, is indexable and is linked

Why

Google can follow canonical chains but strongly recommends pointing every link at a single canonical page. Said at Search Central Live by a community speaker: a canonical loop leaves a group with no leader at all. The leader should carry no noindex, return no error, not be blocked in robots.txt and not redirect; a leader that visitors cannot reach through links asks Google to index a page humans cannot reach.

How

Resolve each canonical target to its final URL when pages are generated, and update canonicals to the new final URLs after migrations or host changes. Never canonicalise to a URL that redirects, returns non-200, carries noindex or is disallowed. Make the canonical target the URL that internal links use, fixing whichever of the two is wrong.

Test

Canonical audit in a crawler: group URLs by the final URL their canonicals and redirects lead to, and flag groups with more than one hop, a loop, several canonicals on one page, or a leader that is noindexed, blocked, non-200, redirecting or without internal links.

Evidence · 8 claims · 3 Google pages
  • DocsSourceD2-C420

    Google's 2009 Search Central blog post introducing rel=canonical says Google's algorithm is lenient and can follow canonical chains, but strongly recommends updating links to point to a single canonical page for optimal canonicalization.

    Search Central blog (12 February 2009)

  • StageConsistent with docsD2-C419

    Canonical chains, where a page's canonical target leads on to yet another URL, are named as improper use in the canonical link specification, according to a community speaker.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C438

    A community speaker advised fixing canonical chains by cleaning up the canonical graph so that every canonical link points directly to the canonical leader.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C432

    The canonical leader of a group should be indexable: it should carry no noindex robots directive, return no error status code and not be blocked in robots.txt.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C424

    Canonical loops, such as two HTML pages whose canonical tags point to each other, are a structural conflict: the group has no canonical leader, so its canonical information cannot be used.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C427

    When internal links point only to page A and the canonical leader is reached only through A's canonical link, the leader is reachable by machines but not by human visitors, a signal conflict that asks the search engine to index a page users cannot reach.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C416

    To audit canonicals at scale, resolve every crawled URL to the final URL its canonical links and redirects lead to, group URLs by that leader, and flag groups with more than one hop, a loop, several canonicals on one page, or a leader that is noindexed, blocked, non-200 or redirecting.

    Ibrahim Anjro (author)

  • AnalysisD2-C423

    After a migration or a subdomain change, update rel=canonical tags to the new final URLs: canonicals still pointing at old, redirecting URLs create chains through server-side redirects like the one in the real-world example shown.

    Ibrahim Anjro (author)

MustDocumentedDEV-CAN-05

Make all canonical signals agree: redirects, rel=canonical, internal links, sitemaps and hreflang

Why

When all signals point to the same URL Google follows the site owner, and when they point in different directions it cannot tell what the owner wants. Redirects and rel=canonical are strong signals, sitemap inclusion is a weak one, and combining them makes them more effective; links to duplicates are consolidated once the preferred URL becomes canonical.

How

Generate canonical tags, sitemap entries, hreflang URLs and internal links from one URL-building function so they cannot drift apart, and check every signal against the preferred URL before a migration or template change.

Test

A crawl finds zero internal links to non-canonical URLs, zero sitemap URLs that are not self-canonical and zero hreflang targets that redirect or canonicalise elsewhere; sampled URLs in URL Inspection show matching user-declared and Google-selected canonicals.

Evidence · 6 claims · 2 Google pages
  • StageConsistent with docsD2-C399

    When all canonical signals point to the same URL, Google follows what the site owner says; when they point in different directions, Google cannot tell what the owner wants.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C391

    Google's canonical guide ranks the ways to signal a preferred canonical by strength: redirects and rel=canonical annotations are strong signals, sitemap inclusion is a weak signal, and combining methods makes them more effective.

    Google Search Central

  • DocsSourceD2-C430

    Google's guide to specifying canonical URLs says a canonical helps consolidate signals for duplicate pages: links to a duplicate URL are consolidated with links to the preferred URL once the preferred URL becomes canonical.

    Google Search Central

  • SlideConsistent with docsD2-C390

    Clear site-owner signals about which URL should be canonical make a big difference to Google's choice; the speaker named redirects, listing only the preferred URL in sitemaps, and rel=canonical, which the speaker said also helps a bit.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C400

    Before a migration or template change, align every canonical signal for the preferred URL (redirects, rel=canonical, internal links, sitemap entries, hreflang and working HTTPS); Google says it follows the site owner only when the signals agree.

    Ibrahim Anjro (author)

  • AnalysisD2-C431

    A leader reachable only through a canonical and a leader with weak link equity share one fix: point internal links, sitemap entries and redirects at the URL chosen as canonical, so the leader is both reachable for users and the strongest page in its group.

    Ibrahim Anjro (author)

AvoidDocumentedDEV-CAN-06

Do not canonicalise one language version to another; connect them with hreflang

Why

A canonical from one language version to another asks Google to keep only one language in the index, so only that version would rank. Google's guide says that on pages using hreflang the canonical should be a page in the same language, or the best substitute language if none exists.

How

Every language or country version keeps a self-referencing canonical and lists its alternates with hreflang (DEV-INT-03). Check that no template falls back to the default-language canonical on translated pages.

Test

Canonical audit: no canonical group contains documents in more than one language, and every localized page's canonical equals its own URL.

Evidence · 4 claims · 2 Google pages
  • StageConsistent with docsD2-C433

    Canonical links between different language versions of a page are a misuse: if a search engine accepted such a canonical group, only one language version would rank.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C434

    Language versions of a page should be connected with hreflang annotations instead of canonical links.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C435

    Google's guide to specifying canonical URLs says that on pages using hreflang, the canonical should be a page in the same language, or the best possible substitute language if no canonical page exists in the same language.

    Google Search Central

  • AnalysisD2-C436

    Check that every canonical group contains a single document language: a group spanning languages means canonicals are joining versions that hreflang should connect, and each language version should keep its own self-referencing canonical.

    Ibrahim Anjro (author)

AvoidDocumentedDEV-CAN-07

Never leave staging, preview or mirror copies of the site publicly crawlable

Why

Google watches for canonical hijacking, where several domains try to be canonical for the same content, including a site owner's own copies such as a staging site. A crawlable copy competes with production for canonical selection and can win.

How

Protect staging, preview and QA environments with HTTP authentication or an IP allowlist, since a 401 or 403 cannot be indexed, and add a site-wide noindex header as a second layer. Never launch with staging's blanket noindex or Disallow: / still in place, and never point production canonicals or hreflang at a staging host.

Test

curl -I on the staging host returns 401 or 403 without credentials; the pre-launch check confirms production has no global noindex and no Disallow: /.

Evidence · 2 claims · 2 Google pages
  • SlideConsistent with docsD2-C388

    Google watches for canonical hijacking, where several domains try to be canonical for the same content, whether accidentally across a site owner's own domains (such as a staging copy) or through third-party domains, maliciously or not, and asks site owners to report cases it gets wrong.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C387

    Google named three considerations when picking a representative URL: hijacking across pages or sites, user experience (the page can load, meta refresh, security) and site-owner signals (redirects, rel=canonical, sitemaps).

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-CAN-08

Make pages that share a URL pattern differ in their main content

Why

Google clusters near matches and structurally similar pages, and fixing a wrong cluster comes down to making the pages sufficiently different; they can stay clustered for up to two weeks after a fix. Said at Search Central Live: when several URLs under one pattern, or city pages with similar stock, show the same content, Google may assume that every URL fitting the pattern shows that content without looking at the page.

How

Give location, variant and category pages distinct main content (local stock, staff, prices, addresses, a unique introduction), not just a different title. Return 404 for empty or invalid combinations and for URLs that no longer exist, and avoid generating many similar-looking URLs that lead to the same content.

Test

The Page indexing report's "Duplicate, Google chose different canonical than user" reason, grouped by URL pattern, stays flat; URL Inspection shows sampled pattern URLs as their own canonical.

Evidence · 6 claims · 2 Google pages
  • DocsSourceD2-C373

    Google's canonicalization troubleshooting guide says fixing a wrong duplicate cluster comes down to making the clustered pages sufficiently different; pages split out faster when the difference is clear and significant, and Google may keep pages in a duplicate cluster for up to two weeks after a fix.

    Google Search Central

  • SlideConsistent with docsD2-C371

    To avoid pattern-based deduplication, Google's speaker recommended not having many unrelated, similar-looking URLs that lead to the same content, and returning error pages for URLs that no longer exist so they are clearly unrelated.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C368

    Google sometimes picks a canonical that looks unrelated because it recognised a URL pattern: if /buy/fax, /buy/typewriter and /office-equipment show the same content, its systems may assume any /buy/ URL, even /buy/seo-service, shows that same office-equipment content without looking at the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C370

    City pages can trigger the same pattern-based deduplication: for a car dealer brand with branches in several cities and similar stock, Google's systems may decide the city name does not matter and canonicalise to one city's page; the slide asked whether a further city page such as /zurich/services would be treated the same way.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C369

    Google's slide on structurally similar content asked whether a new URL that fits a known duplicate pattern, such as /buy/seo-service, even needs to be crawled.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C372

    Make location and variant pages differ in their main content (local stock, staff, prices, addresses) and return 404 for empty or invalid combinations; otherwise every URL that fits the pattern can be folded into one canonical.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-CAN-09

Keep migration redirects in place long term and judge a move in Search Console, not with site: queries

Why

Google keeps the source of a redirect as an alternate name and may still show old URLs to people who search for the old domain or brand; its redirect guide calls this normal and says it fades over time. Removing redirects early breaks those visits and the consolidation of signals. Google's site move guide says to keep redirects generally at least one year so it can transfer all signals, and that small to medium sites take a few weeks for most pages to move. Said at Search Central Live: a site move takes one to three months on average and up to about a year, because Google's slowest signal is recalculated only about once a year (a partly uncertain passage of the recording). Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said sitemaps are not the main tool: decide first what matters from the business's perspective, for example whether to consolidate languages; listing the new URLs in a sitemap cannot hurt.

How

Map every old URL to its final new URL in one hop (DEV-CAN-10), keep the redirects for at least a year and ideally indefinitely, use Search Console's Change of address tool for domain moves, submit the new sitemaps, update internal links to the new URLs so pages do not rely on redirects alone, and follow the new property's indexing and traffic against a pre-migration baseline.

Test

Old URLs still answer 301 to the right targets months after launch; in Search Console the new property's indexed pages rise as the old property's fall.

Evidence · 14 claims · 2 Google pages
  • DocsSourceD2-C357

    Google's redirects guide says Google keeps track of both the source and the target of a redirect: one becomes the canonical, depending on signals such as whether the redirect is permanent or temporary, and the other becomes an alternate name that may appear in results when a query suggests the user trusts the old URL more. After a move to a new domain, old URLs may still show occasionally; the guide calls this normal.

    Google Search Central

  • SlideConsistent with docsD2-C354

    Alternate names are why a site: query for an old domain still shows the old domain's URLs after a site migration, which site owners often misread as a migration that is not working.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C355

    When someone searches for an old domain after a migration, Google shows the old URL as an alternate version because that is what was searched for, and relies on the redirect to take the user to the new site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C358

    After a migration, judge success by the new domain's indexing and traffic in Search Console rather than by a site: query on the old domain, and keep the old domain's redirects in place long term so searches for the old brand still reach the new site.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C351

    Google treats a site migration as deduplication across sites, in which the site owner says the old and the new domain are the same and that Google should pick the new domain, so deduplication also helps Google handle migrations.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD3-C643

    A site move can finish within a few weeks for a small site, takes one to three months on average, and in the worst case about a year.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C644

    Gary Illyes said a site move can take up to about a year in the worst case, because Google's slowest signal is recalculated only about once a year (some words of this passage are uncertain readings).

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C671

    Google's site move guide says to keep redirects for as long as possible, generally at least one year, so that Google can transfer all signals to the new URLs.

    Google Search Central

  • DocsSourceD3-C672

    Google's site move guide says that, as a general rule, a small to medium-sized site can take a few weeks for most pages to move and larger sites take longer, depending on the number of URLs and server speed.

    Google Search Central

  • AnalysisD3-C645

    Keep migration redirects in place for at least a year and judge a site move after one to three months, not days: Google's speaker said its slowest signal needs about a year to be recalculated (a partly uncertain passage of the recording), and Google's site move guide says to keep redirects generally at least one year.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD2-C884

    The technical steps of a community speaker's domain consolidation included submitting new sitemaps, filing a change of address in Search Console and updating internal links so the new pages did not rely on redirects alone.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C912

    Google's site move guide says to submit a Change of Address in Search Console when moving from one domain or subdomain to another, to submit the new sitemap, and to change internal links on the new site from the old URLs to the new ones.

    Google Search Central

  • StageConsistent with docsD2-C888

    After a domain consolidation went live, a community speaker's team checked Search Console every day, ran full site crawls and compared the numbers with the pre-migration baseline for months.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD1-C540

    Asked how to plan a migration that does not leave many URLs unindexed, a Google panelist said the answer is probably not sitemaps: decide what matters from the business's perspective (for example whether to consolidate languages); listing the new URLs in a sitemap is probably a good idea and cannot hurt, but it is not the main tool.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-CAN-10

Redirect an old URL only to a page with the same intent, and return 404 or 410 when there is none

Why

Google's site move guide warns against redirecting many old URLs to one irrelevant destination such as the home page, which can confuse users and might be treated as a soft 404, and says to return 404 or 410 for content that is not moved. Community speakers who ran migrations at the event built them around one redirect map that gives every old URL an approved destination, redirected only where a new page served the same intent and never sent removed URLs to the home page, even those still bringing traffic. A 200 on every new URL proves only that the URLs work, not that the content users came for is still there.

How

Before launch, freeze the inventory (CMS export, crawl, sitemaps, search data and logs) and approve a redirect map with one row per old URL: a same-intent target, or 404 or 410. Where two language or market versions duplicate each other, redirect both old URLs to one shared target. After launch, test the live site against the map.

Test

A script requests every old URL in the map and checks it returns one 301 to the mapped target, or the mapped 404 or 410; no old URL other than the old home page redirects to the new home page; the Page indexing report shows no rise in soft 404s on redirect targets.

Evidence · 7 claims · 2 Google pages
  • DocsSourceD2-C911

    Google's site move guide says not to redirect many old URLs to one irrelevant destination such as the new site's home page, which can confuse users and might be treated as a soft 404, and to return a 404 or 410 for deleted or merged content that is not moved to the new site.

    Google Search Central

  • StageConsistent with docsD2-C882

    In a community case study, an old URL got a 301 redirect only when a new page served the same intent; URLs with no same-intent match were removed with an error status instead of being redirected (the exact code is unclear in the recording).

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C883

    A community speaker's team never redirected removed URLs to the homepage, even those that still brought traffic, because the aim was a clear signal about what the site is rather than keeping every visitor.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C895

    A community speaker said a migration should be run around one document, the redirect map, which gives every old URL an approved destination.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C902

    Before a migration launches, a community speaker advised freezing the inventory (CMS export, crawl, sitemap, search data and logs), saving all content and approving the redirect map.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C903

    A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD1-C479

    Where content was duplicated across languages, Google's own site consolidation redirected two language versions into one, giving both old URLs one target path to redirect to.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

AvoidDocumentedDEV-CAN-11

Never route a crawlable URL's redirects through a URL that robots.txt blocks

Why

Said at Search Central Live by Dave Smart: robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one. His example was a shop that redirected every page through a disallowed /cart/ URL with JavaScript to set the local currency and back, so its pages were reported as blocked by robots.txt. It applies to every kind of redirect, for example through an external authorisation service whose own robots.txt blocks crawlers, through an old intermediate URL that was disallowed years after the content moved past it, or a redirect served only to Googlebot's user agent. Search Console shows the block on the first URL of the chain, which is not itself disallowed; Google's URL Inspection help confirms that the live test follows redirects without saying so or showing the final URL it tested.

How

Set currency, language, consent or session state on the page itself or with a cookie, not with a redirect hop through cart, checkout, login or tracking URLs. Point every redirect straight at its final URL (DEV-CAN-01) and re-check the redirect map whenever robots.txt gains a Disallow rule, so a historical intermediate URL cannot block the target. Do not send crawlers through third-party hosts in a redirect chain, and never redirect Googlebot differently from users (DEV-SPM-03).

Test

When Search Console reports a URL as blocked by robots.txt that the file does not block, follow its redirect chain hop by hop, with a browser user agent and with Googlebot's, and test every hop against its own host's robots.txt with Google's open-source parser; run the same check on one URL per template after each release. JavaScript redirects only show in a rendered check, such as URL Inspection's live test.

Code · Check every hop of a redirect chain against robots.txt

Robots.txt is checked for every URL in a redirect chain, and crawling stops at the first blocked hop, while Search Console shows the block on the first URL. This script follows server-side redirects one hop at a time, with Googlebot's user agent and with a browser's, and tests each hop against its own host's robots.txt. It needs pip install requests protego (Protego is an open-source robots.txt parser with Google-style wildcard and longest-match rules; confirm a doubtful result with Google's own parser, github.com/google/robotstxt). Meta refresh and JavaScript redirects are not followed: check those with URL Inspection's live test.

python
# check_chain.py  -  usage: python check_chain.py https://www.example.com/page [more URLs]
import sys
from urllib.parse import urljoin, urlsplit

import requests
from protego import Protego

AGENTS = {
    "googlebot": "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
    "browser": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36",
}
TOKEN = "Googlebot"  # the robots.txt token the rules are tested for
_robots = {}


def robots_for(url):
    parts = urlsplit(url)
    origin = f"{parts.scheme}://{parts.netloc}"
    if origin not in _robots:
        r = requests.get(origin + "/robots.txt", headers={"User-Agent": AGENTS["googlebot"]}, timeout=10)
        if r.status_code == 200:
            text = r.text
        elif 400 <= r.status_code < 500 and r.status_code != 429:
            text = ""  # 4xx (except 429): Google crawls as if there were no robots.txt
        else:
            text = "User-agent: *\nDisallow: /"  # 5xx or 429: Google stops crawling the host for now
        _robots[origin] = Protego.parse(text)
    return _robots[origin]


def check(url, agent, max_hops=10):
    print(f"[{agent}]")
    for hop in range(max_hops + 1):
        allowed = robots_for(url).can_fetch(url, TOKEN)
        print(f"  {hop}: {'allowed' if allowed else 'BLOCKED by robots.txt'}  {url}")
        if not allowed:
            print("     crawling stops here; Search Console reports the first URL of the chain as blocked")
            return
        r = requests.get(url, headers={"User-Agent": AGENTS[agent]}, allow_redirects=False, timeout=10)
        location = r.headers.get("Location")
        if r.status_code in (301, 302, 303, 307, 308) and location:
            url = urljoin(url, location)
        else:
            print(f"     final status {r.status_code}")
            return
    print(f"     more than {max_hops} redirect hops")


if __name__ == "__main__":
    for start in sys.argv[1:]:
        for agent in AGENTS:  # different chains for the two user agents point to cloaking
            check(start, agent)
Evidence · 4 claims · 3 Google pages
  • StageNot in docsD1-C534

    Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C535

    Dave Smart said Search Console reports such a block only on the first URL of the chain, which is confusing: the URL shown as blocked by robots.txt is not itself disallowed.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C536

    Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C537

    When Search Console says a URL is blocked by robots.txt but the file does not block it, check whether the URL redirects and test every URL in the chain.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

Error pages and soft 404s 4

A page that should be an error but returns 200 is a soft 404: Google keeps it out of the index but keeps crawling it, may cluster many such pages together, and recognises it from the main content rather than from keywords. Every response needs the status code that matches the state of the page.

MustDocumentedDEV-ERR-01

Return 404 or 410 for removed and non-existent URLs

Why

4xx codes other than 429 do not slow crawling, URLs that return them are not indexed, and a 404 is a strong signal not to crawl the URL again. A soft 404 is kept out of the index but keeps being crawled, wasting crawl budget, and soft 404 pages can be clustered together as duplicates.

How

Answer unknown paths, deleted items without a close replacement and invalid parameters with 404 or 410 and a helpful error page (search box, main categories). Redirect only when a genuinely equivalent replacement exists, and never send every missing URL to the homepage.

Test

curl -I https://www.example.com/this-does-not-exist-123 returns 404; the Page indexing report's "Soft 404" count stays near zero.

Evidence · 7 claims · 3 Google pages
  • DocsSourceD1-C072

    4xx status codes other than 429 have no effect on crawl rate.

    Google

  • DocsSourceD1-C073

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

    Google, Google Search Central

  • DocsSourceD1-C109

    Google advises against using noindex to save crawl budget and against using robots.txt to temporarily reallocate budget. Use robots.txt only for pages you never want crawled, and 404 or 410 for removed pages.

    Google

  • StageConfirmed by docsD2-C334

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C367

    Soft 404 pages are another reason Google clusters pages together, an outcome site owners usually do not want.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD1-C351

    404 Not Found and 410 Gone both tell Google there is nothing at the URL, so the URL is not indexable.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C385

    4xx responses do not affect a site's crawl budget, because Google expects pages, content and products to come and go as a natural part of the web.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MustDocumentedDEV-ERR-02

Return a real 404 for unknown routes in single-page apps, or use a documented client-side fallback

Why

In single-page apps the server often returns the app with 200 for every URL and the front-end router shows a "not found" view, so no error is reported. When a server-side 404 is impractical, Google documents two fallbacks: a JavaScript redirect to a URL that returns 404, or a noindex robots meta tag added with JavaScript.

How

Prefer the server: route known paths, look up dynamic ones such as product slugs, and send the app shell with 404 for everything else. Otherwise, in the client, redirect error views to a URL the server answers with 404, or inject <meta name="robots" content="noindex"> on error views.

Test

curl -I on a made-up deep URL returns 404, or the rendered HTML in URL Inspection shows the redirect or the noindex; the Page indexing report shows no soft 404s on app routes.

Code · Real 404 status codes in a single-page app

Fix soft 404s at the source: the server (or hosting rewrite rules) answers known paths with 200 and the app shell, and everything else with 404 and a static error page that loads no router, so it can never redirect again. Where that is impossible, use one of the two client-side fallbacks Google documents: a JavaScript redirect to a URL that returns 404, or a noindex added with JavaScript.

JavaScript
// Server (Express 4 or 5): static files first, then a final handler for every other path
const express = require('express');
const path = require('path');
const app = express();
const shell = path.join(__dirname, 'dist', 'index.html');
const notFound = path.join(__dirname, 'dist', '404.html'); // static error page: no router script
const PAGES = [/^\/$/, /^\/chairs\/$/, /^\/tables\/$/, /^\/about\/$/];
const PRODUCT = /^\/chairs\/([a-z0-9-]+)$/;

app.use(express.static(path.join(__dirname, 'dist'), { index: false }));

app.use(async (req, res) => {
  const product = req.path.match(PRODUCT);
  const found = PAGES.some((re) => re.test(req.path)) || (product !== null && (await productExists(product[1])));
  if (found) res.sendFile(shell);
  else res.status(404).sendFile(notFound); // also answers /not-found, the target of the client fallback
});

async function productExists(slug) {
  return ['oak-dining-chair', 'beech-stool'].includes(slug); // replace with your database or API lookup
}

app.listen(3000);
JavaScript
// Client fallback when the server cannot know: never leave a "not found" view on a 200 URL
fetch(`/api/products/${productId}`)
  .then((response) => response.json())
  .then((product) => {
    if (product.exists) {
      showProductDetails(product);
      return;
    }
    // Option 1: go to a URL whose server response is 404 (and whose page does not redirect again)
    window.location.replace('/not-found');
    // Option 2 (instead of option 1): keep the URL but keep the view out of the index
    // const robots = document.createElement('meta');
    // robots.name = 'robots';
    // robots.content = 'noindex';
    // document.head.appendChild(robots);
  });
Evidence · 7 claims · 2 Google pages
  • StageConfirmed by docsD2-C190

    In single-page apps, a missing page often shows a custom 404 page while the server returns HTTP 200, because the front-end router, not the server, handles the 404.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C291

    In JavaScript single-page apps, soft 404s typically arise because the server returns the app with a 200 status for every URL, so when the app shows a 'not found' message for a URL that does not exist, no error is reported.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C264

    Google's slide defined a soft 404 in a JavaScript application as a page that serves a 'Not Found' message but returns a 200 HTTP status code.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C192

    Google's JavaScript guides say that when client-side routing makes a real 404 status impractical, a single-page app can avoid soft 404s by redirecting with JavaScript to a URL whose server returns 404, or by adding a robots noindex meta tag with JavaScript.

    Google Search Central

  • DocsSourceD2-C294

    For client-side rendered single-page apps, where meaningful status codes can be impossible or impractical, Google's documentation gives two ways to avoid soft 404s: a JavaScript redirect to a URL that returns a 404 status, or a noindex robots meta tag added with JavaScript.

    Google Search Central

  • StageConsistent with docsD2-C191

    The fix for a client-side soft 404 is to make a missing page end with a real HTTP 404 status code instead of 200.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C295

    After moving to real paths, set the server or hosting rewrite rules so a direct request to every valid path returns 200 with the content (ideally server-rendered) and an unknown path returns a 404 status, which removes client-side soft 404s at the source.

    Ibrahim Anjro (author)

AvoidDocumentedDEV-ERR-03

Never serve error states, empty results or failed data loads with a 200 status

Why

Google listed four common causes of soft 404s: pages that look like errors, thin or empty content, server or CMS misconfiguration, and JavaScript-dependent content that fails to load. Said at Search Central Live: detection uses a model that understands page layout, so an error message alone in the main content, such as a database connection error, makes a soft 404 even when header and navigation look normal.

How

Database and backend failures return 503; missing content returns 404 or 410; empty search and filter results return 404 or carry noindex. Client-rendered pages whose data request fails show a retry option and never leave an empty main area on a 200 URL.

Test

In staging, simulate a backend outage: pages answer 503, not 200. Crawl for 200 pages with a very short main content or error wording, and review the Page indexing report's soft 404 list after each release.

Evidence · 10 claims · 3 Google pages
  • SlideConfirmed by docsD2-C336

    Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C339

    For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C213

    A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C376

    Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.

    Search Central blog (24 December 2024)

  • AnalysisD2-C342

    Return real error status codes for error states, 404 or 410 for missing content and 503 for outages such as a failed database connection, also in single-page apps; a 200 page whose main content is only an error message is treated as a soft 404 even when header and navigation look normal.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD2-C334

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD1-C348

    A 204 No Content response is not eligible for indexing, because the server confirms the request but returns no content to index.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C355

    A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C320

    An HTTP 200 OK status only means that the server believes it managed to do what the client asked for.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C369

    Google's status code page says that with a 204 No Content response Google receives no content and cannot process it, and that a 2xx page whose content is empty or an error message shows as a soft 404 in Search Console.

    Google

ShouldDocumentedDEV-ERR-04

Keep temporarily out-of-stock product pages live with 200 and their current availability

Why

Google's guide to pausing a business recommends staying online with limited functionality, such as a disabled cart, and updating Product markup with current availability. Google's answer on out-of-stock pages was that it depends on how important the page is to users, who may wait months or pre-order; a temporary redirect keeps the source URL in results anyway.

How

Keep the page at 200, show the availability (OutOfStock, PreOrder, BackOrder) on the page and in Product markup, offer pre-ordering or a back-in-stock alert, and link to alternatives. Redirect with 301 only when the product is gone for good and a true successor exists; otherwise return 404 or 410 for discontinued items.

Test

The Rich Results Test shows the current availability in Product markup; out-of-stock URLs answer 200 and stay indexed in URL Inspection.

Evidence · 5 claims · 3 Google pages
  • DocsSourceD2-C012

    Google's guide to temporarily pausing an online business recommends that a shop expecting to sell again within weeks or months stays online with limited functionality, such as a disabled cart, and updates its Product structured data to show current availability.

    Google Search Central

  • SlideNot in docsD2-C010

    Google answered on a Q&A slide that whether to keep a temporarily out-of-stock product page or redirect it depends on how important that product page is to users.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C011

    Google's Q&A slide on out-of-stock product pages said users might wait months for some products, or even pre-order them if the site offers pre-ordering.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C013

    Google's redirect documentation says that with a temporary redirect, such as a 302, Google Search shows the source page in search results and does not use the redirect as a signal that the target should be canonical.

    Google Search Central

  • AnalysisD2-C014

    Keep a product page that is out of stock for a few months live with a 200 status, show its availability on the page and in Product structured data, and offer pre-ordering where possible when users would wait for the product. Redirect it only when users would rather switch to a similar product than wait.

    Ibrahim Anjro (author)

International and multilingual sites 12

Google reads a page's language from its content, stores one language per page and uses hreflang to show searchers the right version. Each version needs its own URL, crawlable links to its alternates, valid and reciprocal annotations and content that really differs.

MustDocumentedDEV-INT-01

Give every language and country version its own URL

Why

Switching languages with cookies or by swapping content in place makes the versions very hard for Google to handle. Googlebot mostly crawls from US IP addresses without an Accept-Language header, so content that adapts to the visitor's location or language may never be crawled in its other versions.

How

Give each version its own URL (example.de, de.example.com or example.com/de/), never a ?lang= parameter or a cookie. The URL alone decides the language and region served; a remembered choice may only preselect the language selector, never change the content of a URL.

Test

curl each version's URL from one location with no Accept-Language header and no cookies: each returns its own language and content.

Evidence · 4 claims · 3 Google pages
  • StageConfirmed by docsD2-C560

    Each language version should have its own URL; switching languages with cookies or by updating the content in place makes the versions very difficult for Google to handle.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C561

    Google's documentation says Googlebot mostly crawls from US IP addresses and sends no Accept-Language header, so pages that change content or redirect by the visitor's perceived country or language may not have every version crawled, indexed or ranked; it recommends separate URLs annotated with hreflang.

    Google Search Central

  • SlideConfirmed by docsD2-C562

    Language versions can also be language-plus-region variants: Google's example site had a generic English page (en), a UK English page (en-gb) and a Spanish page (es), each on its own URL.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C595

    Google's table of URL structures for country targeting gives pros and cons for a country-specific domain (example.de), a subdomain (de.example.com) and a subdirectory (example.com/de/) on a generic domain, and marks URL parameters (site.com?loc=de) as not recommended.

    Google Search Central

AvoidDocumentedDEV-INT-02

Do not redirect visitors automatically by IP address or browser language

Why

Google advised against clever geo-redirecting because it very often goes wrong, and its documentation explains why: Googlebot mostly crawls from the US without Accept-Language, so automatic redirects can hide every other version from it.

How

Serve the requested URL as it is and suggest another version with a dismissible banner or the language selector. Point x-default at a language selector or a generic page instead of redirecting from the root.

Test

curl -I on each country URL with IP addresses or Accept-Language headers of other countries: no 3xx to another version.

Evidence · 2 claims · 2 Google pages
  • SlideConfirmed by docsD2-C383

    Google's speaker advised against 'clever' geo-redirecting, because it very often goes wrong.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C561

    Google's documentation says Googlebot mostly crawls from US IP addresses and sends no Accept-Language header, so pages that change content or redirect by the visitor's perceived country or language may not have every version crawled, indexed or ranked; it recommends separate URLs annotated with hreflang.

    Google Search Central

MustDocumentedDEV-INT-03

Make hreflang complete: a self-reference, return links and fully qualified URLs

Why

Missing return links and missing self-references are common hreflang mistakes, and Google ignores annotations that point only one way, since otherwise any site could claim to be a version of someone else's. Alternate URLs must be fully qualified, including https://, and may be on other domains. A brand-only query reveals no language, so Google falls back on the user's location and browser settings and sometimes shows the wrong language version; complete hreflang is what lets it swap in the right one.

How

Every page in a set lists itself and every alternate with an identical block, plus x-default for the fallback or selector page. Point only at canonical, indexable 200 URLs and generate the block for all versions from one data source. If keeping every pair is impractical, at least link new versions both ways with the original or dominant language.

Test

An hreflang audit in a crawler reports, per cluster, no missing return links, no missing self-references, no non-200 or non-canonical targets and no relative URLs.

Code · hreflang link elements: complete, reciprocal, self-referencing

Every language version of a page carries the identical block in its <head>, including a line for itself and an x-default for the fallback, ideally a language selector page. URLs are fully qualified; codes are an ISO 639-1 language, an optional ISO 15924 script (zh-Hant, zh-Hans-CN) and an optional ISO 3166-1 Alpha-2 region.

HTML
<!-- The same block in the <head> of every language version listed here; the x-default target is a language selector page -->
<link rel="alternate" hreflang="en" href="https://www.example.com/en/chairs/">
<link rel="alternate" hreflang="en-GB" href="https://www.example.com/en-gb/chairs/">
<link rel="alternate" hreflang="de-DE" href="https://www.example.com/de-de/stuehle/">
<link rel="alternate" hreflang="de-AT" href="https://www.example.com/de-at/stuehle/">
<link rel="alternate" hreflang="de-CH" href="https://www.example.com/de-ch/stuehle/">
<link rel="alternate" hreflang="sv-SE" href="https://www.example.com/sv-se/stolar/">
<link rel="alternate" hreflang="zh-Hant" href="https://www.example.com/zh-hant/chairs/">
<link rel="alternate" hreflang="x-default" href="https://www.example.com/chairs/choose-language/">

<!-- Wrong values seen in the wild: "se", "dk", "cz" (country codes, not languages: use sv, da, cs),
     "en-UK" (use en-GB), "de-SW" (use de-CH), "en-EU" (no such region), "GB" (a region alone).
     A script subtag is valid: "zh-Hant", "zh-Hans-CN" (language, script, region in that order) -->
Evidence · 12 claims · 2 Google pages
  • SlideConfirmed by docsD2-C565

    Missing return links are a common hreflang mistake: if page X names page Y as a language version, page Y must link back to page X.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C566

    Each page in an hreflang set must also list its own URL, and a missing self-reference is a common mistake Google sees.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C570

    Google ignores hreflang annotations that point only one way, where page A names page B as a language version but page B does not name page A back.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C572

    Google requires hreflang return links because an alternate URL is a full URL that can be on any domain, so without them any site could declare itself a language version of someone else's website.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C573

    Google's hreflang guide says alternate URLs must be fully qualified, including the protocol (https://example.com/foo, not //example.com/foo or /foo), and need not be on the same domain.

    Google Search Central

  • SlideConfirmed by docsD2-C564

    Google's hreflang example lists every language version in the page's HTML head as a link element with rel=alternate, the version's full URL and its hreflang code (en, en-gb, es), including the page's own URL.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C571

    Google's hreflang guide says a site that finds it hard to keep every language pair bidirectional may omit some languages on some pages, because Google still processes the pairs that point to each other; new language versions should at least link both ways with the original or dominant language.

    Google Search Central

  • AnalysisD2-C574

    Audit hreflang per cluster rather than per page: check that every page lists itself and all its alternates, that every alternate links back, and that only one method (HTML head, HTTP header or sitemap) supplies the annotations.

    Ibrahim Anjro (author)

  • StageConsistent with docsD3-C007

    Query language detection works poorly when someone searches only for a brand name, such as Facebook or Google, because the query does not show which language the user wants results in.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C008

    For brand-only queries Google falls back on other information, such as the user's location and browser settings, to work out the language of the results.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C009

    Because a brand-only query does not reveal the user's language, Google sometimes struggles to show the right language version of a page for brand searches.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C010

    Sites with several language versions should not leave brand searches to chance: annotate the versions with hreflang and offer a visible language switcher, because Google may have only the user's location and browser settings to choose a version.

    Ibrahim Anjro (author)

MustDocumentedDEV-INT-04

Use valid hreflang codes: an ISO 639-1 language, an optional ISO 15924 script and an optional ISO 3166-1 Alpha-2 region

Why

Common mistakes are country codes used as languages (se, dk and cz instead of sv, da and cs), wrong regions (UK instead of GB, SW instead of CH, GE instead of DE), regions that are not countries (EU, a city such as LA) and a region on its own (GB). Some mistakes are valid codes for something else (se is Northern Sami, GE is Georgia) and silently target the wrong audience.

How

Store each version as a language plus an optional script and region (sv-SE, de-CH, en-GB, zh-Hant, zh-Hans-CN), validate every value against the ISO 639-1, ISO 15924 and ISO 3166-1 Alpha-2 lists in a build check or crawler rule, and use x-default for everyone else. hreflang targets countries, not groups such as the EU.

Test

An automated check confirms every value matches x-default or language[-Script][-REGION] (case-insensitive), with each part checked against ISO 639-1, ISO 15924 and ISO 3166-1 Alpha-2.

Code · hreflang link elements: complete, reciprocal, self-referencing

Every language version of a page carries the identical block in its <head>, including a line for itself and an x-default for the fallback, ideally a language selector page. URLs are fully qualified; codes are an ISO 639-1 language, an optional ISO 15924 script (zh-Hant, zh-Hans-CN) and an optional ISO 3166-1 Alpha-2 region.

HTML
<!-- The same block in the <head> of every language version listed here; the x-default target is a language selector page -->
<link rel="alternate" hreflang="en" href="https://www.example.com/en/chairs/">
<link rel="alternate" hreflang="en-GB" href="https://www.example.com/en-gb/chairs/">
<link rel="alternate" hreflang="de-DE" href="https://www.example.com/de-de/stuehle/">
<link rel="alternate" hreflang="de-AT" href="https://www.example.com/de-at/stuehle/">
<link rel="alternate" hreflang="de-CH" href="https://www.example.com/de-ch/stuehle/">
<link rel="alternate" hreflang="sv-SE" href="https://www.example.com/sv-se/stolar/">
<link rel="alternate" hreflang="zh-Hant" href="https://www.example.com/zh-hant/chairs/">
<link rel="alternate" hreflang="x-default" href="https://www.example.com/chairs/choose-language/">

<!-- Wrong values seen in the wild: "se", "dk", "cz" (country codes, not languages: use sv, da, cs),
     "en-UK" (use en-GB), "de-SW" (use de-CH), "en-EU" (no such region), "GB" (a region alone).
     A script subtag is valid: "zh-Hant", "zh-Hans-CN" (language, script, region in that order) -->
Evidence · 10 claims · 1 Google page
  • SlideConsistent with docsD2-C575

    hreflang language codes must be ISO 639-1 codes: Google's list of common mistakes marked se, dk and cz (the country codes of Sweden, Denmark and Czechia) as wrong and sv, da and cs (Swedish, Danish, Czech) as right.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C576

    Swedish content for Sweden is marked sv-SE, not se-SE: the language part must be the Swedish language code sv, even though Sweden's country code is SE.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C578

    Czech is a frequent hreflang error: CZ is the country code and ccTLD of Czechia, but the language code for Czech is cs.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C579

    hreflang region codes must be ISO 3166-1 Alpha-2 country codes; frequent mistakes are UK instead of GB for the United Kingdom, SW instead of CH for Switzerland and GE instead of DE for Germany.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C580

    EU is not a valid hreflang region code: sites use it hoping to target the European Union, but hreflang targeting works only by country.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C581

    Google also flagged LA used as a region code for Los Angeles (a city, not a country) and SA used for South Africa as wrong hreflang region codes.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C582

    A region code cannot be used on its own in hreflang: en and en-GB are valid values, but GB alone is not, because the value must name a language.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C583

    Validate hreflang values against the ISO 639-1 and ISO 3166-1 Alpha-2 lists in a crawler rule or build check rather than by eye: reserved codes such as UK and EU are simply ignored, but some plausible mistakes are valid codes for something else (se is Northern Sami, GE Georgia, LA Laos, SA Saudi Arabia) and silently point the annotation at the wrong audience.

    Ibrahim Anjro (author)

  • DocsSourceD2-C830

    Google's hreflang guide says the ISO 639-1 language code can be followed by a script in ISO 15924 format, such as zh-Hant or zh-Hans, and by an optional ISO 3166-1 Alpha-2 region code, as in zh-Hans-US.

    Google Search Central

  • StageConsistent with docsD2-C962

    Google said incorrect hreflang language or region codes are a mistake it finds very often, with the same wrong codes coming up again and again, especially in Europe, because site owners assume the codes are easy.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-INT-05

Supply hreflang through one method only: HTML head, HTTP Link header or sitemap

Why

The three methods are equivalent; combining them brings no benefit in Search and is harder to manage, and Google often sees sites using two or three methods at once with conflicting annotations. The HTTP header suits non-HTML files such as PDFs, and the sitemap suits large sites.

How

Choose one method per site (or per file type), remove leftovers from plug-ins and earlier implementations, and keep the annotations consistent with the canonicals.

Test

A crawl finds annotations in only one source per URL, and the hreflang sets from that source are complete.

Code · hreflang in an XML sitemap

The sitemap method is equivalent to link elements and suits large sites or templates that cannot change the <head>. Use it instead of, not on top of, the HTML or HTTP-header method; every <url> lists all versions, itself included.

XML
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:xhtml="http://www.w3.org/1999/xhtml">
  <url>
    <loc>https://www.example.com/en/chairs/</loc>
    <xhtml:link rel="alternate" hreflang="en" href="https://www.example.com/en/chairs/"/>
    <xhtml:link rel="alternate" hreflang="de-DE" href="https://www.example.com/de-de/stuehle/"/>
    <xhtml:link rel="alternate" hreflang="x-default" href="https://www.example.com/"/>
  </url>
  <url>
    <loc>https://www.example.com/de-de/stuehle/</loc>
    <xhtml:link rel="alternate" hreflang="en" href="https://www.example.com/en/chairs/"/>
    <xhtml:link rel="alternate" hreflang="de-DE" href="https://www.example.com/de-de/stuehle/"/>
    <xhtml:link rel="alternate" hreflang="x-default" href="https://www.example.com/"/>
  </url>
</urlset>
Evidence · 3 claims · 1 Google page
  • StageConsistent with docsD2-C568

    Any one hreflang method is fine, but a site should use only one: Google often sees sites using two or three methods at the same time, with annotations that conflict.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C569

    Google's hreflang guide says its three methods (link elements in the HTML head, an HTTP Link header, which suits non-HTML files such as PDFs, and an XML sitemap) are equivalent; using several at once is allowed but brings no benefit in Search and is harder to manage.

    Google Search Central

  • StageConfirmed by docsD2-C567

    When the HTML head is not suitable for a page, hreflang can be given by other methods instead, such as listing the language versions in a sitemap.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-INT-06

Build language and market selectors as <a href> links to each alternate URL

Why

A selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google and depends on sitemaps to be found, which is slow.

How

Render the selector as a list of <a href> links to the equivalent page in each version, not only to each version's homepage, and keep any visual dropdown on top of those real links.

Test

The rendered HTML of any page contains <a href> links to every version of that page, and a crawl starting from one version reaches all the others.

Evidence · 2 claims · 2 Google pages
  • SlideConsistent with docsD2-C187

    A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C189

    Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-INT-07

Make each page's language unambiguous by translating the template with the main content

Why

Google determines a page's language from its content, not from hreflang or the URL, and stores one language per page, so mixed-language pages make that hard. Pages that translate only the template around untranslated main content still count as duplicates. Detecting the query's language is Google's first step for almost any query, and language and country are the first signals it uses to order candidates already at retrieval, before ranking, so a page whose language is unclear can drop out before ranking starts.

How

Translate the main content and the template (navigation, buttons, footer, legal notices, the text inside structured data) together, and avoid large blocks in a second language. Set the html lang attribute for browsers and assistive technology, knowing that Google relies on the visible content.

Test

Run language detection over the main text and template strings of sampled pages per version, and review pages with untranslated components.

Evidence · 12 claims · 3 Google pages
  • StageConfirmed by docsD2-C585

    Google determines a page's language for indexing from the page content, not from a language code in the URL.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C584

    Google does not use hreflang to determine a page's language when indexing the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C587

    Pages that mix several languages make it very hard for Google to decide which language a page is in, so each page should make its target language obvious and avoid mixing languages.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C589

    Google's hreflang guide says pages that translate only the template (navigation, footer) around main content in one language still count as duplicates, because localized versions are duplicates only if the main content stays untranslated; it recommends hreflang for such pages.

    Google Search Central

  • StageConsistent with docsD2-C588

    Scattered boilerplate in a second language is sometimes fine but still not recommended, partly because mixed languages on one page are odd for users.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C366

    Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C590

    Because Google reads a page's language from its content and stores one language per page, translate templates (navigation, footer, buttons, legal notices) together with the main text; large untranslated blocks risk the page being annotated with the wrong language.

    Ibrahim Anjro (author)

  • StageConsistent with docsD3-C006

    Google's first step in understanding almost any query is to detect its language, which tells Google roughly what content the user wants: a query in German suggests German content, a query in English English content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C079

    To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C080

    At retrieval, Google tries to match results to the user's language wherever possible: someone searching in Spanish does not necessarily want results in Italian.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C082

    Country is the second retrieval signal: a user searching from Switzerland wants cheese from Switzerland, not from Germany, and a user in Spain is poorly served by results targeting a South American country.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C086

    Make each page's language and target country unambiguous (one main language per page, hreflang between versions, local prices, addresses and shipping details), because Google applies language and country already at retrieval, before ranking.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-INT-08

Differentiate same-language country versions beyond boilerplate and mark them with region codes

Why

Same-language content for different countries, such as German for Germany, Austria and Switzerland, is tricky for deduplication; Google tries to use hreflang alternates and recommends hreflang even for small regional variations. Without real differences, Google may cluster the versions and show only one.

How

Use hreflang with region codes (de-DE, de-AT, de-CH) and show country-specific prices and currency, shipping, legal details, addresses, phone numbers and regional vocabulary in the main content.

Test

Main content differs between each pair of country versions beyond the header and footer, and URL Inspection shows each version as its own canonical.

Evidence · 6 claims · 3 Google pages
  • SlideConsistent with docsD2-C381

    Same-language content for different countries is tricky for Google's deduplication, notably German pages for Germany, Austria and Switzerland, and possibly Spanish-language variants.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConfirmed by docsD2-C382

    When localized pages are clustered, Google tries to use hreflang alternates; the speaker's short version of the advice was to use hreflang, which he called really helpful for same-language, different-country content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C628

    Google's documentation recommends hreflang annotations even when a site's versions differ only by small regional variations within one language, for example English-language content targeted to the US, GB and Ireland.

    Google Search Central

  • DocsSourceD2-C385

    Google's canonical guide says that for canonicalization Google prefers URLs that are part of hreflang clusters: if German pages for Germany and Switzerland point to each other with hreflang but not to the Austrian page, the German and Swiss pages are preferred as canonicals.

    Google Search Central

  • DocsSourceD2-C632

    Google's documentation lists the signals it uses to decide which locale a page targets: ccTLDs, hreflang statements, server location, and other signals such as local addresses and phone numbers, local language and currency, links from other local sites and Business Profile signals.

    Google Search Central

  • AnalysisD2-C384

    For German-language sites serving Germany, Austria and Switzerland, add hreflang with region codes (de-DE, de-AT, de-CH) and make the country pages differ in more than boilerplate (prices, shipping, legal details), or expect Google to cluster them and show one.

    Ibrahim Anjro (author)

MayDocumentedDEV-INT-09

Choose ccTLDs, subdomains or subdirectories by business needs, and never URL parameters

Why

A ccTLD is one of the strongest country signals, but subdomains and subdirectories are acceptable too, and Google marks URL parameters such as ?loc=de as not recommended. A separate ccTLD is not necessarily better, and sites should not move their domain for that reason; availability, cost, local regulations and commitment to the market decide.

How

For new markets, subdirectories on the existing domain are usually the cheapest start; ccTLDs make sense where the business is committed and can operate locally. Whatever the choice, use one consistent pattern for every version.

Test

The URL structure decision is recorded before the build, and every version follows the same pattern.

Evidence · 6 claims · 1 Google page
  • StageConsistent with docsD2-C594

    For country versions of a site, a ccTLD, subdomains or subdirectories are all acceptable choices; the right one depends on the site's needs, goals and resources.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C595

    Google's table of URL structures for country targeting gives pros and cons for a country-specific domain (example.de), a subdomain (de.example.com) and a subdirectory (example.com/de/) on a generic domain, and marks URL parameters (site.com?loc=de) as not recommended.

    Google Search Central

  • StageConsistent with docsD2-C596

    A separate ccTLD for each country version is not necessarily better, even though the ccTLD is one of the strongest country signals, and sites should not move their domain for that reason.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C597

    Whether to use ccTLDs depends on whether the domain can be obtained, what it costs, local laws or regulations, and how committed the business is to expanding in that country.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C591

    The ccTLD (country-code top-level domain) is one of the most important and strongest signals Google uses to decide which country a site targets; Google also considers other signals, which matter less.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C598

    For a business testing new countries, subdirectories on the existing domain are usually the cheapest start; a ccTLD's stronger country signal pays off only where domain availability, cost, local rules and long-term commitment justify it, and an established domain should not be moved just for the signal.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-INT-10

Publish machine-translated language versions only after human review and localisation

Why

Google's spam policies count generating many pages mainly to manipulate rankings, with little or no value to users, as scaled content abuse however they are made, and list automated translating of scraped content among the examples. Said at Search Central Live: whether machine-translated content is acceptable depends on the case, after weighing translation quality, local conventions such as date formats and calendars, and cultural adaptation.

How

Route translation plug-in and CMS auto-translate output through a review workflow before publishing: a native speaker checks the text, and dates, calendars, units, currencies and examples are adapted to the market. Keep unreviewed versions unpublished or noindexed, and do not generate a language version for every market by default.

Test

Every published language version has a reviewer and a review date recorded in the CMS, and no auto-translated page is indexable before review.

Evidence · 5 claims · 1 Google page
  • DocsSourceD2-C611

    Google's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.

    Google Search Central

  • StageNot in docsD2-C608

    Whether machine-translated content is acceptable depends on the case and is the site owner's decision, after weighing three things machine translation can miss: translation quality, local conventions and cultural adaptation.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C612

    Treat machine translation as a first draft: have a native speaker review it and adapt dates, calendars and units before publishing, because unreviewed bulk translation that adds little value can also fall under Google's scaled content abuse policy.

    Ibrahim Anjro (author)

  • AnalysisD2-C692

    The lower selection bar in under-served languages is an opening for content written for that market, not for machine translation at scale: index selection also applies spam signals, and Google's spam policies count generating many pages from scraped content through automated transformations such as translating, with little value for users, as scaled content abuse.

    Ibrahim Anjro (author)

  • StageNot in docsD2-C610

    Localisation should account for local conventions such as date formats, which differ between Europe, the UK, the US and other countries, and calendars (in Thailand the current year is 2569), or a date can point users to the wrong day.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-INT-11

Write pages in the words, spellings and script the audience searches with, and never add synonym or variant lists

Why

Google expands almost every query with synonyms, treats spellings with and without diacritics as the same behind the scenes, and recognises an English term and its local-language equivalent as one thing, so a page needs neither every synonym nor both spellings. Blocks of keyword variants read as keyword stuffing, which Google's spam policies prohibit. Said at Search Central Live: users expect text written the way they search, which differs by language and script; Hindi users, for example, search both in Devanagari and in Latin letters.

How

Keep editors' wording as written in the CMS: do not strip diacritics from titles, headings or body text, and do not auto-append synonym, misspelling or transliteration blocks, hidden keyword fields or tag clouds. Where an audience searches in two scripts or spellings, choose the form per market from real query data instead of duplicating the page. Use another name for a term only where readers might not understand otherwise.

Test

Search Console's Performance report shows which spellings and scripts users type for key terms; searching each variant on Google shows whether Google treats them as the same thing; a template audit finds no generated synonym or variant lists.

Evidence · 10 claims · 3 Google pages
  • StageConfirmed by docsD3-C023

    Site owners do not need to list all synonyms of a term on a page, because Google already knows the synonyms and looks for them on pages too.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C042

    Where Google recognises an English term and its local-language equivalent as synonyms, a page does not need to contain both.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C044

    When Google already matches a term's variants, there is no need to add them to pages artificially; mention another name only where visitors might not understand otherwise.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C048

    Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a German 'ü' written as 'ü', as 'ue' or left out.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C050

    Google's recommendation for diacritics, non-English words, product names and spelling variants is to focus first on what users actually search for.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C051

    Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C024

    Write each page in the words its audience uses and drop synonym lists added for search engines: Google adds synonyms at query time, and blocks of keyword variants read as keyword stuffing, which Google's spam policies prohibit.

    Ibrahim Anjro (author)

  • AnalysisD3-C055

    For audiences that type the same words in two scripts or spellings (Hindi in Devanagari and in Latin letters, German with and without umlauts), check which forms appear in Search Console's queries and use those forms in headings and key text.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C965

    Users of non-Latin-script languages do not always search in their own script: the same Persian query may be typed in Persian script or in Latin letters, with the same intent and the same expected results.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C981

    Google's post on multilingual searches (8 September 2023) says that, because of typing difficulty on some keyboards, a person in India might search in Hindi using Latin rather than Devanagari characters and want and receive Hindi results written either way.

    Search Central blog (8 September 2023)

ShouldSaid at Search Central LiveDEV-INT-12

Mark the text direction of right-to-left and mixed-script titles, product names and URLs

Why

Said at Search Central Live by a community speaker: Persian and Arabic run right to left and English left to right, so a title, product name or URL that mixes the two can display in a confusing, unpredictable order even when its content is correct.

How

Set lang and dir="rtl" on the html element of right-to-left pages, wrap embedded left-to-right fragments (brand names, model numbers) in bdi or an element with dir="ltr" or dir="auto", and generate titles so the main language comes first. Keep URL slugs in one script per language.

Test

Review the title, h1 and product names of key right-to-left templates in a browser and in search results: words, numbers and punctuation appear in the intended order.

Evidence · 2 claims
  • StageNot in docsD2-C963

    Persian and Arabic are written right to left and English left to right, so a title that mixes the two scripts can display in a confusing, unpredictable order even when its content is correct.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C964

    The display problem of mixing right-to-left and left-to-right text also affects product titles and URLs, the presenter of the non-Latin-script talk said, showing an example from a large Iranian e-commerce site.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

Page structure and HTML 7

Google parses each page into a DOM, tells header, navigation, main content and footer apart, and weighs words by where they appear. Templates decide whether the main content is easy to find and whether controls in the HTML work as intended.

ShouldDocumentedDEV-HTM-01

Wrap each page's primary content in one clearly delimited main area

Why

Google extracts every element to tell header, navigation and main content apart, treats the main content as very important and clusters pages by it. A page that uses a different template can be reported as "Crawled - currently not indexed" because Google cannot tell where its content is, even when its quality matches indexed pages.

How

Use semantic landmarks: header, nav, one main (with article where it fits), aside and footer. Main content is any part that directly helps the page achieve its purpose: text, images, video, a tool, user-generated content such as comments or reviews, content in tabs, every heading and the visible title; Gary Illyes said it is what Google considers when ranking a page, and that what a site puts in its navigation or header tells Google the site does not particularly care about that content. Keep the h1, the opening paragraph, key images and the facts the page is about inside main, near the top, and use the same structure across templates of one type.

Test

Template validation finds one main element per page, with the page's h1 inside it; the rendered HTML in URL Inspection shows the main text inside main; the DevTools accessibility tree shows the landmarks.

Code · Page template with a clearly delimited main content area

Everything indexing depends on (title, description, canonical, robots rules, main text, links) is in the server HTML. Header, navigation and footer are separated from one <main> element that holds the title, headings, opening text and media, because Google weighs words by where they appear and treats the main content as the most important part.

HTML
<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <title>Oak dining chair with woven seat | Example Shop</title>
  <meta name="description" content="Solid oak dining chair with a hand-woven paper-cord seat. Seat height 45 cm, delivered assembled.">
  <link rel="canonical" href="https://www.example.com/chairs/oak-dining-chair">
  <meta name="robots" content="max-snippet:-1, max-image-preview:large">
</head>
<body>
  <header>
    <a href="/">Example Shop</a>
    <nav aria-label="Main">
      <a href="/chairs/">Chairs</a>
      <a href="/tables/">Tables</a>
    </nav>
  </header>

  <main>
    <article>
      <h1>Oak dining chair with woven seat</h1>
      <p>A solid oak dining chair with a hand-woven paper-cord seat, made for everyday family meals.</p>
      <figure>
        <img src="/img/oak-chair-1200.webp" alt="Oak dining chair with a woven paper-cord seat, seen from the front" width="1200" height="900">
        <figcaption>Natural oak finish, seat height 45 cm.</figcaption>
      </figure>
      <h2>Dimensions and materials</h2>
      <p>Width 46 cm, depth 52 cm, height 80 cm. <strong>Solid European oak</strong>, paper-cord seat.</p>
    </article>
  </main>

  <footer>
    <nav aria-label="Footer">
      <a href="/delivery/">Delivery</a>
      <a href="/returns/">Returns</a>
      <a href="/contact/">Contact</a>
    </nav>
  </footer>
</body>
</html>
Evidence · 14 claims · 3 Google pages
  • SlideConsistent with docsD2-C028

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C309

    A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C310

    Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.

    Google Search Central

  • StageConsistent with docsD2-C713

    If content quality is even across a site, a page reported as 'Crawled – currently not indexed' may be using a different template that keeps Google from understanding where its content is.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C029

    Make the main content of every template easy to separate from the header, navigation and footer, for example as one clearly delimited main area, because Google identifies the main content and treats it as the most important part of the page.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C861

    Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C862

    Gary Illyes defined a page's main content as any part of the page that directly helps the page achieve its purpose, what it was built for.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C863

    Main content is not only text: images, videos, a tool or anything else that helps a page achieve its purpose can be main content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C864

    Content created by other users can be main content: on a user-generated content site, the user-generated content can be the page's main content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C865

    A comment section below a blog post can still be part of the page's main content and can contribute to Google's understanding of the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C867

    A page's main content includes all of its headings and its visible title.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C868

    Gary Illyes said the main content is what Google considers when ranking a page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C870

    Google's Search Quality Rater Guidelines define main content as any part of the page that directly helps it achieve its purpose, including text, images, videos, page features such as calculators and content created by users, and they count the title at the top of the page as part of it.

    Google Search Quality Rater Guidelines (PDF, 11 September 2025)

  • DocsSourceD2-C871

    Google's Search Quality Rater Guidelines say navigation links are a common type of supplementary content, and that content behind tabs and user reviews or comments may count as main content on some pages and as supplementary content on others, depending on the page's purpose.

    Google Search Quality Rater Guidelines (PDF, 11 September 2025)

ShouldDocumentedDEV-HTM-02

Place the words a page should rank for in its main content, not in footers or sidebars

Why

Google weighs words by the part of the page they appear in, and where text sits already contributes quite a bit to ranking. Said at Search Central Live: words in the footer get a lower weight, so footer keyword blocks add little.

How

Give editors template fields for the title, h1, introduction and subheadings inside the main area, and drop footer or sidebar keyword blocks and tag clouds used to add terms.

Test

Per template, the target terms appear in the title, the h1 and the first paragraph, and footers contain no keyword blocks.

Evidence · 5 claims · 2 Google pages
  • StageConsistent with docsD2-C312

    When Google processes a page for indexing, it gives words different weights depending on the part of the page where they appear.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C315

    To make a word count for ranking a page, Gary Illyes said the simplest step is to move it into the main content, because where text sits on a page already contributes quite a bit to ranking.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C313

    Words in the footer of a page get a lower weight, so text placed in the footer is unlikely to contribute much to ranking the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C314

    On Google's example blog page, the post title and opening sentence counted as important because they sit in the main content, in front of the user, while the site tagline, the 'Categories' sidebar and category links such as 'Hugo (7)' counted as less important supplementary text.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C317

    Put the terms a page should rank for in its main content (title, headings, opening paragraphs) rather than only in sidebars, tag lists or footers; footer keyword blocks are unlikely to add ranking value.

    Ibrahim Anjro (author)

MustDocumentedDEV-HTM-03

Give every indexable page a unique, descriptive title element in the server HTML

Why

Indexing analyses a page's text and key tags such as the title element, and snippets come mainly from the page content and sometimes from the meta description. A community speaker showed a site whose rendered meta description contained HTML tags, a sign of a broken process. Google generates title links and snippets from its understanding of the page even when the site provides nothing extra, and it notices a changed title only after recrawling and reprocessing the page, which its title link guide says takes a few days to a few weeks.

How

Generate the title on the server from the page's own data, with no placeholders, no HTML and no duplicates across a template, and keep it descriptive and specific to the page. You should also generate a plain-text meta description per page in the same way: Google builds snippets mainly from the page content and uses the description only sometimes, but a broken one still reaches results.

Test

A crawl finds no missing, duplicate, empty or placeholder titles, and no meta description that contains < or > or a placeholder.

Evidence · 7 claims · 4 Google pages
  • DocsSourceD2-C445

    Google's guide to how Search works says rendering happens during the crawl, and describes indexing as analysing a page's text, key tags and attributes such as title elements and alt attributes, images and videos, and deciding whether the page is a duplicate or the canonical.

    Google Search Central

  • DocsSourceD2-C725

    Google's snippet documentation says snippets are created automatically, primarily from the page content, to preview the part that best relates to the user's specific search, so one page can get different snippets for different searches; sometimes the meta description is used instead.

    Google Search Central

  • StageD2-C142

    One site that builds its whole content by rendering also rendered its meta description with HTML tags inside it, which makes no sense and points to a flawed process.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD3-C316

    Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConsistent with docsD3-C657

    Google estimated that a title update takes as long as a snippet update: 1-2 days on average, up to several weeks to months.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C675

    Google's title link guide says Google has to recrawl and reprocess a page to notice changes to the sources of its title link, which may take a few days to a few weeks.

    Google Search Central

  • StageNot in docsD2-C858

    Search engines such as Google and Bing would ignore HTML tags written inside a meta description, a community speaker said.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-HTM-04

Use semantic, accessible HTML: real buttons, labelled forms and a logical heading order

Why

Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM and the accessibility tree, and recommends semantic HTML; attendees were also told to check ARIA and accessibility, and a community speaker added that an unlabelled div tells an agent nothing about its purpose and that no ARIA is better than bad ARIA. Google has not called accessibility a ranking signal, and its SEO Starter Guide says headings out of order help screen readers less but do not matter to Google Search.

How

Buttons are button elements and links are a elements with href; form fields have labels; headings follow a logical h1 to h3 order; images have alt text; ARIA roles are used only where no native element exists, and only when they are correct (a div acting as a toggle gets role=button, an aria-label and aria-pressed). Avoid layout shifts that move buttons after load. Do not block agents you want to allow (DEV-AIF-04).

Test

An axe or Lighthouse accessibility audit runs in CI on key templates, plus a keyboard-only and screen-reader check of navigation, forms and checkout.

Evidence · 10 claims · 4 Google pages
  • DocsSourceD1-C131

    Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.

    Google Search Central

  • StageConsistent with docsD1-C118

    Attendees were told to check ARIA and accessibility.

    a speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD2-C465

    Gemini in Chrome relies heavily on the screenshot it takes of a page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD1-C119

    Google has not called ARIA a ranking signal. The case for it is that AI agents and assistive tools read pages through the same structure: real buttons, labelled forms and semantic headings.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD1-C274

    A community speaker said AI agents understand a page through a combination of three inputs: a screenshot, the DOM (the HTML plus the changes rendered by JavaScript) and the accessibility tree.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C277

    A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C278

    A community speaker said a block of content without semantic HTML or landmarks is just a div whose purpose an agent cannot tell, and recommended landmark elements (header, nav, main, article for independent sections, footer) plus p and h1-h6 for text.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C279

    A community speaker described ARIA (Accessible Rich Internet Applications) as a set of attributes, not a programming language, that adds accessibility information to HTML: a div used as an 'add to favourites' button can get role=button, an aria-label and aria-pressed set to true or false.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C280

    A community speaker warned that no ARIA is better than bad ARIA: wrong or confusing ARIA attributes do more harm than leaving them out.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C313

    Google's SEO Starter Guide says headings in semantic order help screen readers but do not matter to Google Search if used out of order, and that there is no ideal number of headings for a page.

    Google Search Central

ShouldDocumentedDEV-HTM-05

Write valid HTML and close every element

Why

Valid HTML is technically not a ranking factor, but it matters for controls such as data-nosnippet, where an unclosed element can extend the rule to the rest of the page. Parsing turns the HTML into a DOM tree, and broken markup changes what ends up where.

How

Lint templates in CI with html-validate or the W3C Nu checker, and fix unclosed elements, invalid nesting and stray elements in the head: a parser ends the head at the first element that does not belong there, which pushes later meta tags into the body, and Google stops reading head elements after the first invalid one.

Test

The validator reports no errors on key templates, and the robots meta tag, canonical and hreflang appear inside head in the parsed DOM.

Evidence · 4 claims · 3 Google pages
  • StageConsistent with docsD2-C080

    John Mueller said valid HTML is technically not a ranking factor but does matter for controls such as data-nosnippet.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C078

    Broken HTML can widen data-nosnippet: if a div marked data-nosnippet is not closed properly, the rest of the page can be blocked from the snippet as well.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C027

    The HTML parsing step turns a fetched page's HTML into a Document Object Model (DOM) tree of elements, attributes and text nodes.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C831

    Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.

    Google Search Central

AvoidDocumentedDEV-HTM-06

Never hide text from users, such as white-on-white or off-screen keyword blocks

Why

Hidden text is against Google's spam policies, and sites that violate them may rank lower or not appear at all. Said at Search Central Live: Google records spam metadata such as white-on-white text with a page's tokens, so leftover hidden keyword blocks are a liability rather than neutral clutter.

How

Remove hidden keyword blocks, text in the background colour and text positioned off-screen for search engines. Legitimately hidden interface content (tabs, accordions, screen-reader-only labels) is fine.

Test

A crawl with a style check finds no text whose colour matches its background and no off-screen or display:none blocks outside accessibility helpers and UI components.

Evidence · 3 claims · 1 Google page
  • StageNot in docsD2-C324

    Google also stores spam metadata with the tokens of a page, for example that text was white on a white background, so ranking can use that information.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C343

    Hidden text is recorded at token level: Google stores spam metadata such as white-on-white text with the tokens, so leftover hidden keyword blocks are a liability, not neutral clutter, and should be removed.

    Ibrahim Anjro (author)

  • DocsSourceD2-C667

    Google's spam policies say that sites violating them may rank lower in results or not appear in results at all.

    Google Search Central

MaySaid at Search Central LiveDEV-HTM-07

Use real heading, title and emphasis elements instead of styling alone

Why

Said at Search Central Live: when Google tokenizes a page for its index it attaches metadata to each word, such as whether it appeared in the page header, the main content, a heading or in bold. Text that only looks like a heading or bold text through CSS does not carry that information.

How

Put the page title in the title element, the visible title in an h1 and section titles in h2 to h6 elements, mark important terms with strong or b rather than CSS font-weight on spans, and keep emphasis sparing so it still means something.

Test

Template review or an HTML lint rule finds no headings or emphasis built from styled div or span elements.

Evidence · 2 claims
  • StageNot in docsD2-C323

    When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C720

    Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Structured data 15

Structured data makes pages eligible for rich results, gives Google precise facts such as prices, dates and identifiers more cheaply and accurately than model-based extraction, and the same processed data feeds classic results and AI features. It must describe what users can see.

ShouldDocumentedDEV-SDA-01

Write structured data as JSON-LD by default

Why

Google accepts JSON-LD, microdata and RDFa and interprets them identically, but recommends JSON-LD because one contiguous block is easier to author and people make fewer mistakes with it. Said at Search Central Live: microdata can save payload where page weight is critical.

How

Render JSON-LD on the server in a script type="application/ld+json" block generated from the same data as the visible page, one block per entity or a single @graph. Switch to microdata only for very weight-sensitive templates.

Test

The Rich Results Test on one URL per template detects the expected items without errors, and the JSON-LD is present in the raw HTML (curl).

Code · Product markup with a time-limited sale price

Describe only the product the page is about (not the carousel of related products) and only what the page shows. For a sale, price is the sale price, the regular price is a StrikethroughPrice, and validFrom with priceValidUntil (or validThrough) bound the sale in ISO 8601 with a time zone, so a sale price that lingers in cached markup is not treated as current. The template must still switch to the regular price (and drop the StrikethroughPrice) when the sale ends: Google warns that a listing may not display if priceValidUntil is in the past. Keep the dates aligned with the Merchant Center feed. Shipping and returns point by @id alone to the policies defined once in the organisation block.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Product",
  "@id": "https://www.example.com/chairs/oak-dining-chair#product",
  "name": "Oak dining chair with woven seat",
  "description": "Solid oak dining chair with a hand-woven paper-cord seat.",
  "image": [
    "https://www.example.com/img/oak-chair-1200.webp",
    "https://www.example.com/img/oak-chair-side-1200.webp"
  ],
  "sku": "CH-OAK-01",
  "brand": {
    "@type": "Brand",
    "name": "Example Shop"
  },
  "offers": {
    "@type": "Offer",
    "url": "https://www.example.com/chairs/oak-dining-chair",
    "price": 149.00,
    "priceCurrency": "EUR",
    "availability": "https://schema.org/InStock",
    "itemCondition": "https://schema.org/NewCondition",
    "validFrom": "2026-11-27T00:00:00+01:00",
    "priceValidUntil": "2026-11-30T23:59:59+01:00",
    "priceSpecification": {
      "@type": "UnitPriceSpecification",
      "priceType": "https://schema.org/StrikethroughPrice",
      "price": 199.00,
      "priceCurrency": "EUR"
    },
    "shippingDetails": {
      "@type": "OfferShippingDetails",
      "hasShippingService": {
        "@id": "https://www.example.com/#standard-shipping"
      }
    },
    "hasMerchantReturnPolicy": {
      "@id": "https://www.example.com/#returns"
    }
  }
}
</script>
Evidence · 4 claims · 1 Google page
  • StageConfirmed by docsD2-C495

    Google recommends JSON-LD because it is one contiguous block that is easier to author, so people make fewer mistakes with it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C494

    Google accepts three structured data syntaxes, JSON-LD, microdata and RDFa, which are all valid and are extracted at the very start into the same pipelines, so they are interpreted identically downstream.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C496

    Microdata has an advantage when payload size matters: embedded in the existing HTML, it avoids duplicating the page content in a separate JSON-LD block.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C497

    Default to JSON-LD and switch to microdata only where page weight is critical, since Google interprets all three syntaxes identically and JSON-LD is the one people get wrong least often.

    Ibrahim Anjro (author)

MustDocumentedDEV-SDA-02

Mark up only content that is visible on the page and relevant to it

Why

Google's guidelines say not to mark up content that readers cannot see, even if it is accurate. Irrelevant structured data can be treated as abusive: filters make it ineffective, and egregious cases lead to a manual action that removes rich result eligibility without affecting web ranking. Google said at Search Central Live that once it sees markup it does not trust from a site, it may stop using that site's structured data altogether.

How

Generate markup from the same fields that render the visible content, use it for the machine-precise form of visible facts (an ISO date for a shown date, the homepage url for a named organisation), and remove markup for content a page no longer shows.

Test

For each template, the Rich Results Test output matches what the page visibly shows, and the Manual actions report in Search Console is empty.

Evidence · 5 claims · 2 Google pages
  • DocsSourceD2-C467

    Google's structured data documentation says not to mark up content that is not visible to readers of the page, and not to add structured data about information that users cannot see even if it is accurate.

    Google Search Central

  • StageConfirmed by docsD2-C500

    Structured data that is not relevant to the page's content can be treated as abusive: Google's filters make it ineffective, and egregious cases can lead to a manual action.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C501

    Google's general structured data guidelines say a structured data manual action removes a page's eligibility for rich results but does not affect how the page ranks in web search.

    Google Search Central

  • AnalysisD2-C468

    The 'non-visible metadata' argument is not a licence to mark up hidden content, since Google's structured data guidelines still require markup to describe what users can see; use markup for the machine-precise form of facts the page shows, such as a full ISO date with time zone for a visible event date or a homepage url for a named organisation.

    Ibrahim Anjro (author)

  • StageConsistent with docsD3-C647

    Google may never use structured data from a site it does not trust: once it sees markup it does not trust, it does not touch it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

ShouldDocumentedDEV-SDA-03

Choose the specific types Google documents for its features, and skip generic or irrelevant ones

Why

Google uses schema.org but mostly consumes the subsets its documentation declares, and a very generic type says little and is hard to use. Pick the types relevant to the page instead of every type Google is interested in; extra relevant markup is never penalised, and Google also retires little-used features.

How

Start from Google's Search gallery, map each template to its types and their required and recommended properties, describe each entity with the most specific type, and do not spend effort marking up every semantic detail. Revisit the list when Google adds or retires features.

Test

The Rich Results Test lists the intended features per template, and Search Console shows an enhancement report for each type in use.

Evidence · 7 claims · 4 Google pages
  • SlideConfirmed by docsD2-C490

    Google recommends using the Search gallery in its developer documentation to find the structured data features that suit a site; the gallery shows each feature and how Google uses the markup.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C487

    Google uses schema.org as its markup vocabulary but mostly consumes only the subsets that its developer documentation declares, for the features it builds.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C486

    Marking something up with only a very generic schema.org type says little about it and is very hard for Google to use.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C492

    Pick only the structured data types that are relevant to a page instead of adding every type Google is interested in.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C489

    Describing every semantic detail of a page in markup is probably not worth the effort; focus on the structured data that Google or other consumers actually use.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C488

    Google will never penalise a site just for having more structured data on its pages than Google uses; extra markup does not hurt.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C522

    Google keeps adding structured data features and recommendations but also removes them: in the previous year (2025) it removed several features that brought little benefit and were little used.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-SDA-04

Identify entities with unique identifiers: url, @id and stable IDs

Why

A unique identifier, such as the homepage URL in an organisation's url property, lets Google tell which specific entity is meant instead of reading a bare name. Said at Search Central Live: markup can also carry stable identifiers that the visible text lacks, for example for user-generated content.

How

Give the Organization a url (the homepage) and a logo and reference it by @id from other blocks; give products sku or gtin values; give each entity a stable @id URL with a fragment (https://www.example.com/#organization) and keep it constant across pages and releases.

Test

The Rich Results Test shows url and @id on the Organization and other entities, and the @id values are identical on every page.

Code · Organization-level identity, returns, shipping and loyalty markup

One block, usually on the homepage, identifies the business by its homepage url and a stable @id, and states the policies that apply to most products: a return policy, a shipping service and a loyalty program. Product pages only override what differs.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "OnlineStore",
  "@id": "https://www.example.com/#organization",
  "name": "Example Shop",
  "url": "https://www.example.com/",
  "logo": "https://www.example.com/img/logo-512.png",
  "sameAs": [
    "https://video.example.net/@exampleshop",
    "https://social.example.org/exampleshop"
  ],
  "hasMerchantReturnPolicy": {
    "@type": "MerchantReturnPolicy",
    "@id": "https://www.example.com/#returns",
    "applicableCountry": ["ES", "FR"],
    "returnPolicyCountry": "ES",
    "returnPolicyCategory": "https://schema.org/MerchantReturnFiniteReturnWindow",
    "merchantReturnDays": 30,
    "returnMethod": "https://schema.org/ReturnByMail",
    "returnFees": "https://schema.org/FreeReturn",
    "refundType": "https://schema.org/FullRefund"
  },
  "hasShippingService": {
    "@type": "ShippingService",
    "@id": "https://www.example.com/#standard-shipping",
    "name": "Standard shipping to Spain and France",
    "fulfillmentType": "FulfillmentTypeDelivery",
    "handlingTime": {
      "@type": "ServicePeriod",
      "cutoffTime": "14:00:00+01:00",
      "duration": {
        "@type": "QuantitativeValue",
        "minValue": 0,
        "maxValue": 1,
        "unitCode": "DAY"
      },
      "businessDays": ["Monday", "Tuesday", "Wednesday", "Thursday", "Friday"]
    },
    "shippingConditions": [
      {
        "@type": "ShippingConditions",
        "shippingDestination": [
          { "@type": "DefinedRegion", "addressCountry": "ES" },
          { "@type": "DefinedRegion", "addressCountry": "FR" }
        ],
        "orderValue": {
          "@type": "MonetaryAmount",
          "minValue": 0,
          "maxValue": 99.99,
          "currency": "EUR"
        },
        "shippingRate": {
          "@type": "MonetaryAmount",
          "value": 4.95,
          "currency": "EUR"
        },
        "transitTime": {
          "@type": "ServicePeriod",
          "duration": {
            "@type": "QuantitativeValue",
            "minValue": 2,
            "maxValue": 4,
            "unitCode": "DAY"
          }
        }
      },
      {
        "@type": "ShippingConditions",
        "shippingDestination": [
          { "@type": "DefinedRegion", "addressCountry": "ES" },
          { "@type": "DefinedRegion", "addressCountry": "FR" }
        ],
        "orderValue": {
          "@type": "MonetaryAmount",
          "minValue": 100,
          "currency": "EUR"
        },
        "shippingRate": {
          "@type": "MonetaryAmount",
          "value": 0,
          "currency": "EUR"
        },
        "transitTime": {
          "@type": "ServicePeriod",
          "duration": {
            "@type": "QuantitativeValue",
            "minValue": 2,
            "maxValue": 4,
            "unitCode": "DAY"
          }
        }
      }
    ]
  },
  "hasMemberProgram": {
    "@type": "MemberProgram",
    "name": "Example Club",
    "description": "Free membership: earn points on every order.",
    "url": "https://www.example.com/club/",
    "hasTiers": [
      {
        "@type": "MemberProgramTier",
        "@id": "https://www.example.com/club/#member",
        "name": "Member",
        "hasTierBenefit": ["https://schema.org/TierBenefitLoyaltyPoints"],
        "membershipPointsEarned": 5
      }
    ]
  }
}
</script>
Evidence · 3 claims · 1 Google page
  • StageConfirmed by docsD2-C502

    Use unique identifiers in structured data, for example the homepage URL in an organisation's url property, so Google can tell which specific entity is meant rather than reading just a name string.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C463

    Structured data often carries non-visible metadata that the page text lacks, such as full ISO dates or stable identifiers for user-generated content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C468

    The 'non-visible metadata' argument is not a licence to mark up hidden content, since Google's structured data guidelines still require markup to describe what users can see; use markup for the machine-precise form of facts the page shows, such as a full ISO date with time zone for a visible event date or a homepage url for a named organisation.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-SDA-05

Write dates in ISO 8601 and give every date-time a UTC offset

Why

A full ISO 8601 value in markup disambiguates a visible date whose local time zone Google might not detect correctly, and date formats also differ between countries; this matters for event extraction and wherever the date must be exactly right, such as sale periods. Google accepts a date without a time where the time is not known: its Event guide asks for 2019-08-15 rather than an invented hour.

How

Output values such as 2026-11-27T19:30:00+01:00 for startDate, uploadDate, validFrom, priceValidUntil and similar properties, generated from the same timestamp as the visible date and with the correct UTC offset. Where the time is unknown or irrelevant, write the date alone (2026-11-27).

Test

The Rich Results Test shows the parsed dates, and a unit test asserts that every DateTime value has an offset (+01:00 or Z) and that date-only values are YYYY-MM-DD, used only where the time is unknown or irrelevant.

Code · Product markup with a time-limited sale price

Describe only the product the page is about (not the carousel of related products) and only what the page shows. For a sale, price is the sale price, the regular price is a StrikethroughPrice, and validFrom with priceValidUntil (or validThrough) bound the sale in ISO 8601 with a time zone, so a sale price that lingers in cached markup is not treated as current. The template must still switch to the regular price (and drop the StrikethroughPrice) when the sale ends: Google warns that a listing may not display if priceValidUntil is in the past. Keep the dates aligned with the Merchant Center feed. Shipping and returns point by @id alone to the policies defined once in the organisation block.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Product",
  "@id": "https://www.example.com/chairs/oak-dining-chair#product",
  "name": "Oak dining chair with woven seat",
  "description": "Solid oak dining chair with a hand-woven paper-cord seat.",
  "image": [
    "https://www.example.com/img/oak-chair-1200.webp",
    "https://www.example.com/img/oak-chair-side-1200.webp"
  ],
  "sku": "CH-OAK-01",
  "brand": {
    "@type": "Brand",
    "name": "Example Shop"
  },
  "offers": {
    "@type": "Offer",
    "url": "https://www.example.com/chairs/oak-dining-chair",
    "price": 149.00,
    "priceCurrency": "EUR",
    "availability": "https://schema.org/InStock",
    "itemCondition": "https://schema.org/NewCondition",
    "validFrom": "2026-11-27T00:00:00+01:00",
    "priceValidUntil": "2026-11-30T23:59:59+01:00",
    "priceSpecification": {
      "@type": "UnitPriceSpecification",
      "priceType": "https://schema.org/StrikethroughPrice",
      "price": 199.00,
      "priceCurrency": "EUR"
    },
    "shippingDetails": {
      "@type": "OfferShippingDetails",
      "hasShippingService": {
        "@id": "https://www.example.com/#standard-shipping"
      }
    },
    "hasMerchantReturnPolicy": {
      "@id": "https://www.example.com/#returns"
    }
  }
}
</script>
Evidence · 5 claims · 2 Google pages
  • StageConsistent with docsD2-C466

    Markup can state a date as a full ISO 8601 value, which disambiguates a visible date whose local time zone Google might not detect correctly; this matters for event extraction and wherever the date must be exactly right.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C463

    Structured data often carries non-visible metadata that the page text lacks, such as full ISO dates or stable identifiers for user-generated content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C610

    Localisation should account for local conventions such as date formats, which differ between Europe, the UK, the US and other countries, and calendars (in Thailand the current year is 2569), or a date can point users to the wrong day.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C468

    The 'non-visible metadata' argument is not a licence to mark up hidden content, since Google's structured data guidelines still require markup to describe what users can see; use markup for the machine-precise form of facts the page shows, such as a full ISO date with time zone for a visible event date or a homepage url for a named organisation.

    Ibrahim Anjro (author)

  • DocsSourceD2-C836

    Google's Event structured data guide says to give a date without a time, such as 2019-08-15, when the start hour is not known, and to include the UTC or GMT offset whenever a time is given.

    Google Search Central

ShouldSaid at Search Central LiveDEV-SDA-06

Emit each structured data entity once per page, and audit CMS plug-ins for duplicates

Why

Said at Search Central Live: several plug-ins emitting the same markup type is one of the most common structured data problems, and the duplicates can make an event detail page look like a list of events and change how Google interprets the page.

How

Decide which layer owns each type (theme, SEO plug-in, events plug-in, tag manager), disable the others, and merge the output into one @graph per page.

Test

The Rich Results Test on one URL per template shows each item once, unless the page really lists several entities.

Evidence · 2 claims
  • StageNot in docsD2-C503

    Several plug-ins emitting the same markup type is one of the most common structured data problems: the duplicates can make an event details page look like a list of events and change how Google interprets the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C504

    Audit CMS templates for duplicate markup by running one page per template through the Rich Results Test and checking whether an SEO plug-in and a theme or events plug-in emit the same type twice; 'more markup never hurts' covers relevant, non-duplicated markup only.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-SDA-07

Describe only the page's main product in Product markup, not related-product carousels

Why

Google's merchant listing guide says Product rich results only support pages that focus on a single product, or on several variants of the same product. Said at Search Central Live: structured data also points Google at the pertinent data and stops its systems pulling in extraneous information such as the prices of related products, while Google's own extraction sometimes fails to find a page's main content.

How

Generate Product and Offer markup only for the product the page is about, and give recommendation, recently-viewed and bundle carousels no Product markup.

Test

The Rich Results Test on product pages lists exactly one product (or one product group with its variants).

Evidence · 4 claims · 1 Google page
  • DocsSourceD2-C834

    Google's merchant listing guide says Product rich results only support pages that focus on a single product, or on several variants of the same product.

    Google Search Central

  • SlideNot in docsD2-C473

    Structured data explicitly points to the pertinent data on a page, which reduces noise and stops Google's systems from pulling in extraneous information such as prices of related products.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C474

    Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C475

    On product pages with related-product or recently-viewed carousels, make sure the Product markup describes only the main item and its price; Google's own example of what markup prevents was pulling a price from related products.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-SDA-08

Mark up merchant listings completely, with validity dates for sale prices

Why

Google's merchant listing guide has explained validity dates for sale prices in a Sale duration section since July 2026 (its updates log calls this a clarification). Said at Search Central Live: with the dates in place, merchants need not rush to remove a sale price for fear it shows wrongly in snippets. The template must still switch back to the regular price when the sale ends: Google warns that a listing may not display if priceValidUntil is in the past. Product markup should also show current availability. The product name is required, and unlike product snippets, merchant listing experiences require a price greater than zero; an empty name and a price of zero were among the problems the Rich Results Test flagged in a community demo at the event.

How

Product with a non-empty name, image, description, sku or gtin and brand; Offer with a price above zero, priceCurrency, availability and itemCondition. For a sale, put the sale price in price, the regular price in a StrikethroughPrice UnitPriceSpecification, and bound it with validFrom and priceValidUntil (or validThrough) in ISO 8601. When the sale ends, render the regular price in price and drop the StrikethroughPrice. Keep values and dates identical to the Merchant Center feed.

Test

The Rich Results Test shows a valid merchant listing, the Merchant listings report has no errors, and a scheduled check confirms feed and markup prices match.

Code · Product markup with a time-limited sale price

Describe only the product the page is about (not the carousel of related products) and only what the page shows. For a sale, price is the sale price, the regular price is a StrikethroughPrice, and validFrom with priceValidUntil (or validThrough) bound the sale in ISO 8601 with a time zone, so a sale price that lingers in cached markup is not treated as current. The template must still switch to the regular price (and drop the StrikethroughPrice) when the sale ends: Google warns that a listing may not display if priceValidUntil is in the past. Keep the dates aligned with the Merchant Center feed. Shipping and returns point by @id alone to the policies defined once in the organisation block.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Product",
  "@id": "https://www.example.com/chairs/oak-dining-chair#product",
  "name": "Oak dining chair with woven seat",
  "description": "Solid oak dining chair with a hand-woven paper-cord seat.",
  "image": [
    "https://www.example.com/img/oak-chair-1200.webp",
    "https://www.example.com/img/oak-chair-side-1200.webp"
  ],
  "sku": "CH-OAK-01",
  "brand": {
    "@type": "Brand",
    "name": "Example Shop"
  },
  "offers": {
    "@type": "Offer",
    "url": "https://www.example.com/chairs/oak-dining-chair",
    "price": 149.00,
    "priceCurrency": "EUR",
    "availability": "https://schema.org/InStock",
    "itemCondition": "https://schema.org/NewCondition",
    "validFrom": "2026-11-27T00:00:00+01:00",
    "priceValidUntil": "2026-11-30T23:59:59+01:00",
    "priceSpecification": {
      "@type": "UnitPriceSpecification",
      "priceType": "https://schema.org/StrikethroughPrice",
      "price": 199.00,
      "priceCurrency": "EUR"
    },
    "shippingDetails": {
      "@type": "OfferShippingDetails",
      "hasShippingService": {
        "@id": "https://www.example.com/#standard-shipping"
      }
    },
    "hasMerchantReturnPolicy": {
      "@id": "https://www.example.com/#returns"
    }
  }
}
</script>
Evidence · 7 claims · 2 Google pages
  • DocsSourceD2-C507

    Google's documentation updates log calls the July 2026 sale price change a clarification: a new Sale duration section of the merchant listing guide explains validFrom with validThrough or priceValidUntil, aligned with Merchant Center's sale_price_effective_date attribute.

    Google Search Central

  • StageConsistent with docsD2-C506

    Google's structured data speaker said that in the months before the event Google added support for validity dates on sale prices in product structured data, so merchants no longer need to rush to remove a sale price when the sale ends for fear it shows wrongly in snippets.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C508

    For time-limited sales, give the sale price its validity dates in the product markup (Google's merchant listing documentation describes validFrom, validThrough and priceValidUntil) instead of editing markup by hand when the sale ends, and keep the dates aligned with the Merchant Center feed.

    Ibrahim Anjro (author)

  • DocsSourceD2-C012

    Google's guide to temporarily pausing an online business recommends that a shop expecting to sell again within weeks or months stays online with limited functionality, such as a disabled cart, and updates its Product structured data to show current availability.

    Google Search Central

  • DocsSourceD2-C835

    Google's merchant listing guide warns that a listing may not display if its priceValidUntil property indicates a past date.

    Google Search Central

  • DocsSourceD1-C507

    Google's merchant listing documentation lists the product name as a required property and, unlike product snippets, requires a price greater than zero for merchant listing experiences.

    Google Search Central

  • StageConsistent with docsD1-C502

    In a community demo, the structured-data problems Google's Rich Results Test reported on the test page included an empty name, a breadcrumb problem and a price of zero, with errors shown in pink and warnings in orange (best readings of a largely unintelligible recording; each item is heard in only one of the two recordings).

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

ShouldDocumentedDEV-SDA-09

Declare return, shipping and loyalty policies once at organisation level

Why

Google's shopping markup supports loyalty programs, shipping policies and return policies defined at organisation level, with details at product level only where they differ.

How

On the homepage or a policy page, add an Organization (or OnlineStore) block with hasMerchantReturnPolicy, hasShippingService (a ShippingService with shippingConditions) and hasMemberProgram (a MemberProgram with tiers). Override in the Offer (OfferShippingDetails, member prices with validForMemberTier) only for exceptions. From each Offer, reference the organisation-level policies by @id only, as Google's merchant listing guide shows: shippingDetails with an OfferShippingDetails whose hasShippingService is {"@id": ...}, and hasMerchantReturnPolicy {"@id": ...}.

Test

The Rich Results Test on the homepage detects the organisation details, and the Merchant listings report recognises shipping and returns.

Code · Organization-level identity, returns, shipping and loyalty markup

One block, usually on the homepage, identifies the business by its homepage url and a stable @id, and states the policies that apply to most products: a return policy, a shipping service and a loyalty program. Product pages only override what differs.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "OnlineStore",
  "@id": "https://www.example.com/#organization",
  "name": "Example Shop",
  "url": "https://www.example.com/",
  "logo": "https://www.example.com/img/logo-512.png",
  "sameAs": [
    "https://video.example.net/@exampleshop",
    "https://social.example.org/exampleshop"
  ],
  "hasMerchantReturnPolicy": {
    "@type": "MerchantReturnPolicy",
    "@id": "https://www.example.com/#returns",
    "applicableCountry": ["ES", "FR"],
    "returnPolicyCountry": "ES",
    "returnPolicyCategory": "https://schema.org/MerchantReturnFiniteReturnWindow",
    "merchantReturnDays": 30,
    "returnMethod": "https://schema.org/ReturnByMail",
    "returnFees": "https://schema.org/FreeReturn",
    "refundType": "https://schema.org/FullRefund"
  },
  "hasShippingService": {
    "@type": "ShippingService",
    "@id": "https://www.example.com/#standard-shipping",
    "name": "Standard shipping to Spain and France",
    "fulfillmentType": "FulfillmentTypeDelivery",
    "handlingTime": {
      "@type": "ServicePeriod",
      "cutoffTime": "14:00:00+01:00",
      "duration": {
        "@type": "QuantitativeValue",
        "minValue": 0,
        "maxValue": 1,
        "unitCode": "DAY"
      },
      "businessDays": ["Monday", "Tuesday", "Wednesday", "Thursday", "Friday"]
    },
    "shippingConditions": [
      {
        "@type": "ShippingConditions",
        "shippingDestination": [
          { "@type": "DefinedRegion", "addressCountry": "ES" },
          { "@type": "DefinedRegion", "addressCountry": "FR" }
        ],
        "orderValue": {
          "@type": "MonetaryAmount",
          "minValue": 0,
          "maxValue": 99.99,
          "currency": "EUR"
        },
        "shippingRate": {
          "@type": "MonetaryAmount",
          "value": 4.95,
          "currency": "EUR"
        },
        "transitTime": {
          "@type": "ServicePeriod",
          "duration": {
            "@type": "QuantitativeValue",
            "minValue": 2,
            "maxValue": 4,
            "unitCode": "DAY"
          }
        }
      },
      {
        "@type": "ShippingConditions",
        "shippingDestination": [
          { "@type": "DefinedRegion", "addressCountry": "ES" },
          { "@type": "DefinedRegion", "addressCountry": "FR" }
        ],
        "orderValue": {
          "@type": "MonetaryAmount",
          "minValue": 100,
          "currency": "EUR"
        },
        "shippingRate": {
          "@type": "MonetaryAmount",
          "value": 0,
          "currency": "EUR"
        },
        "transitTime": {
          "@type": "ServicePeriod",
          "duration": {
            "@type": "QuantitativeValue",
            "minValue": 2,
            "maxValue": 4,
            "unitCode": "DAY"
          }
        }
      }
    ]
  },
  "hasMemberProgram": {
    "@type": "MemberProgram",
    "name": "Example Club",
    "description": "Free membership: earn points on every order.",
    "url": "https://www.example.com/club/",
    "hasTiers": [
      {
        "@type": "MemberProgramTier",
        "@id": "https://www.example.com/club/#member",
        "name": "Member",
        "hasTierBenefit": ["https://schema.org/TierBenefitLoyaltyPoints"],
        "membershipPointsEarned": 5
      }
    ]
  }
}
</script>
Evidence · 4 claims · 4 Google pages
  • StageConfirmed by docsD2-C505

    Google's shopping structured data launches of the previous year (2025) added support for merchant loyalty programs and shipping policies, letting merchants define a policy at organisation level and specify details at product level.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C502

    Use unique identifiers in structured data, for example the homepage URL in an organisation's url property, so Google can tell which specific entity is meant rather than reading just a name string.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD3-C383

    In product markup, an offer's shipping and return information can point through a JSON-LD identifier (@id) to shipping and return data defined elsewhere.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C384

    Google's merchant listing documentation recommends defining shipping and return policies once under Organization markup and referencing them from an Offer, even on another page, using only the @id keyword (through hasShippingService for shipping, hasMerchantReturnPolicy for returns).

    Google Search Central

ShouldDocumentedDEV-SDA-10

Put the same structured data on duplicate URLs as on the canonical

Why

Google's guidelines recommend the same structured data on all duplicates of a page, not only on the canonical. Said at Search Central Live: deduplication runs before feature extraction, so markup that exists only on a URL that loses canonical selection may never be extracted.

How

Render markup from the page template rather than per URL, so parameter, print and mobile variants carry identical markup, and make sure the preferred canonical has the complete markup.

Test

The Rich Results Test returns the same items for the canonical and for a duplicate variant.

Evidence · 3 claims · 1 Google page
  • DocsSourceD2-C451

    Google's general structured data guidelines recommend placing the same structured data on all duplicate pages of the same content, not just on the canonical page.

    Google Search Central

  • StageNot in docsD2-C444

    Google runs deduplication before feature extraction, so that the expensive extraction work is spent only on a trimmed-down, manageable set of documents.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C450

    Because Google extracts structured data, images and videos only after deduplication, put markup and media on the URL you want as canonical and keep them identical on its duplicates; markup that exists only on a duplicate that loses canonical selection may never be extracted.

    Ibrahim Anjro (author)

MustDocumentedDEV-SDA-11

Validate markup in the Rich Results Test before it reaches templates, and monitor it after release

Why

Google recommends testing and previewing structured data in the Rich Results Test and testing manually before putting markup into templates. AI-generated markup can contain hallucinations, such as invented properties or broken nesting, and should be fact-checked and validated before publishing.

How

Prototype each type on one page, validate it, then put it in the template. The Rich Results Test takes a public URL, which Google fetches itself, or pasted code: for staging and other private pages paste the rendered HTML, as a community demo's script did at the event, and keep the resources it loads reachable anonymously, because the test cannot load resources behind a firewall or a password. Fix every error (a critical issue makes the item invalid and stops the rich result); warnings are non-critical and only limit how the item can appear. Add a CI step that extracts JSON-LD from rendered templates and checks syntax and required properties, and review the Search Console enhancement reports after each release.

Test

The Rich Results Test passes for one URL per template; CI fails on JSON-LD parse errors or missing required properties.

Evidence · 8 claims · 4 Google pages
  • SlideConfirmed by docsD2-C498

    Google recommends testing and previewing structured data in the Rich Results Test, which shows the rich result features it detected and whether the markup is valid.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C499

    Test structured data manually first and only then put it into the site's templates.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C462

    Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.

    Google Search Central

  • StageNot in docsD2-C460

    In the speaker's own tests, even the latest LLMs asked to generate schema.org markup for a page often invent properties that do not exist, get deeply nested schemas such as complex pricing models wrong, and duplicate content across several fields.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD1-C251

    A community demo treated warnings in Google's structured-data test as not critical, in an example result of 22 items with two errors and two warnings.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C265

    Search Console Help separates critical issues, which make a structured-data item invalid and stop it from appearing as a rich result, from non-critical issues, listed under 'Improve item appearance'; the Rich Results Test reports valid items that have warnings.

    Google Search Console Help

  • StageConsistent with docsD1-C501

    A community demo's script used both input modes of Google's Rich Results Test: the URL mode for public pages, which Google fetches itself, and the code mode, into which the script pasted the page's HTML, for private pages and pages Google cannot fetch (best reading of a largely unintelligible recording; the second kind of page was heard as 'dead', possibly 'dev').

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C506

    Google's Rich Results Test help says the tool tests either a page's full URL or a pasted code snippet (Code instead of URL); all page resources must be reachable by an anonymous user on the internet, so resources behind a firewall or a password are not available to the test unless exposed, for example through a tunnel.

    Google Search Console Help

MaySaid at Search Central LiveDEV-SDA-12

Wire Google's SHACL validation rules into the build once they are published

Why

Said at Search Central Live: Google plans to publish downloadable SHACL rules on each structured data feature guide, to run inside a site's content generation so markup is checked before it is published. They will catch problems such as a missing required field but will not replace Search Console reports, and they were not yet available at the time of the event.

How

Plan a CI or CMS publishing step that runs a SHACL validator (for example pySHACL) against the generated JSON-LD, adopt the rules type by type as Google releases them, and keep the Rich Results Test and Search Console as the reference.

Test

Once rules exist, the build fails when a template drops a required property.

Evidence · 6 claims
  • SlideNot in docsD2-C514

    Google announced server-side structured data validation as coming soon: it will publish downloadable validation rules in SHACL on each structured data feature guide, with other kinds of checks to follow.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideNot in docsD2-C515

    Google's planned validation workflow has five steps: download the rules from the feature guide, generate the JSON or embedded microdata or RDFa, run the rules against the generated server-side markup as a first check, deploy and test in the Rich Results Test, and monitor ongoing performance in Search Console.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C516

    The SHACL rules are meant to run inside a site's content generation, so markup is sanity-checked before it is published and does not silently regress later, a breakage site owners might otherwise discover only through a Search Console report.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C517

    The SHACL rules will not replace Search Console as the canonical place for structured data reports, because some checks use Google's internal libraries and cannot be expressed in SHACL, but they will catch problems such as a missing required field.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C519

    Google plans a SHACL rule set for each of its structured data feature types and will release the rule sets gradually once they have been checked.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C520

    When the SHACL rules appear, wire them into the build or CMS publishing step as an automated test, so a template change that drops a required property fails before deployment instead of surfacing weeks later in Search Console.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-SDA-13

Add Article markup with headline, images, dates and authors to every article template

Why

Article structured data is not required for Top Stories or other Google News features, but Google highly recommends it for all articles: it tells Google explicitly that the content is a news article, what its headline is, who wrote it and when, and it can indicate that content is behind a paywall.

How

Render a NewsArticle, BlogPosting or Article block on the server from the article's own fields: headline, image as the visible lead image in 1x1, 4x3 and 16x9 versions, datePublished and dateModified with a UTC offset (DEV-SDA-05), and author as a list of Person or Organization objects with name and a url to the author page. For paywalled articles add isAccessibleForFree false and a hasPart WebPageElement whose cssSelector names the paywalled section.

Test

The Rich Results Test detects the article on one URL per article template with no errors; the dates match the visible dates; every author url opens a real author page.

Code · Article markup with images, dates and authors

Render one block per article on the server from the article's own fields: the headline, the visible lead image in 1x1, 4x3 and 16x9 versions, both dates with a UTC offset, and each author as a Person with a url to a real author page. Google does not require it for Top Stories but highly recommends it for all articles.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "NewsArticle",
  "headline": "How to care for an oak dining table",
  "image": [
    "https://www.example.com/img/oak-care-1x1.webp",
    "https://www.example.com/img/oak-care-4x3.webp",
    "https://www.example.com/img/oak-care-16x9.webp"
  ],
  "datePublished": "2026-10-02T08:00:00+02:00",
  "dateModified": "2026-10-02T09:20:00+02:00",
  "author": [
    {
      "@type": "Person",
      "name": "Jane Doe",
      "url": "https://www.example.com/authors/jane-doe/"
    },
    {
      "@type": "Person",
      "name": "John Roe",
      "url": "https://www.example.com/authors/john-roe/"
    }
  ]
}
</script>

For a paywalled article, add isAccessibleForFree and a hasPart element whose cssSelector names the section behind the paywall (the class must match the page's HTML).

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "NewsArticle",
  "headline": "How to care for an oak dining table",
  "datePublished": "2026-10-02T08:00:00+02:00",
  "isAccessibleForFree": false,
  "hasPart": {
    "@type": "WebPageElement",
    "isAccessibleForFree": false,
    "cssSelector": ".paywall"
  }
}
</script>
Evidence · 5 claims · 2 Google pages
  • StageConsistent with docsD3-C330

    Article structured data is a broad type with many properties and nested types, which Google uses to learn more about published content, such as its headline, author and publication date.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConsistent with docsD3-C331

    Google's rich results slide said Article structured data is not required to appear in Top Stories but is highly recommended for all articles.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C332

    Google's Article documentation says there is no markup requirement to be eligible for Google News features like Top stories, and that Article markup can tell Google more explicitly that content is a news article, who the author is and what the title is.

    Google Search Central

  • SlideConfirmed by docsD3-C334

    Google's Article slide showed the NewsArticle JSON-LD example from Google's Article documentation: a block in the page head with headline, three image URLs (1x1, 4x3 and 16x9), datePublished and dateModified with time zone offsets, and an author list of Person objects with name and url.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C333

    Google's Article rich results slide said structured data can be used to indicate that content is behind a paywall.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

AvoidDocumentedDEV-SDA-14

Do not mark up self-serving review ratings of your own business

Why

Review snippets show an average star rating, and often the number of reviews, for products, local businesses, films, books and other things. Google's review snippet guide makes ratings of local businesses and organizations eligible only on sites that capture reviews about other businesses; reviews a business shows about itself are self-serving and not eligible.

How

Generate Review and AggregateRating markup only from the reviews visible on the page (DEV-SDA-02), for example on product pages, and remove LocalBusiness or Organization ratings about your own business from the site's markup, including those added by review widgets.

Test

The Rich Results Test shows review snippets only on templates with visible reviews, and a crawl finds no rating on the site's own Organization or LocalBusiness entity.

Evidence · 3 claims · 2 Google pages
  • StageConfirmed by docsD3-C337

    Review structured data lets a site specify how its users rated something; the review snippet shows an average star rating and often the number of reviews of a product, service or piece of content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C339

    Review snippets can be used for many kinds of things, from products to local businesses, movies and books.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C340

    Google's review snippet documentation supports ratings for local businesses and organizations only on sites that capture reviews about other businesses; self-serving reviews of a site's own business are not eligible.

    Google Search Central

ShouldDocumentedDEV-SDA-15

State the preferred site name with WebSite markup on the home page

Why

The site name shown with each text result is an attribution element, generated automatically from the home page and references to the site on the web; site owners can indicate their preferred name with WebSite structured data. Google supports one site name per domain or subdomain, not per subdirectory.

How

Add a WebSite block with name, an optional alternateName and url (the canonical home page) to the home page of every domain and subdomain that should have its own name, and use the same name in the home page title and in Organization markup.

Test

curl the home page of each host and find exactly one WebSite block with the intended name and the canonical home page url.

Code · WebSite markup for the preferred site name

Put one WebSite block on the home page of each domain or subdomain (not a subdirectory) that should have its own site name in Google's results. url is the canonical home page; alternateName is an optional fallback such as an acronym.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "WebSite",
  "name": "Example Shop",
  "alternateName": "ExShop",
  "url": "https://www.example.com/"
}
</script>
Evidence · 2 claims · 1 Google page
  • StageConfirmed by docsD3-C335

    Google described the site name as an attribution feature rather than a rich result, and said site owners can influence which site name Google shows for their site in search results.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C336

    Google's site name documentation says site names are generated automatically from a site's home page and references to it on the web, and that a site can indicate its preferred name with WebSite structured data on its home page.

    Google Search Central

Shopping data (Merchant Center, product markup and Storebot) 4

Product rich results come from structured data, the Shopping tab needs a Merchant Center feed, and Google Shopping's own crawler, Storebot-Google, checks both against the shop. AI shopping experiences ask for richer product data than classic results, delivered first through the feed.

ShouldDocumentedDEV-SHP-01

Use Product markup for rich results and a Merchant Center feed for the Shopping tab, with identical values

Why

Product structured data alone makes pages eligible for product rich results and product annotations in Google Images; Merchant Center is not required for them, and Merchant Center does not require the markup. Google's product guide says providing both maximizes eligibility for shopping experiences and helps Google understand and verify the data, and Merchant Center's automatic item updates can use landing-page markup to fix price and availability mismatches, without replacing regular feed updates. Said at Search Central Live: the Shopping tab needs a Merchant Center feed.

How

Generate the feed and the JSON-LD from the same product database so price, sale price, availability, condition, brand and GTIN never diverge (DEV-SDA-08), update the feed whenever prices or stock change, and turn on automatic item updates as a safety net.

Test

Merchant Center diagnostics show no price or availability mismatches, the Merchant listings and Product snippets reports in Search Console show no errors, and a scheduled job compares feed values with the JSON-LD of sampled product pages.

Code · Product markup with a time-limited sale price

Describe only the product the page is about (not the carousel of related products) and only what the page shows. For a sale, price is the sale price, the regular price is a StrikethroughPrice, and validFrom with priceValidUntil (or validThrough) bound the sale in ISO 8601 with a time zone, so a sale price that lingers in cached markup is not treated as current. The template must still switch to the regular price (and drop the StrikethroughPrice) when the sale ends: Google warns that a listing may not display if priceValidUntil is in the past. Keep the dates aligned with the Merchant Center feed. Shipping and returns point by @id alone to the policies defined once in the organisation block.

HTML
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Product",
  "@id": "https://www.example.com/chairs/oak-dining-chair#product",
  "name": "Oak dining chair with woven seat",
  "description": "Solid oak dining chair with a hand-woven paper-cord seat.",
  "image": [
    "https://www.example.com/img/oak-chair-1200.webp",
    "https://www.example.com/img/oak-chair-side-1200.webp"
  ],
  "sku": "CH-OAK-01",
  "brand": {
    "@type": "Brand",
    "name": "Example Shop"
  },
  "offers": {
    "@type": "Offer",
    "url": "https://www.example.com/chairs/oak-dining-chair",
    "price": 149.00,
    "priceCurrency": "EUR",
    "availability": "https://schema.org/InStock",
    "itemCondition": "https://schema.org/NewCondition",
    "validFrom": "2026-11-27T00:00:00+01:00",
    "priceValidUntil": "2026-11-30T23:59:59+01:00",
    "priceSpecification": {
      "@type": "UnitPriceSpecification",
      "priceType": "https://schema.org/StrikethroughPrice",
      "price": 199.00,
      "priceCurrency": "EUR"
    },
    "shippingDetails": {
      "@type": "OfferShippingDetails",
      "hasShippingService": {
        "@id": "https://www.example.com/#standard-shipping"
      }
    },
    "hasMerchantReturnPolicy": {
      "@id": "https://www.example.com/#returns"
    }
  }
}
</script>
Evidence · 9 claims · 3 Google pages
  • StageConfirmed by docsD3-C343

    Google Search displays product rich results from product structured data; it may also use the product data a merchant provides in Google Merchant Center, but Merchant Center is not required for them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C344

    Google's Product structured data introduction says that providing both on-page structured data and a Merchant Center feed maximizes eligibility for shopping experiences and helps Google understand and verify the data; product snippets may take pricing from the feed when the markup lacks it.

    Google Search Central

  • StageConsistent with docsD3-C345

    Product annotations on Google Images results also come from product structured data and do not need Google Merchant Center.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C346

    To appear in the Shopping tab of Google Search, a product must be submitted in Google Merchant Center; product structured data alone is not enough.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C349

    Product structured data can help Merchant Center in some cases, for example during data validation, but Merchant Center does not require it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C350

    Merchant Center's automatic item updates use structured data and other product data found on landing pages to fix price, sale price, availability and condition mismatches in a merchant's product data; Google says they do not replace regular feed updates.

    Google Merchant Center Help

  • StageConfirmed by docsD3-C354

    Product snippets in Google Search are powered by both shopping feeds and schema.org structured data.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C402

    Basic, well-specified product attributes such as price, brand and availability remain as important as ever for Google Shopping.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C351

    For an online shop, Product structured data is enough for product rich results and Google Images product annotations, but the Shopping tab needs a Merchant Center feed too; keep feed and markup values identical, since Google can use the markup to validate the feed.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-SHP-02

Let Storebot-Google crawl product, cart and checkout pages, through robots.txt and bot protection

Why

Google Shopping crawls with its own user agent, Storebot-Google, because it needs fresh prices, availability and shipping details. It visits product, cart and checkout pages, can fill in checkout forms and records price, shipping, availability, coupons and payment methods to verify the data merchants send. robots.txt rules addressed to Storebot-Google affect all Google Shopping surfaces, and a crawler without its own group follows the * group, so a shop that disallows /cart/ and /checkout/ for every crawler also blocks these checks.

How

Give Storebot-Google its own robots.txt group that repeats the general rules but leaves cart and checkout paths open; verify it like Googlebot, as one of Google's common crawlers (DEV-SRV-01), and exempt it from rate limits, CAPTCHAs and bot challenges on product, cart and checkout flows. Block it only if the shop deliberately stays out of Google Shopping.

Test

Google's open-source robots.txt parser, run with the Storebot-Google token, returns allowed for a product URL, /cart/ and /checkout/; server logs show Storebot-Google requests answered with 200, not 403, 429 or challenge pages.

Code · robots.txt that keeps cart and checkout open for Storebot-Google

A crawler follows only the most specific group that names it, so Storebot-Google, Google Shopping's crawler, uses its own group and ignores *. Repeat the general rules there but leave cart and checkout open, so it can verify prices, shipping and availability; every other crawler still stays out of them.

Text
# https://www.example.com/robots.txt
User-agent: *
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
Disallow: /api/
Allow: /api/products/

# Google Shopping's crawler: same rules, but cart and checkout stay crawlable
User-agent: Storebot-Google
Disallow: /search?
Disallow: /api/
Allow: /api/products/

Sitemap: https://www.example.com/sitemap.xml
Evidence · 6 claims · 2 Google pages
  • StageConsistent with docsD3-C385

    Google Shopping uses its own crawler, Storebot (Storebot-Google), instead of Googlebot because it needs fresh product information, prices, availability and shipping details and therefore crawls much more often.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C386

    Storebot-Google is used to validate merchant feeds: when a feed gives a price, Google double-checks it on the merchant's site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C387

    Storebot-Google also does deep crawls, for example of checkout pages, which do not belong in the search index; this is another reason it is kept separate from Googlebot.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C388

    Google's Merchant Center help says the StoreBot crawler goes through product detail, cart and checkout pages, can fill in checkout forms, and records price, shipping, availability, coupons and payment methods to verify the data merchants share in Merchant Center.

    Google Merchant Center Help

  • DocsSourceD3-C389

    Google's crawler list says robots.txt rules addressed to the Storebot-Google user agent affect all surfaces of Google Shopping, such as the Shopping tab in Google Search.

    Google

  • AnalysisD3-C390

    Treat Storebot-Google separately from Googlebot in robots.txt and bot protection: a shop that blocks it on product, cart or checkout paths can undermine Shopping price and availability checks, even when Googlebot is allowed.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-SHP-03

Send conversational product data in the Merchant Center feed: Q&A, product details, highlights and related products

Why

In 2026 Google added Merchant Center attributes intended mainly for conversational experiences such as AI Mode (question_and_answer, document_link, related_product and popularity_rank) and more flexible variant data (item_group_title and variant_option), and it now stresses the existing product_highlight and product_detail attributes for AI. In Google's testing with one brand, the conversational attributes it submitted were used in 50% of relevant product recommendations in AI Mode.

How

Extend the feed generator: up to 30 question_and_answer pairs per product (question and answer up to 1,000 characters each, 10,000 in total, with no prices, shipping, dates or company name), up to 100 product_detail specifications as section, attribute name and value, 2 to 100 product_highlight lines of up to 150 characters without promotional text, up to five document_link URLs of PDFs you hold the rights to, related_product for accessories and parts, and popularity_rank as a 0-100 rank against your own inventory. Source the answers from real customer questions, and keep the core attributes (price, availability, brand, GTIN) accurate first.

Test

Merchant Center reports no errors for the new attributes, and a sampled product's Q&A, details and highlights in the feed match its product page.

Code · Merchant Center feed item with Q&A, product details and highlights

An XML (RSS 2.0) feed item that adds the attributes Google stresses for AI shopping experiences to the core data: product_highlight lines, product_detail specifications as section, name and value, and question_and_answer pairs (no prices, shipping, dates or company name in them). Keep price, availability, brand and GTIN identical to the product page's markup.

XML
<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:g="http://base.google.com/ns/1.0">
  <channel>
    <title>Example Shop</title>
    <link>https://www.example.com/</link>
    <description>Example Shop product feed</description>
    <item>
      <g:id>KT-17-STEEL</g:id>
      <g:title>Steel kettle 1.7 l with temperature control</g:title>
      <g:description>Stainless steel kettle with five temperature settings and a keep-warm mode.</g:description>
      <g:link>https://www.example.com/kettles/steel-kettle-17</g:link>
      <g:image_link>https://www.example.com/img/steel-kettle-17-1200.jpg</g:image_link>
      <g:price>59.00 EUR</g:price>
      <g:availability>in_stock</g:availability>
      <g:condition>new</g:condition>
      <g:brand>Example Home</g:brand>
      <g:gtin>4006381333931</g:gtin>
      <g:product_highlight>Five temperature settings from 70 to 100 degrees Celsius</g:product_highlight>
      <g:product_highlight>Keeps water warm for up to 30 minutes</g:product_highlight>
      <g:product_detail>
        <g:section_name>General</g:section_name>
        <g:attribute_name>Capacity</g:attribute_name>
        <g:attribute_value>1.7 l</g:attribute_value>
      </g:product_detail>
      <g:product_detail>
        <g:section_name>Power</g:section_name>
        <g:attribute_name>Wattage</g:attribute_name>
        <g:attribute_value>2200 W</g:attribute_value>
      </g:product_detail>
      <g:question_and_answer>
        <g:question>Does the kettle switch off when it boils dry?</g:question>
        <g:answer>Yes. Boil-dry protection switches it off when there is no water in it.</g:answer>
      </g:question_and_answer>
      <g:question_and_answer>
        <g:question>Can I set it to 80 degrees for green tea?</g:question>
        <g:answer>Yes. The 80 degree setting is one of the five presets.</g:answer>
      </g:question_and_answer>
    </item>
  </channel>
</rss>
Evidence · 13 claims · 9 Google pages
  • StageConfirmed by docsD3-C365

    Google added six Merchant Center feed attributes for AI shopping experiences: question and answer, documents, related products, item group title and variant option for variants, and popularity rank.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C367

    Google's Merchant Center help marks the question_and_answer, document_link, related_product and popularity_rank attributes as primarily intended for conversational experiences such as AI Mode in Google Search.

    Google Merchant Center Help

  • StageConfirmed by docsD3-C368

    The question and answer feed attribute lets a merchant or brand upload questions and answers relevant to a product, whether they come from the manufacturer or from consumers.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C369

    Merchant Center's question_and_answer attribute takes up to 30 question-and-answer pairs per product, each question and answer up to 1,000 characters and 10,000 characters in total, and must not contain prices, shipping, dates or the company name.

    Google Merchant Center Help

  • DocsSourceD3-C371

    Merchant Center's document_link attribute takes up to five URLs of PDF documents about a product, such as manuals, user guides or assembly instructions, and the merchant must own the content licensing rights.

    Google Merchant Center Help

  • StageConfirmed by docsD3-C372

    The related products feed attribute describes accessories, parts and replacement products, answering frequent questions such as what else to buy with a printer.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C373

    Variant data in Google's feed specification used to be limited to basics such as colour and size; the new item group title and variant option attributes make it much more flexible.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C375

    Merchant Center's popularity_rank is a 0-100 value the merchant assigns to rank a product's popularity, based on its recent sales, against the rest of its own inventory; it does not reflect user ratings.

    Google Merchant Center Help

  • StageConfirmed by docsD3-C392

    Google now emphasises two existing Merchant Center attributes for AI, product highlights (a short bulleted list) and product detail (detailed specifications), and updated its Help Center guidance to say they help AI systems.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C393

    Merchant Center's product_highlight attribute takes 2 to 100 highlights of up to 150 characters each that describe only the product itself, without promotional text, keywords or search terms.

    Google Merchant Center Help

  • StageConfirmed by docsD3-C394

    Instead of adding feed attributes for individual specifications such as connectivity, memory, camera settings or box contents, Google asks merchants to send them as key-value pairs in product detail, because merchants know best what matters.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C395

    Merchant Center's product_detail attribute takes up to 100 specifications, each a section name, attribute name and attribute value, and Google's help says clean key-value specifications improve product details on Google Shopping and AI-driven surfaces.

    Google Merchant Center Help

  • DocsSourceD3-C408

    Google's 16 September 2026 holiday shopping post says that in testing with lululemon, conversational attributes submitted by the brand were used 50% of the time in relevant product recommendations in AI Mode.

    Google blog (16 September 2026)

MaySaid at Search Central LiveDEV-SHP-04

Treat product Q&A and specification markup in schema.org as an extra on top of the feed

Why

Said at Search Central Live: most of the new conversational feed attributes already exist in schema.org and Google can use them from product markup too, product Q&A can be marked up as Question items, each with an acceptedAnswer, linked to the Product through subjectOf, and a recent schema.org release added specification and valueGroup for product details. As of 3 October 2026 neither Search Central nor the Merchant Center attribute pages document such markup (only item_group_title maps, to ProductGroup.name), so the feed is the documented route, and the extra markup makes pages heavier.

How

Keep the Merchant Center feed as the source (DEV-SHP-03). If you add markup, generate it from the same Q&A and specification data the page visibly shows (DEV-SDA-02), attach Question items with Answer objects to the Product through subjectOf, and watch the HTML size (DEV-PRF-04).

Test

The JSON-LD parses, the Rich Results Test still reports the product without errors, and every marked-up question and answer is visible on the page.

Evidence · 6 claims · 8 Google pages
  • StageNot in docsD3-C391

    Product Q&A can be marked up in schema.org by linking Question items, each with an acceptedAnswer of type Answer, to the product through the subjectOf property, because the product is the subject of the Q&A.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C380

    Most of the new conversational feed attributes were already available in schema.org, so Google ties them back to structured data and can use them from product markup as well as from feeds.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C399

    A recent schema.org release added two properties, specification and valueGroup, so that the Merchant Center product detail attribute can be expressed in schema.org, with product specs nested under specification.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C406

    Google Shopping named the downside of adding Q&As and documentation to product markup: the page's markup becomes somewhat bigger.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C381

    As of 3 October 2026 Google's Merchant Center help lists no schema.org property for question_and_answer, document_link, related_product, popularity_rank, variant_option or product_detail (only item_group_title maps, to ProductGroup.name), and Search Central does not document such markup: reading these attributes from schema.org is a stage statement, the feed is the documented route.

    Ibrahim Anjro (author)

  • AnalysisD3-C400

    Schema.org's release notes show version 30.1 (16 September 2026) added specification, isOftenBoughtWith and consumerNotice for Product, valueGroup for PropertyValue and itemPopularity for Offer as vocabulary for retail feed data; Google's Search Central docs did not yet document these properties on 3 October 2026.

    Ibrahim Anjro (author)

Images 8

Google finds images through img elements in the HTML and through image sitemaps, understands them through alt text and the surrounding text, and can show them in web search, Google Images, Discover and AI features.

MustDocumentedDEV-IMG-01

Put every image that should be found in an <img> element with a src attribute

Why

Google finds images with an HTML parser that looks for img elements, and src is the most important attribute: without it Google does not know where the image is. The picture element is supported only through the img it contains, and CSS background images are not extracted.

How

Product, article and other content images use img with src, inside picture where art direction or modern formats are needed. Keep CSS backgrounds for decoration. Lazy-loaded images keep a real src with loading="lazy", not only a data-src filled in by a script.

Test

In the rendered HTML (URL Inspection) key images are img elements with a src; a template search finds no background-image on content images.

Code · Indexable, responsive image with a modern-format fallback

Google extracts images from <img src>; a <picture> element counts only through the <img> inside it, and CSS background images are not extracted. Serve AVIF or WebP through <source> elements, keep a widely supported file in the img src, describe the image in alt and put a caption or explanatory text next to it.

HTML
<figure>
  <picture>
    <source type="image/avif"
            srcset="/img/oak-chair-800.avif 800w, /img/oak-chair-1600.avif 1600w"
            sizes="(max-width: 800px) 100vw, 800px">
    <source type="image/webp"
            srcset="/img/oak-chair-800.webp 800w, /img/oak-chair-1600.webp 1600w"
            sizes="(max-width: 800px) 100vw, 800px">
    <img src="/img/oak-chair-800.jpg"
         srcset="/img/oak-chair-800.jpg 800w, /img/oak-chair-1600.jpg 1600w"
         sizes="(max-width: 800px) 100vw, 800px"
         alt="Oak dining chair with a woven paper-cord seat, seen from the front"
         width="800" height="600" loading="lazy">
  </picture>
  <figcaption>The oak dining chair in a natural oak finish, seat height 45 cm.</figcaption>
</figure>

<!-- Not indexable as an image: keep CSS backgrounds for decoration only -->
<div class="hero" style="background-image: url('/img/oak-chair-hero.jpg')"></div>
Evidence · 6 claims · 1 Google page
  • StageConsistent with docsD2-C527

    One of the most reliable ways to make sure Google finds an image is to include it in an img element in the HTML; Gary Illyes named image sitemaps as the other method.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C528

    Google does not extract CSS background images.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C526

    Google supports the picture element only because a picture element must contain an img element, and that img element is what Google extracts and passes to its media indexer.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C532

    The src attribute is the most important img attribute, because without it Google does not know where the image bytes are and cannot index the image.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C524

    Google extracts images and videos from the page's document object model (DOM), and finds images with a fairly standard HTML parser that looks for img elements.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C538

    Images that should rank need an img element with a real src; hero or product images set as CSS backgrounds are invisible to Google Images and image features, so keep CSS backgrounds for decorative images you do not need indexed.

    Ibrahim Anjro (author)

MustDocumentedDEV-IMG-02

Describe every meaningful image in its alt attribute

Why

Google calls alt text the most important attribute for image metadata and uses it, with computer vision and the page content, to understand the image. Alt text should describe the image for someone who cannot see it, and the image can rank for what it describes.

How

Make alt text a required CMS field for content images, write specific descriptions rather than file names or keyword lists, and give decorative images an empty alt="". If alt text is generated with an LLM, as a community speaker showed through a crawler's custom JavaScript, first run it on a few chosen URLs and review the output, and give the model the page title as well as the image so the text carries the page's intent.

Test

A crawl finds no content image with a missing alt, an empty alt or an alt equal to the file name; the axe or Lighthouse image-alt check passes.

Evidence · 5 claims · 1 Google page
  • DocsSourceD2-C534

    Google's image SEO guide calls alt text the most important attribute for providing more metadata about an image, and says Google uses it together with computer vision and the page content to understand the image.

    Google Search Central

  • StageConfirmed by docsD2-C535

    Alt text should describe the image in words for someone who cannot see it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C536

    Google uses the words in an alt attribute to understand the image, may attach them to the image at serving time, and the image can rank for concepts the alt text describes.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C949

    To test LLM-generated alt text before a rollout, a community speaker recommended running the crawler's custom JavaScript in Screaming Frog's List mode on a few chosen URLs and reviewing the generated alt text, instead of crawling the whole site in Spider mode.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C950

    Because not every generated alt text can be checked by hand, a community speaker's safer script adds the page title to the LLM's image description, so the alt text carries the page's intent and not only what the image shows.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-IMG-03

Place images next to text that explains them

Why

The text around an image is critical: Google uses it as context to understand and rank the image, so alt text alone is not enough. An extracted image can then appear almost anywhere Google shows results.

How

Put images in the main content near the paragraph they illustrate, add a figcaption or explanatory sentence to key images, and avoid image-only galleries without text.

Test

Template review: key images sit inside main with a caption or adjacent descriptive text.

Code · Indexable, responsive image with a modern-format fallback

Google extracts images from <img src>; a <picture> element counts only through the <img> inside it, and CSS background images are not extracted. Serve AVIF or WebP through <source> elements, keep a widely supported file in the img src, describe the image in alt and put a caption or explanatory text next to it.

HTML
<figure>
  <picture>
    <source type="image/avif"
            srcset="/img/oak-chair-800.avif 800w, /img/oak-chair-1600.avif 1600w"
            sizes="(max-width: 800px) 100vw, 800px">
    <source type="image/webp"
            srcset="/img/oak-chair-800.webp 800w, /img/oak-chair-1600.webp 1600w"
            sizes="(max-width: 800px) 100vw, 800px">
    <img src="/img/oak-chair-800.jpg"
         srcset="/img/oak-chair-800.jpg 800w, /img/oak-chair-1600.jpg 1600w"
         sizes="(max-width: 800px) 100vw, 800px"
         alt="Oak dining chair with a woven paper-cord seat, seen from the front"
         width="800" height="600" loading="lazy">
  </picture>
  <figcaption>The oak dining chair in a natural oak finish, seat height 45 cm.</figcaption>
</figure>

<!-- Not indexable as an image: keep CSS backgrounds for decoration only -->
<div class="hero" style="background-image: url('/img/oak-chair-hero.jpg')"></div>
Evidence · 3 claims · 1 Google page
  • StageConsistent with docsD2-C537

    The text around an image is critical: Google uses it as context to understand the image and to rank it, so an alt attribute alone is not enough.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C525

    An image Google has extracted can appear almost anywhere Google shows results, including Discover, image search, web search and AI features, so its potential reach is immense.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C539

    Give important images a caption or explanatory sentence next to them rather than relying on alt text alone, since Gary Illyes ranked src above alt and called the surrounding text critical for ranking the image.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-IMG-04

List important images in an image sitemap

Why

Besides img elements, image sitemaps tell Google about the images on each page, which helps for images that JavaScript loads or that are otherwise hard to find.

How

Add image:image entries with image:loc under each page's url in the sitemap, or in a separate image sitemap, up to 1,000 images per page, and list only images you want indexed.

Test

The Sitemaps report reads the file, and the image URLs in it return 200 and are not disallowed.

Code · Image sitemap

Image sitemaps list, under each page's <loc>, the images that page uses, which helps Google find images that are loaded by JavaScript or are otherwise hard to discover. Only image:image and image:loc are used (caption, title, geo location and license tags are deprecated); up to 1,000 images per page.

XML
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
        xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
  <url>
    <loc>https://www.example.com/chairs/oak-dining-chair</loc>
    <image:image>
      <image:loc>https://www.example.com/img/oak-chair-1200.webp</image:loc>
    </image:image>
    <image:image>
      <image:loc>https://www.example.com/img/oak-chair-side-1200.webp</image:loc>
    </image:image>
  </url>
</urlset>
Evidence · 2 claims · 1 Google page
  • StageConfirmed by docsD2-C540

    Besides img elements, image sitemaps tell Google about images: they are XML sitemaps that list, under a page's loc entry, the locations of the images on that page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C527

    One of the most reliable ways to make sure Google finds an image is to include it in an img element in the HTML; Gary Illyes named image sitemaps as the other method.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-IMG-05

Serve modern, compact image formats with a widely supported file in the img src

Why

Google supports pretty much all popular image formats and recommends high-quality modern formats such as WebP with small files. Its documentation lists AVIF as supported, but Gary Illyes said at Search Central Live that AVIF currently has hiccups and Google may have problems ingesting it.

How

Offer AVIF and WebP through source elements inside picture, keep a WebP or JPEG file in the img src, set width and height, compress the files and serve responsive sizes with srcset.

Test

Lighthouse format and sizing checks pass; the img src of key images is WebP or JPEG; key images appear in Google Images after launch.

Code · Indexable, responsive image with a modern-format fallback

Google extracts images from <img src>; a <picture> element counts only through the <img> inside it, and CSS background images are not extracted. Serve AVIF or WebP through <source> elements, keep a widely supported file in the img src, describe the image in alt and put a caption or explanatory text next to it.

HTML
<figure>
  <picture>
    <source type="image/avif"
            srcset="/img/oak-chair-800.avif 800w, /img/oak-chair-1600.avif 1600w"
            sizes="(max-width: 800px) 100vw, 800px">
    <source type="image/webp"
            srcset="/img/oak-chair-800.webp 800w, /img/oak-chair-1600.webp 1600w"
            sizes="(max-width: 800px) 100vw, 800px">
    <img src="/img/oak-chair-800.jpg"
         srcset="/img/oak-chair-800.jpg 800w, /img/oak-chair-1600.jpg 1600w"
         sizes="(max-width: 800px) 100vw, 800px"
         alt="Oak dining chair with a woven paper-cord seat, seen from the front"
         width="800" height="600" loading="lazy">
  </picture>
  <figcaption>The oak dining chair in a natural oak finish, seat height 45 cm.</figcaption>
</figure>

<!-- Not indexable as an image: keep CSS backgrounds for decoration only -->
<div class="hero" style="background-image: url('/img/oak-chair-hero.jpg')"></div>
Evidence · 5 claims · 2 Google pages
  • StageConsistent with docsD2-C544

    Use high-quality modern image formats such as WebP for a good balance of quality and compression, and keep image files small so people can enjoy them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C541

    Google supports pretty much all of the most popular image formats on the web.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C543

    Google's image SEO guide lists AVIF among the supported image formats (with BMP, GIF, JPEG, PNG, WebP and SVG), and Google announced in August 2024 that AVIF files need nothing special to be indexed.

    Google Search Central, Search Central blog (30 August 2024)

  • StageNot in docsD2-C542

    Gary Illyes said the AVIF image format currently has hiccups and Google may have problems ingesting it, although it should technically be supported.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C545

    While the AVIF ingestion hiccup Gary Illyes mentioned lasts, serve AVIF only through source elements inside picture and keep a WebP or JPEG file in the img src, which is the URL Google extracts, even though Google's documentation lists AVIF as supported; then check in Google Images that key images appear.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-IMG-06

Keep images out of Search with robots.txt or X-Robots-Tag, not with CSS tricks

Why

Google's documented ways to keep a site's images out of search results are a robots.txt disallow (for example for Googlebot-Image) or a noindex X-Robots-Tag header, with the Removals tool for emergencies. A CSS background image, suggested on stage, only avoids extraction from that page; the same image URL used in an img element elsewhere or listed in a sitemap can still be indexed.

How

Disallow the image paths for Googlebot-Image, or send X-Robots-Tag: noindex for them, but not both for the same URLs, because the header is not seen on a disallowed URL.

Test

Google's open-source robots.txt parser (github.com/google/robotstxt), run with the Googlebot-Image token, returns disallowed for the image paths, or curl -I on an image shows the X-Robots-Tag header; after recrawl the images no longer appear in Google Images.

Code · X-Robots-Tag header for PDFs, images and other non-HTML files

Files that cannot carry a meta tag get their robots rules as an HTTP response header. The same rules as the meta tag apply (noindex, nosnippet, max-snippet...), and the URL must stay crawlable for Google to see the header.

HTTP
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex
nginx
# Inside the server block: office documents out of the index.
# An add_header in a location replaces every add_header set at server level for those responses:
# repeat your HSTS and CSP lines in each such location (or use the headers-more module).
location ~* \.(pdf|docx?|xlsx?)$ {
    add_header X-Robots-Tag "noindex" always;
    add_header Strict-Transport-Security "max-age=31536000" always;  # repeated from the server block
}

# Images that must not appear in Google Images (they still display on your pages)
location ^~ /internal-images/ {
    add_header X-Robots-Tag "noindex" always;
    add_header Strict-Transport-Security "max-age=31536000" always;  # repeated from the server block
}
Apache
<FilesMatch "\.(pdf|docx?|xlsx?)$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>
Evidence · 5 claims · 1 Google page
  • DocsSourceD2-C529

    Google's documented ways to keep a site's images out of search results are a robots.txt disallow rule (for example for Googlebot-Image) or a noindex X-Robots-Tag HTTP header, with the Removals tool for emergencies.

    Google Search Central

  • AnalysisD2-C530

    Hiding an image as a CSS background is a fragile way to keep it out of Google: it only stops extraction from that page, so the same image URL used in an img element elsewhere or listed in a sitemap can still be indexed; the documented robots.txt or noindex X-Robots-Tag methods are the reliable route.

    Ibrahim Anjro (author)

  • StageNot in docsD2-C826

    Gary Illyes suggested a div with a CSS background image as a way to keep an image from being picked up by Google.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C931

    Googlebot-Image and Googlebot-Video are the Google crawlers that fetch images and videos, and robots.txt rules addressed to their user agent tokens control how Google indexes a site's images and videos.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C932

    If robots.txt disallows the location of an image or video file, Google does not index that image or video.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-IMG-07

Give every article a large, relevant lead image and declare it in og:image or markup

Why

For Discover, Google recommends large images at least 1,200 px wide, with more than 300,000 pixels and a 16x9 aspect ratio, enabled by max-image-preview:large (DEV-IDX-04); a small thumbnail lowers click-through and a large one raises it. Avoid generic images such as the site logo, text-heavy images and misleading or exaggerated ones. The image named in schema.org markup or og:image can influence the thumbnail Discover chooses.

How

Make a lead image at least 1,200 px wide a required field of article templates, output it in og:image and in the Article markup's image list (DEV-SDA-13), crop 16x9 versions that keep the important details, and fall back to no image rather than the logo.

Test

A crawl of article URLs finds on each an og:image at least 1,200 px wide that is not the logo; Search Console's Discover report shows impressions for article templates.

Code · Large-image preview and lead image for Discover

Allow large image previews and name a large, relevant lead image (at least 1,200 px wide, 16x9, not the logo and not text-heavy) in the head of every article template; the same image belongs in the Article markup's image list.

HTML
<head>
  <meta name="robots" content="max-snippet:-1, max-image-preview:large">
  <meta property="og:image" content="https://www.example.com/img/oak-care-16x9.webp">
  <meta property="og:image:width" content="1600">
  <meta property="og:image:height" content="900">
  <meta property="og:image:alt" content="Hand applying oil to an oak table top">
</head>
Evidence · 6 claims · 1 Google page
  • SlideConfirmed by docsD3-C224

    Google's Discover slide recommended large images at least 1,200 px wide, with more than 300,000 total pixels and a 16x9 aspect ratio, enabled by the max-image-preview:large setting.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C225

    Google's Discover slide said to avoid generic images, such as a site logo, and text-heavy images.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConsistent with docsD3-C226

    Google's Discover slide said images are subject to the same policy requirements as the content and that misleading, exaggerated or outrageous images should be avoided.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C227

    Google's Discover documentation says to specify the large image with schema.org markup or the og:image meta tag, which can influence the thumbnail Discover chooses, and that a vertical image cropped to 16x9 must keep its important details.

    Google Search Central

  • SlideConsistent with docsD3-C221

    Google's Discover slide said that for eligible content image quality decides click-through: a small thumbnail lowers the click-through rate and a large, high-resolution image raises it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C230

    For Discover, set max-image-preview:large on article templates, give every article a relevant image at least 1,200 px wide in og:image or schema.org markup that is neither the site logo nor text-heavy, and show clear bylines and dates.

    Ibrahim Anjro (author)

MayDocumentedDEV-IMG-08

Keep the IPTC and C2PA metadata that marks AI-generated images

Why

Whether a site uses AI-generated images is up to its owner, Google said at the event, but they must work for users and be checked for errors such as garbled text, which image models typically render poorly. Google Images supports the IPTC Digital Source Type values for algorithmically created images, such as trainedAlgorithmicMedia, and can show C2PA details in "About this image".

How

Keep the IPTC DigitalSourceType and any C2PA manifest from the generating tool when images are resized or compressed (many image pipelines strip metadata by default), and review AI images before publishing (DEV-SPM-05).

Test

exiftool -XMP-iptcExt:DigitalSourceType on a published AI-generated image shows the source type that the original file had.

Evidence · 4 claims · 2 Google pages
  • DocsSourceD2-C943

    Google's image metadata guide says Google Images supports the IPTC Digital Source Type values for algorithmically created images, such as trainedAlgorithmicMedia, and can show C2PA details in 'About this image', such as whether an image was created or edited with AI tools.

    Google Search Central

  • StageConsistent with docsD2-C937

    To the frequent question about AI-generated images on a site, Google's answer is that it is up to the site owner.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C941

    Sites that use AI-generated images or videos should make sure they work for users, check them for hallucinations and regenerate them where needed.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C939

    The diffusion models that generate images were built to generate images, not text, so they are typically poor at rendering text inside an image.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Video 6

Google extracts videos from the page together with the data around them and passes them to its media indexing. A dedicated watch page, a video element that loads without interaction and VideoObject markup make a video eligible for video features such as key moments.

MustDocumentedDEV-VID-01

Embed each video in a video, iframe, embed or object element that loads without user action

Why

Google finds videos referenced in these elements and extracts the video with the data around it. A video that loads only after a click, swipe or typing is not found.

How

Output the player element in the HTML at load (a poster frame is fine), at its real size and position and not hidden behind other elements, and for lazy players keep the element and its src or embed URL in the DOM. The Video indexing report flags "Cannot determine video position and size" when the player is not on the page at load. Said at Search Central Live: without a video container in the HTML Google treats the page as having no video, and the video should be embedded prominently, above the fold.

Test

The rendered HTML in URL Inspection contains the video element or iframe with its URL, and the Video indexing report lists the page.

Code · Video watch page: video element plus VideoObject with key moments

Put the video in a <video> (or <iframe>/<embed>/<object>) element that loads without a click, on a page where it is the main content, and describe it with VideoObject markup. Clip parts enable key moments; their url points to the same watch page at that second.

HTML
<video controls preload="metadata" width="1280" height="720"
       poster="https://www.example.com/video/chair-assembly-1280.jpg">
  <source src="https://www.example.com/video/chair-assembly.mp4" type="video/mp4">
</video>

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "VideoObject",
  "name": "How to assemble the oak dining chair",
  "description": "Step-by-step assembly of the oak dining chair in four minutes, with the tools you need.",
  "thumbnailUrl": ["https://www.example.com/video/chair-assembly-1280.jpg"],
  "uploadDate": "2026-09-15T09:00:00+02:00",
  "duration": "PT4M12S",
  "contentUrl": "https://www.example.com/video/chair-assembly.mp4",
  "embedUrl": "https://www.example.com/embed/chair-assembly",
  "hasPart": [
    {
      "@type": "Clip",
      "name": "Unpacking the parts",
      "startOffset": 0,
      "endOffset": 35,
      "url": "https://www.example.com/videos/chair-assembly?t=0"
    },
    {
      "@type": "Clip",
      "name": "Attaching the legs",
      "startOffset": 35,
      "endOffset": 140,
      "url": "https://www.example.com/videos/chair-assembly?t=35"
    },
    {
      "@type": "Clip",
      "name": "Fitting the seat",
      "startOffset": 140,
      "endOffset": 252,
      "url": "https://www.example.com/videos/chair-assembly?t=140"
    }
  ]
}
</script>
Evidence · 5 claims · 2 Google pages
  • StageConsistent with docsD2-C448

    For videos, Google's feature extraction takes the video itself and the data around it, to get a better sense of what happens in the video, and passes them to Google's media indexing engine.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C551

    Google strongly suggests describing a video's metadata with JSON-LD structured data, on top of placing the video in an HTML element on the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C918

    Gary Illyes called a video container in the page's HTML critical: without one, Google treats the page as having no video.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C942

    Google's Video indexing report help says only videos on a watch page are eligible for indexing, and flags 'Cannot determine video position and size' when the player is not on the page at load, for example behind a click-to-play image, asking for the player to load at its real size and position without user interaction.

    Google Search Console Help

  • StageNot in docsD2-C924

    For a video to be discovered it must be embedded prominently, above the fold; Gary Illyes said a video placed below the fold is not going to be indexed.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-VID-02

Give each important video its own watch page

Why

A video must be embedded on an indexed watch page to be eligible for video features, and Google recommends a dedicated watch page per video where it makes sense. A Google slide listed a dedicated watch page, compelling titles and descriptions, relevant thumbnails, markup, fast pages and sitemap inclusion as key factors.

How

Give each important video one URL where it is the main content, with a title and description specific to that video, a transcript or summary, and links from related articles. Gary Illyes said a video without a dedicated watch page is not going to be indexed, and that descriptive text around the video helps Google rank and retrieve it.

Test

The Video indexing report shows the watch pages indexed with a detected video and no "video is not the main content" issue.

Evidence · 6 claims · 2 Google pages
  • DocsSourceD2-C554

    Google's video SEO guide says a video must be embedded on an indexed watch page to be eligible for video features and recommends a dedicated watch page per video where it makes sense for the business; a non-watch page with the video can still appear as a text result or a Google Images result with a video badge.

    Google Search Central

  • SlideConsistent with docsD2-C553

    Google's video SEO slide said each video must be on its own dedicated HTML watch page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • SlideConsistent with docsD2-C552

    Google's slide listed seven key factors for video SEO success: high-quality video content, a dedicated watch page, compelling titles and descriptions, relevant thumbnails, video markup, fast-loading pages and sitemap inclusion.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C557

    Give each important video its own watch page where the video is the main content, add VideoObject markup and list the video in a video sitemap; a video buried in a long article lacks the dedicated watch page that Google listed as a key factor.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD2-C916

    Gary Illyes said a video without a dedicated watch page is not going to be indexed.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C927

    Descriptive text around a video helps Google rank and retrieve the video.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-VID-03

Describe videos with VideoObject markup and mark key moments with Clip or SeekToAction

Why

Google strongly suggests describing a video's metadata with JSON-LD in addition to the HTML element. Video features such as key moments and previews build on it: Clip markup gives exact segments, and SeekToAction tells Google how timestamps appear in your URLs.

How

VideoObject with name, description, thumbnailUrl, uploadDate in ISO 8601, duration, and contentUrl or embedUrl; hasPart Clip items with name, startOffset, endOffset and a url pointing at the same watch page at that second. Keep thumbnails crawlable and make them show the video's real content: said at Search Central Live, Google's experiments show a very high drop-out at the start of videos whose thumbnail is wrong.

Test

The Rich Results Test detects the video and its clips, and the Video indexing report shows it as indexed.

Code · Video watch page: video element plus VideoObject with key moments

Put the video in a <video> (or <iframe>/<embed>/<object>) element that loads without a click, on a page where it is the main content, and describe it with VideoObject markup. Clip parts enable key moments; their url points to the same watch page at that second.

HTML
<video controls preload="metadata" width="1280" height="720"
       poster="https://www.example.com/video/chair-assembly-1280.jpg">
  <source src="https://www.example.com/video/chair-assembly.mp4" type="video/mp4">
</video>

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "VideoObject",
  "name": "How to assemble the oak dining chair",
  "description": "Step-by-step assembly of the oak dining chair in four minutes, with the tools you need.",
  "thumbnailUrl": ["https://www.example.com/video/chair-assembly-1280.jpg"],
  "uploadDate": "2026-09-15T09:00:00+02:00",
  "duration": "PT4M12S",
  "contentUrl": "https://www.example.com/video/chair-assembly.mp4",
  "embedUrl": "https://www.example.com/embed/chair-assembly",
  "hasPart": [
    {
      "@type": "Clip",
      "name": "Unpacking the parts",
      "startOffset": 0,
      "endOffset": 35,
      "url": "https://www.example.com/videos/chair-assembly?t=0"
    },
    {
      "@type": "Clip",
      "name": "Attaching the legs",
      "startOffset": 35,
      "endOffset": 140,
      "url": "https://www.example.com/videos/chair-assembly?t=35"
    },
    {
      "@type": "Clip",
      "name": "Fitting the seat",
      "startOffset": 140,
      "endOffset": 252,
      "url": "https://www.example.com/videos/chair-assembly?t=140"
    }
  ]
}
</script>
Evidence · 3 claims · 2 Google pages
  • StageConsistent with docsD2-C551

    Google strongly suggests describing a video's metadata with JSON-LD structured data, on top of placing the video in an HTML element on the page.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C550

    Google has video-specific features, such as key moments and previews, that help users interact with videos more easily.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C917

    Google knows from experiments that when a video's thumbnail is wrong, the share of viewers who drop out at the start of the video is extremely high.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-VID-04

List videos in a video sitemap

Why

Including videos in a sitemap helps Google discover all of a site's video content. Said at Search Central Live: video sitemaps are not critical but good to have, because Google ingests them much more often than it can process HTML pages, and they tell it which URLs to check for video instead of checking every page.

How

Add video:video entries (thumbnail, title, description, and content or player location) for each watch page, in the page sitemap or a separate video sitemap.

Test

The Sitemaps report reads the video sitemap, and its video count matches the Video indexing report.

Evidence · 3 claims · 1 Google page
  • SlideConfirmed by docsD2-C556

    Including videos in a sitemap helps Google discover all of a site's video content.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C925

    Video sitemaps are not critical but good to have, because Google ingests video sitemaps much more often than it can process HTML pages.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C926

    A video sitemap or video feed tells Google which URLs carry videos so it can visit those URLs to double-check; without one, Google has to check every page individually.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MayDocumentedDEV-VID-05

Allow video previews with max-video-preview:-1

Why

max-video-preview sets the maximum number of seconds of a video preview in Search: 0 allows only a static image and -1 means no limit. Google said it does not expect -1 to produce an hour-long preview.

How

Add max-video-preview:-1 to the robots meta tag of watch pages unless licensing requires a limit.

Test

curl the watch page and check the robots meta content.

Code · Robots meta tags for common page types

Put robots rules in the <head> of the HTML the server sends, one variant per page type below; never add, change or remove them with JavaScript. A noindex only works if the URL is not disallowed in robots.txt, because Google has to fetch the page to see it.

HTML
<!-- Indexable templates: allow full-length snippets and large image and video previews -->
<meta name="robots" content="max-snippet:-1, max-image-preview:large, max-video-preview:-1">

<!-- Keep a page out of Search (internal admin, thin or temporary pages) -->
<meta name="robots" content="noindex">

<!-- The same rule for Google only -->
<meta name="googlebot" content="noindex">

<!-- Time-limited page (event, offer, job ad): drop it from results after the end date -->
<meta name="robots" content="unavailable_after: 2026-12-31T23:59:59+01:00">
Evidence · 2 claims · 2 Google pages
  • SlideConfirmed by docsD2-C090

    The max-video-preview rule sets the maximum number of seconds of a video preview in Search; 0 allows only a static image and -1 means no time limit.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C091

    John Mueller said he thinks setting max-video-preview to -1 will not make Google show an hour-long video preview in search results.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-VID-06

Serve videos fast in a common container with standard codecs, and never block the files

Why

Google processes the most popular video file types, MP4 and WebM among them, and the server streaming the video needs enough resources to support crawling. Said at Search Central Live by Gary Illyes: what matters is that the video itself loads fast, a CDN that serves it faster than the site's own server is a win, MP4 is recommended for its compatibility, and unusual codecs leave some people unable to play it. If robots.txt disallows a video file, Google does not index the video, and unlike a blocked web page it does not show the bare URL.

How

Encode in MP4 (H.264 video, AAC audio) or WebM, serve the files and thumbnails from a CDN with range requests, and keep the video, thumbnail and player URLs allowed in robots.txt for Googlebot and Googlebot-Video. Alternatively host on a video platform such as YouTube and embed it.

Test

ffprobe shows a standard container and codecs; the video starts quickly on a throttled mobile connection; Google's open-source robots.txt parser, run with the Googlebot-Video token, returns allowed for the video and thumbnail URLs.

Evidence · 8 claims · 2 Google pages
  • StageConfirmed by docsD2-C921

    Google supports the most popular video file formats on the web.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C919

    Gary Illyes called the 'fast-loading pages' factor for video SEO a misnomer: what matters is that the video itself loads fast, because people no longer have the patience to wait for videos to load.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C920

    Serving videos from a content delivery network (CDN) that loads them faster than the site's own server is a win for video SEO, Gary Illyes said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C922

    Google recommends the MP4 container for videos because of its browser and device compatibility.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C923

    Encode videos with standard codecs, because some people will not be able to play videos that use unusual ones.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C932

    If robots.txt disallows the location of an image or video file, Google does not index that image or video.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C933

    Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C928

    Gary Illyes suggested considering video hosting platforms such as YouTube or Vimeo, because they solved video search years ago.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Performance and server capacity 4

Speed matters to crawling as much as to users: Google adjusts how much it crawls a host to its connect time, time to first byte and error rate, and that capacity is shared by every Google crawler. Fast rendering keeps content from timing out.

ShouldDocumentedDEV-PRF-01

Keep connect time and time to first byte low and stable under crawler load

Why

The crawl rate limit (hostload) is a host-wide limit shared by all Google crawlers, including Search, Ads, Shopping and Images; it applies per host, so a CDN host, an app host, subdomains and the www host may each have their own. When connect time, time to first byte or the share of 429 and 5xx responses rises, Google lowers it and crawls less, while better performance increases the crawl budget. Google's crawl budget guide does not express this limit in requests per second: it counts the parallel connections Google holds open to the server and how long they last.

How

Cache rendered HTML at the edge, keep origin response times low for uncached pages, size servers for crawler bursts, give bots exactly what an anonymous first-time visitor gets (no personalisation, the default test variant), never content or URLs that no user can get, and do not route Googlebot to a slower backend.

Test

Search Console Crawl stats shows a stable average response time and healthy host status; server logs compare time to first byte for verified Googlebot with users. An abrupt step down in crawling, or fewer connections opened by Googlebot, points to a lowered capacity limit rather than lower demand.

Evidence · 13 claims · 3 Google pages
  • SlideConsistent with docsD1-C092

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideConfirmed by docsD1-C091

    Crawl rate limit, or hostload, is a host-wide metric shared across all Google crawlers.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C132

    Google's crawl budget guide says each crawler has its own crawl demand, but the crawl capacity limit (hostload) is shared across all crawlers, so high demand from one crawler can reduce the capacity left for others.

    Google

  • StageConsistent with docsD1-C096

    Crawl budget was described as the attention span Google gives a website, and better performance increases it.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideConsistent with docsD1-C066

    Google Search, Ads, Shopping and Images all request through one centralised crawling infrastructure, whose primary mandate is to fetch from the internet while strictly preventing the overloading of external servers.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideNot in docsD2-C024

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD1-C141

    Google's guide to A/B testing says never to show one set of URLs to Googlebot and a different set to humans: that is cloaking, which is against Google's spam policies whether or not a test is running.

    Google Search Central

  • StageConsistent with docsD1-C374

    Crawl rate limit, or hostload, is a proxy for how many requests per second the connection to a site can handle before Google hits the limit.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C376

    Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C440

    If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its crawling of that site again.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C493

    There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C494

    To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot opens to the site: if it has dropped, the capacity limit changed.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C547

    Google's crawl budget guide does not express the crawl capacity limit in requests per second: it limits the total time a server spends holding connections open for Google, counting both the number of parallel connections and their duration. That fits the panel's advice to watch how many connections Googlebot opens (D1-C494): fewer connections is the documented form of a lower capacity limit.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-PRF-02

Make JavaScript-generated content render quickly

Why

Pages wait in a render queue and rendering does not go on forever. Said at Search Central Live: when a Gemini user asks about a specific page it is read at that moment, which takes extra time, and Google's answer is fast, efficient JavaScript rendering or server-side rendering.

How

Ship less JavaScript on content templates, avoid sequential API waterfalls before the main content, stream or server-render above-the-fold content, and lazy-load only non-essential widgets.

Test

Lighthouse or PageSpeed Insights on key templates; a headless render under CPU throttling shows the main content quickly; the URL Inspection screenshot shows the full content.

Evidence · 5 claims · 1 Google page
  • StageConsistent with docsD2-C124

    Google's simple solution for the extra time of live page reads is to make JavaScript-generated content render quickly and efficiently, or to use server-side rendering.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C172

    Google's JavaScript SEO basics guide says a page may wait in the render queue for a few seconds but that it can take longer, and it gives no upper limit.

    Google Search Central

  • StageNot in docsD2-C120

    When a Gemini user asks about a specific web page (for example, whether it says anything about the ruby HTML tag), the page is read at that moment rather than during crawling and may be used to ground the answer, which takes extra time.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C125

    For pages people are likely to ask an AI assistant about (product, pricing, documentation and policy pages), server-render the main content so a live, user-triggered read does not depend on client-side JavaScript finishing quickly.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C117

    AI Overviews and AI Mode typically do not ground their answers by reading pages live, unlike Gemini when a user asks about a specific page, Google said.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MaySaid at Search Central LiveDEV-PRF-03

Load per-session-billed third-party tools only after a user interaction

Why

Said at Search Central Live by a community speaker: third-party systems billed per session (chat, personalisation, testing, session replay) also run when bots render pages and can cost a high-traffic site thousands of euros a month.

How

Initialise such tools on the first user interaction (click, keypress, scroll) when they add no content that needs indexing, and compare the sessions they bill with real user sessions.

Test

Vendor-billed session counts match analytics sessions, and a headless render of a page starts no vendor session before an interaction.

Evidence · 2 claims
  • StageNot in docsD2-C163

    Third-party systems billed per session that also run when bots render a page can cost a high-traffic site thousands of euros per month, a community speaker said, adding that he sees this regularly.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C165

    Compare the session counts that per-session-billed tools (chat, personalisation, testing, session replay) charge for with real user sessions, and load such tools only after a user interaction where they add no content that needs to be indexed.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-PRF-04

Keep each HTML response well under 2 MB and put the main content and JSON-LD early

Why

Googlebot fetches only the first 2MB of an HTML URL, HTTP headers included: bytes past that cutoff are not fetched, rendered or indexed. Google warns that bloated inline base64 images, large blocks of inline CSS or JavaScript and megabytes of menus can push a page's text or structured data past the limit, a real risk for single-page apps that inline large hydration state. Scripts, stylesheets and API responses the page loads each have their own limit.

How

Move base64 images, large inline scripts and styles and serialised application state (hydration JSON) to external files or endpoints, keep navigation markup compact, and place the meta tags, title, canonical, hreflang and essential JSON-LD high in the head, with the main content before long secondary blocks.

Test

For every template, curl -s <url> | wc -c stays under about 1.5 MB, and the main heading and the JSON-LD appear within the first 2 MB of the response (curl -s <url> | head -c 2000000 | grep -c 'application/ld+json').

Evidence · 2 claims · 2 Google pages
  • DocsSourceD1-C136

    Google's Inside Googlebot post (March 2026) says Googlebot currently fetches only the first 2MB of each URL, HTTP headers included (64MB for PDFs); bytes past that cutoff are not fetched, rendered or indexed, and each resource the page loads has its own separate limit.

    Search Central blog (31 March 2026)

  • DocsSourceD1-C137

    Google's Inside Googlebot post warns that bloated inline base64 images, large blocks of inline CSS or JavaScript, or megabytes of menus can push a page's text or structured data past Googlebot's 2MB cutoff, and advises moving heavy CSS and JavaScript to external files and placing meta tags, the title, the canonical and essential structured data high in the HTML.

    Search Central blog (31 March 2026)

AI features (AI Overviews, AI Mode, Gemini) 6

AI Overviews and AI Mode use the same crawling and the same index as Search, and query fan-out sends generated queries to that index and gets documents back with their snippets. There is no separate AI index, markup or file to build for; the controls for AI features are covered in the indexing controls above.

MustDocumentedDEV-AIF-01

Keep pages indexable and snippet-eligible, the whole technical requirement for AI features

Why

To appear as a supporting link in AI Overviews or AI Mode a page must be indexed and eligible to be shown with a snippet, with no additional technical requirements. AI features are built on Google's core ranking systems and its index, so what works for web results works for them.

How

Apply the rest of this kit, do not block snippets (DEV-IDX-05), and keep the Search Console generative AI setting on include unless the site is excluded on purpose (DEV-IDX-10). When a page never shows up as a source, check indexing, snippet rules and that setting first.

Test

URL Inspection shows the page indexed, its robots rules contain no nosnippet, and the Generative AI performance report shows impressions.

Evidence · 5 claims · 2 Google pages
  • DocsSourceD2-C074

    Google's AI features guide says a page must be indexed and eligible to be shown in Google Search with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements.

    Google Search Central

  • SlideConsistent with docsD1-C038

    AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • SlideConfirmed by docsD1-C051

    Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD2-C601

    Sites need no extra work to appear in AI Overviews and AI Mode: both work on top of the existing web results, so what works for web results also works for the AI features.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C730

    There is no separate AI index to optimise for: when a page never shows up as a source in AI Overviews or AI Mode, first check that it is indexed, that no nosnippet rule blocks its snippet and that the site is not excluded in Search Console's generative AI setting, the eligibility conditions Google lists.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-AIF-02

Spend no effort on llms.txt, AI text files or AI-specific markup for Google

Why

Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in Search, which ignores them, so they neither help nor harm. Structured data is not required for generative AI search either, though it remains worth using for rich results, and the same processed markup feeds classic results and AI answers. Google repeated it on serving day: AI Overviews and AI Mode are standard search features, not rich results, and need no structured data to function.

How

Do not build llms.txt or AI-only endpoints for Google; invest instead in documented structured data types (DEV-SDA-03) and crawlable, server-rendered content. If llms.txt is published for other tools, make it answer 200 (Lighthouse's optional llms.txt audit flags a server error), and do not publish markdown copies of HTML pages for agents: a community speaker warned they are duplicates, possibly even a cloaking risk. The same speaker cited a study of almost 40,000 sites in which 97% of llms.txt files got no AI agent hits in a month; Ahrefs' published study (June 2026) reports the same 97%.

Test

The backlog has no tickets for llms.txt or AI-only markup aimed at Google.

Evidence · 15 claims · 3 Google pages
  • DocsSourceD1-C061

    Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in Search, and that such files neither harm nor help visibility because Google Search ignores them.

    Google Search Central

  • SlideConfirmed by docsD1-C054

    Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD2-C480

    Google's guide to optimizing for generative AI features lists 'overfocusing on structured data' among the things site owners don't need to do: structured data is not required for generative AI search and no special schema.org markup is needed, though it remains worth using because it helps pages become eligible for rich results.

    Google Search Central

  • StageConsistent with docsD2-C476

    The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C479

    There is no separate markup for AI features: by Google's account the same processed markup feeds classic results and AI answers, and raw schema.org is generally not passed into model context, so invest in the types Google documents for its features rather than in extra markup written for AI.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD3-C325

    AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to function and work with the normal text results from Google's index.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C326

    Google's AI features guide says there are no additional requirements, and no special schema.org structured data, needed to appear in AI Overviews or AI Mode; a page must be indexed and eligible to be shown with a snippet.

    Google Search Central

  • AnalysisD3-C328

    Day 2 said the structured data Google processes also feeds AI Overviews and AI Mode (D2-C476), while Day 3 said these features do not need structured data to function: markup is not a condition for appearing in AI features, but data taken from it can still reach them.

    Ibrahim Anjro (author)

  • StageConsistent with docsD1-C271

    A community speaker said llms.txt files are not necessary, citing a third-party study from summer 2026 (heard as Ahrefs') that looked at almost 40,000 sites over one month: 97% had no AI agent hits on the file, and those that had any got about two hits a month.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C272

    A community speaker advised against publishing markdown copies of HTML pages for AI agents: the copy is a duplicate (which the speaker also called a possible source of cloaking, an uncertain word in the recordings), and the models are trained to read HTML, CSS and JavaScript.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C454

    Gary Illyes said llms.txt does not matter for Google Search right now and that he does not expect it to, while adding that he had been wrong before.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C455

    Gary Illyes said llms.txt matters to some people, which is why tools such as Lighthouse add checks for it, so both sides of the debate are right in their own context.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C460

    Chrome's Lighthouse documentation lists an llms.txt audit among its agentic browsing audits: it flags a server error when llms.txt is fetched and marks the audit not applicable when the file is missing, because providing the file is optional for now.

    Chrome for Developers

  • StageConsistent with docsD1-C473

    Files such as cats.txt have no importance for Google Search: if a site links to one, Googlebot will find and crawl it, but it has no other effect.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C505

    The llms.txt study cited on stage matches Ahrefs' June 2026 study of 137,210 domains: 28% (about 38,000, the 'almost 40,000' sites heard on stage) published an llms.txt file and 97% of those files received no requests at all in May 2026; the files that were fetched got about 22,000 requests across some 1,100 domains, around 20 each, not the 'two hits a month' heard on stage.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-AIF-03

Structure content for readers instead of splitting it into small chunks for AI

Why

Google's guidance says there is no need to chop content for AI systems. Chunking matters at the level of a model's context window, and Gemini's holds a whole smaller book; said at Search Central Live, once chunk size is thought of in millions of tokens, chunking has perhaps lost its meaning anyway. Separate pages for every query variation, made to manipulate rankings or AI answers, count as scaled content abuse. Fan-out queries differ from system to system and, said at Search Central Live, are not added to Search Console; Google advises knowing that fan-out happens without overfocusing on individual fan-out queries, so do not build pages for fan-out queries copied from LLM tools.

How

Keep complete pages with clear headings and full explanations, model content in the CMS by topic rather than by fragment length, and do not auto-generate thin pages per question or query variation.

Test

Content-model review finds no template that splits articles into many thin URLs.

Evidence · 9 claims · 2 Google pages
  • SlideConfirmed by docsD1-C054

    Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C129

    Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.

    Google Search Central

  • StageConsistent with docsD2-C332

    Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C825

    Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C333

    Do not rewrite pages into short, self-contained chunks for AI systems; Google says Gemini reads context windows of millions of tokens, so structure content for readers, with clear headings and complete explanations.

    Ibrahim Anjro (author)

  • StageNot in docsD3-C063

    Google said fan-out queries are not added to Search Console, because Google considers them part of its infrastructure.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C065

    Because every system runs fan-out differently, Google advised understanding that fan-out happens but not overfocusing on individual fan-out queries or on how to rank for them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C066

    Do not build pages for fan-out queries copied from Gemini or other LLM tools: they differ from system to system and Search Console does not report them. Cover a topic's real subquestions on solid pages instead.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C869

    Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldSaid at Search Central LiveDEV-AIF-04

Do not block the AI agents you want as visitors

Why

Said at Search Central Live: Google's duplication talk closed with the advice not to block agents, and AI agents browsing for users run into the same bot walls that sites build against scrapers and may go elsewhere, for example to buy the product from another shop.

How

Decide per agent which ones to allow, tune bot protection to challenge abusive patterns rather than every automated browser, and if a challenge is needed send it with 503 (DEV-SRV-02).

Test

WAF and bot-management logs show no blocks of agents the team decided to allow.

Evidence · 3 claims · 1 Google page
  • SlideNot in docsD2-C397

    Google's duplication talk closed with the advice not to block agents, which the speaker said are sometimes really cool.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C377

    AI agents that browse the web for users run into the same bot walls that sites put up against scrapers, and may give up and go to another site, for example to buy the product elsewhere.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C378

    Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.

    Ibrahim Anjro (author)

MayDocumentedDEV-AIF-05

Set robots.txt policy per AI crawler by its own token, and know that user-triggered fetchers ignore it

Why

Mainstream crawlers from Google, other search engines and AI companies try to follow robots.txt, and nearly all AI crawlers use their own user agents, so a site can allow search crawling and refuse AI training per company; Google's token for that is Google-Extended (DEV-IDX-11), and said at Search Central Live, other companies have similar "extended" tokens, such as Apple's Applebot-Extended, which lets a site allow Applebot for search while opting out of Apple's AI use. User-triggered fetchers, which fetch a page because a user asked, and AI agents acting for a user generally ignore robots.txt.

How

Write one robots.txt group per AI crawler token you want to restrict, and keep Googlebot allowed. Avoid a default-deny file that blocks every crawler and allows only chosen ones: a Google panelist called it a bad pattern, because you do not know what you are blocking. When logs show an unknown crawler fetching a lot, look up its user agent to find the token that controls it. To restrict user-triggered fetchers or agents, use authentication or bot management, not robots.txt.

Test

Google's open-source robots.txt parser returns the intended result for Googlebot, Google-Extended and each listed AI token; logs show the restricted crawlers stop after robots.txt is refreshed.

Code · robots.txt with crawl controls and a rendering carve-out

Disallow only URLs that should never be crawled (internal search, cart and checkout actions, filter parameters), keep every script, style and API path that pages need for rendering crawlable, and list the sitemap. Google reads only user-agent, allow, disallow and sitemap; each host (www, api, cdn) needs its own file at its root. Robots.txt is public, so never list secret paths in it.

Text
# https://www.example.com/robots.txt
User-agent: *
# Internal search results and cart/checkout actions (adapt to your URL patterns)
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
# (Shops in Google Shopping: give Storebot-Google its own group that leaves cart and checkout open.
#  A named group replaces this * group for that crawler, so repeat in it every rule that should still apply.)
# Filter and sort parameters of faceted navigation, as the first parameter (?) or a later one (&);
# /*?*size= would also block ?pagesize= and /*?*color= would block ?bgcolor=
Disallow: /*?color=
Disallow: /*&color=
Disallow: /*?size=
Disallow: /*&size=
Disallow: /*?sort=
Disallow: /*&sort=
# API: blocked in general, but the endpoints pages render from stay crawlable
# (the longer, more specific Allow rule wins)
Disallow: /api/
Allow: /api/products/
Allow: /api/reviews/
# Never disallow /static/, /assets/ or other JS and CSS folders

# Optional: keep content out of Gemini training and grounding (no effect on Google Search)
User-agent: Google-Extended
Disallow: /

Sitemap: https://www.example.com/sitemap.xml
Evidence · 13 claims · 4 Google pages
  • StageConfirmed by docsD1-C428

    Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C482

    Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C485

    To opt out of Google's AI training, a site can use the Google-Extended token in robots.txt.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C487

    A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C488

    When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C434

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C435

    AI agents acting on a user's request fall into the same category as user-initiated fetchers: an agent told to look at a site reads the page without checking robots.txt.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C273

    A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C424

    Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas many AI systems crawl a site with no understanding of it and take everything.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C425

    Crawling for AI model training differs from search crawling because training mainly needs a very large number of tokens, and it matters little which pages they come from.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C524

    Other companies have similar 'extended' robots.txt tokens; Google named Apple's Applebot-Extended.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C523

    The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

    Ibrahim Anjro (author)

  • StageD1-C489

    A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

MayDocumentedDEV-AIF-06

Try WebMCP to declare the actions a site offers to AI agents

Why

A community speaker presented WebMCP, a proposed web standard in Chrome, as a way to tell AI agents which actions a site offers, such as booking an appointment, so they do not have to scrape the interface. It has declarative tools, annotations on standard HTML forms, and imperative tools defined in JavaScript; at the event it was not supported by all browsers but could be enabled for testing. It plays no documented role in Google Search.

How

Pilot it on one task flow: annotate an existing form for the declarative API or register a JavaScript tool for booking, filtering or adding to cart, through Chrome's origin trial (Chrome 149 and later) or the chrome://flags/#enable-webmcp-testing flag. Keep the normal HTML flow working for users and agents without WebMCP.

Test

With the flag enabled, Chrome's tooling lists the site's tools, and the flow still completes without WebMCP.

Evidence · 4 claims · 1 Google page
  • StageConfirmed by docsD1-C281

    A community speaker presented WebMCP (Web Model Context Protocol) as a way for a site to declare the actions it offers to AI agents, for example booking an appointment in a calendar, so an agent does not have to scrape the interface; the speaker called it faster, easier and cheaper.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C282

    A community speaker said that at the time of the event (30 September 2026) WebMCP was not supported by all browsers but could be enabled for testing through browser flags.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C283

    A community speaker said WebMCP has two kinds of tools: declarative ones, mostly HTML annotations such as the fields of a contact form, and imperative ones for other actions such as booking, filtering a catalogue, adding products to a cart, getting product specs or reordering.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C316

    Chrome's WebMCP documentation describes WebMCP as a proposed web standard that site owners can test through an origin trial starting in Chrome 149 or through a Chrome flag for local development.

    Chrome for Developers

  • WebMCP Chrome for Developers · checked 3 October 2026

Spam policies and AI-generated content 5

Google's spam policies cover practices that deceive users or manipulate Search, including its generative AI responses, and several of them are built in code: history manipulation, cloaking, generated page sets and hacked pages. AI is a tool; how its output is used decides whether content is helpful or spam.

AvoidDocumentedDEV-SPM-01

Never manipulate browser history so that Back does not return users to where they came from

Why

Back button hijacking, interfering with browser navigation so users cannot use Back to return immediately to the page they came from, is an explicit violation of Google's malicious practices spam policy, enforced since 15 June 2026 with manual spam actions or automated demotions. Google added it after many complaints about landing on auto-generated pages or staying on the same site after pressing Back twice, and says the cause can be a site's included libraries or advertising platform.

How

Call history.pushState only for real in-app navigation that the user triggered (DEV-URL-03), never on page load, on scroll or to insert interstitial, recommended or exit pages, and do not intercept popstate to keep users on the site. Audit scripts from ad, engagement and recommendation vendors, and remove or disable any code, import or configuration that adds history entries.

Test

An automated browser test opens each template from another page, waits and interacts, presses Back once and lands on the previous page; a code search of the built bundles finds no pushState or replaceState outside the router.

Code · Automated check that Back leaves the page

A Playwright test opens each template from another page, interacts with it the way users do (many history hijacks wait for a key press, a scroll or a timer), presses Back once and expects to be on the previous page again. Run it in CI for every template and after adding any ad, engagement or recommendation script.

JavaScript
// back-button.spec.js: npm i -D @playwright/test, then npx playwright test
const { test, expect } = require('@playwright/test');

const START = 'https://www.example.com/'; // stands in for the page the user came from, such as a search results page
const PAGES = [
  'https://www.example.com/blog/oak-table-care/',
  'https://www.example.com/chairs/oak-dining-chair',
];

for (const url of PAGES) {
  test(`one Back press leaves ${url}`, async ({ page }) => {
    await page.goto(START);
    await page.goto(url);
    await page.keyboard.press('End'); // a real user input that also scrolls to the bottom
    await page.waitForTimeout(5000);  // give delayed scripts time to run
    await page.goBack();
    await expect(page).toHaveURL(START);
  });
}

Then search the built bundles for history calls outside the router and review each one.

Shell
grep -rnE "history\.(pushState|replaceState)|addEventListener\(.popstate" dist/
Evidence · 5 claims · 2 Google pages
  • StageConfirmed by docsD3-C207

    Google's quality talk listed recent updates to Google's spam policies: back button hijacking, scaled content abuse and manipulating AI responses.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C208

    Google's quality talk said back button hijacking was added after many user complaints about pressing the back button and landing on an auto-generated page, or staying on the same site after pressing it twice.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C209

    Google announced on 13 April 2026 that back button hijacking is an explicit violation of its malicious practices spam policy, with enforcement from 15 June 2026; pages that do it may get manual spam actions or automated demotions.

    Search Central blog (13 April 2026), Google Search Central

  • DocsSourceD3-C210

    Google's back button hijacking announcement says some instances may come from a site's included libraries or advertising platform, and asks site owners to remove or disable any code, imports or configurations responsible.

    Search Central blog (13 April 2026)

  • AnalysisD3-C211

    For development teams: test that the browser back button returns users from your pages to where they came from, including search results; audit scripts that call the History API, insert interstitial or 'recommended' pages, or come from ad and engagement vendors.

    Ibrahim Anjro (author)

AvoidDocumentedDEV-SPM-02

Do not generate large sets of low-value pages from templates, feeds or LLM output

Why

Scaled content abuse, many pages generated mainly to manipulate rankings rather than help users, is among the spam policies Google recently updated, whether the pages come from generative AI, scraped feeds, stitched sources or keyword text. Said at Search Central Live: churning out pages with LLMs repeats the script-generated spam of the early 2000s, Google counts AI slop as scaled content abuse, and this type, rather than link spam, is the one worth talking about today. Google's guidance says using generative AI to produce lots of text without manual oversight represents little to no effort.

How

Put programmatic page types (location by service, product by keyword, one page per question or per fan-out query) behind a review: each page needs unique, useful main content, otherwise do not create it or keep it noindexed. Do not auto-publish LLM drafts or feed-derived pages in bulk; cap automated publishing and require an editor's approval per page.

Test

A crawl of each programmatic template finds no clusters of near-identical pages; publishing logs show an editor's approval for generated pages; Search Console's Manual actions report is empty.

Evidence · 7 claims · 3 Google pages
  • StageConfirmed by docsD3-C207

    Google's quality talk listed recent updates to Google's spam policies: back button hijacking, scaled content abuse and manipulating AI responses.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C284

    Reviewing the spam slide, the speaker said link spam would probably be replaced with scaled content abuse as the spam type worth talking about today.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C694

    Scaled content abuse is becoming a problem again: in the early 2000s pages were churned out with Perl or PHP scripts, and now the same is done with LLMs.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C695

    The cheaper tokens become, the more AI slop is created, and Google counts AI slop as scaled content abuse.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C176

    Google's helpful-content guidance lists effort, originality, talent or skill and accuracy as the main-content attributes search quality raters are trained to evaluate, and says using generative AI to produce large amounts of text without manual oversight or curation represents little to no effort.

    Google Search Central

  • AnalysisD3-C213

    Tactics sold as AI visibility or GEO, such as hidden text aimed at AI answers, mass-produced pages for fan-out queries or keyword-stuffed passages, are judged under the same spam policies as classic ranking manipulation.

    Ibrahim Anjro (author)

  • AnalysisD3-C300

    Weight spam audits toward scaled content abuse: a Googler said that type would replace link spam on the slide, so review mass-produced pages, including AI-generated ones, before older link-building patterns.

    Ibrahim Anjro (author)

AvoidDocumentedDEV-SPM-03

Never serve Googlebot different content from users, and do not build doorway pages

Why

Cloaking, presenting different content to users and to search engines to manipulate rankings and mislead users, and doorways, sites or pages created to rank for specific, similar queries that lead users to intermediate pages less useful than the final destination, are spam types in Google's spam policies and on its Day 3 spam slide.

How

Never branch on the Googlebot user agent or Google's IP ranges to change text, links or redirects. Do not generate near-identical city or keyword landing pages that funnel visitors to one destination; merge them into one useful page per real offering.

Test

The HTML served to a Googlebot user agent and to a browser have the same main content, links and redirects, and URL Inspection's rendered HTML matches what users see.

Evidence · 4 claims · 1 Google page
  • SlideConfirmed by docsD3-C277

    Google's slide 'Spam, spam and spam!' listed five spam types, cloaking, doorways, scraped content, link spam and hacked content, and pointed to goo.gle/spam-policy for more.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C278

    Google's slide defined cloaking as presenting different content to users and to search engines with the intent to manipulate search rankings and mislead users.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C279

    Google's slide defined doorways as sites or pages created to rank for specific, similar search queries that lead users to intermediate pages less useful than the final destination.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C301

    For developers, Google's five listed spam types map to concrete checks: serve Googlebot the same content as users (cloaking), avoid near-identical location or keyword pages (doorways), add value to reused content (scraped content), and patch and monitor the CMS against injected pages and links (hacked content).

    Ibrahim Anjro (author)

MustDocumentedDEV-SPM-04

Keep the CMS and plug-ins patched and watch for injected pages and links

Why

Hacked content, placed on a site without permission through security vulnerabilities, is one of the spam types on Google's list; cleaning it up is hard and prevention is key. If it leads to a manual action, Google's help says most reconsideration reviews take several days or weeks; said at Search Central Live, removal takes one to two weeks on average, and for some sites the action stays because the owners never actually clean up.

How

Automate CMS, theme and plug-in updates, remove unused plug-ins and accounts, require two-factor login for editors, and alert on new URLs, outbound links or sitemap entries that no editor created. After a clean-up, file the reconsideration request only once every injected page and link is gone.

Test

A weekly comparison of the sitemap and a crawl with the CMS's own page list finds no unknown URLs; Search Console's Security issues and Manual actions reports are empty.

Evidence · 5 claims · 2 Google pages
  • SlideConsistent with docsD3-C282

    Google's slide defined hacked content as content placed on a site without permission through security vulnerabilities, said cleaning it up is hard and prevention is key, and linked to goo.gle/hack-tips.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C277

    Google's slide 'Spam, spam and spam!' listed five spam types, cloaking, doorways, scraped content, link spam and hacked content, and pointed to goo.gle/spam-policy for more.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConsistent with docsD3-C659

    Google estimated that removing a manual action after a reconsideration request takes 1-2 weeks on average, at least 2-3 days, and 4-6 weeks or much longer for dormant sites.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C674

    Google's Manual actions report help says most reconsideration reviews take several days or weeks, and that some, such as link-related requests, may take longer.

    Google Search Console Help

  • StageConsistent with docsD3-C660

    For some sites a manual action is almost never removed, because the owners do not make the effort to actually clean up the site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

ShouldDocumentedDEV-SPM-05

Publish AI-generated text, images and markup only after human review, and show only real authors

Why

Any AI model can hallucinate. Google treats quality problems as quality issues, not as AI against human content: the tool is not the problem, but how it was used and what for. Its rater guidelines treat fake owner or creator profiles, such as made-up authors with AI-generated headshots, as deception that earns the lowest rating; raters cannot penalise a site, but their ratings become labels that improve Google's algorithms, so unchecked AI output can harm a site indirectly.

How

Build the CMS workflow so AI drafts, AI images and AI-generated markup cannot publish without a named editor's fact check; link every byline to a real author page with a real photo and verifiable profiles, and never generate author personas or headshots.

Test

Publishing logs show an editor's approval for every AI-assisted item, and an audit of author pages finds no persona without a real person behind it.

Evidence · 11 claims · 4 Google pages
  • StageConsistent with docsD3-C686

    Hallucinations can happen with any AI model, and with current training methods there is no way to get rid of them.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C690

    Quality raters cannot give a site a penalty or a manual action; their ratings are converted into labels that Google uses to improve its algorithms.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C691

    Publishing AI output without checking it can indirectly harm a site's standing in Search, because deceptive content rated lowest by raters feeds the labels used to improve the algorithms.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C692

    Google sees fake online personas as a growing type of deception, for example a professional author bio whose headshot is an AI-generated image of a non-existent person, and pages doing this get the lowest rater rating.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C697

    Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem, but how it was used and what for.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConsistent with docsD3-C700

    Google's closing slide said to use AI responsibly because AI hallucinates, and, especially when creating content briefs with AI, to make sure not to add to the sea of AI slop already flooding the internet.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C710

    Google's Search Quality Rater Guidelines (September 2025) list fake owner or content creator profiles, such as made-up author profiles with AI-generated images, as deception, and say pages using deception of any type should be rated Lowest.

    Google Search Quality Rater Guidelines (PDF, 11 September 2025)

  • StageConfirmed by docsD3-C190

    Google's quality talk said quality problems should be treated as quality issues, not as AI versus human content, because a lot of good AI-assisted or AI-written content exists.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C693

    Show only real authors with real photos and verifiable profiles; never invent author personas or use AI-generated headshots, which Google's raters treat as deception.

    Ibrahim Anjro (author)

  • StageConsistent with docsD2-C937

    To the frequent question about AI-generated images on a site, Google's answer is that it is up to the site owner.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConsistent with docsD2-C941

    Sites that use AI-generated images or videos should make sure they work for users, check them for hallucinations and regenerate them where needed.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

Testing and monitoring 11

Search Console shows what Google crawled, rendered, chose as canonical and indexed. Render one URL per template in a headless browser before every release, check it in URL Inspection right after, and watch the reports that reveal crawling, rendering and index selection problems.

MustDocumentedDEV-MON-01

Verify every production host in Search Console before launch

Why

Debugging what Google can see on a site, with URL Inspection, the Page indexing report or Crawl stats, starts with verified access to the site in Search Console.

How

Add a Domain property (DNS verification) that covers every subdomain and protocol, plus URL-prefix properties for sections or hosts that need separate access, and keep the verification record in infrastructure code so it survives DNS or CMS changes.

Test

Search Console lists the Domain property as verified, and every team that ships code has access.

Evidence · 1 claim · 1 Google page
  • StageConsistent with docsD2-C306

    To use Search Console to debug what Google can see on a site and why, Erin Sparling said the first step is to get access to the site by verifying ownership (the speaker's words were authorized domains).

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

MustDocumentedDEV-MON-02

Render every template in a headless browser before release and check it in URL Inspection right after

Why

The rendered HTML in Google's testing tools shows what Google can index, and the live test also shows page resources, JavaScript console messages and a screenshot. A valid live test only confirms that Google can access the page; indexing still depends on other conditions, and the live result can differ from the indexed version. URL Inspection tests only public URLs of a verified property, so it cannot test a staging host that answers 401 or 403 (DEV-CAN-07).

How

Keep a list of one representative URL per template (home, category, product, article, landing pages, each language version). Before release, render each template on staging with headless Chrome (Puppeteer or Playwright): compare the raw and rendered DOM, find the main content, prices and internal links, collect console errors, and run the structured data through the Rich Results Test's code input. Right after release, run the URL Inspection live test on one production URL per template and check the rendered HTML, resources and console messages. Request indexing for key URLs after fixing signals calculated at indexing time, such as language or robots rules.

Test

The release checklist records a headless render result for each template before release and a URL Inspection live test result for each template after release.

Evidence · 7 claims · 3 Google pages
  • StageConfirmed by docsD2-C176

    To check whether JavaScript content is affected, open the URL Inspection tool in Search Console or the Rich Results Test and look for that content in the rendered HTML tab.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C178

    Google's URL Inspection help says the tool's live test examines a URL in real time, so its results can differ from Google's indexed version of the page.

    Google Search Console Help

  • DocsSourceD2-C308

    Google's URL Inspection help says a valid live test only confirms that Google can access a page for indexing; the page must still meet other conditions to be indexed, such as having no manual action, not being a duplicate and being of high enough quality.

    Google Search Console Help

  • StageConsistent with docsD2-C215

    To find rendering issues, compare a page's raw source code with its rendered versions, including the rendered HTML that Google's testing tools show; the differences point to the problems.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C179

    Check rendering per template rather than per URL: inspect one URL of each page type (product, category, article) in URL Inspection and confirm that the main content, prices and internal links appear in the rendered HTML.

    Ibrahim Anjro (author)

  • AnalysisD2-C158

    When something is missing from a rendered page, check both robots.txt and the Content Security Policy: the URL Inspection live test shows the page resources, the JavaScript console output and a screenshot of the rendered page.

    Ibrahim Anjro (author)

  • AnalysisD2-C677

    Language, country, SafeSearch and spam signals, and some freshness signals, are calculated when a page is indexed, so a fix such as correcting a page's language or removing content that triggers SafeSearch only counts once Google recrawls and reprocesses the page; request recrawling of the most important URLs after the fix.

    Ibrahim Anjro (author)

ShouldDocumentedDEV-MON-03

Watch the Page indexing report and treat its not-indexed reasons as a work list

Why

The Page indexing report is where index selection problems show. "Discovered - currently not indexed" means Google knows the URL but has not crawled it yet, and "Crawled - currently not indexed" means Google processed the page but did not keep it, which was described at Search Central Live as most of the time a quality issue; such pages need no resubmission. Soft 404s are also looked up in this report, while other crawl issues are debugged in the Crawl stats report (DEV-MON-04).

How

Export not-indexed URLs by reason and template regularly. For Discovered, first rule out slow responses and server errors; for Crawled, compare with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences. Do not resubmit the same URLs again and again.

Test

A trend of indexed and not-indexed URLs per sitemap and template is reviewed, and spikes after releases are investigated.

Evidence · 10 claims · 3 Google pages
  • StageConsistent with docsD2-C717

    Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageConfirmed by docsD2-C718

    Not-indexed reasons in the Page indexing report include pages excluded by a noindex rule and 'Alternate page with proper canonical tag'.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • DocsSourceD2-C707

    Google's Page indexing report help says a 'Discovered – currently not indexed' page was found but not crawled yet, typically because Google wanted to crawl it but expected the crawl to overload the site, so it rescheduled the crawl.

    Google Search Console Help

  • DocsSourceD2-C712

    Google's Page indexing report help says a 'Crawled – currently not indexed' page was crawled but not indexed, may or may not be indexed in the future, and does not need to be resubmitted for crawling.

    Google Search Console Help

  • StageConsistent with docsD2-C711

    'Crawled – currently not indexed' in Search Console is an index selection decision: Google crawled and processed the page but decided not to keep it in the index.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C706

    'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C714

    'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C708

    The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.

    Ibrahim Anjro (author)

  • AnalysisD2-C716

    Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD1-C508

    Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

ShouldDocumentedDEV-MON-04

Monitor Crawl stats and server logs for verified Googlebot traffic

Why

Firewalls, DNS problems and robots.txt errors are blind spots, because Google only sees failing requests, and response times and 429 or 5xx rates drive how much Google crawls. Logs show what Googlebot actually fetches, which matters on sites with filter parameters or complicated URLs, where crawlers can easily wander off into URLs that make no sense for the site, and client-side analytics may not see Googlebot at all, because Google's renderer may skip reporting and error requests.

How

Keep access logs with status code, response time and user agent (logs alone do not show where Google's crawl limit lies, so read them with Crawl stats: total crawling, average response time, the response and file-type breakdowns and the discovery versus refresh split), verify Googlebot against Google's published crawler IP ranges or by reverse DNS to crawl-*.googlebot.com when analysing them, and alert on spikes of 5xx or 429 responses, robots.txt fetch failures and DNS errors. Review the Crawl stats host status regularly.

Test

A dashboard shows verified Googlebot requests by status code and template, and the Crawl stats host status has no failures.

Code · Verify that a request really comes from Googlebot

Before a firewall, CDN or bot-protection rule blocks or challenges a "Googlebot" request, verify it: the user-agent string alone can be faked. In WAF rules, prefer matching the IP against the ranges Google publishes as common-crawlers.json and special-crawlers.json on its page on verifying its crawlers; for log analysis, a reverse DNS lookup followed by a forward lookup works too. Never allowlist a whole googleusercontent.com host: Google Cloud customer VMs resolve there too, and only *.gae.googleusercontent.com belongs to Google's user-triggered fetchers, which are not Googlebot.

Shell
# 1. Reverse DNS: Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com
#    or geo-crawl-*.geo.googlebot.com; special-case crawlers to rate-limited-proxy-*.google.com.
#    A host such as 81.59.117.34.bc.googleusercontent.com is a Google Cloud customer, not Googlebot.
host 66.249.66.1
# -> 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.

# 2. Forward DNS: that host name must resolve back to the same IP
host crawl-66-249-66-1.googlebot.com
# -> crawl-66-249-66-1.googlebot.com has address 66.249.66.1

The published IP ranges are easier to keep in sync with WAF and CDN allowlists than DNS lookups:

Shell
# Common crawlers (Googlebot and others) and special-case crawlers; refresh the lists regularly
curl -s https://developers.google.com/static/crawling/ipranges/common-crawlers.json
curl -s https://developers.google.com/static/crawling/ipranges/special-crawlers.json
Evidence · 13 claims · 5 Google pages
  • StageConsistent with docsD1-C069

    DNS issues and firewalls are blind spots: Google does not know a firewall is blocking it, it only sees that requests fail.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • DocsSourceD1-C126

    Google treats network timeouts, connection resets and DNS errors like 5xx server errors: crawling slows down immediately, and already indexed URLs that stay unreachable are removed from Google's index within days.

    Google

  • SlideConsistent with docsD1-C092

    Hostload is driven by changes in connect time, changes in time to first byte, and HTTP 429 or 5xx status codes. If these increase, hostload is adjusted and crawling slows down.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • AnalysisD1-C078

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

    Ibrahim Anjro (author)

  • DocsSourceD1-C138

    Google's page on verifying its crawlers says Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com or geo-crawl-*.geo.googlebot.com host names, special-case crawlers to rate-limited-proxy-*.google.com and user-triggered fetchers to *.gae.googleusercontent.com or google-proxy-*.google.com, and it publishes each group's IP ranges as JSON files such as common-crawlers.json and special-crawlers.json.

    Google

  • DocsSourceD2-C164

    Google says Googlebot and its Web Rendering Service identify resources that do not contribute to essential page content, such as reporting and error requests, and may not fetch them, so client-side analytics may not give a full or accurate picture of their activity.

    Google Search Central

  • StageConfirmed by docsD1-C339

    To debug crawl issues, Google points site owners to the Crawl Stats report in Search Console, found in the Settings page under Crawl stats.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C340

    Search Console's Crawl Stats report shows how much Google crawls from a site, the errors Google received, and the content types fetched by the different crawlers.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C393

    To check whether a site has a crawl budget problem, use Search Console's crawl report (Crawl Stats), which breaks crawl requests down by response and by file type and shows crawl problems Google finds.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C394

    The file-type breakdown of crawl requests in Search Console is useful for spotting anomalies, such as most crawling going to images on a site that has no images worth crawling.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C396

    Search Console's crawl data separates discovery (fetches of URLs Google has not seen before) from refresh (fetches of known URLs), and lets you drill down to the problematic URLs.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C442

    Log files alone cannot show crawl inefficiency because they do not reveal where Google's crawl limit for the site lies; many similar URLs being crawled may not matter if the site is nowhere near that limit.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C443

    To judge crawl efficiency, check Search Console's Crawl Stats report: whether total crawling has plateaued over time, and whether the average response time shows the server is fast enough or is limiting Googlebot.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

ShouldDocumentedDEV-MON-05

Run recurring rendering checks: raw versus rendered HTML, placeholders and console errors

Why

Sites that create content by rendering need at least basic monitoring that their rendered pages come out right. Rendering failures such as timeouts, unresolved placeholders or pages that render blank because of server overload or bad scripts are intermittent, and invisible if only the raw HTML is checked.

How

Schedule a JavaScript-rendering crawl of sample URLs per template, besides spot checks in Chrome DevTools, diff raw and rendered HTML for main content, links, prices and meta tags, flag unresolved placeholders, render the same URLs more than once, and archive the HTML and resources of failing pages for analysis.

Test

The monitoring job runs on schedule, alerts on differences, and reports zero placeholder hits.

Evidence · 7 claims · 1 Google page
  • StageConsistent with docsD2-C143

    Sites that create their content by rendering should have at least basic monitoring to check that the rendered pages come out right, a community speaker advised.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C159

    A page can render completely blank while its raw HTML looks fine, so a team checking only the raw HTML thinks all is well; causes include server misconfiguration, overload and bad scripts.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C151

    Unresolved placeholders caused by rendering timeouts are hard to catch because the problem moves around: it is not always the same page that is broken.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • StageNot in docsD2-C153

    To catch intermittent rendering problems, archive the full HTML and the resources of each page during a site audit and analyse them yourself, a community speaker advised.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C154

    Add a check for unresolved template placeholders (such as {{brand}}, undefined or null) in rendered titles, meta descriptions and visible text to recurring crawls, and render the same URLs more than once, because such failures are intermittent.

    Ibrahim Anjro (author)

  • AnalysisD2-C148

    Diff the raw and the rendered HTML of each key template for prices, offer counts, links and meta tags, and treat any difference in commercial data as a bug, because bots and users may get different versions.

    Ibrahim Anjro (author)

  • StageNot in docsD2-C860

    Besides using Chrome DevTools more often, a community speaker advised checking rendering at scale with a modern crawler.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-MON-06

Audit canonical groups with a crawler on a schedule

Why

Broken canonical tags can make the wrong pages show up in search results, and a broken canonical setup does not fail visibly: pages load normally while the choice of URL passes to the search engine. Checking canonicals pair by pair misses chains, loops and leaders that nobody links to.

How

Crawl with a tool that groups URLs by canonical leader and flag chains, loops, multiple canonicals, leaders that are non-200, noindexed, blocked, redirecting or unlinked, and groups that mix languages. Compare user-declared and Google-selected canonicals in URL Inspection for samples.

Test

The scheduled crawl report shows zero canonical errors, and the "Duplicate, Google chose different canonical than user" reason in the Page indexing report trends down.

Evidence · 5 claims · 2 Google pages
  • StageConfirmed by docsD2-C401

    Broken canonical tags can make the wrong pages of a site show up in search results.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C404

    A broken canonical setup does not fail visibly: pages still load normally while the choice of URL passes to the search engine's own heuristics, so canonical tags need a regular audit with a crawler and the URL Inspection tool.

    Ibrahim Anjro (author)

  • StageNot in docsD2-C398

    Google's speaker suggested checking rel=canonical links with a crawler such as Screaming Frog to make sure they are reasonable.

    Google at Search Central Live Deep Dive Europe 2026 (Day 2)

  • AnalysisD2-C416

    To audit canonicals at scale, resolve every crawled URL to the final URL its canonical links and redirects lead to, group URLs by that leader, and flag groups with more than one hop, a loop, several canonicals on one page, or a leader that is noindexed, blocked, non-200 or redirecting.

    Ibrahim Anjro (author)

  • StageD2-C408

    A community speaker advised auditing canonicals as whole canonical groups rather than pair by pair, paying special attention to each group's canonical leader.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

MayDocumentedDEV-MON-07

Track AI feature visibility in the Generative AI performance report

Why

Google recommends the Generative AI performance report in Search Console for measuring how content performs in generative AI features on Search and Discover. It shows impressions, pages, countries, devices and dates, but no clicks or queries. The regular Performance report still counts traffic from AI Overviews and AI Mode within the Web search type, so analytics setups that split out LLM referrals do not isolate it. Since 24 September 2026 the Performance report also has a multimodal search type for searches with Google Lens, Circle to Search, uploaded images and Chrome's image search; it has no query dimension.

How

Add the report to the regular dashboard and compare countries and languages with native-language spot checks in AI Overviews and AI Mode. Add the multimodal search type to the same review and judge it by page, country and device.

Test

The report is reviewed monthly alongside the Performance report.

Evidence · 11 claims · 5 Google pages
  • DocsSourceD1-C062

    Google's guide for generative AI features recommends the Generative AI performance report in Search Console for measuring how content performs in generative AI features on Google Search and Discover.

    Google Search Central

  • DocsSourceD1-C124

    Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.

    Search Central blog (3 June 2026)

  • AnalysisD2-C607

    Check AI Overviews and AI Mode with native-language queries in each target market, not with translated English keywords, and compare with the country breakdown of Search Console's generative AI performance report; topics where competitors are cited and you are not point to missing or weak local content.

    Ibrahim Anjro (author)

  • StageConfirmed by docsD3-C438

    The regular Search Console Performance report still shows all traffic including AI features, while the generative AI view shows only impressions from AI surfaces.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD3-C439

    Search Console's generative AI report shows impressions broken down by pages, countries and devices.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C441

    Search Console's generative AI performance reports count impressions only, with no clicks and no query dimension (Google documents the Search report for AI Overviews and AI Mode, plus a separate Discover report), so measure visits that follow AI answers with your own analytics.

    Ibrahim Anjro (author)

  • DocsSourceD3-C498

    Google's AI features guide says traffic from sites appearing in AI features such as AI Overviews and AI Mode is included in the overall search traffic in Search Console and reported in the Performance report within the Web search type.

    Google Search Central

  • StageConfirmed by docsD3-C458

    Multimodal searches reported in Search Console include searches with Google Lens, Circle to Search on Android, images uploaded to Google Search and Chrome's right-click search on an image.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C459

    To see multimodal traffic in the Search Console Performance report, select the multimodal search type instead of the default text-based one.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • AnalysisD3-C462

    Google's Performance report help says the queries dimension is not available for the multimodal search type, because these searches mostly use images, so judge image-led traffic by page, country and device, and make product and article images distinctive enough for Lens to match.

    Ibrahim Anjro (author)

  • DocsSourceD3-C465

    Google's announcement of 24 September 2026 says multimodal search data appears both in the Performance report for Search results and in the generative AI performance report, can be exported, and began rolling out globally that day.

    Search Central blog (24 September 2026)

ShouldDocumentedDEV-MON-08

Set up the Search Console bulk data export to BigQuery on large sites

Why

The Search Console interface exports at most 1,000 rows, and the Search Analytics API and the Looker Studio connector return up to 50,000 rows per day per site per search type. The bulk data export to BigQuery is the most complete query dataset available. Anonymized queries are always left out of report tables and are left out whenever a filter is applied, so filtered rows do not add up to the chart totals; in the export they stay as rows with an empty query.

How

Turn on the daily bulk data export in Search Console settings to a BigQuery project the team owns, set table expiration to fit the budget, and query the searchdata_site_impression and searchdata_url_impression tables. Keep anonymized rows in totals but exclude them (is_anonymized_query, or query = '') when ranking top queries, and tell any LLM or dashboard that analyses the data to do the same.

Test

New daily rows appear in the searchconsole dataset each day, and the summed clicks for a recent day are close to the Performance report's total for that day.

Code · BigQuery: top queries per page from the Search Console bulk export

Run against the searchdata_url_impression table of the bulk data export (default dataset searchconsole; replace the project name). Anonymized queries have an empty query and is_anonymized_query set, so they are left out of the ranking here but should stay in totals.

sql
-- Top non-anonymized queries per page, web search, last 28 days
SELECT
  url,
  query,
  SUM(clicks) AS clicks,
  SUM(impressions) AS impressions,
  SAFE_DIVIDE(SUM(clicks), SUM(impressions)) AS ctr,
  SAFE_DIVIDE(SUM(sum_position), SUM(impressions)) + 1 AS avg_position
FROM `my-project.searchconsole.searchdata_url_impression`
WHERE search_type = 'WEB'
  AND data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 28 DAY)
  AND NOT is_anonymized_query
GROUP BY url, query
ORDER BY clicks DESC
LIMIT 1000;
Evidence · 7 claims · 3 Google pages
  • StageConsistent with docsD3-C471

    Nik Vujic said it is super important to export all of a site's Search Console data, and his agency's workflow starts from Search Console's export to BigQuery.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C474

    A community slide said the Search Console interface export is limited to 1,000 rows.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConfirmed by docsD3-C475

    A community slide said the Search Console bulk data export to BigQuery gives the most complete query dataset available, with anonymized queries excluded.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C480

    A Search Central blog post on Search Console data limits (October 2022) says the interface exports at most 1,000 rows, while the Search Analytics API and the Looker Studio connector return up to 50,000 rows per day per site per search type.

    Search Central blog (19 October 2022)

  • DocsSourceD3-C481

    Google's blog post on Search Console data limits says anonymized queries are always left out of the report tables but counted in chart totals unless you filter by query, and are left out whenever a filter is applied, so filtered rows do not add up to the chart totals.

    Search Central blog (19 October 2022)

  • AnalysisD3-C476

    In the bulk export, 'anonymized queries excluded' means the query text, not the traffic: Google's table reference and query guidelines show those rows kept with an empty query field, often as the single most common 'query'. Tell an LLM to skip empty queries when ranking top queries, but keep them in totals.

    Ibrahim Anjro (author)

  • AnalysisD3-C502

    Looker Studio lifts Search Console's 1,000-row interface limit but keeps the connector's cap of 50,000 rows per day per site per search type; only the bulk export to BigQuery is free of the daily row limit, so large sites that need every non-anonymized query should use it.

    Ibrahim Anjro (author)

MayDocumentedDEV-MON-09

Add the brand's YouTube, Instagram, TikTok and X accounts to Search Console as platform properties

Why

Platform properties show the traffic Google Search, Discover and Google News send to a verified Instagram, TikTok, X or YouTube account, with its top content, query groups and countries, but not how often content is seen on the platform itself. Verification runs inside Search Console through a permission pop-up from the platform, with no token, and accounts of a claimed Search profile are added automatically.

How

Add one property per account or channel, record who owns each login, re-verify when the external login expires (access pauses until then), and compare formats by URL pattern in comparison mode, for example /watch against /shorts/ on YouTube.

Test

Every official account appears as a platform property with data, and the launch checklist names the owner of each login.

Evidence · 7 claims · 2 Google pages
  • StageConfirmed by docsD3-C443

    Platform properties let a user verify ownership of an Instagram, TikTok, X or YouTube account in Search Console and see the traffic Google sends to that account.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C446

    A platform property is verified entirely from within Search Console: a pop-up from the platform, for example Instagram, asks the user to grant Google permission to identify the account and its handle, and completing it proves ownership without a token.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C466

    Google's help page on platform properties says ownership is checked periodically and, if the external login expires, access pauses until the account is re-verified; each account or channel needs its own property, and data takes a few days to appear.

    Google Search Console Help

  • DocsSourceD3-C467

    Google's help page says platform properties show only how content performs on Google Search, Discover and Google News, not how often it is seen on the social or video platform itself.

    Google Search Console Help

  • DocsSourceD3-C468

    Google's social and video performance guide says that if you already claimed your Search profile, all of its verified accounts are added automatically as platform properties in Search Console.

    Google Search Central

  • StageConfirmed by docsD3-C451

    In a YouTube platform property, long-form and short-form videos can be compared in the Performance report's comparison mode by URL pattern, /watch versus /shorts/; Instagram posts (/p/) and reels can be compared the same way.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConfirmed by docsD2-C960

    A community speaker said publishers can also verify their social accounts in Search Console, which makes it easier to track the performance of content repurposed for social platforms.

    a community speaker at Search Central Live Deep Dive Europe 2026 (Day 2)

ShouldDocumentedDEV-MON-10

Plan releases and fixes around Google's processing times, and judge results only after them

Why

Said at Search Central Live, as estimates from Google's internal logs meant to inform rather than guarantee: robots.txt changes apply in about 24 hours, JavaScript-added content reaches indexing within hours and at worst weeks, a 404 or noindex drops a page within one to three weeks, a canonical change shows in one to three weeks, structured data updates in hours to two weeks, titles and snippets in one or two days up to several weeks, and a site move takes one to three months on average. Each step waits for the one before, so nothing changes until the page is recrawled. Google's documentation gives compatible figures, such as up to two weeks to split a duplicate cluster and a few days to a few weeks for title link changes.

How

Put the expected window into release plans and stakeholder updates, schedule robots.txt changes a day ahead, request indexing for the key changed URLs after a release, and review the effect only once the window for that kind of change has passed.

Test

Each SEO-relevant ticket records its expected processing window, and the post-release review is dated after it.

Evidence · 11 claims · 5 Google pages
  • SlideConfirmed by docsD3-C613

    Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C627

    Content that JavaScript adds to a page is typically seen by Google's indexing system within a few hours, and at worst within weeks.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C639

    After Google processes a page that now returns a 404 or a noindex, it usually removes the page from its serving index within one to three weeks, sometimes much sooner.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C640

    A change of canonical URL usually takes one to three weeks to show, though it can happen in seconds because Google allocates enough resources to canonicalization.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageNot in docsD3-C646

    Google usually picks up structured data updates within hours to one or two weeks, sometimes immediately.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • SlideConsistent with docsD3-C657

    Google estimated that a title update takes as long as a snippet update: 1-2 days on average, up to several weeks to months.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C643

    A site move can finish within a few weeks for a small site, takes one to three months on average, and in the worst case about a year.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • StageConsistent with docsD3-C665

    Google's processes depend on each other: crawling happens first, then indexing, then serving, so a change such as a new canonical cannot be processed until the page has been recrawled.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

  • DocsSourceD3-C669

    Google's guide to fixing canonicalization issues says that even after content issues are fixed, Google might hold pages in a duplicate cluster for up to two weeks, and that pages split out faster when they differ clearly and significantly.

    Google Search Central

  • DocsSourceD3-C675

    Google's title link guide says Google has to recrawl and reprocess a page to notice changes to the sources of its title link, which may take a few days to a few weeks.

    Google Search Central

  • StageD3-C598

    Google presented its processing times for crawling, indexing and serving as an experiment: estimates from its internal logs, meant to be informative and not something to obsess about, with feedback from the audience invited.

    Google at Search Central Live Deep Dive Europe 2026 (Day 3)

ShouldDocumentedDEV-MON-11

Measure the time from publishing a URL to Googlebot's first crawl of it

Why

Said at Search Central Live by Gary Illyes: the reliable check that new content is crawled fast enough is the time between publishing a URL and Googlebot's first fetch of it; it should stay flat or fall, and only a rising trend calls for investigation, starting with URLs Google crawls and recrawls a lot without use. How much it rises matters: for fresh news stories he called something like two hours probably not great. The Crawl stats split by purpose (discovery versus refresh) gives a rougher view; John Mueller said that if about half of a site's crawling goes to discovery, crawling of new pages is not the problem. News sites rarely need to worry about crawl budget, because Google crawls them aggressively.

How

Log each URL's publish time in the CMS, join it with the first verified Googlebot request in the access logs, and chart the median delay per section each week. When it rises, review the most-crawled URLs in the logs and block or remove useless ones (DEV-URL-08).

Test

The weekly chart exists per section and alerts when the median publish-to-first-crawl delay rises for several weeks in a row.

Evidence · 7 claims · 2 Google pages
  • StageNot in docsD1-C467

    To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a URL and Googlebot's first crawl of it: ideally it stays flat or falls, and only a rising trend calls for investigation.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C468

    When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google crawls and recrawls a lot and judging whether they are useful.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C492

    John Mueller said the Crawl Stats report's split of crawling by purpose, discovery versus refresh, gives a rough view of whether discovery crawling is in a reasonable range, though measuring publish-to-first-crawl time is more accurate.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C466

    Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConfirmed by docsD1-C396

    Search Console's crawl data separates discovery (fetches of URLs Google has not seen before) from refresh (fetches of known URLs), and lets you drill down to the problematic URLs.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageNot in docsD1-C543

    Gary Illyes said that if the time from publishing to Googlebot's first crawl starts rising, how much it rises matters: for the crawling of fresh stories, something like two hours is probably not great (a smaller example figure is unclear in the recording).

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

  • StageConsistent with docsD1-C544

    John Mueller said that if the Crawl Stats report shows about half of a site's crawling going to discovering new pages, that is a lot of discovery crawling, which suggests that crawling of new pages is not the problem.

    Google at Search Central Live Deep Dive Europe 2026 (Day 1)

Working together

From checklist to an agent that guards every site

This developer kit is free and stays open. If you manage many websites, we can build it into your workflow: an agentic system that checks every site against these requirements on each release, flags what breaks with the claim and the Google page behind it, and keeps launches clean.

Open for collaboration on this community version: Ibram & Dawwa GmbH · LinkedIn

Take it with you