Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Developer kit

Pre-launch checklist

56 items for the release review. Tick a MUST item when it is in place and an AVOID item when you have confirmed the site does not do it.

Server, status codes and crawling 10

  • Test

    Search Console Crawl stats report: host status shows no robots.txt fetch, DNS resolution or server connectivity failures, and no Search Console alert for DNS errors, which can get a site removed from Google Search very aggressively and, with it, from every feature that depends on Search. URL Inspection live test succeeds for one URL per template. WAF/CDN logs show no 403, 429 or challenge responses to IPs that verify as Googlebot.

    Why and how
  • Test

    Request a few URLs with a Googlebot user agent from an unverified IP: any challenge comes back as 503 or 429, never 200. The Page indexing report shows no sudden rise of "Duplicate" or "Soft 404" on unrelated URLs.

    Why and how
  • Test

    In a maintenance drill, curl -I on any page returns 503 while curl -I on /robots.txt returns 200. Afterwards, the 5xx share in Crawl stats falls back to zero.

    Why and how
  • Test

    curl -I https://<each-host>/robots.txt returns 200 or 404, never 5xx, 401/403 or a redirect to a login page. Search Console's robots.txt report (Settings) shows each host's file as fetched.

    Why and how
  • Test

    Check the file's fetch status, warnings and errors in Search Console's robots.txt report and test it with an RFC 9309 parser such as Google's open-source robots.txt library (github.com/google/robotstxt); check the size is well under 500 KiB; spot-check important URLs (homepage with tracking parameters, JS bundles, API endpoints, filter URLs) against the rules.

    Why and how
  • Test

    Fetch the live file (curl https://www.example.com/robots.txt) and check that every run of user-agent lines is uninterrupted; test it in Bing Webmaster Tools' robots.txt tester and with Google's open-source robots.txt parser (github.com/google/robotstxt). The robots.txt report's latest fetched version matches the file the team deployed.

    Why and how
  • Test

    curl -I --http1.1 https://www.example.com/ returns 200, and curl -I --http2 https://www.example.com/ returns the same status, canonical, robots and caching headers (Date and similar headers may differ), unless the site deliberately answers 421 to opt out of HTTP/2 crawling; Crawl stats shows normal response times.

    Why and how
  • Test

    Crawl stats and server logs show no 401, 402 or 403 responses to verified Googlebot on indexable templates, and URL Inspection's live test returns 200 for one URL per template.

    Why and how
  • Test

    For each named group, run Google's open-source robots.txt parser (github.com/google/robotstxt) with that crawler's token against a list of sample URLs covering every path the * group disallows: each result is the one intended for that crawler. For the quiz case, /staging/preview/ is allowed and /staging/other/ disallowed for Googlebot.

    Why and how
  • Test

    The live robots.txt names no admin, backup, export or internal paths, and curl -I on each private path returns 401 or 403 without credentials.

    Why and how
  • Test

    In URL Inspection (rendered HTML) and in a crawler with JavaScript off and on, every navigation item appears as <a href="...">; a crawl from the homepage reaches all key templates through links; search the rendered HTML for onclick=, javascript:, routerLink and href="#".

    Why and how
  • Test

    Search templates and the rendered HTML (DevTools Elements panel, URL Inspection) for onclick navigation, javascript:, routerLink without href, href on span or div, and href="#/"; a JavaScript-rendering crawler finds no orphaned templates.

    Why and how
  • Test

    Open a deep URL directly in a fresh browser session and with curl: it returns 200 and the right content. The rendered HTML in URL Inspection contains no href="#/..." routes.

    Why and how
  • Test

    Crawl from the homepage following only <a href> links and compare the result with the sitemap: sitemap URLs missing from the crawl are orphans. Check the click depth of key templates.

    Why and how
  • Test

    For pages 1 to 3 of a listing, curl the HTML: the canonical is self-referencing and the next-page <a href> is present without JavaScript. URL Inspection shows later pages as indexable.

    Why and how
  • Test

    With JavaScript disabled, follow the "more" link through every page of a listing. In URL Inspection the rendered HTML of page 1 contains a crawlable link to page 2, and server logs show Googlebot fetching the paginated URLs.

    Why and how
  • Test

    A crawl that respects robots.txt finds a bounded number of listing URLs; server logs and Crawl stats show Googlebot not spending fetches on parameter URLs; Google's open-source robots.txt parser (github.com/google/robotstxt), run with the Googlebot token, returns disallowed for sample filter URLs and allowed for item pages, and URL Inspection shows item pages as Crawl allowed.

    Why and how

Rendering and JavaScript 4

  • Test

    In URL Inspection (rendered HTML) or the Rich Results Test, find text from every tab and accordion panel. In the DevTools Network panel, no content request waits for a click.

    Why and how
  • Test

    The rendered HTML in URL Inspection contains below-the-fold images and list items. A code search finds no addEventListener('scroll') or onscroll handler that loads content.

    Why and how
  • Test

    URL Inspection live test, page resources: nothing needed for content is "Blocked by robots.txt". The DevTools Network panel (Fetch/XHR) lists each API host, and each host's robots.txt allows those paths.

    Why and how
  • Test

    The robots meta tag in curl output and in URL Inspection's rendered HTML is identical on every indexable template. A search of the client bundle finds no code that writes meta[name="robots"] except error views.

    Why and how

Indexing and snippet controls 3

  • Test

    For each noindexed URL: URL Inspection shows Crawl allowed: Yes, curl shows the noindex, and URL Inspection and the Page indexing report list it as excluded by the noindex tag.

    Why and how
  • Test

    Server logs show few or no Googlebot fetches of parameter and action URLs; robots.txt has no crawl-delay meant for Google; Crawl stats shows fewer fetches of low-value URLs after fixes.

    Why and how
  • Test

    Search the codebase, CMS settings and CDN transforms for nosnippet and max-snippet; curl key templates to confirm none is present unless intended.

    Why and how

Canonicals, redirects and duplicates 9

  • Test

    curl -sIL on an old URL shows one 301 or 308 hop ending in 200; a crawler's redirect-chain report is empty; after recrawl, URL Inspection of the new URL shows it as the Google-selected canonical.

    Why and how
  • Test

    curl -I on http://example.com, http://www.example.com and https://example.com each returns one 301 to the https://www. URL; a TLS check (openssl s_client or an online scanner) shows a valid certificate on every host name.

    Why and how
  • Test

    A crawl finds no indexable page with zero or several canonicals, or with a relative or placeholder value; raw and rendered HTML carry the same canonical; URL Inspection shows the user-declared and Google-selected canonical as equal on key templates.

    Why and how
  • Test

    Canonical audit in a crawler: group URLs by the final URL their canonicals and redirects lead to, and flag groups with more than one hop, a loop, several canonicals on one page, or a leader that is noindexed, blocked, non-200, redirecting or without internal links.

    Why and how
  • Test

    A crawl finds zero internal links to non-canonical URLs, zero sitemap URLs that are not self-canonical and zero hreflang targets that redirect or canonicalise elsewhere; sampled URLs in URL Inspection show matching user-declared and Google-selected canonicals.

    Why and how
  • Test

    Canonical audit: no canonical group contains documents in more than one language, and every localized page's canonical equals its own URL.

    Why and how
  • Test

    curl -I on the staging host returns 401 or 403 without credentials; the pre-launch check confirms production has no global noindex and no Disallow: /.

    Why and how
  • Test

    A script requests every old URL in the map and checks it returns one 301 to the mapped target, or the mapped 404 or 410; no old URL other than the old home page redirects to the new home page; the Page indexing report shows no rise in soft 404s on redirect targets.

    Why and how
  • Test

    When Search Console reports a URL as blocked by robots.txt that the file does not block, follow its redirect chain hop by hop, with a browser user agent and with Googlebot's, and test every hop against its own host's robots.txt with Google's open-source parser; run the same check on one URL per template after each release. JavaScript redirects only show in a rendered check, such as URL Inspection's live test.

    Why and how

Error pages and soft 404s 3

  • Test

    curl -I https://www.example.com/this-does-not-exist-123 returns 404; the Page indexing report's "Soft 404" count stays near zero.

    Why and how
  • Test

    curl -I on a made-up deep URL returns 404, or the rendered HTML in URL Inspection shows the redirect or the noindex; the Page indexing report shows no soft 404s on app routes.

    Why and how
  • Test

    In staging, simulate a backend outage: pages answer 503, not 200. Crawl for 200 pages with a very short main content or error wording, and review the Page indexing report's soft 404 list after each release.

    Why and how

International and multilingual sites 5

  • Test

    curl each version's URL from one location with no Accept-Language header and no cookies: each returns its own language and content.

    Why and how
  • Test

    curl -I on each country URL with IP addresses or Accept-Language headers of other countries: no 3xx to another version.

    Why and how
  • Test

    An hreflang audit in a crawler reports, per cluster, no missing return links, no missing self-references, no non-200 or non-canonical targets and no relative URLs.

    Why and how
  • Test

    An automated check confirms every value matches x-default or language[-Script][-REGION] (case-insensitive), with each part checked against ISO 639-1, ISO 15924 and ISO 3166-1 Alpha-2.

    Why and how
  • Test

    The rendered HTML of any page contains <a href> links to every version of that page, and a crawl starting from one version reaches all the others.

    Why and how

Page structure and HTML 2

  • Test

    A crawl finds no missing, duplicate, empty or placeholder titles, and no meta description that contains < or > or a placeholder.

    Why and how
  • Test

    A crawl with a style check finds no text whose colour matches its background and no off-screen or display:none blocks outside accessibility helpers and UI components.

    Why and how

Structured data 3

  • Test

    For each template, the Rich Results Test output matches what the page visibly shows, and the Manual actions report in Search Console is empty.

    Why and how
  • Test

    The Rich Results Test passes for one URL per template; CI fails on JSON-LD parse errors or missing required properties.

    Why and how
  • Test

    The Rich Results Test shows review snippets only on templates with visible reviews, and a crawl finds no rating on the site's own Organization or LocalBusiness entity.

    Why and how

Images 2

  • Test

    In the rendered HTML (URL Inspection) key images are img elements with a src; a template search finds no background-image on content images.

    Why and how
  • Test

    A crawl finds no content image with a missing alt, an empty alt or an alt equal to the file name; the axe or Lighthouse image-alt check passes.

    Why and how

Video 1

  • Test

    The rendered HTML in URL Inspection contains the video element or iframe with its URL, and the Video indexing report lists the page.

    Why and how

AI features (AI Overviews, AI Mode, Gemini) 1

  • Test

    URL Inspection shows the page indexed, its robots rules contain no nosnippet, and the Generative AI performance report shows impressions.

    Why and how

Spam policies and AI-generated content 4

  • Test

    An automated browser test opens each template from another page, waits and interacts, presses Back once and lands on the previous page; a code search of the built bundles finds no pushState or replaceState outside the router.

    Why and how
  • Test

    A crawl of each programmatic template finds no clusters of near-identical pages; publishing logs show an editor's approval for generated pages; Search Console's Manual actions report is empty.

    Why and how
  • Test

    The HTML served to a Googlebot user agent and to a browser have the same main content, links and redirects, and URL Inspection's rendered HTML matches what users see.

    Why and how
  • Test

    A weekly comparison of the sitemap and a crawl with the CMS's own page list finds no unknown URLs; Search Console's Security issues and Manual actions reports are empty.

    Why and how

Testing and monitoring 2

  • Test

    Search Console lists the Domain property as verified, and every team that ships code has access.

    Why and how
  • Test

    The release checklist records a headless render result for each template before release and a URL Inspection live test result for each template after release.

    Why and how

Ticks are kept in this browser only. Print the page for a paper copy.

Looking after many sites? This kit can run as an agent on every release.