Developer kit
Pre-launch checklist
56 items for the release review. Tick a MUST item when it is in place and an AVOID item when you have confirmed the site does not do it.
Server, status codes and crawling 10
- TestWhy and how
Search Console Crawl stats report: host status shows no robots.txt fetch, DNS resolution or server connectivity failures, and no Search Console alert for DNS errors, which can get a site removed from Google Search very aggressively and, with it, from every feature that depends on Search. URL Inspection live test succeeds for one URL per template. WAF/CDN logs show no 403, 429 or challenge responses to IPs that verify as Googlebot.
- TestWhy and how
Request a few URLs with a Googlebot user agent from an unverified IP: any challenge comes back as 503 or 429, never 200. The Page indexing report shows no sudden rise of "Duplicate" or "Soft 404" on unrelated URLs.
- TestWhy and how
In a maintenance drill, curl -I on any page returns 503 while curl -I on /robots.txt returns 200. Afterwards, the 5xx share in Crawl stats falls back to zero.
- TestWhy and how
curl -I https://<each-host>/robots.txt returns 200 or 404, never 5xx, 401/403 or a redirect to a login page. Search Console's robots.txt report (Settings) shows each host's file as fetched.
- TestWhy and how
Check the file's fetch status, warnings and errors in Search Console's robots.txt report and test it with an RFC 9309 parser such as Google's open-source robots.txt library (github.com/google/robotstxt); check the size is well under 500 KiB; spot-check important URLs (homepage with tracking parameters, JS bundles, API endpoints, filter URLs) against the rules.
- TestWhy and how
Fetch the live file (curl https://www.example.com/robots.txt) and check that every run of user-agent lines is uninterrupted; test it in Bing Webmaster Tools' robots.txt tester and with Google's open-source robots.txt parser (github.com/google/robotstxt). The robots.txt report's latest fetched version matches the file the team deployed.
- TestWhy and how
curl -I --http1.1 https://www.example.com/ returns 200, and curl -I --http2 https://www.example.com/ returns the same status, canonical, robots and caching headers (Date and similar headers may differ), unless the site deliberately answers 421 to opt out of HTTP/2 crawling; Crawl stats shows normal response times.
- TestWhy and how
Crawl stats and server logs show no 401, 402 or 403 responses to verified Googlebot on indexable templates, and URL Inspection's live test returns 200 for one URL per template.
- TestWhy and how
For each named group, run Google's open-source robots.txt parser (github.com/google/robotstxt) with that crawler's token against a list of sample URLs covering every path the * group disallows: each result is the one intended for that crawler. For the quiz case, /staging/preview/ is allowed and /staging/other/ disallowed for Googlebot.
- TestWhy and how
The live robots.txt names no admin, backup, export or internal paths, and curl -I on each private path returns 401 or 403 without credentials.
URLs, links and discovery 7
- TestWhy and how
In URL Inspection (rendered HTML) and in a crawler with JavaScript off and on, every navigation item appears as <a href="...">; a crawl from the homepage reaches all key templates through links; search the rendered HTML for onclick=, javascript:, routerLink and href="#".
- TestWhy and how
Search templates and the rendered HTML (DevTools Elements panel, URL Inspection) for onclick navigation, javascript:, routerLink without href, href on span or div, and href="#/"; a JavaScript-rendering crawler finds no orphaned templates.
- TestWhy and how
Open a deep URL directly in a fresh browser session and with curl: it returns 200 and the right content. The rendered HTML in URL Inspection contains no href="#/..." routes.
- TestWhy and how
Crawl from the homepage following only <a href> links and compare the result with the sitemap: sitemap URLs missing from the crawl are orphans. Check the click depth of key templates.
- TestWhy and how
For pages 1 to 3 of a listing, curl the HTML: the canonical is self-referencing and the next-page <a href> is present without JavaScript. URL Inspection shows later pages as indexable.
- TestWhy and how
With JavaScript disabled, follow the "more" link through every page of a listing. In URL Inspection the rendered HTML of page 1 contains a crawlable link to page 2, and server logs show Googlebot fetching the paginated URLs.
- TestWhy and how
A crawl that respects robots.txt finds a bounded number of listing URLs; server logs and Crawl stats show Googlebot not spending fetches on parameter URLs; Google's open-source robots.txt parser (github.com/google/robotstxt), run with the Googlebot token, returns disallowed for sample filter URLs and allowed for item pages, and URL Inspection shows item pages as Crawl allowed.
Rendering and JavaScript 4
- TestWhy and how
In URL Inspection (rendered HTML) or the Rich Results Test, find text from every tab and accordion panel. In the DevTools Network panel, no content request waits for a click.
- TestWhy and how
The rendered HTML in URL Inspection contains below-the-fold images and list items. A code search finds no addEventListener('scroll') or onscroll handler that loads content.
- TestWhy and how
URL Inspection live test, page resources: nothing needed for content is "Blocked by robots.txt". The DevTools Network panel (Fetch/XHR) lists each API host, and each host's robots.txt allows those paths.
- TestWhy and how
The robots meta tag in curl output and in URL Inspection's rendered HTML is identical on every indexable template. A search of the client bundle finds no code that writes meta[name="robots"] except error views.
Indexing and snippet controls 3
- TestWhy and how
For each noindexed URL: URL Inspection shows Crawl allowed: Yes, curl shows the noindex, and URL Inspection and the Page indexing report list it as excluded by the noindex tag.
- TestWhy and how
Server logs show few or no Googlebot fetches of parameter and action URLs; robots.txt has no crawl-delay meant for Google; Crawl stats shows fewer fetches of low-value URLs after fixes.
- TestWhy and how
Search the codebase, CMS settings and CDN transforms for nosnippet and max-snippet; curl key templates to confirm none is present unless intended.
Canonicals, redirects and duplicates 9
- TestWhy and how
curl -sIL on an old URL shows one 301 or 308 hop ending in 200; a crawler's redirect-chain report is empty; after recrawl, URL Inspection of the new URL shows it as the Google-selected canonical.
- TestWhy and how
curl -I on http://example.com, http://www.example.com and https://example.com each returns one 301 to the https://www. URL; a TLS check (openssl s_client or an online scanner) shows a valid certificate on every host name.
- TestWhy and how
A crawl finds no indexable page with zero or several canonicals, or with a relative or placeholder value; raw and rendered HTML carry the same canonical; URL Inspection shows the user-declared and Google-selected canonical as equal on key templates.
- TestWhy and how
Canonical audit in a crawler: group URLs by the final URL their canonicals and redirects lead to, and flag groups with more than one hop, a loop, several canonicals on one page, or a leader that is noindexed, blocked, non-200, redirecting or without internal links.
- TestWhy and how
A crawl finds zero internal links to non-canonical URLs, zero sitemap URLs that are not self-canonical and zero hreflang targets that redirect or canonicalise elsewhere; sampled URLs in URL Inspection show matching user-declared and Google-selected canonicals.
- TestWhy and how
Canonical audit: no canonical group contains documents in more than one language, and every localized page's canonical equals its own URL.
- TestWhy and how
curl -I on the staging host returns 401 or 403 without credentials; the pre-launch check confirms production has no global noindex and no Disallow: /.
- TestWhy and how
A script requests every old URL in the map and checks it returns one 301 to the mapped target, or the mapped 404 or 410; no old URL other than the old home page redirects to the new home page; the Page indexing report shows no rise in soft 404s on redirect targets.
- TestWhy and how
When Search Console reports a URL as blocked by robots.txt that the file does not block, follow its redirect chain hop by hop, with a browser user agent and with Googlebot's, and test every hop against its own host's robots.txt with Google's open-source parser; run the same check on one URL per template after each release. JavaScript redirects only show in a rendered check, such as URL Inspection's live test.
Error pages and soft 404s 3
- TestWhy and how
curl -I https://www.example.com/this-does-not-exist-123 returns 404; the Page indexing report's "Soft 404" count stays near zero.
- TestWhy and how
curl -I on a made-up deep URL returns 404, or the rendered HTML in URL Inspection shows the redirect or the noindex; the Page indexing report shows no soft 404s on app routes.
- TestWhy and how
In staging, simulate a backend outage: pages answer 503, not 200. Crawl for 200 pages with a very short main content or error wording, and review the Page indexing report's soft 404 list after each release.
International and multilingual sites 5
- TestWhy and how
curl each version's URL from one location with no Accept-Language header and no cookies: each returns its own language and content.
- TestWhy and how
curl -I on each country URL with IP addresses or Accept-Language headers of other countries: no 3xx to another version.
- TestWhy and how
An hreflang audit in a crawler reports, per cluster, no missing return links, no missing self-references, no non-200 or non-canonical targets and no relative URLs.
- TestWhy and how
An automated check confirms every value matches x-default or language[-Script][-REGION] (case-insensitive), with each part checked against ISO 639-1, ISO 15924 and ISO 3166-1 Alpha-2.
- TestWhy and how
The rendered HTML of any page contains <a href> links to every version of that page, and a crawl starting from one version reaches all the others.
Page structure and HTML 2
- TestWhy and how
A crawl finds no missing, duplicate, empty or placeholder titles, and no meta description that contains < or > or a placeholder.
- TestWhy and how
A crawl with a style check finds no text whose colour matches its background and no off-screen or display:none blocks outside accessibility helpers and UI components.
Structured data 3
- TestWhy and how
For each template, the Rich Results Test output matches what the page visibly shows, and the Manual actions report in Search Console is empty.
- TestWhy and how
The Rich Results Test passes for one URL per template; CI fails on JSON-LD parse errors or missing required properties.
- TestWhy and how
The Rich Results Test shows review snippets only on templates with visible reviews, and a crawl finds no rating on the site's own Organization or LocalBusiness entity.
Images 2
- TestWhy and how
In the rendered HTML (URL Inspection) key images are img elements with a src; a template search finds no background-image on content images.
- TestWhy and how
A crawl finds no content image with a missing alt, an empty alt or an alt equal to the file name; the axe or Lighthouse image-alt check passes.
Video 1
- TestWhy and how
The rendered HTML in URL Inspection contains the video element or iframe with its URL, and the Video indexing report lists the page.
AI features (AI Overviews, AI Mode, Gemini) 1
- TestWhy and how
URL Inspection shows the page indexed, its robots rules contain no nosnippet, and the Generative AI performance report shows impressions.
Spam policies and AI-generated content 4
- TestWhy and how
An automated browser test opens each template from another page, waits and interacts, presses Back once and lands on the previous page; a code search of the built bundles finds no pushState or replaceState outside the router.
- TestWhy and how
A crawl of each programmatic template finds no clusters of near-identical pages; publishing logs show an editor's approval for generated pages; Search Console's Manual actions report is empty.
- TestWhy and how
The HTML served to a Googlebot user agent and to a browser have the same main content, links and redirects, and URL Inspection's rendered HTML matches what users see.
- TestWhy and how
A weekly comparison of the sitemap and a crawl with the CMS's own page list finds no unknown URLs; Search Console's Security issues and Manual actions reports are empty.
Testing and monitoring 2
- TestWhy and how
Search Console lists the Domain property as verified, and every team that ships code has access.
- TestWhy and how
The release checklist records a headless render result for each template before release and a URL Inspection live test result for each template after release.
Ticks are kept in this browser only. Print the page for a paper copy.
Looking after many sites? This kit can run as an agent on every release.