Developer kit
Code snippets
35 snippets to copy into templates and server configuration. Replace example.com and the sample values with your own.
Verify that a request really comes from Googlebot
Before a firewall, CDN or bot-protection rule blocks or challenges a "Googlebot" request, verify it: the user-agent string alone can be faked. In WAF rules, prefer matching the IP against the ranges Google publishes as common-crawlers.json and special-crawlers.json on its page on verifying its crawlers; for log analysis, a reverse DNS lookup followed by a forward lookup works too. Never allowlist a whole googleusercontent.com host: Google Cloud customer VMs resolve there too, and only *.gae.googleusercontent.com belongs to Google's user-triggered fetchers, which are not Googlebot.
# 1. Reverse DNS: Googlebot and Google's other common crawlers resolve to crawl-*.googlebot.com
# or geo-crawl-*.geo.googlebot.com; special-case crawlers to rate-limited-proxy-*.google.com.
# A host such as 81.59.117.34.bc.googleusercontent.com is a Google Cloud customer, not Googlebot.
host 66.249.66.1
# -> 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
# 2. Forward DNS: that host name must resolve back to the same IP
host crawl-66-249-66-1.googlebot.com
# -> crawl-66-249-66-1.googlebot.com has address 66.249.66.1The published IP ranges are easier to keep in sync with WAF and CDN allowlists than DNS lookups:
# Common crawlers (Googlebot and others) and special-case crawlers; refresh the lists regularly
curl -s https://developers.google.com/static/crawling/ipranges/common-crawlers.json
curl -s https://developers.google.com/static/crawling/ipranges/special-crawlers.jsonUsed by DEV-SRV-01 Let verified Google crawlers through firewalls, CDNs and bot protection, DEV-MON-04 Monitor Crawl stats and server logs for verified Googlebot traffic
Planned maintenance with 503 (robots.txt stays available)
During a short outage every page answers 503 Service Unavailable, never a 200 "we'll be back" page and never a 404. Keep it to a day or two, and keep serving robots.txt normally: a 5xx on robots.txt stops crawling of the whole site.
HTTP/1.1 503 Service Unavailable
Content-Type: text/html; charset=utf-8
Retry-After: 3600
Cache-Control: no-storeserver {
listen 443 ssl;
server_name www.example.com;
ssl_certificate /etc/ssl/www.example.com.crt;
ssl_certificate_key /etc/ssl/www.example.com.key;
error_page 503 /maintenance.html;
location = /maintenance.html {
root /var/www/static;
internal;
# add_header here replaces server-level add_header lines: repeat HSTS or CSP if you set them
add_header Retry-After 3600 always;
add_header Cache-Control "no-store" always;
}
# robots.txt keeps answering 200 during maintenance
location = /robots.txt {
root /var/www/static;
}
location / {
return 503;
}
}Used by DEV-SRV-03 Answer planned maintenance and short outages with 503, for a day or two at most
robots.txt with crawl controls and a rendering carve-out
Disallow only URLs that should never be crawled (internal search, cart and checkout actions, filter parameters), keep every script, style and API path that pages need for rendering crawlable, and list the sitemap. Google reads only user-agent, allow, disallow and sitemap; each host (www, api, cdn) needs its own file at its root. Robots.txt is public, so never list secret paths in it.
# https://www.example.com/robots.txt
User-agent: *
# Internal search results and cart/checkout actions (adapt to your URL patterns)
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
# (Shops in Google Shopping: give Storebot-Google its own group that leaves cart and checkout open.
# A named group replaces this * group for that crawler, so repeat in it every rule that should still apply.)
# Filter and sort parameters of faceted navigation, as the first parameter (?) or a later one (&);
# /*?*size= would also block ?pagesize= and /*?*color= would block ?bgcolor=
Disallow: /*?color=
Disallow: /*&color=
Disallow: /*?size=
Disallow: /*&size=
Disallow: /*?sort=
Disallow: /*&sort=
# API: blocked in general, but the endpoints pages render from stay crawlable
# (the longer, more specific Allow rule wins)
Disallow: /api/
Allow: /api/products/
Allow: /api/reviews/
# Never disallow /static/, /assets/ or other JS and CSS folders
# Optional: keep content out of Gemini training and grounding (no effect on Google Search)
User-agent: Google-Extended
Disallow: /
Sitemap: https://www.example.com/sitemap.xmlUsed by DEV-SRV-04 Keep robots.txt reachable on every host: 200 with rules, or 404 if there are none, DEV-SRV-05 Rely only on the fields Google supports (user-agent, allow, disallow, sitemap) and keep robots.txt under 500 KiB, DEV-URL-08 Keep faceted navigation out of the crawl with robots.txt rules or URL fragments, DEV-REN-04 Do not block JavaScript, CSS or API endpoints that pages need for rendering in robots.txt, DEV-IDX-11 Use the Google-Extended robots.txt token to control Gemini training and grounding, DEV-AIF-05 Set robots.txt policy per AI crawler by its own token, and know that user-triggered fetchers ignore it
ETag and 304 Not Modified for crawlers
Send a validator (ETag, which Google's crawling team prefers, or Last-Modified) and answer a matching conditional request with 304 and no body, so unchanged pages cost almost nothing to recrawl. Express does this for you when it sends the body, but its ETag hashes the whole response: if the HTML embeds a CSP nonce, a CSRF token or a timestamp, compute the ETag from the content data instead (for example a hash of the product record and the template version) and set it with res.set('ETag', ...).
GET /chairs/oak-dining-chair HTTP/1.1
Host: www.example.com
If-None-Match: "a3f9c2e1"
HTTP/1.1 304 Not Modified
ETag: "a3f9c2e1"
Cache-Control: no-cache// Express: strong ETags, and an automatic 304 when If-None-Match matches
const express = require('express');
const app = express();
app.set('etag', 'strong');
app.get('/chairs/:slug', async (req, res) => {
const html = await renderChairPage(req.params.slug); // your server-side renderer
res.set('Cache-Control', 'no-cache'); // may be stored, but must be revalidated
res.type('html').send(html); // res.send() compares the ETag and answers 304 when fresh
});
// Pages with a per-request nonce or CSRF token: an ETag from the data, not from the body
const crypto = require('crypto');
const TEMPLATE_VERSION = '2026-10-02';
app.get('/products/:slug', async (req, res) => {
const product = await loadProduct(req.params.slug); // your data access
const etag = '"' + crypto.createHash('sha256').update(JSON.stringify(product) + TEMPLATE_VERSION).digest('hex').slice(0, 16) + '"';
res.set('ETag', etag);
res.set('Cache-Control', 'no-cache');
if (req.fresh) return res.status(304).end(); // If-None-Match matches the data ETag
res.type('html').send(renderProductPage(product, res.locals.cspNonce));
});Used by DEV-SRV-08 Support conditional requests with ETag or Last-Modified and answer 304 when nothing changed
Complete robots.txt groups for named crawlers
A crawler obeys only the most specific group that names it and ignores the * group, so groups are not additive: a named group must repeat every shared rule that should still apply to that crawler. One group can name several crawlers, and Google combines all groups that name the same user agent into one.
# WRONG: the googlebot group replaces the * group instead of adding to it,
# so Googlebot may crawl /goats/ and /cows/ and only /dogs/ is closed to it
User-agent: *
Disallow: /goats/
Disallow: /cows/
User-agent: googlebot
Disallow: /dogs/Right: the shared rules sit in the * group, and the named group repeats them before adding its own. To open /staging/preview/ to Googlebot while the rest of /staging/ stays closed to it, the Googlebot group needs both the disallow and the allow; an allow in the * group, or a Googlebot group with only the allow, does not do it.
# Shared rules for every crawler without a group of its own
User-agent: *
Disallow: /goats/
Disallow: /cows/
Disallow: /staging/
# Googlebot: the shared rules again, plus its own
User-agent: googlebot
Disallow: /goats/
Disallow: /cows/
Disallow: /dogs/
Disallow: /staging/
Allow: /staging/preview/
# Several crawlers can share one group: stack their user-agent lines with nothing in between
User-agent: bingbot
User-agent: applebot
Disallow: /goats/
Disallow: /cows/
Disallow: /staging/
Sitemap: https://www.example.com/sitemap.xmlUsed by DEV-SRV-10 Give a crawler its own robots.txt group only when needed, and make that group complete
Crawlable links versus links Google cannot rely on
Google extracts links from <a> elements with an href that holds a real URL, both from the server HTML and from the rendered page. Click handlers, javascript: URLs, routerLink without href, href on other elements and # routes are not dependable.
<!-- Crawlable: <a> with a real relative or absolute URL -->
<a href="/chairs/">Chairs</a>
<a href="https://www.example.com/chairs/oak-dining-chair">Oak dining chair</a>
<!-- Crawlable and still client-side routed: keep the href, intercept the click in JavaScript -->
<a href="/chairs/" data-link>Chairs</a>
<!-- Not dependable: do not use for anything that should be discovered -->
<a onclick="goTo('/chairs/')">Chairs</a>
<a routerLink="/chairs/">Chairs</a>
<a href="javascript:goTo('chairs')">Chairs</a>
<a href="#/chairs">Chairs</a>
<span data-href="/chairs/" onclick="location.href=this.dataset.href">Chairs</span>
<button type="button" onclick="location.href='/chairs/'">Chairs</button>Used by DEV-URL-01 Make every navigational link an <a> element whose href holds a real URL, DEV-URL-02 Do not navigate with onclick handlers, javascript: URLs, routerLink without href or # pseudo-links
Single-page app routing with the History API instead of #/ routes
Every view gets a real path in a real <a href>; JavaScript intercepts the click, updates the address with pushState and renders the view, so users skip full reloads while Google can follow the links. The server must answer a direct request for each path with the full page (ideally server-rendered) and unknown paths with 404. The not-found view is a route of its own, so a shell served with 404 renders it instead of redirecting again.
<nav>
<a href="/" data-link>Home</a>
<a href="/chairs/" data-link>Chairs</a>
<a href="/tables/" data-link>Tables</a>
</nav>
<main id="app"></main>
<script>
const routes = {
'/': () => '<h1>Example Shop</h1>',
'/chairs/': () => '<h1>Chairs</h1>',
'/tables/': () => '<h1>Tables</h1>',
'/not-found': () => '<h1>Page not found</h1>', // the server answers this URL with 404
};
function renderRoute(path) {
const view = routes[path];
if (!view) {
// Unknown route on a 200 URL: go to the URL the server answers with 404 (once; never loop)
if (path !== '/not-found') window.location.replace('/not-found');
return;
}
document.getElementById('app').innerHTML = view();
}
document.addEventListener('click', (event) => {
const link = event.target.closest('a[data-link]');
if (!link || link.origin !== location.origin) return;
if (event.button !== 0 || event.metaKey || event.ctrlKey || event.shiftKey || event.altKey) return;
event.preventDefault();
history.pushState({}, '', link.pathname + link.search);
renderRoute(link.pathname);
});
window.addEventListener('popstate', () => renderRoute(location.pathname));
renderRoute(location.pathname);
</script>Used by DEV-URL-03 Route single-page apps with the History API and real paths, not # fragments
XML sitemap with canonical URLs and real lastmod dates
List only canonical, indexable URLs that answer 200 (no redirects, no noindex, no parameters you canonicalise away), with a lastmod that changes only for significant updates (the main content, structured data or links), never for footer, copyright or timestamp changes; Google uses lastmod only when it is consistently accurate. Split large sites with a sitemap index (at most 50,000 URLs or 50 MB uncompressed per file), reference it in robots.txt and submit it in Search Console.
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemaps/products-1.xml</loc>
<lastmod>2026-09-30T06:00:00+00:00</lastmod>
</sitemap>
</sitemapindex><?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://www.example.com/chairs/oak-dining-chair</loc>
<lastmod>2026-09-28T10:15:00+00:00</lastmod>
</url>
<url>
<loc>https://www.example.com/chairs/beech-stool</loc>
<lastmod>2026-09-12T08:00:00+00:00</lastmod>
</url>
</urlset>Used by DEV-URL-05 Publish XML sitemaps listing only canonical, indexable URLs with accurate lastmod
Paginated listing with crawlable page links
Each page of a series has its own URL (?page=n), its own self-referencing canonical and plain <a href> links to the next page and back to the first page; the pages may share one title and description, so a page number in the title is optional. Do not canonicalise page 2 and later to page 1, and do not put page numbers in # fragments; Google no longer uses rel=next/rel=prev.
<!-- https://www.example.com/chairs/?page=2 -->
<head>
<title>Chairs, page 2 | Example Shop</title> <!-- the page number is optional -->
<link rel="canonical" href="https://www.example.com/chairs/?page=2">
</head>
<body>
<main>
<h1>Chairs</h1>
<ul>
<li><a href="/chairs/oak-dining-chair">Oak dining chair</a></li>
<li><a href="/chairs/beech-stool">Beech stool</a></li>
</ul>
<nav aria-label="Pagination">
<a href="/chairs/">1</a>
<a href="/chairs/?page=2" aria-current="page">2</a>
<a href="/chairs/?page=3">3</a>
<a href="/chairs/?page=3">Next page</a>
</nav>
</main>
</body>Used by DEV-URL-06 Give each page of a paginated series its own URL, its own canonical and an <a href> to the next page
Lazy loading that works without scrolling
Googlebot does not scroll; it renders with a very tall viewport. Use native loading="lazy" for images and iframes, and load list chunks with an IntersectionObserver on a sentinel element, never with a scroll event listener. The plain <a href> links to the paginated URLs stay in the page whatever the script does, so the rendered HTML always links to the next page.
<!-- Images and iframes: native lazy loading, real src in the HTML -->
<img src="/img/oak-chair-800.webp" alt="Oak dining chair with a woven paper-cord seat" width="800" height="600" loading="lazy">
<!-- Lists: page 1 of /chairs/, a sentinel for the observer, and pagination links the script never removes -->
<ul id="products">
<li><a href="/chairs/oak-dining-chair">Oak dining chair</a></li>
</ul>
<div id="load-more-sentinel" style="height: 1px"></div>
<nav aria-label="Pagination">
<a href="/chairs/" aria-current="page">1</a>
<a href="/chairs/?page=2">2</a>
<a href="/chairs/?page=3">3</a>
<a href="/chairs/?page=2">Next page</a>
</nav>
<!-- Avoid: window.addEventListener('scroll', loadMoreProducts) never runs for Googlebot -->
<script>
const list = document.getElementById('products');
const sentinel = document.getElementById('load-more-sentinel');
let next = '/chairs/?page=2'; // written by the server: the next page's URL, empty on the last page
const observer = new IntersectionObserver(async ([entry]) => {
if (!entry.isIntersecting || !next) return;
observer.unobserve(sentinel);
const url = new URL(next, location.href);
const res = await fetch('/fragments' + url.pathname + url.search); // returns the <li> items of that page
if (!res.ok) { observer.disconnect(); return; } // keep the links; never insert an error page
list.insertAdjacentHTML('beforeend', await res.text());
history.replaceState({}, '', url.pathname + url.search); // the URL follows the chunk once it is in place
next = res.headers.get('X-Next-Page') || ''; // e.g. "/chairs/?page=3", absent on the last page
if (next) observer.observe(sentinel); else observer.disconnect();
});
if (next) observer.observe(sentinel);
</script>Used by DEV-URL-07 Back infinite scroll and load-more buttons with paginated URLs, DEV-REN-03 Lazy-load with native lazy loading or an IntersectionObserver, never with scroll events
Redirect URL variants to the one URL the application expects
Build the expected URL from the route and the parameters the page really uses, in a fixed order, and answer any other spelling with one 301. Generate internal links with the same function so the site never links to a variant. The example is Express middleware; lowercasing the path assumes the site's routes are all lowercase.
const ORIGIN = 'https://www.example.com';
// Parameters each route actually uses, in the order they must appear.
const PARAMS = { '/chairs': ['colour', 'page'], '/search': ['q'] };
function expectedUrl(pathname, searchParams) {
let path = pathname.toLowerCase().replace(/\/{2,}/g, '/');
if (path.length > 1 && path.endsWith('/')) path = path.slice(0, -1);
const kept = new URLSearchParams();
for (const name of PARAMS[path] || []) {
const value = searchParams.get(name);
if (value) kept.set(name, value); // unused and empty parameters are dropped
}
const query = kept.toString(); // spaces always come out as +
return path + (query ? '?' + query : '');
}
app.set('trust proxy', true); // so req.protocol is right behind a CDN or load balancer
app.use((req, res, next) => {
const requested = new URL(ORIGIN + req.originalUrl); // not new URL(path, base): '//x' would parse as a host
const expected = expectedUrl(requested.pathname, requested.searchParams);
const wrongOrigin = req.protocol !== 'https' || req.hostname !== 'www.example.com';
if (wrongOrigin || expected !== requested.pathname + requested.search) {
return res.redirect(301, ORIGIN + expected); // one hop, straight to the final URL
}
next();
});Used by DEV-URL-11 Compute the expected URL for every request and redirect variants to it
Page template with a clearly delimited main content area
Everything indexing depends on (title, description, canonical, robots rules, main text, links) is in the server HTML. Header, navigation and footer are separated from one <main> element that holds the title, headings, opening text and media, because Google weighs words by where they appear and treats the main content as the most important part.
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Oak dining chair with woven seat | Example Shop</title>
<meta name="description" content="Solid oak dining chair with a hand-woven paper-cord seat. Seat height 45 cm, delivered assembled.">
<link rel="canonical" href="https://www.example.com/chairs/oak-dining-chair">
<meta name="robots" content="max-snippet:-1, max-image-preview:large">
</head>
<body>
<header>
<a href="/">Example Shop</a>
<nav aria-label="Main">
<a href="/chairs/">Chairs</a>
<a href="/tables/">Tables</a>
</nav>
</header>
<main>
<article>
<h1>Oak dining chair with woven seat</h1>
<p>A solid oak dining chair with a hand-woven paper-cord seat, made for everyday family meals.</p>
<figure>
<img src="/img/oak-chair-1200.webp" alt="Oak dining chair with a woven paper-cord seat, seen from the front" width="1200" height="900">
<figcaption>Natural oak finish, seat height 45 cm.</figcaption>
</figure>
<h2>Dimensions and materials</h2>
<p>Width 46 cm, depth 52 cm, height 80 cm. <strong>Solid European oak</strong>, paper-cord seat.</p>
</article>
</main>
<footer>
<nav aria-label="Footer">
<a href="/delivery/">Delivery</a>
<a href="/returns/">Returns</a>
<a href="/contact/">Contact</a>
</nav>
</footer>
</body>
</html>Used by DEV-REN-01 Server-render the main content and everything indexing depends on (SSR, static generation or hybrid), DEV-HTM-01 Wrap each page's primary content in one clearly delimited main area
Tabs and accordions whose content is in the DOM from the start
Googlebot does not click, so content fetched only when a tab is clicked never exists for it. Ship every panel in the HTML and only hide the inactive ones with the hidden attribute or CSS: hidden content in the DOM can be indexed, absent content cannot.
<div class="tabs">
<div role="tablist" aria-label="Product details">
<button type="button" role="tab" id="tab-specs" aria-controls="panel-specs" aria-selected="true">Specifications</button>
<button type="button" role="tab" id="tab-care" aria-controls="panel-care" aria-selected="false">Care</button>
</div>
<section role="tabpanel" id="panel-specs" aria-labelledby="tab-specs">
<p>Seat height 45 cm, solid oak frame, paper-cord seat.</p>
</section>
<section role="tabpanel" id="panel-care" aria-labelledby="tab-care" hidden>
<p>Wipe with a damp cloth; re-oil the frame once a year.</p>
</section>
</div>
<!-- Avoid: tab.onclick = () => fetch('/api/specs') ... content that exists only after a click -->
<script>
const tabs = document.querySelectorAll('[role="tab"]');
tabs.forEach((tab) => {
tab.addEventListener('click', () => {
tabs.forEach((other) => {
const selected = other === tab;
other.setAttribute('aria-selected', String(selected));
document.getElementById(other.getAttribute('aria-controls')).hidden = !selected;
});
});
});
</script>Used by DEV-REN-02 Have all indexable content in the DOM after load, without clicks, typing or scrolling
Feature detection with a fallback for unsupported or permission-based APIs
Google's renderer does not grant permission prompts (location, payments, camera) and does not run WebGL well. Render the indexable content first and treat such features as optional enhancements behind feature detection.
<main>
<h1>Our stores</h1>
<ul id="stores">
<li>Barcelona, Carrer Example 1</li>
<li>Madrid, Calle Example 2</li>
</ul>
<button type="button" id="near-me" hidden>Sort by distance</button>
<canvas class="hero-effect" width="1200" height="400"></canvas>
</main>
<script>
// Location: optional, only after a user click; the full list is already in the page
if ('geolocation' in navigator) {
const button = document.getElementById('near-me');
button.hidden = false;
button.addEventListener('click', () => {
navigator.geolocation.getCurrentPosition(
(position) => sortStoresByDistance(position.coords),
() => {} // declined or unavailable: keep the default order
);
});
}
function sortStoresByDistance(coords) {
console.log('Sort stores near', coords.latitude, coords.longitude);
}
// WebGL: decoration only; without it the server-rendered text and images stay as they are
const canvas = document.querySelector('canvas.hero-effect');
const gl = canvas.getContext('webgl');
if (gl) {
gl.clearColor(0.1, 0.3, 0.5, 1.0);
gl.clear(gl.COLOR_BUFFER_BIT);
} else {
canvas.remove();
}
</script>Used by DEV-REN-08 Feature-detect permission-based APIs and WebGL, and render the content without them
Robots meta tags for common page types
Put robots rules in the <head> of the HTML the server sends, one variant per page type below; never add, change or remove them with JavaScript. A noindex only works if the URL is not disallowed in robots.txt, because Google has to fetch the page to see it.
<!-- Indexable templates: allow full-length snippets and large image and video previews -->
<meta name="robots" content="max-snippet:-1, max-image-preview:large, max-video-preview:-1">
<!-- Keep a page out of Search (internal admin, thin or temporary pages) -->
<meta name="robots" content="noindex">
<!-- The same rule for Google only -->
<meta name="googlebot" content="noindex">
<!-- Time-limited page (event, offer, job ad): drop it from results after the end date -->
<meta name="robots" content="unavailable_after: 2026-12-31T23:59:59+01:00">Used by DEV-IDX-01 Use noindex to keep a page out of Search, and leave that URL crawlable, DEV-IDX-04 Allow full snippets and large previews with max-snippet:-1 and max-image-preview:large, DEV-IDX-08 Use unavailable_after on pages with a known end date, DEV-VID-05 Allow video previews with max-video-preview:-1
X-Robots-Tag header for PDFs, images and other non-HTML files
Files that cannot carry a meta tag get their robots rules as an HTTP response header. The same rules as the meta tag apply (noindex, nosnippet, max-snippet...), and the URL must stay crawlable for Google to see the header.
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex# Inside the server block: office documents out of the index.
# An add_header in a location replaces every add_header set at server level for those responses:
# repeat your HSTS and CSP lines in each such location (or use the headers-more module).
location ~* \.(pdf|docx?|xlsx?)$ {
add_header X-Robots-Tag "noindex" always;
add_header Strict-Transport-Security "max-age=31536000" always; # repeated from the server block
}
# Images that must not appear in Google Images (they still display on your pages)
location ^~ /internal-images/ {
add_header X-Robots-Tag "noindex" always;
add_header Strict-Transport-Security "max-age=31536000" always; # repeated from the server block
}<FilesMatch "\.(pdf|docx?|xlsx?)$">
Header set X-Robots-Tag "noindex"
</FilesMatch>Used by DEV-IDX-03 Send robots rules for PDFs, images and other non-HTML files in the X-Robots-Tag header, DEV-IMG-06 Keep images out of Search with robots.txt or X-Robots-Tag, not with CSS tricks
data-nosnippet for one detail instead of the whole snippet
data-nosnippet keeps a specific passage out of search snippets (and so out of the snippet-based inputs of AI features) while the page can still rank for it. It works only on span, div and section, must be in the server HTML (not toggled by JavaScript) and needs valid, closed markup: an unclosed element can hide the rest of the page.
<p>
Call our workshop on <span data-nosnippet>+1 555-0100</span> for a repair quote.
</p>
<section data-nosnippet>
<h2>Member prices</h2>
<p>Log in to see your personal discount.</p>
</section>Used by DEV-IDX-07 Use data-nosnippet on a span, div or section to keep one detail out of snippets
Permanent redirects to one host, one protocol, one hop
Send http:// and the other host name to the preferred HTTPS host with a single 301, and send moved pages straight to their final URL (301 or 308). Keep migration redirects in place long term; avoid chains, meta refresh and JavaScript redirects for permanent moves.
# http:// on both host names -> https://www. in one hop
server {
listen 80;
server_name example.com www.example.com;
return 301 https://www.example.com$request_uri;
}
# https:// on the bare domain -> https://www.
server {
listen 443 ssl;
server_name example.com;
ssl_certificate /etc/ssl/example.com.crt;
ssl_certificate_key /etc/ssl/example.com.key;
return 301 https://www.example.com$request_uri;
}
server {
listen 443 ssl;
http2 on; # nginx 1.25.1+; older versions use "listen 443 ssl http2;"
server_name www.example.com;
ssl_certificate /etc/ssl/www.example.com.crt;
ssl_certificate_key /etc/ssl/www.example.com.key;
# Moved pages: permanent and straight to the final URL
location = /old-chairs/ {
return 301 https://www.example.com/chairs/;
}
location ^~ /shop/chairs/ {
rewrite ^/shop/chairs/(.*)$ https://www.example.com/chairs/$1 permanent;
}
root /var/www/example;
}Used by DEV-CAN-01 Use permanent server-side redirects (301 or 308) for moved URLs, straight to the final URL, DEV-CAN-02 Serve one host over HTTPS with a valid certificate and redirect every other origin to it
One rel=canonical per page, pointing straight at the final URL
Exactly one absolute canonical in the server-sent <head>, self-referencing on the preferred URL and identical on its duplicates (tracking parameters, sort orders, print views). The target must answer 200, be indexable, not redirect and be in the same language. Non-HTML files can declare it in an HTTP Link header.
<!-- On https://www.example.com/chairs/oak-dining-chair
and on https://www.example.com/chairs/oak-dining-chair?utm_source=newsletter -->
<head>
<link rel="canonical" href="https://www.example.com/chairs/oak-dining-chair">
</head>HTTP/1.1 200 OK
Content-Type: application/pdf
Link: <https://www.example.com/guides/chair-care.pdf>; rel="canonical"Used by DEV-CAN-03 Put exactly one absolute rel=canonical in the server-rendered head of every indexable page
Check every hop of a redirect chain against robots.txt
Robots.txt is checked for every URL in a redirect chain, and crawling stops at the first blocked hop, while Search Console shows the block on the first URL. This script follows server-side redirects one hop at a time, with Googlebot's user agent and with a browser's, and tests each hop against its own host's robots.txt. It needs pip install requests protego (Protego is an open-source robots.txt parser with Google-style wildcard and longest-match rules; confirm a doubtful result with Google's own parser, github.com/google/robotstxt). Meta refresh and JavaScript redirects are not followed: check those with URL Inspection's live test.
# check_chain.py - usage: python check_chain.py https://www.example.com/page [more URLs]
import sys
from urllib.parse import urljoin, urlsplit
import requests
from protego import Protego
AGENTS = {
"googlebot": "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)",
"browser": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0 Safari/537.36",
}
TOKEN = "Googlebot" # the robots.txt token the rules are tested for
_robots = {}
def robots_for(url):
parts = urlsplit(url)
origin = f"{parts.scheme}://{parts.netloc}"
if origin not in _robots:
r = requests.get(origin + "/robots.txt", headers={"User-Agent": AGENTS["googlebot"]}, timeout=10)
if r.status_code == 200:
text = r.text
elif 400 <= r.status_code < 500 and r.status_code != 429:
text = "" # 4xx (except 429): Google crawls as if there were no robots.txt
else:
text = "User-agent: *\nDisallow: /" # 5xx or 429: Google stops crawling the host for now
_robots[origin] = Protego.parse(text)
return _robots[origin]
def check(url, agent, max_hops=10):
print(f"[{agent}]")
for hop in range(max_hops + 1):
allowed = robots_for(url).can_fetch(url, TOKEN)
print(f" {hop}: {'allowed' if allowed else 'BLOCKED by robots.txt'} {url}")
if not allowed:
print(" crawling stops here; Search Console reports the first URL of the chain as blocked")
return
r = requests.get(url, headers={"User-Agent": AGENTS[agent]}, allow_redirects=False, timeout=10)
location = r.headers.get("Location")
if r.status_code in (301, 302, 303, 307, 308) and location:
url = urljoin(url, location)
else:
print(f" final status {r.status_code}")
return
print(f" more than {max_hops} redirect hops")
if __name__ == "__main__":
for start in sys.argv[1:]:
for agent in AGENTS: # different chains for the two user agents point to cloaking
check(start, agent)Used by DEV-CAN-11 Never route a crawlable URL's redirects through a URL that robots.txt blocks
Real 404 status codes in a single-page app
Fix soft 404s at the source: the server (or hosting rewrite rules) answers known paths with 200 and the app shell, and everything else with 404 and a static error page that loads no router, so it can never redirect again. Where that is impossible, use one of the two client-side fallbacks Google documents: a JavaScript redirect to a URL that returns 404, or a noindex added with JavaScript.
// Server (Express 4 or 5): static files first, then a final handler for every other path
const express = require('express');
const path = require('path');
const app = express();
const shell = path.join(__dirname, 'dist', 'index.html');
const notFound = path.join(__dirname, 'dist', '404.html'); // static error page: no router script
const PAGES = [/^\/$/, /^\/chairs\/$/, /^\/tables\/$/, /^\/about\/$/];
const PRODUCT = /^\/chairs\/([a-z0-9-]+)$/;
app.use(express.static(path.join(__dirname, 'dist'), { index: false }));
app.use(async (req, res) => {
const product = req.path.match(PRODUCT);
const found = PAGES.some((re) => re.test(req.path)) || (product !== null && (await productExists(product[1])));
if (found) res.sendFile(shell);
else res.status(404).sendFile(notFound); // also answers /not-found, the target of the client fallback
});
async function productExists(slug) {
return ['oak-dining-chair', 'beech-stool'].includes(slug); // replace with your database or API lookup
}
app.listen(3000);// Client fallback when the server cannot know: never leave a "not found" view on a 200 URL
fetch(`/api/products/${productId}`)
.then((response) => response.json())
.then((product) => {
if (product.exists) {
showProductDetails(product);
return;
}
// Option 1: go to a URL whose server response is 404 (and whose page does not redirect again)
window.location.replace('/not-found');
// Option 2 (instead of option 1): keep the URL but keep the view out of the index
// const robots = document.createElement('meta');
// robots.name = 'robots';
// robots.content = 'noindex';
// document.head.appendChild(robots);
});Used by DEV-ERR-02 Return a real 404 for unknown routes in single-page apps, or use a documented client-side fallback
hreflang link elements: complete, reciprocal, self-referencing
Every language version of a page carries the identical block in its <head>, including a line for itself and an x-default for the fallback, ideally a language selector page. URLs are fully qualified; codes are an ISO 639-1 language, an optional ISO 15924 script (zh-Hant, zh-Hans-CN) and an optional ISO 3166-1 Alpha-2 region.
<!-- The same block in the <head> of every language version listed here; the x-default target is a language selector page -->
<link rel="alternate" hreflang="en" href="https://www.example.com/en/chairs/">
<link rel="alternate" hreflang="en-GB" href="https://www.example.com/en-gb/chairs/">
<link rel="alternate" hreflang="de-DE" href="https://www.example.com/de-de/stuehle/">
<link rel="alternate" hreflang="de-AT" href="https://www.example.com/de-at/stuehle/">
<link rel="alternate" hreflang="de-CH" href="https://www.example.com/de-ch/stuehle/">
<link rel="alternate" hreflang="sv-SE" href="https://www.example.com/sv-se/stolar/">
<link rel="alternate" hreflang="zh-Hant" href="https://www.example.com/zh-hant/chairs/">
<link rel="alternate" hreflang="x-default" href="https://www.example.com/chairs/choose-language/">
<!-- Wrong values seen in the wild: "se", "dk", "cz" (country codes, not languages: use sv, da, cs),
"en-UK" (use en-GB), "de-SW" (use de-CH), "en-EU" (no such region), "GB" (a region alone).
A script subtag is valid: "zh-Hant", "zh-Hans-CN" (language, script, region in that order) -->Used by DEV-INT-03 Make hreflang complete: a self-reference, return links and fully qualified URLs, DEV-INT-04 Use valid hreflang codes: an ISO 639-1 language, an optional ISO 15924 script and an optional ISO 3166-1 Alpha-2 region
hreflang in an XML sitemap
The sitemap method is equivalent to link elements and suits large sites or templates that cannot change the <head>. Use it instead of, not on top of, the HTML or HTTP-header method; every <url> lists all versions, itself included.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
xmlns:xhtml="http://www.w3.org/1999/xhtml">
<url>
<loc>https://www.example.com/en/chairs/</loc>
<xhtml:link rel="alternate" hreflang="en" href="https://www.example.com/en/chairs/"/>
<xhtml:link rel="alternate" hreflang="de-DE" href="https://www.example.com/de-de/stuehle/"/>
<xhtml:link rel="alternate" hreflang="x-default" href="https://www.example.com/"/>
</url>
<url>
<loc>https://www.example.com/de-de/stuehle/</loc>
<xhtml:link rel="alternate" hreflang="en" href="https://www.example.com/en/chairs/"/>
<xhtml:link rel="alternate" hreflang="de-DE" href="https://www.example.com/de-de/stuehle/"/>
<xhtml:link rel="alternate" hreflang="x-default" href="https://www.example.com/"/>
</url>
</urlset>Used by DEV-INT-05 Supply hreflang through one method only: HTML head, HTTP Link header or sitemap
Product markup with a time-limited sale price
Describe only the product the page is about (not the carousel of related products) and only what the page shows. For a sale, price is the sale price, the regular price is a StrikethroughPrice, and validFrom with priceValidUntil (or validThrough) bound the sale in ISO 8601 with a time zone, so a sale price that lingers in cached markup is not treated as current. The template must still switch to the regular price (and drop the StrikethroughPrice) when the sale ends: Google warns that a listing may not display if priceValidUntil is in the past. Keep the dates aligned with the Merchant Center feed. Shipping and returns point by @id alone to the policies defined once in the organisation block.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Product",
"@id": "https://www.example.com/chairs/oak-dining-chair#product",
"name": "Oak dining chair with woven seat",
"description": "Solid oak dining chair with a hand-woven paper-cord seat.",
"image": [
"https://www.example.com/img/oak-chair-1200.webp",
"https://www.example.com/img/oak-chair-side-1200.webp"
],
"sku": "CH-OAK-01",
"brand": {
"@type": "Brand",
"name": "Example Shop"
},
"offers": {
"@type": "Offer",
"url": "https://www.example.com/chairs/oak-dining-chair",
"price": 149.00,
"priceCurrency": "EUR",
"availability": "https://schema.org/InStock",
"itemCondition": "https://schema.org/NewCondition",
"validFrom": "2026-11-27T00:00:00+01:00",
"priceValidUntil": "2026-11-30T23:59:59+01:00",
"priceSpecification": {
"@type": "UnitPriceSpecification",
"priceType": "https://schema.org/StrikethroughPrice",
"price": 199.00,
"priceCurrency": "EUR"
},
"shippingDetails": {
"@type": "OfferShippingDetails",
"hasShippingService": {
"@id": "https://www.example.com/#standard-shipping"
}
},
"hasMerchantReturnPolicy": {
"@id": "https://www.example.com/#returns"
}
}
}
</script>Used by DEV-SDA-01 Write structured data as JSON-LD by default, DEV-SDA-05 Write dates in ISO 8601 and give every date-time a UTC offset, DEV-SDA-08 Mark up merchant listings completely, with validity dates for sale prices, DEV-SHP-01 Use Product markup for rich results and a Merchant Center feed for the Shopping tab, with identical values
Organization-level identity, returns, shipping and loyalty markup
One block, usually on the homepage, identifies the business by its homepage url and a stable @id, and states the policies that apply to most products: a return policy, a shipping service and a loyalty program. Product pages only override what differs.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "OnlineStore",
"@id": "https://www.example.com/#organization",
"name": "Example Shop",
"url": "https://www.example.com/",
"logo": "https://www.example.com/img/logo-512.png",
"sameAs": [
"https://video.example.net/@exampleshop",
"https://social.example.org/exampleshop"
],
"hasMerchantReturnPolicy": {
"@type": "MerchantReturnPolicy",
"@id": "https://www.example.com/#returns",
"applicableCountry": ["ES", "FR"],
"returnPolicyCountry": "ES",
"returnPolicyCategory": "https://schema.org/MerchantReturnFiniteReturnWindow",
"merchantReturnDays": 30,
"returnMethod": "https://schema.org/ReturnByMail",
"returnFees": "https://schema.org/FreeReturn",
"refundType": "https://schema.org/FullRefund"
},
"hasShippingService": {
"@type": "ShippingService",
"@id": "https://www.example.com/#standard-shipping",
"name": "Standard shipping to Spain and France",
"fulfillmentType": "FulfillmentTypeDelivery",
"handlingTime": {
"@type": "ServicePeriod",
"cutoffTime": "14:00:00+01:00",
"duration": {
"@type": "QuantitativeValue",
"minValue": 0,
"maxValue": 1,
"unitCode": "DAY"
},
"businessDays": ["Monday", "Tuesday", "Wednesday", "Thursday", "Friday"]
},
"shippingConditions": [
{
"@type": "ShippingConditions",
"shippingDestination": [
{ "@type": "DefinedRegion", "addressCountry": "ES" },
{ "@type": "DefinedRegion", "addressCountry": "FR" }
],
"orderValue": {
"@type": "MonetaryAmount",
"minValue": 0,
"maxValue": 99.99,
"currency": "EUR"
},
"shippingRate": {
"@type": "MonetaryAmount",
"value": 4.95,
"currency": "EUR"
},
"transitTime": {
"@type": "ServicePeriod",
"duration": {
"@type": "QuantitativeValue",
"minValue": 2,
"maxValue": 4,
"unitCode": "DAY"
}
}
},
{
"@type": "ShippingConditions",
"shippingDestination": [
{ "@type": "DefinedRegion", "addressCountry": "ES" },
{ "@type": "DefinedRegion", "addressCountry": "FR" }
],
"orderValue": {
"@type": "MonetaryAmount",
"minValue": 100,
"currency": "EUR"
},
"shippingRate": {
"@type": "MonetaryAmount",
"value": 0,
"currency": "EUR"
},
"transitTime": {
"@type": "ServicePeriod",
"duration": {
"@type": "QuantitativeValue",
"minValue": 2,
"maxValue": 4,
"unitCode": "DAY"
}
}
}
]
},
"hasMemberProgram": {
"@type": "MemberProgram",
"name": "Example Club",
"description": "Free membership: earn points on every order.",
"url": "https://www.example.com/club/",
"hasTiers": [
{
"@type": "MemberProgramTier",
"@id": "https://www.example.com/club/#member",
"name": "Member",
"hasTierBenefit": ["https://schema.org/TierBenefitLoyaltyPoints"],
"membershipPointsEarned": 5
}
]
}
}
</script>Used by DEV-SDA-04 Identify entities with unique identifiers: url, @id and stable IDs, DEV-SDA-09 Declare return, shipping and loyalty policies once at organisation level
Article markup with images, dates and authors
Render one block per article on the server from the article's own fields: the headline, the visible lead image in 1x1, 4x3 and 16x9 versions, both dates with a UTC offset, and each author as a Person with a url to a real author page. Google does not require it for Top Stories but highly recommends it for all articles.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "NewsArticle",
"headline": "How to care for an oak dining table",
"image": [
"https://www.example.com/img/oak-care-1x1.webp",
"https://www.example.com/img/oak-care-4x3.webp",
"https://www.example.com/img/oak-care-16x9.webp"
],
"datePublished": "2026-10-02T08:00:00+02:00",
"dateModified": "2026-10-02T09:20:00+02:00",
"author": [
{
"@type": "Person",
"name": "Jane Doe",
"url": "https://www.example.com/authors/jane-doe/"
},
{
"@type": "Person",
"name": "John Roe",
"url": "https://www.example.com/authors/john-roe/"
}
]
}
</script>For a paywalled article, add isAccessibleForFree and a hasPart element whose cssSelector names the section behind the paywall (the class must match the page's HTML).
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "NewsArticle",
"headline": "How to care for an oak dining table",
"datePublished": "2026-10-02T08:00:00+02:00",
"isAccessibleForFree": false,
"hasPart": {
"@type": "WebPageElement",
"isAccessibleForFree": false,
"cssSelector": ".paywall"
}
}
</script>Used by DEV-SDA-13 Add Article markup with headline, images, dates and authors to every article template
WebSite markup for the preferred site name
Put one WebSite block on the home page of each domain or subdomain (not a subdirectory) that should have its own site name in Google's results. url is the canonical home page; alternateName is an optional fallback such as an acronym.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "WebSite",
"name": "Example Shop",
"alternateName": "ExShop",
"url": "https://www.example.com/"
}
</script>Used by DEV-SDA-15 State the preferred site name with WebSite markup on the home page
robots.txt that keeps cart and checkout open for Storebot-Google
A crawler follows only the most specific group that names it, so Storebot-Google, Google Shopping's crawler, uses its own group and ignores *. Repeat the general rules there but leave cart and checkout open, so it can verify prices, shipping and availability; every other crawler still stays out of them.
# https://www.example.com/robots.txt
User-agent: *
Disallow: /search?
Disallow: /cart/
Disallow: /checkout/
Disallow: /api/
Allow: /api/products/
# Google Shopping's crawler: same rules, but cart and checkout stay crawlable
User-agent: Storebot-Google
Disallow: /search?
Disallow: /api/
Allow: /api/products/
Sitemap: https://www.example.com/sitemap.xmlUsed by DEV-SHP-02 Let Storebot-Google crawl product, cart and checkout pages, through robots.txt and bot protection
Merchant Center feed item with Q&A, product details and highlights
An XML (RSS 2.0) feed item that adds the attributes Google stresses for AI shopping experiences to the core data: product_highlight lines, product_detail specifications as section, name and value, and question_and_answer pairs (no prices, shipping, dates or company name in them). Keep price, availability, brand and GTIN identical to the product page's markup.
<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:g="http://base.google.com/ns/1.0">
<channel>
<title>Example Shop</title>
<link>https://www.example.com/</link>
<description>Example Shop product feed</description>
<item>
<g:id>KT-17-STEEL</g:id>
<g:title>Steel kettle 1.7 l with temperature control</g:title>
<g:description>Stainless steel kettle with five temperature settings and a keep-warm mode.</g:description>
<g:link>https://www.example.com/kettles/steel-kettle-17</g:link>
<g:image_link>https://www.example.com/img/steel-kettle-17-1200.jpg</g:image_link>
<g:price>59.00 EUR</g:price>
<g:availability>in_stock</g:availability>
<g:condition>new</g:condition>
<g:brand>Example Home</g:brand>
<g:gtin>4006381333931</g:gtin>
<g:product_highlight>Five temperature settings from 70 to 100 degrees Celsius</g:product_highlight>
<g:product_highlight>Keeps water warm for up to 30 minutes</g:product_highlight>
<g:product_detail>
<g:section_name>General</g:section_name>
<g:attribute_name>Capacity</g:attribute_name>
<g:attribute_value>1.7 l</g:attribute_value>
</g:product_detail>
<g:product_detail>
<g:section_name>Power</g:section_name>
<g:attribute_name>Wattage</g:attribute_name>
<g:attribute_value>2200 W</g:attribute_value>
</g:product_detail>
<g:question_and_answer>
<g:question>Does the kettle switch off when it boils dry?</g:question>
<g:answer>Yes. Boil-dry protection switches it off when there is no water in it.</g:answer>
</g:question_and_answer>
<g:question_and_answer>
<g:question>Can I set it to 80 degrees for green tea?</g:question>
<g:answer>Yes. The 80 degree setting is one of the five presets.</g:answer>
</g:question_and_answer>
</item>
</channel>
</rss>Used by DEV-SHP-03 Send conversational product data in the Merchant Center feed: Q&A, product details, highlights and related products
Indexable, responsive image with a modern-format fallback
Google extracts images from <img src>; a <picture> element counts only through the <img> inside it, and CSS background images are not extracted. Serve AVIF or WebP through <source> elements, keep a widely supported file in the img src, describe the image in alt and put a caption or explanatory text next to it.
<figure>
<picture>
<source type="image/avif"
srcset="/img/oak-chair-800.avif 800w, /img/oak-chair-1600.avif 1600w"
sizes="(max-width: 800px) 100vw, 800px">
<source type="image/webp"
srcset="/img/oak-chair-800.webp 800w, /img/oak-chair-1600.webp 1600w"
sizes="(max-width: 800px) 100vw, 800px">
<img src="/img/oak-chair-800.jpg"
srcset="/img/oak-chair-800.jpg 800w, /img/oak-chair-1600.jpg 1600w"
sizes="(max-width: 800px) 100vw, 800px"
alt="Oak dining chair with a woven paper-cord seat, seen from the front"
width="800" height="600" loading="lazy">
</picture>
<figcaption>The oak dining chair in a natural oak finish, seat height 45 cm.</figcaption>
</figure>
<!-- Not indexable as an image: keep CSS backgrounds for decoration only -->
<div class="hero" style="background-image: url('/img/oak-chair-hero.jpg')"></div>Used by DEV-IMG-01 Put every image that should be found in an <img> element with a src attribute, DEV-IMG-03 Place images next to text that explains them, DEV-IMG-05 Serve modern, compact image formats with a widely supported file in the img src
Image sitemap
Image sitemaps list, under each page's <loc>, the images that page uses, which helps Google find images that are loaded by JavaScript or are otherwise hard to discover. Only image:image and image:loc are used (caption, title, geo location and license tags are deprecated); up to 1,000 images per page.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
<url>
<loc>https://www.example.com/chairs/oak-dining-chair</loc>
<image:image>
<image:loc>https://www.example.com/img/oak-chair-1200.webp</image:loc>
</image:image>
<image:image>
<image:loc>https://www.example.com/img/oak-chair-side-1200.webp</image:loc>
</image:image>
</url>
</urlset>Used by DEV-IMG-04 List important images in an image sitemap
Large-image preview and lead image for Discover
Allow large image previews and name a large, relevant lead image (at least 1,200 px wide, 16x9, not the logo and not text-heavy) in the head of every article template; the same image belongs in the Article markup's image list.
<head>
<meta name="robots" content="max-snippet:-1, max-image-preview:large">
<meta property="og:image" content="https://www.example.com/img/oak-care-16x9.webp">
<meta property="og:image:width" content="1600">
<meta property="og:image:height" content="900">
<meta property="og:image:alt" content="Hand applying oil to an oak table top">
</head>Used by DEV-IMG-07 Give every article a large, relevant lead image and declare it in og:image or markup
Video watch page: video element plus VideoObject with key moments
Put the video in a <video> (or <iframe>/<embed>/<object>) element that loads without a click, on a page where it is the main content, and describe it with VideoObject markup. Clip parts enable key moments; their url points to the same watch page at that second.
<video controls preload="metadata" width="1280" height="720"
poster="https://www.example.com/video/chair-assembly-1280.jpg">
<source src="https://www.example.com/video/chair-assembly.mp4" type="video/mp4">
</video>
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "VideoObject",
"name": "How to assemble the oak dining chair",
"description": "Step-by-step assembly of the oak dining chair in four minutes, with the tools you need.",
"thumbnailUrl": ["https://www.example.com/video/chair-assembly-1280.jpg"],
"uploadDate": "2026-09-15T09:00:00+02:00",
"duration": "PT4M12S",
"contentUrl": "https://www.example.com/video/chair-assembly.mp4",
"embedUrl": "https://www.example.com/embed/chair-assembly",
"hasPart": [
{
"@type": "Clip",
"name": "Unpacking the parts",
"startOffset": 0,
"endOffset": 35,
"url": "https://www.example.com/videos/chair-assembly?t=0"
},
{
"@type": "Clip",
"name": "Attaching the legs",
"startOffset": 35,
"endOffset": 140,
"url": "https://www.example.com/videos/chair-assembly?t=35"
},
{
"@type": "Clip",
"name": "Fitting the seat",
"startOffset": 140,
"endOffset": 252,
"url": "https://www.example.com/videos/chair-assembly?t=140"
}
]
}
</script>Used by DEV-VID-01 Embed each video in a video, iframe, embed or object element that loads without user action, DEV-VID-03 Describe videos with VideoObject markup and mark key moments with Clip or SeekToAction
Automated check that Back leaves the page
A Playwright test opens each template from another page, interacts with it the way users do (many history hijacks wait for a key press, a scroll or a timer), presses Back once and expects to be on the previous page again. Run it in CI for every template and after adding any ad, engagement or recommendation script.
// back-button.spec.js: npm i -D @playwright/test, then npx playwright test
const { test, expect } = require('@playwright/test');
const START = 'https://www.example.com/'; // stands in for the page the user came from, such as a search results page
const PAGES = [
'https://www.example.com/blog/oak-table-care/',
'https://www.example.com/chairs/oak-dining-chair',
];
for (const url of PAGES) {
test(`one Back press leaves ${url}`, async ({ page }) => {
await page.goto(START);
await page.goto(url);
await page.keyboard.press('End'); // a real user input that also scrolls to the bottom
await page.waitForTimeout(5000); // give delayed scripts time to run
await page.goBack();
await expect(page).toHaveURL(START);
});
}Then search the built bundles for history calls outside the router and review each one.
grep -rnE "history\.(pushState|replaceState)|addEventListener\(.popstate" dist/Used by DEV-SPM-01 Never manipulate browser history so that Back does not return users to where they came from
BigQuery: top queries per page from the Search Console bulk export
Run against the searchdata_url_impression table of the bulk data export (default dataset searchconsole; replace the project name). Anonymized queries have an empty query and is_anonymized_query set, so they are left out of the ranking here but should stay in totals.
-- Top non-anonymized queries per page, web search, last 28 days
SELECT
url,
query,
SUM(clicks) AS clicks,
SUM(impressions) AS impressions,
SAFE_DIVIDE(SUM(clicks), SUM(impressions)) AS ctr,
SAFE_DIVIDE(SUM(sum_position), SUM(impressions)) + 1 AS avg_position
FROM `my-project.searchconsole.searchdata_url_impression`
WHERE search_type = 'WEB'
AND data_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 28 DAY)
AND NOT is_anonymized_query
GROUP BY url, query
ORDER BY clicks DESC
LIMIT 1000;Used by DEV-MON-08 Set up the Search Console bulk data export to BigQuery on large sites
Looking after many sites? This kit can run as an agent on every release.