Google’s documentation 2
Documented in
Google's Search Central post on CDNs says that if Google cannot recognise an error message served with a 200 status as an error, all pages showing the same message may be dropped from the index as duplicates, and recovery can be slow because Google has little incentive to recrawl duplicates. For bot-verification interstitials it recommends sending crawlers a 503 status.
Search Central blog (24 December 2024) · Day 2 · Handling web duplication
Google's documentation calls a ccTLD a strong signal that a site is meant for a certain country and still lists server location as a possible signal of a site's audience, though not a definitive one because sites use CDNs or are hosted abroad; it does not rank the signals.
Google Search Central · Day 2 · Focusing on Internationalisation and Localisation
Said at the event 9
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Consistent with docs 6
When Search Console reports network errors, the first step Google recommends is to ask the hosting provider, or the CDN, whether they changed anything.
Gary Illyes · Day 1 · How crawling errors affect Search
Network timeouts are usually caused close to the site, very often by a firewall, a CDN, the hosting provider or the DNS provider, and Google has no visibility into them.
Gary Illyes · Day 1 · How crawling errors affect Search
For DNS and network errors caused by a firewall or CDN, Google advises checking whether new firewall rules were set recently and otherwise asking in the CDN's forum, as Google itself cannot see or help with these errors.
Gary Illyes · Day 1 · How crawling errors affect Search
A rise in HTTP errors in crawl reports can come from a CDN throttling crawlers by injecting 429 or 503 responses on the network path between Google's crawler and the site.
Gary Illyes · Day 1 · How crawling errors affect Search
CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200 status, which indexing then classifies as a soft 404, reported as an error in Search Console.
Gary Illyes · Day 1 · How crawling errors affect Search
A growing cause of wrong clustering is bot protection: when a CDN or another protection system shows Googlebot the same 'you look like a bot, solve this puzzle' page on many unrelated URLs or domains, Google can cluster those pages as duplicates.
Day 2 · Handling web duplication
Confirmed by docs 1
Hostload applies per host, not necessarily per site: depending on how a site is configured, its CDN, its app, its subdomains and its main www host may be different hosts.
Day 1 · How Google thinks about crawl budget
Press and analysis 7
Day 1 said a CDN captcha page served with HTTP 200 becomes a soft 404, while Day 2 (D2-C374, D2-C375) said such challenge pages are hard to recognise as errors and can be clustered as duplicates; Google's CDN post describes both outcomes, and in both the real pages drop out of Search.
Ibrahim Anjro · Day 1 · How crawling errors affect Search
Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files. Never place any non-standard directive between user-agent lines; put it after a group's rules and test the file with each search engine's tools, such as Bing's robots.txt tester and Google's open-source robots.txt parser.
Ibrahim Anjro · Day 1 · How Google interprets robots.txt
This does not mean avoiding HTTP/3, which benefits browsers. The rule is never to run a host that only answers over HTTP/3, and to check that CDN or firewall rules written for HTTP/3 traffic do not break HTTP/1.1 and HTTP/2.
Ibrahim Anjro · Day 1 · session not recorded
Set max-snippet:-1 and max-image-preview:large on every indexable template unless licensing requires otherwise, and check that no CMS, plug-in or CDN setting adds lower snippet or image preview limits by default.
Ibrahim Anjro · Day 2 · Controlling indexing
An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
Check robots.txt for Disallow rules covering JavaScript bundles, build folders or API endpoints the page calls while rendering, including on separate API or CDN hostnames with their own robots.txt, then confirm in URL Inspection's live test that the rendered HTML contains the main content.
Ibrahim Anjro · Day 2 · What is Google friendly JavaScript
Check what your CDN or bot protection serves to verified Googlebot and to the AI agents you want to allow; if crawlers must get a challenge page, return it with a 503 status as Google recommends, never as a 200 page across many URLs, which Google may cluster as duplicates.
Ibrahim Anjro · Day 2 · Handling web duplication
Built on these claims 9
Kit items about CDN: their own words name it, or several of the claims they rest on do.
Developer requirements 6
1 more
Also inglossary terms ccTLD, Crawl rate limit (hostload), Soft 404
Connected things 18
Most often named with it
Things named in the same claim, with the number of claims they share.
2 more things
Topics that feature it