Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
Thing · Directive
Disallow
The robots.txt rule that tells a crawler not to fetch the URLs under a path.
- Claims
- 18
- In Google’s docs
- 4
- Said at the event
- 9
- Not in docs
- 3
- Kit items
- 9
Google’s documentation 4
Documented in
Google's faceted navigation guide prefers prevention: disallow filter URLs in robots.txt and keep crawlable only item pages plus one unfiltered listing page, or use URL fragments, which Google generally does not crawl. rel=canonical and nofollow are weaker, slower options.
Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.
Google Search Central, Search Central blog (10 September 2019) · Day 1 · How Google thinks about crawl budget
Google's documented ways to keep a site's images out of search results are a robots.txt disallow rule (for example for Googlebot-Image) or a noindex X-Robots-Tag HTTP header, with the Removals tool for emergencies.
Google Search Central · Day 2 · Using images to your advantage and Engaging Search users with videos
Said at the event 9
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Not in docs 3
Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.
Consistent with docs 6
Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.
A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.
In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.
To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.
In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.
If an API folder must stay blocked in robots.txt, allow the endpoints that rendering needs, for example Disallow: /api/ together with Allow: /api/products/.
Rebecca Yu · Day 2 · Lightning session D: Rendering and JavaScript
Press and analysis 5
A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.
Ibrahim Anjro · Day 1 · How Google interprets robots.txt
Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.
Ibrahim Anjro · Day 2 · Welcome to indexing day!
A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
Check robots.txt for Disallow rules covering JavaScript bundles, build folders or API endpoints the page calls while rendering, including on separate API or CDN hostnames with their own robots.txt, then confirm in URL Inspection's live test that the rendered HTML contains the main content.
Ibrahim Anjro · Day 2 · What is Google friendly JavaScript
Built on these claims 9
Kit items about Disallow: their own words name it, or several of the claims they rest on do.
Developer requirements 5
Rests on 20 claims, 6 of them naming Disallow; its own words name Disallow
Use noindex to keep a page out of Search, and leave that URL crawlable
Rests on 12 claims, 3 of them naming Disallow; its own words name Disallow
Give a crawler its own robots.txt group only when needed, and make that group complete
Rests on 5 claims, 2 of them naming Disallow
Keep images out of Search with robots.txt or X-Robots-Tag, not with CSS tricks
Rests on 5 claims, 1 of them naming Disallow; its own words name Disallow
Keep faceted navigation out of the crawl with robots.txt rules or URL fragments
Rests on 12 claims, 1 of them naming Disallow; its own words name Disallow
Also inglossary terms Faceted navigation, nofollow, User-agent group, Google-Extended
Connected things 16
Relations
- Part of robots.txt structure, no claim needed
Most often named with it
Things named in the same claim, with the number of claims they share.
- robots.txt 14
- Allow 9
- noindex 4
- CDN 2
- Googlebot 2
- nofollow 2
- Rendering 2
- Sitemaps 2
- 5xx 1
- Faceted navigation 1
- JavaScript 1
- Main content 1
- rel=canonical 1
- URL Inspection 1
- X-Robots-Tag 1