Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
Thing · Directive
Allow
The robots.txt rule that lets a crawler fetch a path inside a disallowed one; the most specific matching rule wins.
- Claims
- 11
- In Google’s docs
- 1
- Said at the event
- 7
- Not in docs
- 0
- Kit items
- 6
Google’s documentation 1
Documented in
Said at the event 7
Slide and stage claims that name it, the ones Google’s documentation does not cover first.
Consistent with docs 6
Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.
A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.
In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.
To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.
In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.
If an API folder must stay blocked in robots.txt, allow the endpoints that rendering needs, for example Disallow: /api/ together with Allow: /api/products/.
Rebecca Yu · Day 2 · Lightning session D: Rendering and JavaScript
Confirmed by docs 1
Everything on a site is implicitly allowed, so an allow rule is only needed to re-open a specific path inside a disallowed one.
Press and analysis 3
Under the example file, Googlebot may crawl /, /politics/eu-vote and /sports/live/, and is blocked from /?utm_source=x, /index.html, /politics (no trailing slash), /live/ and /sports/live-score. An 'allow: /$' rule does not cover the homepage with tracking parameters.
Ibrahim Anjro · Day 1 · How Google interprets robots.txt
A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.
Ibrahim Anjro · Day 1 · How Google interprets robots.txt
A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.
Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript
Built on these claims 6
Kit items about Allow: their own words name it, or several of the claims they rest on do.
Developer requirements 5
Rests on 20 claims, 6 of them naming Allow; its own words name Allow
Give a crawler its own robots.txt group only when needed, and make that group complete
Rests on 5 claims, 2 of them naming Allow; its own words name Allow
Rests on 13 claims, 0 of them naming Allow; its own words name Allow
Allow full snippets and large previews with max-snippet:-1 and max-image-preview:large
Rests on 8 claims, 0 of them naming Allow; its own words name Allow
Allow video previews with max-video-preview:-1
Rests on 2 claims, 0 of them naming Allow; its own words name Allow
Also inglossary term User-agent group
Connected things 6
Relations
- Part of robots.txt structure, no claim needed
Most often named with it
Things named in the same claim, with the number of claims they share.