Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Thing · Directive

Disallow

The robots.txt rule that tells a crawler not to fetch the URLs under a path.

Directive

Open in Reef mapOpen in Graph

Narrative see the topic robots.txt rules

Claims
18
In Google’s docs
4
Said at the event
9
Not in docs
3
Kit items
9

Google’s documentation 4

Documented in

DocsSourceD1-C127

Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

Google Search Central, Search Central blog (10 September 2019) · Day 1 · How Google thinks about crawl budget

Said at the event 9

Slide and stage claims that name it, the ones Google’s documentation does not cover first.

Not in docs 3

StageNot in docsD2-C844

Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.

Day 2 · Welcome to indexing day!

Consistent with docs 6

StageConsistent with docsD1-C516

Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C519

A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C520

In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C531

In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

Day 1 · How Google interprets robots.txt

Press and analysis 5

AnalysisD1-C532

A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

Ibrahim Anjro · Day 1 · How Google interprets robots.txt

AnalysisD2-C023

Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.

Ibrahim Anjro · Day 2 · Welcome to indexing day!

Built on these claims 9

Kit items about Disallow: their own words name it, or several of the claims they rest on do.

Developer requirements 5

Also inglossary terms Faceted navigation, nofollow, User-agent group, Google-Extended

Connected things 16

Relations

Most often named with it

Things named in the same claim, with the number of claims they share.

Topics that feature it