Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Thing · Standard

robots.txt

The text file at the root of a host that tells crawlers which paths they may fetch, standardised by the IETF in 2022 as the Robots Exclusion Protocol (RFC 9309).

StandardAlso calledRobots Exclusion Protocol, RFC 9309, REP

Open in Reef mapOpen in Graph

Narrative see the topics robots.txt rules, robots.txt group parsing differences

Claims
113
In Google’s docs
17
Said at the event
78
Not in docs
16
Kit items
36

Glossary · robots.txt

A file in the root of each host (example.com/robots.txt) that tells crawlers which URLs they may fetch. It controls crawling, not indexing, applies only to its own host and is not a security measure: anyone can read it.

Google’s documentation 17

Documented in

DocsSourceD1-C343

Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.

Google · Day 1 · How crawling works

DocsSourceD1-C127

Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.

Google Search Central, Search Central blog (10 September 2019) · Day 1 · How Google thinks about crawl budget

DocsSourceD2-C022

Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.

Google Search Central · Day 2 · Welcome to indexing day!

DocsSourceD2-C297

Google's introduction to robots.txt says unimportant image, script or style files may be blocked, but resources whose absence makes a page harder for Google to understand should not be blocked.

Google Search Central · Day 2 · What is Google friendly JavaScript

Said at the event 78

Slide and stage claims that name it, the ones Google’s documentation does not cover first.

Not in docs 16

SlideNot in docsD1-C123

Bing closes the robots.txt group at an unknown line placed between two user-agent lines, so in the slide's example bingbot ends up not blocked while Googlebot is; the slide warned that different crawlers behave differently.

Day 1 · How Google interprets robots.txt

StageNot in docsD1-C089

The robots.txt talk pointed to the robots.txt file of thebestfriedchickenever.com to show why comments in robots.txt are useful; the live file is mostly a large ASCII-art drawing written as # comment lines, followed by a single user-agent group (checked 2026-10-04).

Day 1 · How Google interprets robots.txt

StageNot in docsD1-C526

The real trouble is an unknown line placed between two user-agent lines: while writing RFC 9309, Google asked people whether the first crawler should then inherit the rules that follow, and opinions split roughly 50-50 (how the talk said the question was settled is unclear in the recording).

Day 1 · How Google interprets robots.txt

StageNot in docsD1-C534

Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.

Dave Smart · Day 1 · Lightning session B: Robots.txt

StageNot in docsD1-C536

Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.

Dave Smart · Day 1 · Lightning session B: Robots.txt

StageNot in docsD1-C432

A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

Day 1 · Q&A

StageNot in docsD1-C487

A Google panelist said Apple offers a robots.txt control for Applebot similar to Google-Extended, so a site can allow Applebot for search while opting out of Apple's AI use.

Day 1 · Q&A

StageNot in docsD1-C488

When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.

Day 1 · Q&A

StageNot in docsD2-C844

Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.

Day 2 · Welcome to indexing day!

SlideNot in docsD2-C328

Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.

Gary Illyes · Day 2 · Understanding what's on a page

Consistent with docs 25

SlideConsistent with docsD1-C037

For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.

Day 1 · How Search works and where's AI?

SlideConsistent with docsD1-C064

The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

Gary Illyes · Day 1 · How crawling works

StageConsistent with docsD1-C513

The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C514

Robots.txt is not a security measure: the file sits at a predictable public address, so anyone can read the paths it disallows. Protect a secret folder with authentication, or do not put it on the internet.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C516

Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C522

Google said a site can keep a directory or the whole site out of Gemini model training by disallowing it for the Google-Extended token in robots.txt, and that Google respects that policy; the speaker believed Google was the first to offer such an opt-out from model training.

Day 1 · How Google interprets robots.txt

StageConsistent with docsD1-C430

Google follows robots.txt partly in its own interest: crawling an infinite URL space such as a calendar that robots.txt blocks would waste Google's crawling time as well as the site's resources.

Day 1 · Q&A

StageConsistent with docsD1-C435

AI agents acting on a user's request fall into the same category as user-initiated fetchers: an agent told to look at a site reads the page without checking robots.txt.

Day 1 · Q&A

StageConsistent with docsD1-C469

Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

Gary Illyes · Day 1 · Q&A

StageConsistent with docsD1-C482

Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

Day 1 · Q&A

StageConsistent with docsD2-C846

Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

Day 2 · Welcome to indexing day!

Confirmed by docs 26

StageConfirmed by docsD1-C510

Google's robots.txt talk recalled that robots.txt began in 1994, when bots were crashing servers: Martijn Koster proposed a text file in the root of a site with rules for how automated clients may access it.

Day 1 · How Google interprets robots.txt

StageConfirmed by docsD1-C512

Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

Day 1 · How Google interprets robots.txt

StageConfirmed by docsD1-C518

Comments in robots.txt start with #, and a line without the # that a parser cannot read is ignored anyway, because the standard requires parsers to skip lines they cannot parse, so it acts like a comment.

Day 1 · How Google interprets robots.txt

StageConfirmed by docsD1-C533

Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

Dave Smart · Day 1 · Lightning session B: Robots.txt

StageConfirmed by docsD1-C428

Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.

Day 1 · Q&A

StageConfirmed by docsD1-C434

User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

Day 1 · Q&A

SlideConfirmed by docsD2-C021

Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

Day 2 · Welcome to indexing day!

StageConfirmed by docsD2-C050

John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

John Mueller · Day 2 · Controlling indexing

StageConfirmed by docsD2-C851

When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

John Mueller · Day 2 · Controlling indexing

StageConfirmed by docsD3-C614

Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

Gary Illyes · Day 3 · How long does it take to..?

Nothing to verify 11

StageD1-C427

An audience question asked how to make sure bots behave well beyond robots.txt.

From the audience · Day 1 · Q&A

StageD1-C429

A Google panelist said crawlers that ignore robots.txt and cause a nuisance are better treated as a scraping problem than as a crawling problem.

Day 1 · Q&A

StageD1-C436

A Google panelist argued that it makes business sense for agents not to follow robots.txt: an agent sent to buy something that obeyed a robots.txt block could not complete the purchase, a bad experience for the user and lost revenue for the shop.

Day 1 · Q&A

StageD1-C437

A Google panelist said much of robots.txt is written with search engines in mind: a search engine should never add items to a cart or check out, but an agent acting for a user probably should be able to.

Day 1 · Q&A

StageD1-C483

A Google panelist doubted that setting different robots.txt policies per AI crawler makes practical sense yet, because nobody knows how these systems will develop.

Day 1 · Q&A

StageD1-C489

A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.

Day 1 · Q&A

SlideD2-C015

An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.

From the audience · Day 2 · Welcome to indexing day!

SlideD2-C019

An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.

From the audience · Day 2 · Welcome to indexing day!

StageD2-C212

In the speaker's product page example, three of the mistakes combine: product links behind onclick handlers keep the product detail pages hidden, the product API is blocked in robots.txt, and some content waits for a user interaction.

Rebecca Yu · Day 2 · Lightning session D: Rendering and JavaScript

Press and analysis 18

AnalysisD1-C088

Content-Signal lines are now common because a large CDN provider adds them to its managed robots.txt files. Never place any non-standard directive between user-agent lines; put it after a group's rules and test the file with each search engine's tools, such as Bing's robots.txt tester and Google's open-source robots.txt parser.

Ibrahim Anjro · Day 1 · How Google interprets robots.txt

AnalysisD1-C523

The 'first' needs care: OpenAI documented how to block its GPTBot crawler in robots.txt in August 2023, before Google announced Google-Extended on 28 September 2023. What Google-Extended added was a token that controls AI-training use of pages fetched by Google's existing crawlers, separately from Search; Apple later introduced a similar token, Applebot-Extended.

Ibrahim Anjro · Day 1 · How Google interprets robots.txt

AnalysisD1-C532

A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

Ibrahim Anjro · Day 1 · How Google interprets robots.txt

AnalysisD1-C470

Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.

Ibrahim Anjro · Day 1 · Q&A

AnalysisD2-C023

Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.

Ibrahim Anjro · Day 2 · Welcome to indexing day!

AnalysisD2-C849

Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

Ibrahim Anjro · Day 2 · Welcome to indexing day!

AnalysisD2-C245

Make sure robots.txt does not block the URLs from which the browser loads a fragment's assets (the gateway paths on the host page's origin, or the endpoint's own host if assets load from there), because Google does not render JavaScript from blocked files and robots.txt rules apply per host.

Ibrahim Anjro · Day 2 · Lightning session D: Rendering and JavaScript

Built on these claims 36

Kit items about robots.txt: their own words name it, or several of the claims they rest on do.

Developer requirements 20

15 more

Facts 1

Also inglossary terms robots.txt, RFC 9309 (Robots Exclusion Protocol), User-triggered fetchers, Contractual crawlers (special-case crawlers), Googlebot-Image and Googlebot-Video, Redirect chain, X-Robots-Tag, Faceted navigation, Google-Extended, Grounding, nofollow, noindex, robots.txt report, Storebot-Google, User-agent group

Connected things 40

Relations

  • Has part Allow structure, no claim needed
  • Has part Disallow structure, no claim needed

Most often named with it

Things named in the same claim, with the number of claims they share.

22 more things

Topics that feature it