Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Crawling

robots.txt rules

A crawler obeys only the most specific group that names it; within the group the longest matching path wins, and ties go to the less restrictive rule. An audio recording of Google's Day 1 robots.txt talk added the history and the basics. Robots.txt began in 1994, when Martijn Koster proposed a text file of access rules because bots were crashing servers; Google has supported it since it started crawling in 1996, and it is now an IETF standard, RFC 9309, the Robots Exclusion Protocol, which Google follows because anyone who wants to opt out of crawling should be able to. Google stressed that the protocol only controls which automated clients may access what, not how the content is used, and that it is not a security measure: the file always sits at the root of the host, where anyone can read the paths it disallows, so a secret folder needs authentication. Everything is implicitly allowed, so an allow rule only re-opens a path inside a disallowed one; * matches any characters and $ ends the match, so user-agent: *, disallow: / and allow: /$ let unnamed crawlers fetch only the homepage. Comments start with #, and the talk pointed to the ASCII-art robots.txt of thebestfriedchickenever.com to show why they are useful. Google called the format extremely forgiving: lines a parser cannot read are skipped, a typo in a path blocks only the wrong path and a misspelt rule name makes Google ignore that line (not in Google's docs). Search Console's robots.txt report shows the file as Google last fetched it, with a version history and the errors and successes of each fetch; Google said the report uses its open-source parser and helps catch CDNs that change robots.txt without the owner knowing and hosts that cloak the file, which happens more often than people think (not in Google's docs). In Lightning session B Dave Smart added that robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one: a page that redirected through a disallowed /cart/ URL to set the local currency was reported as blocked, and Search Console names only the first URL of the chain, which is not itself disallowed; the same happens with an external authorisation service blocked by its own robots.txt, content that moved through several URLs and redirects served only to Googlebot (the per-URL check is not in Google's docs). Google supports only user-agent, allow, disallow and sitemap, and Day 2 confirmed that a sitemap line can be picked up by any crawler, while a sitemap not listed there has to be submitted, for example in Search Console. A disallow is not a noindex: Google does not index a disallowed page's content, but its URL can still appear in results. Google said it will not treat a disallow as noindex because some very important sites block their most important pages by accident (said at the event, not in Google's docs). Day 2 also showed that disallowing JavaScript or API endpoints a page needs for rendering leaves its content missing, and that rules apply per host, so an API or CDN on its own hostname has its own robots.txt. Day 3 added timing and Shopping: Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for and its robots.txt guide confirms, though delays happen, and that requesting a recrawl in Search Console's robots.txt report refreshes it sooner. Rules addressed to Storebot-Google affect all Google Shopping surfaces, such as the Shopping tab. The second recordings added Google's own stance. Gary Illyes said all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling, except contractual crawlers that crawl by agreement; on Day 2 Google repeated that site owners should be able to opt out, for legal reasons or for crawl budget, and in the Day 1 Q&A said it follows robots.txt partly in its own interest, since crawling a blocked infinite URL space would waste its time too. Mainstream crawlers from search engines and AI companies try to follow robots.txt, so correct rules are the control; user-initiated fetchers and agents acting for a user generally do not check it, and a panelist noted that much of robots.txt is written for search engines, which should never add items to a cart, while an agent probably should. A panelist also mentioned crawler best practices in progress that would exempt research, malware-scanning and privacy crawlers, apparently because they sometimes need to ignore robots.txt or probe URLs other crawlers would not touch. Day 2's opening Q&A said listing a sitemap in robots.txt is fine but makes the sitemap public, and explained why a disallow is not a noindex with a national tax authority that blocks very important PDFs: Google can show their URLs but cannot index their content; very few disallowed URLs are in the index, and John Mueller added that robots meta rules only work if robots.txt lets Google fetch the page. For media, rules for Googlebot-Image and Googlebot-Video control image and video indexing, and a video blocked by robots.txt is not shown by its bare URL at all (said at the event, not in Google's docs). Author’s view: the best practices mentioned match the public IETF draft 'Crawler best practices', not Google documentation; and Google's docs name links from elsewhere as the reason a disallowed URL can still be indexed, so keep an important page crawlable with noindex if it must stay out of Search. Author’s view: Google's open-source parser in fact accepts common misspellings of disallow and user-agent, though not of allow, and other crawlers may be stricter, so spell rule names correctly.

What to do

  • Never disallow JavaScript, CSS or API endpoints that rendering needs; if an API folder must stay blocked, allow only what rendering uses (for example Disallow: /api/ with Allow: /api/products/).
  • To keep a page out of Search, allow crawling and use noindex; a disallow alone can leave the bare URL in results.
  • Test each important URL pattern against the group that applies to each crawler, for Google with its open-source robots.txt parser.
  • Check robots.txt on every host a page loads data from, including API and CDN subdomains.
  • List your XML sitemaps in robots.txt so every crawler can find them.
  • Check that the homepage with tracking parameters is not blocked by an 'allow /$' pattern.
  • Plan robots.txt changes a day ahead, and request a refresh in Search Console when a change is urgent.
  • Leave a sitemap out of robots.txt if you do not want it public, and submit it in Search Console instead.
  • Use Googlebot-Image and Googlebot-Video groups, or disallow the file locations, to control which images and videos Google indexes.
  • Protect private paths with authentication, not robots.txt: anyone can read the file and the paths it disallows.
  • Use allow rules only to re-open a path inside a disallowed one; everything else is allowed by default.
  • When Search Console reports a URL as blocked by robots.txt that the file does not block, test every URL in its redirect chain.
  • Check the version history in Search Console's robots.txt report after CDN or hosting changes, to catch rule changes you did not make.
  • Spell rule names correctly: Google ignores or tolerates some typos, but other crawlers may be stricter.

Day 1: Crawling 52

Shown on screen 3

SlideConsistent with docsD1-C064

The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

“Ensure we don't break the internet”

Wording checked against the slide or recording

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence slide photo, transcript

  • Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
  • Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
  • Extended by D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
  • Extended by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
SlideD1-C079

The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence slide photo, transcript

  • Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
  • Extended by D1-C521 Day 1: To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the…
  • Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
SlideConsistent with docsD1-C105

URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.

Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript

Used byrequirement DEV-IDX-02fact F-025

  • Extended by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…

Said on stage 38

StageConfirmed by docsD1-C273

A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-AIF-05glossary term User-triggered fetchers

  • Repeated by D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…
  • Repeated by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
StageConsistent with docsD1-C329

Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

  • Extended by D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
  • Repeated by D2-C848 Day 2: Google said it strongly believes site owners should be able to opt out of crawling and control how their site…
StageConsistent with docsD1-C330

The only Google-owned crawlers that do not obey robots.txt are contractual crawlers, which crawl a site whose owner has agreed that Google may crawl it however it likes.

Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript

Used byglossary term Contractual crawlers (special-case crawlers)

StageNot in docsD1-C089

The robots.txt talk pointed to the robots.txt file of thebestfriedchickenever.com to show why comments in robots.txt are useful; the live file is mostly a large ASCII-art drawing written as # comment lines, followed by a single user-agent group (checked 2026-10-04).

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence notes, transcript

StageConfirmed by docsD1-C510

Google's robots.txt talk recalled that robots.txt began in 1994, when bots were crashing servers: Martijn Koster proposed a text file in the root of a site with rules for how automated clients may access it.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

StageConfirmed by docsD1-C512

Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)

  • Extends D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
  • Repeated by D2-C050 Day 2: John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes…
StageConsistent with docsD1-C513

The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byglossary term RFC 9309 (Robots Exclusion Protocol)

  • Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
StageConsistent with docsD1-C516

Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
StageConfirmed by docsD1-C518

Comments in robots.txt start with #, and a line without the # that a parser cannot read is ignored anyway, because the standard requires parsers to skip lines they cannot parse, so it acts like a comment.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)

StageConsistent with docsD1-C519

A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConsistent with docsD1-C520

In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript, slide photo

Used byrequirement DEV-SRV-05

  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
StageConsistent with docsD1-C521

To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript, slide photo

Used byrequirement DEV-SRV-05

  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extends D1-C082 Day 1: In robots.txt, * matches zero or more of any character and $ marks the end of the URL.
StageNot in docsD1-C525

Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-05

  • Extended by D1-C532 Day 1: A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts…
StageConfirmed by docsD1-C528

Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirements DEV-SRV-05, DEV-SRV-06glossary term robots.txt report

  • Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
  • Extended by D3-C667 Day 3: Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours…
StageConsistent with docsD1-C531

In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript

Used byrequirement DEV-SRV-10

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageConfirmed by docsD1-C533

Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript

Used byrequirement DEV-SRV-10glossary term User-agent group

  • Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
StageNot in docsD1-C534

Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.

Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript

Used byrequirement DEV-CAN-11glossary term Redirect chain

StageConsistent with docsD1-C535

Dave Smart said Search Console reports such a block only on the first URL of the chain, which is confusing: the URL shown as blocked by robots.txt is not itself disallowed.

Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript

Used byrequirement DEV-CAN-11glossary term Redirect chain

StageNot in docsD1-C536

Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.

Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript

Used byrequirement DEV-CAN-11

StageConfirmed by docsD1-C428

Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C427 Day 1: An audience question asked how to make sure bots behave well beyond robots.txt.
StageNot in docsD1-C432

A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

  • Extended by D1-C433 Day 1: The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices'…
StageConfirmed by docsD1-C434

User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05glossary term User-triggered fetchers

  • Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
  • Extended by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
StageConsistent with docsD1-C469

Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-URL-08

  • Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
  • Extended by D1-C470 Day 1: Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless…
StageConsistent with docsD1-C482

Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
  • Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
StageNot in docsD1-C488

When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
StageD1-C489

A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.

“Personally, I think that's a bad pattern, because you don't know what you're blocking.”

Wording checked against the slide or recording

Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript

Used byrequirement DEV-AIF-05

  • Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…

What Google's documentation says 8

DocsSourceD1-C343

Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.

Publisher GoogleAnnotates Day 1, 14:05 · How crawling works

Used byglossary term Contractual crawlers (special-case crawlers)

DocsSourceD1-C074

If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.

Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search

Used byrequirement DEV-SRV-04

  • Extended by D2-C210 Day 2: An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that…
DocsSourceD1-C080

Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirements DEV-SRV-05, DEV-SRV-10glossary term User-agent group

  • Extended by D1-C519 Day 1: A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several…
  • Extended by D1-C527 Day 1: The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another…
  • Extended by D1-C531 Day 1: In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of…
  • Extended by D1-C533 Day 1: Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that…
DocsSourceD1-C081

When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

“In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”

Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-SRV-05

  • Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
  • Extended by D2-C208 Day 2: A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/…
DocsSourceD1-C084

Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-SRV-05glossary term robots.txt

  • Extended by D1-C516 Day 1: Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was…
  • Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
  • Extended by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
DocsSourceD1-C085

Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-SRV-05

  • Extended by D3-C613 Day 3: Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an…
  • Extended by D3-C614 Day 3: Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for…
DocsSourceD1-C140

Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

Publisher Google Search Console HelpAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-SRV-05

  • Extended by D1-C528 Day 1: Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version…
  • Extended by D1-C530 Day 1: Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
  • Extended by D3-C615 Day 3: A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

Analysis by the author 3

AnalysisD1-C532

A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

Author Ibrahim AnjroAnnotates Day 1, 15:10 · How Google interprets robots.txt

Used byrequirement DEV-SRV-05

  • Extends D1-C525 Day 1: Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule…
AnalysisD1-C433

The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.

Author Ibrahim AnjroAnnotates Day 1, 16:35 · Q&A

  • Extends D1-C432 Day 1: A Google panelist said they were working on a set of crawler best practices and offering research…

Day 2: Indexing 37

Shown on screen 10

SlideD2-C015

An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.

From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

  • Answered by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
  • Answered by D2-C018 Day 2: Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in…
  • Answered by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
  • Answered by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
SlideConfirmed by docsD2-C017

Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

“If it's included in the robots.txt file, any crawler can pick your sitemaps up”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
  • Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
  • Extended by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
SlideD2-C019

An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.

From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

  • Answered by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
  • Answered by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
  • Answered by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
  • Answered by D2-C845 Day 2: Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no…
  • Answered by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
SlideNot in docsD2-C020

Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

“Simply put, it's because of sites that are extremely important and like to disallow their most important pages, either accidentally or out of ignorance.”

Wording checked against the slide or recording

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
  • Extended by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
SlideConfirmed by docsD2-C021

Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D1-C105 Day 1: URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
  • Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
  • Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
  • Extended by D2-C933 Day 2: Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt…
SlideConfirmed by docsD2-C204

Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

“If the renderer can't fetch it, the renderer can't run it.”

Wording checked against the slide or recording

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

Used byrequirement DEV-REN-04

  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Repeated by D2-C265 Day 2: Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the…
  • Repeated by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
SlideConsistent with docsD2-C205

Disallowing a script folder such as /static/js/ or a generic /api/ folder in robots.txt can block the endpoints that supply a page's content, leaving blank modules and missing content after rendering.

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

SlideConsistent with docsD2-C207

If an API folder must stay blocked in robots.txt, allow the endpoints that rendering needs, for example Disallow: /api/ together with Allow: /api/products/.

“Carve out only what rendering needs.”

Wording checked against the slide or recording

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript

Used byrequirement DEV-REN-04

SlideConsistent with docsD2-C265

Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the crawling of critical JavaScript files or API endpoints.

“robots.txt disallowing crawling of critical .js or API endpoints.”

Wording checked against the slide or recording

Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence slide photo, transcript

Used byrequirement DEV-REN-04

  • Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…

Said on stage 16

StageConfirmed by docsD2-C842

Google said listing the sitemap in robots.txt is fine, as many websites do.

“you can include it in robots.txt. No problem whatsoever.”

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
StageConsistent with docsD2-C843

Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-URL-05

  • Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
  • Extends D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
StageNot in docsD2-C844

Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
StageConsistent with docsD2-C846

Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

Used byrequirement DEV-IDX-01

  • Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
  • Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
StageConsistent with docsD2-C848

Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.

Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript

  • Repeats D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
StageConfirmed by docsD2-C050

John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript

  • Repeats D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
StageConfirmed by docsD2-C850

John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.

Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript

Used byrequirement DEV-IDX-01

  • Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
StageConfirmed by docsD2-C851

When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript

Used byrequirement DEV-IDX-01

  • Repeats D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
  • Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
StageConfirmed by docsD2-C157

Never block a resource that a page needs for rendering with robots.txt, a community speaker said, calling this common sense.

Speaker Sören BendigIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

StageConsistent with docsD2-C181

For each suspect script or API request in Chrome DevTools, check whether it runs, whether something blocks it and whether robots.txt disallows it.

Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript

StageConfirmed by docsD2-C296

If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript

Used byrequirement DEV-REN-04

  • Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
  • Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
StageConsistent with docsD2-C432

The canonical leader of a group should be indexable: it should carry no noindex robots directive, return no error status code and not be blocked in robots.txt.

Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript

Used byrequirement DEV-CAN-04

StageConfirmed by docsD2-C931

Googlebot-Image and Googlebot-Video are the Google crawlers that fetch images and videos, and robots.txt rules addressed to their user agent tokens control how Google indexes a site's images and videos.

Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript

Used byrequirement DEV-IMG-06glossary term Googlebot-Image and Googlebot-Video

StageConfirmed by docsD2-C932

If robots.txt disallows the location of an image or video file, Google does not index that image or video.

Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript

Used byrequirements DEV-IMG-06, DEV-VID-06glossary term Googlebot-Image and Googlebot-Video

StageNot in docsD2-C933

Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.

Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript

Used byrequirement DEV-VID-06

  • Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…

What Google's documentation says 4

DocsSourceD2-C022

Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.

Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!

Used byrequirement DEV-IDX-01glossary term noindex

  • Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
DocsSourceD2-C069

Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.

Publisher Google Search CentralAnnotates Day 2, 10:30 · Controlling indexing

Used byrequirements DEV-IDX-01, DEV-IDX-03glossary term X-Robots-Tag

  • Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
DocsSourceD2-C122

Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

Publisher GoogleAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-IDX-11

  • Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
  • Extends D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…

Analysis by the author 7

AnalysisD2-C849

Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!

  • Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
AnalysisD2-C208

A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.

Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

  • Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
AnalysisD2-C210

An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.

Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

Used byrequirement DEV-SRV-04

  • Extends D1-C074 Day 1: If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the…
AnalysisD2-C245

Make sure robots.txt does not block the URLs from which the browser loads a fragment's assets (the gateway paths on the host page's origin, or the endpoint's own host if assets load from there), because Google does not render JavaScript from blocked files and robots.txt rules apply per host.

Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript

AnalysisD2-C298

Check robots.txt for Disallow rules covering JavaScript bundles, build folders or API endpoints the page calls while rendering, including on separate API or CDN hostnames with their own robots.txt, then confirm in URL Inspection's live test that the rendered HTML contains the main content.

Author Ibrahim AnjroAnnotates Day 2, 11:15 · What is Google friendly JavaScript

Day 3: Serving: Ranking, Search Console, and Performance 7

Shown on screen 1

SlideConfirmed by docsD3-C613

Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.

Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript

Used byrequirements DEV-MON-10, DEV-SRV-05

  • Extends D1-C085 Day 1: Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

Said on stage 2

StageConfirmed by docsD3-C614

Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C085 Day 1: Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
StageConfirmed by docsD3-C615

A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript

Used byrequirement DEV-SRV-05

  • Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…

What Google's documentation says 2

DocsSourceD3-C389

Google's crawler list says robots.txt rules addressed to the Storebot-Google user agent affect all surfaces of Google Shopping, such as the Shopping tab in Google Search.

Publisher GoogleAnnotates Day 3, 13:40 · Shopping on Search: Beyond the blue links

Used byrequirement DEV-SHP-02glossary term Storebot-Google

DocsSourceD3-C667

Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.

Publisher Google Search Central, Google Search Console HelpAnnotates Day 3, 15:45 · How long does it take to..?

Used byrequirement DEV-SRV-05

  • Extends D1-C528 Day 1: Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version…

Analysis by the author 2

Across days and sessions 50

  1. Analysis D1-C433 Day 1 · Q&A

    The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.

    extends
    Stage D1-C432 Day 1 · Q&A

    A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).

  2. Analysis D1-C470 Day 1 · Q&A

    Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.

    extends
    Stage D1-C469 Day 1 · Q&A

    Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.

  3. Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

    extends
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

  4. Stage D1-C513 Day 1 · How Google interprets robots.txt

    The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.

    extends
    Docs D1-C086 Day 1 · How Google interprets robots.txt

    Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled content may be used to train future Gemini models and to ground Gemini apps and Vertex AI.

  5. Stage D1-C516 Day 1 · How Google interprets robots.txt

    Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  6. Stage D1-C519 Day 1 · How Google interprets robots.txt

    A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  7. Stage D1-C520 Day 1 · How Google interprets robots.txt

    In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.

    extends
    Slide D1-C079 Day 1 · How Google interprets robots.txt

    The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

  8. Stage D1-C520 Day 1 · How Google interprets robots.txt

    In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.

    extends
    Docs D1-C081 Day 1 · How Google interprets robots.txt

    When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

  9. Stage D1-C521 Day 1 · How Google interprets robots.txt

    To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.

    extends
    Slide D1-C079 Day 1 · How Google interprets robots.txt

    The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

  10. Stage D1-C521 Day 1 · How Google interprets robots.txt

    To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.

    extends
    Docs D1-C082 Day 1 · How Google interprets robots.txt

    In robots.txt, * matches zero or more of any character and $ marks the end of the URL.

  11. Stage D1-C527 Day 1 · How Google interprets robots.txt

    The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another group for that crawler. Google's spec combines all groups that name the same user agent into one.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  12. Stage D1-C528 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

    extends
    Docs D1-C140 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

  13. Stage D1-C530 Day 1 · How Google interprets robots.txt

    Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.

    extends
    Docs D1-C140 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

  14. Stage D1-C531 Day 1 · How Google interprets robots.txt

    In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  15. Analysis D1-C532 Day 1 · How Google interprets robots.txt

    A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.

    extends
    Stage D1-C525 Day 1 · How Google interprets robots.txt

    Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.

  16. Stage D1-C533 Day 1 · Lightning session B: Robots.txt

    Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.

    extends
    Docs D1-C080 Day 1 · How Google interprets robots.txt

    Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.

  17. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Slide D1-C079 Day 1 · How Google interprets robots.txt

    The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.

  18. Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  19. Slide D2-C020 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  20. Slide D2-C021 Day 2 · Welcome to indexing day!

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

    extends
    Slide D1-C105 Day 1 · How Google thinks about crawl budget

    URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.

  21. Docs D2-C022 Day 2 · Welcome to indexing day!

    Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  22. Slide D2-C024 Day 2 · How is HTML interpreted

    A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the fetch result, the connect time and the time to first byte in milliseconds, the robots policies that apply to the fetch, and the raw HTTP response with the page's HTML.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  23. Slide D2-C025 Day 2 · How is HTML interpreted

    In Google's example fetch record, the robots policies that apply to a fetch, shown as the value opt-out-google-extended, travel with the fetched page into processing.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  24. Analysis D2-C067 Day 2 · Controlling indexing

    nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.

    extends
    Stage D1-C482 Day 1 · Q&A

    Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.

  25. Docs D2-C069 Day 2 · Controlling indexing

    Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.

    extends
    Slide D1-C106 Day 1 · How Google thinks about crawl budget

    The noindex rule consumes crawl budget, because Google must fetch the page to see it.

  26. Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    extends
    Stage D1-C434 Day 1 · Q&A

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

  27. Slide D2-C204 Day 2 · Lightning session D: Rendering and JavaScript

    Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  28. Analysis D2-C208 Day 2 · Lightning session D: Rendering and JavaScript

    A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.

    extends
    Docs D1-C081 Day 1 · How Google interprets robots.txt

    When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.

  29. Analysis D2-C210 Day 2 · Lightning session D: Rendering and JavaScript

    An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.

    extends
    Docs D1-C074 Day 1 · How crawling errors affect Search

    If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.

  30. Stage D2-C296 Day 2 · What is Google friendly JavaScript

    If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

    extends
    Slide D1-C064 Day 1 · How crawling works

    The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.

  31. Stage D2-C842 Day 2 · Welcome to indexing day!

    Google said listing the sitemap in robots.txt is fine, as many websites do.

    extends
    Docs D1-C084 Day 1 · How Google interprets robots.txt

    Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.

  32. Stage D2-C843 Day 2 · Welcome to indexing day!

    Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.

    extends
    Slide D2-C017 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.

  33. Stage D2-C844 Day 2 · Welcome to indexing day!

    Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.

    extends
    Slide D2-C020 Day 2 · Welcome to indexing day!

    Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.

  34. Stage D2-C846 Day 2 · Welcome to indexing day!

    Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  35. Stage D2-C846 Day 2 · Welcome to indexing day!

    Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.

    extends
    Slide D2-C021 Day 2 · Welcome to indexing day!

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

  36. Analysis D2-C849 Day 2 · Welcome to indexing day!

    Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.

    extends
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  37. Stage D2-C933 Day 2 · Using images to your advantage and Engaging Search users with videos

    Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.

    extends
    Slide D2-C021 Day 2 · Welcome to indexing day!

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

  38. Slide D3-C613 Day 3 · How long does it take to..?

    Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an end point of 25 hours on the slide.

    extends
    Docs D1-C085 Day 1 · How Google interprets robots.txt

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

  39. Stage D3-C614 Day 3 · How long does it take to..?

    Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for, though delays happen now and then.

    extends
    Docs D1-C085 Day 1 · How Google interprets robots.txt

    Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.

  40. Stage D3-C615 Day 3 · How long does it take to..?

    A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.

    extends
    Docs D1-C140 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.

  41. Docs D3-C667 Day 3 · How long does it take to..?

    Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours, and that the Request a recrawl option in Search Console's robots.txt report refreshes it faster.

    extends
    Stage D1-C528 Day 1 · How Google interprets robots.txt

    Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.

  42. Stage D1-C434 Day 1 · Q&A

    User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.

    repeats
    Stage D1-C273 Day 1 · Lightning session A: Automation and AI

    A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

  43. Stage D2-C050 Day 2 · Controlling indexing

    John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.

    repeats
    Stage D1-C512 Day 1 · How Google interprets robots.txt

    Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.

  44. Docs D2-C122 Day 2 · Lightning session D: Rendering and JavaScript

    Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.

    repeats
    Stage D1-C273 Day 1 · Lightning session A: Automation and AI

    A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.

  45. Slide D2-C265 Day 2 · What is Google friendly JavaScript

    Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the crawling of critical JavaScript files or API endpoints.

    repeats
    Slide D2-C204 Day 2 · Lightning session D: Rendering and JavaScript

    Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

  46. Stage D2-C296 Day 2 · What is Google friendly JavaScript

    If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.

    repeats
    Slide D2-C204 Day 2 · Lightning session D: Rendering and JavaScript

    Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.

  47. Stage D2-C848 Day 2 · Welcome to indexing day!

    Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.

    repeats
    Stage D1-C329 Day 1 · How crawling works

    Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.

  48. Stage D2-C850 Day 2 · Controlling indexing

    John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.

    repeats
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  49. Stage D2-C851 Day 2 · Controlling indexing

    When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

    repeats
    Analysis D1-C110 Day 1 · How Google thinks about crawl budget

    A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.

  50. Stage D2-C851 Day 2 · Controlling indexing

    When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.

    repeats
    Slide D2-C021 Day 2 · Welcome to indexing day!

    Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.

Built on these claims 22

Developer requirements 21

Facts 1

Sources 27