Day 1: Crawling 52
Shown on screen 3
The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce robots.txt policies. Fetched data is sent for indexing.
“Ensure we don't break the internet”
Wording checked against the slide or recording
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence slide photo, transcript
- Extended by D2-C024 Day 2: A Google slide titled Data from Crawling showed an example of what the crawler passes on for processing: the…
- Extended by D2-C025 Day 2: In Google's example fetch record, the robots policies that apply to a fetch, shown as the value…
- Extended by D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
- Extended by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
The robots.txt session worked through an example file with three groups: a default group that disallows everything except the bare homepage, a shared bingbot and googlebot group that also allows /politics/ and /*/live/, and a google-extended group that disallows /subscriptions/.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence slide photo, transcript
- Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
- Extended by D1-C521 Day 1: To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the…
- Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-IDX-02fact F-025
- Extended by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
Said on stage 38
A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which fetch a page because a user asked for it, generally ignore robots.txt rules.
Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript
Used byrequirement DEV-AIF-05glossary term User-triggered fetchers
- Repeated by D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…
- Repeated by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as the way site owners opt out of its crawling.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Extended by D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
- Repeated by D2-C848 Day 2: Google said it strongly believes site owners should be able to opt out of crawling and control how their site…
The only Google-owned crawlers that do not obey robots.txt are contractual crawlers, which crawl a site whose owner has agreed that Google may crawl it however it likes.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byglossary term Contractual crawlers (special-case crawlers)
The robots.txt talk pointed to the robots.txt file of thebestfriedchickenever.com to show why comments in robots.txt are useful; the live file is mostly a large ASCII-art drawing written as # comment lines, followed by a single user-agent group (checked 2026-10-04).
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence notes, transcript
Google's robots.txt talk recalled that robots.txt began in 1994, when bots were crashing servers: Martijn Koster proposed a text file in the root of a site with rules for how automated clients may access it.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Google said it already supported robots.txt when it started crawling in 1996 as a Stanford PhD project, and has supported it ever since.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it follows it because it is the right thing to do: anyone who wants to opt out of crawling should be able to.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)
- Extends D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
- Repeated by D2-C050 Day 2: John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes…
The standard's official name is the Robots Exclusion Protocol (REP). Google stressed that it was designed in 1994 only to control which automated clients may access what on a site, and has nothing to do with how the content is used.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byglossary term RFC 9309 (Robots Exclusion Protocol)
- Extends D1-C086 Day 1: Google-Extended is a control token, not a crawler with its own user agent string. It decides whether crawled…
Robots.txt is not a security measure: the file sits at a predictable public address, so anyone can read the paths it disallows. Protect a secret folder with authentication, or do not put it on the internet.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-11glossary term robots.txt
The robots.txt file always sits in the root of the host; it cannot be placed anywhere else.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-04glossary term robots.txt
Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was extended as websites grew more complex, but it still has only three rules: user-agent (a named crawler, or * for every crawler), disallow and allow.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-05
- Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
Everything on a site is implicitly allowed, so an allow rule is only needed to re-open a specific path inside a disallowed one.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-05
Comments in robots.txt start with #, and a line without the # that a parser cannot read is ignored anyway, because the standard requires parsers to skip lines they cannot parse, so it acts like a comment.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-05glossary term RFC 9309 (Robots Exclusion Protocol)
A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several crawlers (the talk's example: Googlebot and Bingbot); give named crawlers their own group only when they need rules the * group should not grant to every crawler.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-10glossary term User-agent group
- Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches any number of characters, and an allow /science/live/ rule re-opens that one path.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript, slide photo
Used byrequirement DEV-SRV-05
- Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
- Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the match, so /$ means only the root path and /cats$ means exactly /cats.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript, slide photo
Used byrequirement DEV-SRV-05
- Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
- Extends D1-C082 Day 1: In robots.txt, * matches zero or more of any character and $ marks the end of the URL.
Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule name such as disallow makes Google ignore that line, and the rest of the file is still used.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-05
- Extended by D1-C532 Day 1: A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts…
Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version history and the errors and successes of each fetch.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirements DEV-SRV-05, DEV-SRV-06glossary term robots.txt report
- Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
- Extended by D3-C667 Day 3: Google's robots.txt guide says its crawlers update their cached copy of a site's robots.txt every 24 hours…
Google said the robots.txt report helps catch CDNs that update a site's robots.txt without the owner's knowledge, and hosts that cloak the robots.txt file, which happens more often than people think.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-06
Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byglossary term robots.txt report
- Extends D1-C140 Day 1: Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain…
In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of /staging/ closed to it was a user-agent: googlebot group with both rules, disallow /staging/ and allow /staging/preview/, rather than an allow in the * group or a Googlebot group with only the allow.
Speaker GoogleIn Day 1, 15:10 · How Google interprets robots.txtEvidence transcript
Used byrequirement DEV-SRV-10
- Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that group's rules apply, so a googlebot group that blocks /dogs/ leaves Googlebot free to crawl the /goats/ and /cows/ that the * group blocks.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-SRV-10glossary term User-agent group
- Extends D1-C080 Day 1: Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most…
Dave Smart said robots.txt is checked for every URL in a redirect chain and crawling stops at the first blocked one; in his example a site redirected through /cart/ with JavaScript to set the local currency and back, and because /cart/ was disallowed the page was reported as blocked.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11glossary term Redirect chain
Dave Smart said Search Console reports such a block only on the first URL of the chain, which is confusing: the URL shown as blocked by robots.txt is not itself disallowed.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11glossary term Redirect chain
Dave Smart said this applies to all redirects, not only JavaScript ones; his examples: a redirect through an external authorisation service that is blocked by its own robots.txt, content that moved through several URLs over the years with one of them later blocked, and unexpected redirects, such as one served only to Googlebot's user agent.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11
When Search Console says a URL is blocked by robots.txt but the file does not block it, check whether the URL redirects and test every URL in the chain.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11
An audience question asked how to make sure bots behave well beyond robots.txt.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C428 Day 1: Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so…
- Answered by D1-C429 Day 1: A Google panelist said crawlers that ignore robots.txt and cause a nuisance are better treated as a scraping…
Mainstream crawlers from Google, other large search engines and AI companies try to follow robots.txt, so implementing robots.txt correctly is the way to stop them doing something specific on a site.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
- Answers D1-C427 Day 1: An audience question asked how to make sure bots behave well beyond robots.txt.
A Google panelist said crawlers that ignore robots.txt and cause a nuisance are better treated as a scraping problem than as a crawling problem.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
- Answers D1-C427 Day 1: An audience question asked how to make sure bots behave well beyond robots.txt.
Google follows robots.txt partly in its own interest: crawling an infinite URL space such as a calendar that robots.txt blocks would waste Google's crawling time as well as the site's resources.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
A Google panelist described research, malware-scanning and privacy crawlers as a separate category of crawler that is generally harmless and useful to the web.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
A Google panelist said they were working on a set of crawler best practices and offering research, malware-scanning and privacy crawlers an exemption from following them, apparently because such crawlers sometimes need to ignore robots.txt or probe URLs that other crawlers would not touch (the reason is a best reading: the audio has 'don't need', which would not explain an exemption).
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
- Extended by D1-C433 Day 1: The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices'…
User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a different category from crawlers and generally do not check robots.txt, because a user asked for the fetch.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05glossary term User-triggered fetchers
- Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
- Extended by D2-C122 Day 2: Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for…
A Google panelist said much of robots.txt is written with search engines in mind: a search engine should never add items to a cart or check out, but an agent acting for a user probably should be able to.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
- Extended by D1-C470 Day 1: Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless…
Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate policy for each; at the event OpenAI was said to have three or four, Anthropic a few and Microsoft some.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
- Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
- Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
When logs show an unknown crawler fetching a lot, search for its user agent online to find the robots.txt token that controls it; the mainstream crawlers that send the most traffic can all be controlled this way.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
- Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
A Google panelist called a default-deny robots.txt, which blocks every crawler and allows only chosen ones, a bad pattern because the site owner does not know what is being blocked.
“Personally, I think that's a bad pattern, because you don't know what you're blocking.”
Wording checked against the slide or recording
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
- Answers D1-C481 Day 1: An audience member asked whether Google plans dedicated user agents for AI, as OpenAI has several, so that…
What Google's documentation says 8
Google's crawler documentation says special-case crawlers serve specific Google products where the crawled site and the product have an agreement about the crawl process, so they may ignore robots.txt rules; AdsBot, for example, ignores the global (*) user agent with the ad publisher's permission.
Publisher GoogleAnnotates Day 1, 14:05 · How crawling works
Used byglossary term Contractual crawlers (special-case crawlers)
If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the cached copy for up to 30 days.
Publisher GoogleAnnotates Day 1, 14:35 · How crawling errors affect Search
Used byrequirement DEV-SRV-04
- Extended by D2-C210 Day 2: An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that…
Rules can be grouped for several crawlers by repeating user-agent lines. A crawler follows only the most specific group that names it.
Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirements DEV-SRV-05, DEV-SRV-10glossary term User-agent group
- Extended by D1-C519 Day 1: A user-agent line with its allow and disallow rules forms a user-agent group, and one group can name several…
- Extended by D1-C527 Day 1: The talk's fix for giving one crawler both an extra rule and the rules of another group: simply add another…
- Extended by D1-C531 Day 1: In the talk's quiz, the answer for letting Googlebot crawl /staging/preview/ while keeping the rest of…
- Extended by D1-C533 Day 1: Dave Smart said robots.txt groups are not additive: if a group matches a crawler's user agent, only that…
When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses the least restrictive one.
“In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.”
Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-SRV-05
- Extended by D1-C520 Day 1: In the talk's worked example, disallow /*/live/ blocks /science/live/ and /sports/live/, because * matches…
- Extended by D2-C208 Day 2: A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/…
In robots.txt, * matches zero or more of any character and $ marks the end of the URL.
Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-SRV-05
- Extended by D1-C521 Day 1: To let unnamed crawlers fetch only the homepage, use user-agent: *, disallow: / and allow: /$; the $ ends the…
Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-SRV-05glossary term robots.txt
- Extended by D1-C516 Day 1: Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was…
- Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
- Extended by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
Google generally caches robots.txt for up to 24 hours and reads only the first 500 KiB of the file.
Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-SRV-05
- Extended by D3-C613 Day 3: Google estimated that a robots.txt update is picked up in about 24 hours, with a minimum of seconds and an…
- Extended by D3-C614 Day 3: Google said its service level objective is to refresh robots.txt every 24 hours, as RFC 9309 calls for…
Search Console's robots.txt report shows the robots.txt files Google found for the top 20 hosts of a Domain property, with their fetch status, warnings and errors; to test whether a specific URL is blocked, Google's help page points to the URL Inspection tool and to Google's open-source robots.txt library.
Publisher Google Search Console HelpAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-SRV-05
- Extended by D1-C528 Day 1: Search Console's robots.txt report shows the robots.txt file as Google last fetched it, with a version…
- Extended by D1-C530 Day 1: Google said the robots.txt report uses the parser Google open-sourced at github.com/google/robotstxt.
- Extended by D3-C615 Day 3: A site owner can submit robots.txt in Search Console to force Google to refresh it sooner.
Analysis by the author 3
Under the example file, Googlebot may crawl /, /politics/eu-vote and /sports/live/, and is blocked from /?utm_source=x, /index.html, /politics (no trailing slash), /live/ and /sports/live-score. An 'allow: /$' rule does not cover the homepage with tracking parameters.
Author Ibrahim AnjroAnnotates Day 1, 15:10 · How Google interprets robots.txt
A typo in a rule name is not always ignored by Google: its open-source robots.txt parser deliberately accepts common misspellings of disallow (such as dissallow, dissalow and disalow) and of user-agent (useragent, user agent), but not of allow. Google's spec page does not mention typos, and other crawlers may be stricter, so spell rule names correctly.
Author Ibrahim AnjroAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-SRV-05
- Extends D1-C525 Day 1: Google called robots.txt extremely forgiving: a typo in a path only blocks the wrong path, a typo in a rule…
The crawler best practices mentioned on stage match the public IETF Internet-Draft 'Crawler best practices' (draft-illyes-aipref-cbcp-00, July 2025), which says self-declared research crawlers, including privacy and malware discovery crawlers, may exempt themselves from any of its practices with a rationale; it is a draft, not Google documentation.
Author Ibrahim AnjroAnnotates Day 1, 16:35 · Q&A
- Extends D1-C432 Day 1: A Google panelist said they were working on a set of crawler best practices and offering research…
Day 2: Indexing 37
Shown on screen 10
An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites include it but that it was missing from an earlier robots.txt slide by a Google speaker, which the question called 'Gary's slide'.
From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
- Answered by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
- Answered by D2-C018 Day 2: Google's Q&A slide said that a sitemap not listed in robots.txt has to be submitted instead, for example in…
- Answered by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
- Answered by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler, not only by Google's.
“If it's included in the robots.txt file, any crawler can pick your sitemaps up”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
Used byrequirement DEV-URL-05
- Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
- Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
- Extends D1-C079 Day 1: The robots.txt session worked through an example file with three groups: a default group that disallows…
- Extended by D2-C843 Day 2: Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a…
An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
- Answered by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
- Answered by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Answered by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
- Answered by D2-C845 Day 2: Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no…
- Answered by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
“Simply put, it's because of sites that are extremely important and like to disallow their most important pages, either accidentally or out of ignorance.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
- Extended by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D1-C105 Day 1: URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
- Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
- Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
- Extended by D2-C933 Day 2: Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt…
Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer cannot fetch a resource, it cannot run it.
“If the renderer can't fetch it, the renderer can't run it.”
Wording checked against the slide or recording
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
Used byrequirement DEV-REN-04
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Repeated by D2-C265 Day 2: Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the…
- Repeated by D2-C296 Day 2: If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that…
Disallowing a script folder such as /static/js/ or a generic /api/ folder in robots.txt can block the endpoints that supply a page's content, leaving blank modules and missing content after rendering.
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
If an API folder must stay blocked in robots.txt, allow the endpoints that rendering needs, for example Disallow: /api/ together with Allow: /api/products/.
“Carve out only what rendering needs.”
Wording checked against the slide or recording
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo, transcript
Used byrequirement DEV-REN-04
robots.txt rules apply per host, so the rules in example.com/robots.txt do not apply at all to an API served from api.example.com.
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence slide photo
Used byrequirement DEV-SRV-04glossary term robots.txt
Blocked resources, one of Google's four common JavaScript indexing problems, means robots.txt disallowing the crawling of critical JavaScript files or API endpoints.
“robots.txt disallowing crawling of critical .js or API endpoints.”
Wording checked against the slide or recording
Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence slide photo, transcript
Used byrequirement DEV-REN-04
- Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
Said on stage 16
Google said listing the sitemap in robots.txt is fine, as many websites do.
“you can include it in robots.txt. No problem whatsoever.”
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-URL-05
- Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
- Extends D1-C084 Day 1: Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as…
Google pointed out the trade-off of listing a sitemap in robots.txt: anyone can then see the sitemap, which a site may not want.
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-URL-05
- Answers D2-C015 Day 2: An audience member asked whether the sitemap link should be included in robots.txt, noting that many websites…
- Extends D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no figure given).
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
Google said it strongly believes site owners should be able to opt out of crawling and control how their site is crawled, which can matter for legal reasons or for crawl budget.
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
- Repeats D1-C329 Day 1: Apart from contractual crawlers, all of Google's automated crawlers obey robots.txt, which Google treats as…
John Mueller noted that robots.txt is defined by an internet standard, RFC 9309, which Google's Gary Illyes co-authored.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
- Repeats D1-C512 Day 1: Robots.txt is now an IETF standard, RFC 9309, supported by basically all major crawlers; Google said it…
John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-01
- Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-01
- Repeats D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
Never block a resource that a page needs for rendering with robots.txt, a community speaker said, calling this common sense.
Speaker Sören BendigIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
For each suspect script or API request in Chrome DevTools, check whether it runs, whether something blocks it and whether robots.txt disallows it.
Speaker Rebecca YuIn Day 2, 10:40 · Lightning session D: Rendering and JavaScriptEvidence transcript
If robots.txt blocks JavaScript that client-side code needs to render, Google cannot render the content that the JavaScript would produce.
Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript
Used byrequirement DEV-REN-04
- Extends D1-C064 Day 1: The crawler has multiple tasks: fetch from the internet, ensure it doesn't break the internet, and enforce…
- Repeats D2-C204 Day 2: Blocking JavaScript or API resources in robots.txt is a classic rendering mistake: if Google's renderer…
The canonical leader of a group should be indexable: it should carry no noindex robots directive, return no error status code and not be blocked in robots.txt.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-CAN-04
Googlebot-Image and Googlebot-Video are the Google crawlers that fetch images and videos, and robots.txt rules addressed to their user agent tokens control how Google indexes a site's images and videos.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-IMG-06glossary term Googlebot-Image and Googlebot-Video
If robots.txt disallows the location of an image or video file, Google does not index that image or video.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirements DEV-IMG-06, DEV-VID-06glossary term Googlebot-Image and Googlebot-Video
Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-VID-06
- Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
What Google's documentation says 4
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-IDX-01glossary term noindex
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.
Publisher Google Search CentralAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirements DEV-IDX-01, DEV-IDX-03glossary term X-Robots-Tag
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Google says user-triggered fetchers, which fetch a URL because a user asked for it in a Google product (for example Gemini Notebook fetching URLs users add as sources, or Google-Agent acting on a user's request), generally ignore robots.txt rules.
Publisher GoogleAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirement DEV-IDX-11
- Repeats D1-C273 Day 1: A community speaker pointed out that, according to Google's documentation, user-triggered fetchers, which…
- Extends D1-C434 Day 1: User-initiated fetchers, such as a translation service fetching a page a user asked to translate, are a…
Google's introduction to robots.txt says unimportant image, script or style files may be blocked, but resources whose absence makes a page harder for Google to understand should not be blocked.
Publisher Google Search CentralAnnotates Day 2, 11:15 · What is Google friendly JavaScript
Used byrequirement DEV-REN-04
Analysis by the author 7
Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-IDX-01
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
When something is missing from a rendered page, check both robots.txt and the Content Security Policy: the URL Inspection live test shows the page resources, the JavaScript console output and a screenshot of the rendered page.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirements DEV-MON-02, DEV-REN-09
A robots.txt carve-out works because Google applies the most specific matching rule, so Allow: /api/products/ beats Disallow: /api/ for product endpoints only; re-test a rendered page after every robots.txt change to script or API paths.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
- Extends D1-C081 Day 1: When matching rules to a URL, Google uses the most specific rule by path length. If rules conflict, it uses…
An API or CDN on its own host needs its own robots.txt check: a blanket Disallow there, or a robots.txt that returns 5xx errors, can stop Google fetching the data a page renders from.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Used byrequirement DEV-SRV-04
- Extends D1-C074 Day 1: If robots.txt returns a 5xx error, Google stops crawling the site for the first 12 hours, then uses the…
Make sure robots.txt does not block the URLs from which the browser loads a fragment's assets (the gateway paths on the host page's origin, or the endpoint's own host if assets load from there), because Google does not render JavaScript from blocked files and robots.txt rules apply per host.
Author Ibrahim AnjroAnnotates Day 2, 10:40 · Lightning session D: Rendering and JavaScript
Check robots.txt for Disallow rules covering JavaScript bundles, build folders or API endpoints the page calls while rendering, including on separate API or CDN hostnames with their own robots.txt, then confirm in URL Inspection's live test that the rendered HTML contains the main content.
Author Ibrahim AnjroAnnotates Day 2, 11:15 · What is Google friendly JavaScript