Day 1: Crawling 10
Shown on screen 4
URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-IDX-02fact F-025
- Extended by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-IDX-02fact F-025
- Extended by D2-C022 Day 2: Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a…
- Extended by D2-C069 Day 2: Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a…
- Extended by D2-C697 Day 2: Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing…
The nofollow rule can still consume crawl budget: Google does not crawl through the nofollow link itself, but it still crawls the linked page when it finds that page through other links.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirement DEV-IDX-02
The non-standard crawl-delay rule is not processed by Googlebot, so it does not save any crawl budget.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byrequirements DEV-IDX-02, DEV-SRV-05
Said on stage 1
Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in robots.txt shifts crawl budget to the rest of the site.
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
- Extended by D1-C470 Day 1: Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless…
What Google's documentation says 3
Google supports only the user-agent, allow, disallow and sitemap fields in robots.txt. Other fields such as crawl-delay are not supported.
Publisher GoogleAnnotates Day 1, 15:10 · How Google interprets robots.txt
Used byrequirement DEV-SRV-05glossary term robots.txt
- Extended by D1-C516 Day 1: Robots.txt was made extremely simple by design so that anyone could implement and understand it; it was…
- Extended by D2-C017 Day 2: Google answered on a Q&A slide that a sitemap listed in the robots.txt file can be picked up by any crawler…
- Extended by D2-C842 Day 2: Google said listing the sitemap in robots.txt is fine, as many websites do.
Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed, but the linked pages may still be crawled if Google finds them through sitemaps or other links. To stop Google fetching URLs on your own site, use a robots.txt disallow rule.
Publisher Google Search Central, Search Central blog (10 September 2019)Annotates Day 1, 16:00 · How Google thinks about crawl budget
Used byrequirements DEV-IDX-02, DEV-IDX-06glossary term nofollow
- Extended by D2-C067 Day 2: nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from…
Google advises against using noindex to save crawl budget and against using robots.txt to temporarily reallocate budget. Use robots.txt only for pages you never want crawled, and 404 or 410 for removed pages.
Publisher GoogleAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Used byrequirements DEV-ERR-01, DEV-IDX-02
Analysis by the author 2
A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Used byrequirement DEV-IDX-01
- Extended by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
- Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
- Extended by D2-C849 Day 2: Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can…
- Extended by D2-C058 Day 2: John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use…
- Repeated by D2-C850 Day 2: John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
- Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
Google's crawl budget guide says crawl budget freed by robots.txt blocks is not shifted to other pages unless the site already hits its crawl capacity limit, and advises against robots.txt for temporary reallocation; so block only sections you never want crawled, and expect a shift only on capacity-limited sites.
Author Ibrahim AnjroAnnotates Day 1, 16:35 · Q&A
- Extends D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…
Day 2: Indexing 22
Shown on screen 3
An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
- Answered by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
- Answered by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Answered by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
- Answered by D2-C845 Day 2: Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no…
- Answered by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some extremely important sites disallow their most important pages, by accident or out of ignorance.
“Simply put, it's because of sites that are extremely important and like to disallow their most important pages, either accidentally or out of ignorance.”
Wording checked against the slide or recording
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
- Extended by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will still find the URLs of those pages in Search.
Speaker Gary Illyes, Cherry PrommawinIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D1-C105 Day 1: URLs disallowed through robots.txt do not affect crawl budget, because they are not fetched.
- Extended by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
- Repeated by D2-C851 Day 2: When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John…
- Extended by D2-C933 Day 2: Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt…
Said on stage 12
Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site, such as a national tax authority, that blocks very important PDF files with robots.txt: Google cannot index the PDFs' content but can at least show their URLs in search results.
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
“if a URL is important, then it might get indexed even if it's disallowed by robots.txt. So the URL gets indexed, not the content.”
Speaker not identifiedIn Day 2, 10:15 · Welcome to indexing day!Evidence transcript
Used byrequirement DEV-IDX-01
- Answers D2-C019 Day 2: An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers…
- Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byglossary term robots.txt
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page, which John Mueller called a weird and very broad rule that makes the page stand on its own.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-06glossary term nofollow
- Extended by D2-C821 Day 2: John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would…
John Mueller recommends rel=nofollow on individual links instead of the page-level nofollow robots rule, so a site can choose which links are useful and which are not.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-06
John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would probably not be part of them.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
- Extends D2-C063 Day 2: The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page…
John Mueller said AI crawlers do not really know what to do with nofollow links, because they look at the content rather than building a link graph; he did not say whether he meant Google's AI systems, other AI crawlers or both.
“AI crawlers don't really know what to do with a nofollow link, because they're looking at the content”
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-01
- Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
When robots.txt blocks a page, the page's content is not indexable but Google can still index its URL, John Mueller said.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-01
- Repeats D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
Gary Illyes suggested a div with a CSS background image as a way to keep an image from being picked up by Google.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-IMG-06
If robots.txt disallows the location of an image or video file, Google does not index that image or video.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirements DEV-IMG-06, DEV-VID-06glossary term Googlebot-Image and Googlebot-Video
Unlike a disallowed web page, whose bare URL can still appear in web results, a video blocked by robots.txt is not shown by its URL, because Google sees no good reason to show a video URL alone.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-VID-06
- Extends D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
What Google's documentation says 2
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-IDX-01glossary term noindex
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Google's documented ways to keep a site's images out of search results are a robots.txt disallow rule (for example for Googlebot-Image) or a noindex X-Robots-Tag HTTP header, with the Removals tool for emergencies.
Publisher Google Search CentralAnnotates Day 2, 13:50 · Using images to your advantage and Engaging Search users with videos
Used byrequirements DEV-IDX-03, DEV-IMG-06glossary term X-Robots-Tag
Analysis by the author 5
Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-IDX-01
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
Audit robots meta tags set by templates and plug-ins: an explicit all rule does nothing and can go, while a page-level nofollow strips link signals from every link on the page and is better replaced by qualifying only the specific links that need it.
Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirement DEV-IDX-06
nofollow is a link-graph hint for search engines, not an access control: to keep AI crawlers away from content, use robots.txt rules for the specific crawlers; at Google, Google-Extended covers Gemini training and grounding, and the Search generative AI control covers AI Overviews and AI Mode.
Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing
- Extends D1-C127 Day 1: Google treats nofollow as a hint for crawling, not a block: nofollow links will generally not be followed…
- Extends D1-C482 Day 1: Nearly all mainstream AI crawlers and AI systems use their own user agents, so robots.txt can set a separate…
Hiding an image as a CSS background is a fragile way to keep it out of Google: it only stops extraction from that page, so the same image URL used in an img element elsewhere or listed in a sitemap can still be indexed; the documented robots.txt or noindex X-Robots-Tag methods are the reliable route.
Author Ibrahim AnjroAnnotates Day 2, 13:50 · Using images to your advantage and Engaging Search users with videos
Used byrequirement DEV-IMG-06