Day 2: Indexing 42
Shown on screen 11
An audience member noted that a robots.txt block does not mean noindex and that explaining this to developers is a constant struggle, and asked why Google does not treat a robots.txt block as a noindex.
From the audienceIn Day 2, 10:15 · Welcome to indexing day!Evidence slide photo, transcript
- Answered by D2-C020 Day 2: Google answered on a Q&A slide that it does not treat a robots.txt disallow as a noindex because some…
- Answered by D2-C021 Day 2: Google's Q&A slide said Google does not index the content of pages disallowed in robots.txt, but users will…
- Answered by D2-C844 Day 2: Google illustrated why it does not treat a robots.txt disallow as noindex with an extremely important site…
- Answered by D2-C845 Day 2: Google said very few URLs disallowed by robots.txt are in its index, compared with the index as a whole (no…
- Answered by D2-C846 Day 2: Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the…
John Mueller opened with a true-or-false quiz slide asking whether robots meta tags can make a page more visible in Search results than having none, and later answered that the statement is true.
“You can use robots meta tags to be more visible in Search results than without robots meta tags.”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
Used byrequirement DEV-IDX-04
The robots meta tag is a piece of HTML code in the head section of a page that gives search engine crawlers page-specific instructions.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
Used byglossary term Robots meta tag
The primary function of the robots meta tag is to control how a page is indexed and shown in search results.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
Used byglossary term Robots meta tag
The robots meta tag is written as <meta name="robots" content="rule1,rule2">, or with the name googlebot instead of robots, with several rules separated by commas.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo
Used byglossary term Robots meta tag
The robots meta rule all is the default value and does nothing: it places no restrictions on indexing the page or following its links, the same as having no robots meta tag, and it is not an instruction that search engines must index the page.
“This is the default value - it does nothing”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
Used byrequirement DEV-IDX-06
The noindex robots rule tells Google not to show the page in search results.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
Used byrequirement DEV-IDX-01glossary term noindex
A slide named internal admin pages, temporary landing pages and thin content as typical use cases for the noindex rule.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo
John Mueller's slide said Search Console has very few indexing-like settings and listed the Search generative AI control as the one control in his talk that is not a meta tag.
“Very few "indexing-like" settings are in Search Console”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
A slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
Used byrequirement DEV-REN-05
- Extended by D2-C853 Day 2: John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding…
Removing a robots restriction such as noindex with JavaScript does not work, a slide said.
“But... it takes more time, and removing restrictions (like "noindex") doesn't work.”
Wording checked against the slide or recording
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence slide photo, transcript
Used byrequirement DEV-REN-05
- Extended by D2-C852 Day 2: When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript…
Said on stage 18
The robots meta element is also extracted when Google processes a page's HTML, and the speaker called it probably one of the most important extracted elements, or one the audience is probably interested in.
Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence transcript
John Mueller said there is currently no RFC that standardises robots meta tags, unlike robots.txt.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
John Mueller encouraged attendees to help write a specification for robots meta tags, saying internet standards groups are made up of ordinary people and need only passion and a willingness to join the discussions.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
- Repeated by D2-C217 Day 2: Natalia Venditto opened by saying that Gary had said earlier at the event that everyone can write a standard…
John Mueller said robots.txt does not control indexing, so robots meta tags are what site owners have to use to control it.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byglossary term robots.txt
- Extends D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page, which John Mueller called a weird and very broad rule that makes the page stand on its own.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-06glossary term nofollow
- Extended by D2-C821 Day 2: John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would…
John Mueller recommends rel=nofollow on individual links instead of the page-level nofollow robots rule, so a site can choose which links are useful and which are not.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-06
John Mueller suspected that if robots meta tags were reinvented today, the page-level nofollow rule would probably not be part of them.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
- Extends D2-C063 Day 2: The page-level nofollow robots rule tells search engines not to pass signals to any of the links on the page…
The robots rule none is equivalent to noindex plus nofollow, so a page that carries it will not show up in Search, provided Google can see the tag.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-06
The noimageindex rule tells Google not to index any of the images on the page, and John Mueller said he could not see why a site would want that.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-09
- Repeated by D2-C934 Day 2: The noimageindex robots meta tag tells Google not to index the images on the page.
The unavailable_after rule lets a page drop out of search results after a set date and time, which suits time-bound pages, though John Mueller said most sites do not use it.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-08glossary term unavailable_after
- Extended by D2-C698 Day 2: Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index…
John Mueller said robots meta rules only work if robots.txt allows Google to fetch the page.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-IDX-01
- Repeats D1-C110 Day 1: A URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google…
When Google finds a noindex rule in a page's HTML, it drops the page without even processing its JavaScript, so a script cannot switch the page back to indexable, John Mueller said.
“we will see the noindex and say, oh, we will get rid of this page; we won't even process the JavaScript”
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-REN-05
- Extends D2-C108 Day 2: Removing a robots restriction such as noindex with JavaScript does not work, a slide said.
John Mueller recommended putting robots meta tags directly in a page's HTML exactly as intended, and adding them with JavaScript only where that is not possible, as in a JavaScript web app.
Speaker John MuellerIn Day 2, 10:30 · Controlling indexingEvidence transcript
Used byrequirement DEV-REN-05
- Extends D2-C106 Day 2: A slide said robots meta rules can be added to a page's HTML with JavaScript, but doing so takes more time.
The canonical leader of a group should be indexable: it should carry no noindex robots directive, return no error status code and not be blocked in robots.txt.
Speaker Tobias SchwarzIn Day 2, 12:05 · Lightning session E: Managing Duplicates and Site MovesEvidence transcript
Used byrequirement DEV-CAN-04
The noimageindex robots meta tag tells Google not to index the images on the page.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-IDX-09glossary term noimageindex
- Repeats D2-C099 Day 2: The noimageindex rule tells Google not to index any of the images on the page, and John Mueller said he could…
Gary Illyes said noimageindex also affects videos on the page, because Google has to index a video's thumbnail, which is an image.
Speaker Gary IllyesIn Day 2, 13:50 · Using images to your advantage and Engaging Search users with videosEvidence transcript
Used byrequirement DEV-IDX-09glossary term noimageindex
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Index selection checks for the robots unavailable_after rule and its date, and drops a page from the index once that date is reached.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-IDX-08glossary term unavailable_after
- Extends D2-C100 Day 2: The unavailable_after rule lets a page drop out of search results after a set date and time, which suits…
What Google's documentation says 7
Google's noindex documentation says the noindex rule works only if the page is not blocked by robots.txt: a crawler that cannot fetch the page never sees the rule, and the page can still appear in search results, for example if other pages link to it.
Publisher Google Search CentralAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-IDX-01glossary term noindex
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Google's page on valid page metadata says that once Google detects an invalid element in the head, it assumes the head has ended and stops reading further elements there; only title, meta, link, script, style, base, noscript and template elements belong in the head.
Publisher Google Search CentralAnnotates Day 2, 10:25 · How is HTML interpreted
Used byrequirements DEV-CAN-03, DEV-HTM-05
Google's robots meta tag specification tells site owners to place the tag in the head section but notes that Google Search does not enforce that placement and also respects robots meta tags in the body of an HTML document.
Publisher Google Search CentralAnnotates Day 2, 10:30 · Controlling indexing
Google's robots meta tag specification says robots meta tags and X-Robots-Tag headers are found only when a URL is crawled, so the rules on a URL disallowed in robots.txt are never seen and are ignored.
Publisher Google Search CentralAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirements DEV-IDX-01, DEV-IDX-03glossary term X-Robots-Tag
- Extends D1-C106 Day 1: The noindex rule consumes crawl budget, because Google must fetch the page to see it.
Google's robots meta tag specification says Googlebot considerably decreases the crawl rate of a URL after the date and time set in its unavailable_after rule.
Publisher Google Search CentralAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirement DEV-IDX-08
Google's meta tags documentation recommends avoiding JavaScript to inject or change meta tags whenever possible and testing the implementation thoroughly when it must be used.
Publisher Google Search CentralAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirement DEV-REN-05
Google's JavaScript SEO basics guide says that when Google encounters a noindex rule it may skip rendering and JavaScript execution, so using JavaScript to change or remove a noindex robots meta tag may not work as expected.
“it may skip rendering and JavaScript execution”
Publisher Google Search CentralAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirement DEV-REN-05
Analysis by the author 6
Explain robots.txt and noindex to developers as two separate controls: robots.txt controls crawling, noindex controls indexing. To keep a page out of Search, let Google crawl it and serve noindex; a robots.txt disallow alone can leave the bare URL in results.
Author Ibrahim AnjroAnnotates Day 2, 10:15 · Welcome to indexing day!
Used byrequirement DEV-IDX-01
Audit robots meta tags set by templates and plug-ins: an explicit all rule does nothing and can go, while a page-level nofollow strips link signals from every link on the page and is better replaced by qualifying only the specific links that need it.
Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirement DEV-IDX-06
Ship robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may skip rendering a page that arrives with noindex, so the page can stay out of the index even if a script removes the tag later.
Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing
Used byrequirement DEV-REN-05
- Extended by D2-C855 Day 2: On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex)…
On stage the rule was absolute (Google won't even process the JavaScript of a page served with noindex), while Google's guide says only that it may skip rendering; either way, a noindex in the served HTML must never be one that JavaScript is expected to lift.
Author Ibrahim AnjroAnnotates Day 2, 10:30 · Controlling indexing
- Extends D2-C109 Day 2: Ship robots rules in the HTML the server sends and never rely on JavaScript to lift a noindex: Google may…
Do not set noimageindex on video watch pages: Gary Illyes said it also stops the video, through its thumbnail, and although the robots meta tag specification mentions only images, the Video indexing report treats a missing or blocked thumbnail as a reason a video is not indexed.
Author Ibrahim AnjroAnnotates Day 2, 13:50 · Using images to your advantage and Engaging Search users with videos
Use the unavailable_after robots rule on pages with a known end date, such as event pages, time-limited offers or job ads, so that index selection drops them automatically when the date passes instead of leaving expired pages in search results.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-IDX-08