Day 1: Crawling 4
Said on stage 4
Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
- Extended by D1-C508 Day 1: Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing…
- Extended by D2-C717 Day 2: Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed…
Google said soft 404s, like some other problem categories, are looked up in Search Console's page indexing report, while other crawl issues are debugged in the Crawl Stats report.
Speaker Gary IllyesIn Day 1, 14:35 · How crawling errors affect SearchEvidence transcript
Used byrequirement DEV-MON-03
- Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…
- Extends D1-C367 Day 1: CDN captcha challenges often return HTTP 200; Googlebot does not solve them and sees only content with a 200…
Dave Smart said Search Console reports such a block only on the first URL of the chain, which is confusing: the URL shown as blocked by robots.txt is not itself disallowed.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11glossary term Redirect chain
When Search Console says a URL is blocked by robots.txt but the file does not block it, check whether the URL redirects and test every URL in the chain.
Speaker Dave SmartIn Day 1, 15:30 · Lightning session B: Robots.txtEvidence transcript
Used byrequirement DEV-CAN-11
Day 2: Indexing 17
Said on stage 11
To use Search Console to debug what Google can see on a site and why, Erin Sparling said the first step is to get access to the site by verifying ownership (the speaker's words were authorized domains).
Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript
Used byrequirement DEV-MON-01
Erin Sparling showed a Search Console result saying a URL would be indexed only under certain conditions, one of which had not been met, and contrasted it with the result for an available URL whose content loads.
Speaker Erin SparlingIn Day 2, 11:15 · What is Google friendly JavaScriptEvidence transcript
Google's duplication talk described three related parts of deduplication: building clusters, localization, and selecting the representative URL, which is the canonicalization site owners see in Search Console.
Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence transcript
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
“The first one is kind of nastier.”
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed
- Extends D1-C097 Day 1: Google's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over…
- Extends D1-C065 Day 1: The scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler.…
Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
'Crawled – currently not indexed' in Search Console is an index selection decision: Google crawled and processed the page but decided not to keep it in the index.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03glossary term Crawled – currently not indexed
If content quality is even across a site, a page reported as 'Crawled – currently not indexed' may be using a different template that keeps Google from understanding where its content is.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-HTM-01
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
“most of the time it is actually a quality issue”
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
The first thing to check for pages reported as 'Crawled – currently not indexed' is whether their quality is on par with the parts of the site that Google does index.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03
- Extends D1-C368 Day 1: Search Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its…
Not-indexed reasons in the Page indexing report include pages excluded by a noindex rule and 'Alternate page with proper canonical tag'.
Speaker GoogleIn Day 2, 15:40 · Deciding what goes in the index?Evidence transcript
Used byrequirement DEV-MON-03
What Google's documentation says 3
Google's URL Inspection help says a valid live test only confirms that Google can access a page for indexing; the page must still meet other conditions to be indexed, such as having no manual action, not being a duplicate and being of high enough quality.
Publisher Google Search Console HelpAnnotates Day 2, 11:15 · What is Google friendly JavaScript
Used byrequirement DEV-MON-02
Google's Page indexing report help says a 'Discovered – currently not indexed' page was found but not crawled yet, typically because Google wanted to crawl it but expected the crawl to overload the site, so it rescheduled the crawl.
Publisher Google Search Console HelpAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-MON-03glossary term Discovered – currently not indexed
Google's Page indexing report help says a 'Crawled – currently not indexed' page was crawled but not indexed, may or may not be indexed in the future, and does not need to be resubmitted for crawling.
Publisher Google Search Console HelpAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-MON-03glossary term Crawled – currently not indexed
Analysis by the author 3
The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Repeatedly submitting 'Discovered – currently not indexed' URLs does not change why they wait, because the status reflects a crawl-scheduling decision; raise the site's demonstrated quality instead, for example by improving or removing weak pages that are already indexed.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
Author Ibrahim AnjroAnnotates Day 2, 15:40 · Deciding what goes in the index?
Used byrequirement DEV-MON-03
- Extends D1-C098 Day 1: On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.