Day 1: Crawling 20
Shown on screen 1
Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence slide photo, transcript
Used byglossary term Crawl demand
- Extended by D1-C377 Day 1: The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual…
- Extended by D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
- Extended by D2-C695 Day 2: Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most…
- Extended by D2-C709 Day 2: Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and…
- Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
Said on stage 15
Google may visit hub pages, such as a homepage or category pages, more often than other pages, because they usually link out to new or updated pages.
Speaker Cherry PrommawinIn Day 1, 11:45 · How Search works and where's AI?Evidence transcript
Used byrequirement DEV-URL-04glossary term Hub pages
- Extends D1-C040 Day 1: URL discovery works through links: a homepage links to section pages, which link to further pages.
- Extended by D2-C820 Day 2: Google's Q&A slide advised large sites to rely on hub pages, such as category pages, that link out to their…
Google's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Extended by D3-C612 Day 3: Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower…
The crawl scheduler logs how often each page changes and crawls frequently changing pages first: a news site's homepage, which changes very often, is prioritised over its terms of service page, which may change once a year.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
Used byglossary term Crawl scheduler
Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site and its content, while Ads wants to check every publisher page that wants to appear in Google Ads, so it schedules those URLs as they come in.
Speaker Gary IllyesIn Day 1, 14:05 · How crawling worksEvidence transcript
- Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual page.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
- Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
Frequent crawling is usually a consequence, not a cause: Google visiting a site more often tends to signal that people are interested in the site and that its content is of high quality.
Speaker Cherry PrommawinIn Day 1, 16:00 · How Google thinks about crawl budgetEvidence transcript
- Answers D1-C390 Day 1: An audience member asked, in a question submitted before the event, whether crawl frequency affects the…
Google's search crawling works to keep content fresh and to understand which pages change frequently, whereas many AI systems crawl a site with no understanding of it and take everything.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-AIF-05
- Answers D1-C422 Day 1: An audience member asked, in a question submitted at registration, what the biggest difference is between…
An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as sports sites, whose content changes constantly and is largely near-duplicate across seasons.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
- Answered by D1-C440 Day 1: If a site's server cannot cope with the extra crawling that higher crawl demand brings, Google reduces its…
Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to users, Google raises crawl demand and tries to raise its crawl capacity for the site, the same logic it applies to small sites.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
- Answers D1-C438 Day 1: An audience member asked how Google prioritises crawling and indexing for very large real-time sites, such as…
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
- Extended by D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
- Extended by D3-C623 Day 3: When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the…
Google crawls new URLs, such as plugin-generated calendar pages, because it wants to see what they are; its automated systems learn that such URLs are useless only from large samples, which takes time.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-URL-08
- Extended by D1-C539 Day 1: A Google panelist said useless plugin-generated parameter URLs can be handled at the web server, for example…
An audience member asked how a large news site can tell whether crawl budget is limiting how fast new articles are discovered (within minutes), and which statistics in Search Console and the logs show this.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
- Answered by D1-C467 Day 1: To check that new content is crawled fast enough, Gary Illyes advised measuring the time between publishing a…
- Answered by D1-C468 Day 1: When new URLs are picked up more slowly, Gary Illyes advised checking the log files for URLs that Google…
- Answered by D1-C469 Day 1: Gary Illyes said disallowing a section that should not be crawled, such as /ads, in a Googlebot group in…
Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively given their constant flow of new content and new URLs.
Speaker Gary IllyesIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-MON-11
- Answers D1-C465 Day 1: An audience member asked how a large news site can tell whether crawl budget is limiting how fast new…
- Repeats D1-C381 Day 1: Not every site needs to worry about crawl budget, and this has always been so: a site with fewer than a few…
- Extended by D3-C606 Day 3: For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading…
An audience member asked, for very large sites of about 100 million pages, what signs show that a site is limited by crawl budget, and how to tell a crawl capacity limit problem from a crawl demand problem.
From the audienceIn Day 1, 16:35 · Q&AEvidence transcript
- Answered by D1-C493 Day 1: There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower…
- Answered by D1-C494 Day 1: To check whether Google lowered a site's crawl capacity limit, look at the number of connections Googlebot…
- Answered by D1-C546 Day 1: A Google panelist illustrated a crawl capacity limit drop as a step down in the site's hostload, for example…
There is no easy way to tell whether a drop in Google's crawling comes from lower crawl demand or a lower crawl capacity limit; most of the time a capacity-limit drop is an abrupt step down.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
- Answers D1-C545 Day 1: An audience member asked, for very large sites of about 100 million pages, what signs show that a site is…
- Extended by D3-C617 Day 3: Google's crawl chart puts a crawl capacity update at seconds when backing off, typically 4 hours or 1-2…
A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search, Google Ads or another Google product, Google raises the site's crawl demand and crawls up to that level of demand as far as the site's crawl capacity allows.
Speaker not identifiedIn Day 1, 16:35 · Q&AEvidence transcript
Used byglossary term Crawl demand
- Extends D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
- Extends D1-C334 Day 1: Different Google teams prioritise crawling differently: web search cares a lot about the quality of a site…
- Repeated by D3-C624 Day 3: Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate…
What Google's documentation says 1
Google's crawl budget guide says each crawler has its own crawl demand, but the crawl capacity limit (hostload) is shared across all crawlers, so high demand from one crawler can reduce the capacity left for others.
“high demand from one crawler can reduce the capacity available for others”
Publisher GoogleAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Used byrequirement DEV-PRF-01glossary term Crawl rate limit (hostload)
Analysis by the author 3
A page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
Author Ibrahim AnjroAnnotates Day 1, 14:05 · How crawling works
Used byrequirements DEV-URL-04, DEV-URL-05
- Extended by D2-C189 Day 2: Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script…
- Extended by D2-C427 Day 2: When internal links point only to page A and the canonical leader is reached only through A's canonical link…
On smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
- Extended by D2-C708 Day 2: The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on…
- Extended by D2-C714 Day 2: 'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages…
- Extended by D2-C716 Day 2: Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the…
The stage remark that crawl demand follows the quality of the site as a whole sits beside the same talk's slide (D1-C094), which falls back to the parent path's aggregate only when a URL's own quality is unknown, and Google's crawl budget guide lists page quality among the demand factors; read it as site quality setting the baseline while known URL-level signals still count.
Author Ibrahim AnjroAnnotates Day 1, 16:00 · How Google thinks about crawl budget
Day 3: Serving: Ranking, Search Console, and Performance 10
Shown on screen 2
Google estimated that refreshing (recrawling) a known URL takes about 30 days on average, with a minimum of seconds and an end point of weeks to never.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript
Google estimated that a crawl demand update driven by Search takes about 20 hours on average, with a minimum of minutes and an end point of weeks to months.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence slide photo, transcript
Said on stage 7
Google's index drops URLs that have not been recrawled for a very long time, which is the 'never' end of the refresh estimate.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
For news sites in particular, Google said the refresh of a known URL can literally take seconds (the reading 'news sites' is likely but not certain).
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
- Extends D1-C466 Day 1: Gary Illyes said news sites rarely need to worry about crawl budget, because Google crawls them aggressively…
An audience member reported recrawl intervals on one very large, popular client site: the homepage about 5 times a day, first-level pages about every 1.5 days (uncertain reading), and pages clustered as soft 404s every 160 to 190 days, adding that other sites will differ.
From the audienceIn Day 3, 15:45 · How long does it take to..?Evidence transcript
An audience member said that on the large site they observed, recrawl frequency followed the site hierarchy, demand and how well each page is internally linked.
From the audienceIn Day 3, 15:45 · How long does it take to..?Evidence transcript
Increases in crawl capacity take longer than decreases, within one to three weeks, because Google first needs to know that the higher demand will last.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
Used byrequirement DEV-SRV-03
When Google notices the web getting excited about a few URLs on a site, it allocates more crawl demand to the whole site so it does not miss content useful to future searchers.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
- Extends D1-C093 Day 1: Crawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on…
- Extends D1-C377 Day 1: The quality that drives crawl demand is the quality of the site as a whole, not the quality of an individual…
- Extends D1-C439 Day 1: Very large, constantly changing sites get no special handling: when URLs change frequently or are useful to…
Crawl demand can also come from other Google products, such as Shopping, and the roughly 20-hour estimate covers only demand from Search.
Speaker Gary IllyesIn Day 3, 15:45 · How long does it take to..?Evidence transcript
- Extends D1-C375 Day 1: Googlebot as a single standalone crawler is a historical idea: Google crawls through a centralised crawling…
- Repeats D1-C538 Day 1: A Google panelist said crawl demand is not only Search's: when enough of a site's URLs are wanted by Search…
Analysis by the author 1
With a typical refresh of about 30 days and deep or soft-404-like pages recrawled only every few months, link important deep pages from strong hub pages and fix pages that look like errors; then check recrawl intervals by site depth in server logs or the Crawl Stats report.
Author Ibrahim AnjroAnnotates Day 3, 15:45 · How long does it take to..?