Day 1: Crawling 7
Said on stage 4
Google's crawling talk opened with how the internet works (TCP/IP, internet service providers, IP addresses, URLs, DNS and HTTP), because crawling can only be understood with that wider scope.
Speaker Cherry PrommawinIn Day 1, 14:05 · How crawling worksEvidence transcript
- Repeated by D2-C847 Day 2: Opening Day 2, Google recapped that Day 1 began with how Search works and then covered crawling, including…
A URL states how a resource is requested (the protocol, HTTP or HTTPS), where (the host, meaning which computer on the network) and what (the path to the exact page or file).
Speaker Cherry PrommawinIn Day 1, 14:05 · How crawling worksEvidence transcript
DNS works as the internet's address book: it is consulted for each request and tells the client at which IP address a host name can be reached.
Speaker Cherry PrommawinIn Day 1, 14:05 · How crawling worksEvidence transcript
Googlebot does not support HTTP/3 today.
In Day 1 · session not recordedEvidence notes
Used byrequirement DEV-SRV-07
What Google's documentation says 2
Google's crawlers support HTTP/1.1 and HTTP/2, use whichever gives the best crawling performance, and may switch between sessions. HTTP/2 can save server resources but gives no ranking benefit.
Publisher GoogleAnnotates Day 1 · session not recorded
Used byrequirement DEV-SRV-07
Google's crawler overview says HTTP/1.1 is the default protocol of Google's crawlers and that a site can opt out of crawling over HTTP/2 by answering Google's HTTP/2 requests with a 421 status code.
Publisher GoogleAnnotates Day 1 · session not recorded
Used byrequirement DEV-SRV-07
Analysis by the author 1
This does not mean avoiding HTTP/3, which benefits browsers. The rule is never to run a host that only answers over HTTP/3, and to check that CDN or firewall rules written for HTTP/3 traffic do not break HTTP/1.1 and HTTP/2.
Author Ibrahim AnjroAnnotates Day 1 · session not recorded
Used byrequirement DEV-SRV-07