Area
The index and its signals
The signals Google calculates while indexing, how it decides what enters the index, what the index looks like, and how Search Console reports it.
Topics in this area 5
25 claims · 7 sessions
Signals calculated during indexing
Of the many signals Google calculates during indexing, it singled out country, language, freshness, SafeSearch and spam as having large effects, and said country and language have been among its most important signals since its early days. Google stores one language per page and weights language and country in index selection so the index is not dominated by English or by the largest countries (said at the event, not in Google's docs). Some freshness signals are calculated at indexing but used in ranking for queries that deserve fresh results, and Google said freshness plays no big role in whether a page is indexed (not in its docs). SafeSearch signals are calculated at indexing because ranking runs online with no time for them, need the whole parsed page plus its inbound and outbound links, and, Google said, do not change a page's chance of being indexed (the timing and indexing points are not in its docs); Google's guidelines advise keeping explicit content on a separate domain or subdomain so the whole site is not filtered. Each indexed document carries its calculated signals for index selection to use; Google said many are documented under other names, but that creating content users like is a better use of time than hunting for signals. Day 3 showed where these signals are used next: at retrieval Google orders candidates with per-document signals collected during indexing, language first and then country. Google added that its slowest signal is recalculated only about once a year, which is why a site move can take up to a year (a partly uncertain passage of the recording). Day 1's second recording added that the signals calculated during indexing are stored and used both for index selection and later for ranking.
Day 1Day 2Day 3
50 claims · 10 sessions
Spam detection and SpamBrain
Google said spam is stopped at indexing: index selection loads most spam signals and keeps documents detected as spam out of the index, and Google's spam policies say violating sites may rank lower or not appear at all. Google has used statistical models against spam for over 20 years and now uses more and more AI; it said SpamBrain is built on Gemini and fine-tuned for spam, which is not in its documentation. Figures said on stage differ from the record: Google's 2021 webspam report gives 2018 as SpamBrain's launch year (2022 was mentioned on stage), and its 2022 report's '5 times more spam sites' compares 2022 with 2021, with 200 times more than at launch, rather than comparing SpamBrain with older algorithms. Google said its testing shows more than 99% of visits from Search are spam-free, a figure in its 2022 webspam report. Google said spam metadata such as white-on-white text is stored with a page's tokens (not in its docs), and its scaled content abuse policy covers automatically translated scraped content. Day 3 put the scale in numbers: Google discovers about 40 billion spammy pages a day, the figure in its webspam report for 2020 and its How Search Works page (a Day 3 slide headed 'In 2023' repeated it, and the quality talk said tens of billions), and its spam systems remove a very large amount before results are served. Google said AI has changed how it builds spam updates, letting it evaluate more candidates, catch new types of spam and launch faster; four spam updates had shipped in 2026 by 2 October, the September 2026 one still rolling out. Google repeated its aim of keeping more than 99% of results spam-free, which its webspam reports measure as visits from Search. Keyword stuffing remains a policy violation, and Google counts AI slop, mass-produced LLM content, as scaled content abuse that it is working hard to remove. Day 1's second recording added that the crawl scheduler very likely deprioritises URLs or sites known to be historically spammy. On Day 2 the non-Latin-script talk cited the October 2023 spam update, whose post says it improved coverage in many languages, naming Turkish, Vietnamese, Indonesian, Hindi and Chinese, particularly against cloaking, hacked, auto-generated and scraped spam. Author’s view: the post names neither Persian nor Arabic and does not mention link spam, so it does not show that the bought links the presenter saw working in those languages were addressed.
Day 1Day 2Day 3
50 claims · 8 sessions
Index selection: deciding what gets indexed
Google's index is immense but finite, so index selection, the last step before it, sets thresholds and keeps only documents likely to be useful to users now or later; Google said the system is predictive and relies heavily on machine learning (not in its docs). Quality is ultimately what decides; Google added that a site's or document's overall importance plays a significant role, as with news sites, and that a site or section already satisfying users gets more forgiving treatment for its new URLs (not in its docs). Language and country are weighted so the index is not dominated by English or by the largest countries, and where Google lacks content in a language such as Basque, lower-quality documents are more likely to be selected, which Google presented as an opening for new publishers (said at the event, not in Google's docs). Index selection also applies blocking signals: a noindex rule (a likely reading of the recording), an expired unavailable_after date, soft 404s, non-canonical duplicates, spam signals and other policies, while freshness and explicit content play no big role in the decision. In Search Console, 'Crawled – currently not indexed' is an index selection decision and, Google said, most of the time a quality issue rather than a technical one: the pages are low quality or useless for the index. Day 3 repeated that Google does not index every URL, which is why it has to rank better and better; Google's indexing chart said end-to-end indexing can take months or never happen, for reasons of quality, and Google said URLs not recrawled for a very long time are dropped from the index (not in its docs). Author’s view: set against the hundreds of trillions of URLs Google said it knows and the hundreds of billions of pages its ranking guide puts in the index, only around one known URL in a thousand is indexed, if both figures hold. A second recording of Day 1 added Cherry Prommawin's version: signals calculated during indexing are stored and used both to decide whether a page is indexed and later for ranking, index selection runs after signals are collected and duplicates dropped, and a selected page takes its whole duplicate cluster into the index with it. Day 2's opening Q&A added that very few robots.txt-disallowed URLs are in the index, and that an important disallowed URL might still be indexed without its content. Author’s view: Google's robots.txt guide names links from elsewhere as the reason, so keep important pages crawlable with noindex if they must stay out of Search.
Day 1Day 2Day 3
22 claims · 6 sessions
Indexing statuses in Search Console
Google described the two not-indexed statuses as different stages: 'Discovered – currently not indexed' is a crawl-scheduling state in which Google knows the URL but does not want to crawl it yet, and 'Crawled – currently not indexed' is an index selection decision taken after processing. On stage Google called the first the nastier one and said owners can influence it by getting other URLs indexed and showing that the site's content is useful, while the Page indexing help explains it by expected server overload. For 'Crawled – currently not indexed', Google said the cause is most of the time quality rather than a technical fault: check whether the pages match the quality of the indexed parts of the site, or whether a different template hides where their content is; the help page adds that such pages need no resubmission. Other not-indexed reasons include noindex and 'Alternate page with proper canonical tag', the canonical choice of deduplication that site owners see in Search Console, and Google suggested the report's reasons for checking index selection issues when testing changes. Google's URL Inspection help says a valid live test only confirms that Google can access a page; indexing still depends on other conditions, such as not being a duplicate and being of high enough quality. On Day 3 a community speaker said analysing Search Console query data with an LLM helped his agency spot developer mistakes such as unwanted pages being indexed, which Search Console also shows directly. On Day 1 Gary Illyes had already pointed to the report's breakdown of reasons as a way to find patterns in how a site's content is crawled and served. An audio recording of Day 1 added that soft 404s are looked up in this report rather than in Crawl Stats, and Dave Smart warned in Lightning session B that a URL reported as blocked by robots.txt may not be disallowed itself: Search Console reports a block anywhere in a redirect chain on the chain's first URL, so check whether the URL redirects and test every URL in the chain.
Day 1Day 2Day 3
31 claims · 6 sessions
What the index looks like
Google's Search index, described on Day 1 as big but not limitless, holds the tokens of each page with their metadata and the signals calculated for the document, not the full page content. Retrieval works through posting lists, kept for most tokens, that list the URLs containing each token: Google intersects the lists of the important query words to get an unranked set of candidates, and its public How Search Works explainer likens the index to the index at the back of a book. Google can also retrieve documents through vector embeddings, where closeness to the query's embedding decides what comes back, and both methods work from the page's content. Snippets are rebuilt from the stored tokens and their positions, and AI Overviews and AI Mode use the same index and snippets: fan-out queries go to the Search index and the returned snippets are the material for the AI answer, so a page that forbids snippets cannot be used for them. The token storage, the posting-list details and snippet reconstruction were said at the event and are not in Google's documentation. Day 3 repeated the posting-list model at retrieval: a document can be retrieved only if it contains the query's words or, for embeddings, related concepts. For scale, Google's first index in 1998 held 26 million pages, Google's own figure (25 million was said on stage). Cherry Prommawin said on Day 1 that Google's index, printed on paper, would reach the Moon and back twelve times (said at the event).
Day 1Day 2Day 3
Across days 66
- Stage D2-C073 Day 2 · Controlling indexing
John Mueller said Google uses the snippet as a way of building AI Overviews and AI Mode answers, so if a page forbids a snippet, Google cannot use that snippet for them.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
extendsStage D1-C070 Day 1 · How crawling errors affect SearchSoft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.
- Stage D2-C344 Day 2 · Handling web duplication
Google deduplicates pages because many sites have very many pages and Google's index does not have room for everything.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Analysis D2-C612 Day 2 · Focusing on Internationalisation and Localisation
Treat machine translation as a first draft: have a native speaker review it and adapt dates, calendars and units before publishing, because unreviewed bulk translation that adds little value can also fall under Google's scaled content abuse policy.
extendsAnalysis D1-C048 Day 1 · How Search works and where's AI?The answer is not a claim that Google detects AI text. It says ranking favours text that reads as natural to people. The risk with AI content is scale without value, which falls under Google's scaled content abuse policy, not the tool itself.
- Stage D2-C648 Day 2 · Calculating (some) signals
Among the many signals Google calculates during indexing, the ones singled out as having large effects on search results were country, language, freshness, SafeSearch and spam.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C656 Day 2 · Calculating (some) signals
Freshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking event hits a city, such as possible closure of Barcelona's airport, users want really fresh results, not results from two weeks ago.
extendsD1-C045 Day 1 · How Search works and where's AI?Ranking signals differ by result type: web pages (text, links, passages), images (resolution, colour, associated text), news (freshness, originality, diversity), local (location, type, rating, reviews, hours) and videos (language, text from speech).
- Stage D2-C668 Day 2 · Calculating (some) signals
Google uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is highly accurate.
extendsD1-C041 Day 1 · How Search works and where's AI?Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.
- Stage D2-C669 Day 2 · Calculating (some) signals
SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.
extendsD1-C039 Day 1 · How Search works and where's AI?Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
- Stage D2-C669 Day 2 · Calculating (some) signals
SpamBrain, Google's AI-based spam detection system, is nowadays built on Gemini and fine-tuned specifically for finding spam.
extendsD1-C041 Day 1 · How Search works and where's AI?Statistical models have been used at Google for over 20 years, for catching spam and originally for the 'Did you mean' feature.
- Stage D2-C682 Day 2 · Deciding what goes in the index?
Google's index selection system calculates thresholds and decides which documents are kept and which are thrown out; a URL that does not meet the thresholds is not indexed.
extendsStage D1-C211 Day 1 · How Search works and where's AI?Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
- Stage D2-C684 Day 2 · Deciding what goes in the index?
Index selection is a predictive AI system that relies heavily on machine learning.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C685 Day 2 · Deciding what goes in the index?
Index selection uses what Google already knows about a site: if the site, or even a section of it, satisfies users' needs well, new URLs from it are treated more forgivingly.
extendsD1-C094 Day 1 · How Google thinks about crawl budgetIf the quality or popularity of a URL is unknown, the aggregate quality or popularity of its parent path is used, then that path's parent, and so on.
- Analysis D2-C686 Day 2 · Deciding what goes in the index?
Launch new pages under sections that Google already indexes well, and improve or remove weak sections, because index selection judges new URLs partly by what it knows about the site and the section they sit in.
extendsAnalysis D1-C095 Day 1 · How Google thinks about crawl budgetNew content inherits its starting crawl demand from the folder it sits in. Put new high-value content under sections Google already rates well, not under weak ones.
- Stage D2-C687 Day 2 · Deciding what goes in the index?
Index selection uses the signals calculated earlier in indexing for each document it has to select or discard.
extendsStage D1-C205 Day 1 · How Search works and where's AI?Signals calculated for a page during indexing are stored in the index and used both to decide whether the page gets indexed and, later, for ranking.
- Stage D2-C688 Day 2 · Deciding what goes in the index?
Index selection is the last step before documents enter Google's index.
extendsStage D1-C211 Day 1 · How Search works and where's AI?Index selection runs after signals are collected and duplicates are dropped, and decides what goes into Google's index, which is big but not limitless.
- Stage D2-C695 Day 2 · Deciding what goes in the index?
Page quality is ultimately what decides whether a document is indexed, so focusing on quality is the most reliable way to get pages into Google's index.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D2-C697 Day 2 · Deciding what goes in the index?
Index selection drops a document carrying a noindex rule if it was not already dropped earlier in processing (noindex is a likely but not certain reading of the transcript, supported by the later mention of noindex among the Page indexing report reasons).
extendsD1-C106 Day 1 · How Google thinks about crawl budgetThe noindex rule consumes crawl budget, because Google must fetch the page to see it.
- Stage D2-C700 Day 2 · Deciding what goes in the index?
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
extendsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C701 Day 2 · Deciding what goes in the index?
When Google already has duplicate information for a document, for example when reprocessing it, index selection uses it to drop non-canonical duplicates from further processing, so that, as the speaker put it, only canonicals end up in search results (a simplification; see D2-C703).
extendsDocs D1-C128 Day 1 · session not recordedGoogle's canonicalization guide says some duplicate content on a site is normal and not a violation of its spam policies; Google clusters duplicate pages, picks the most representative one as canonical and crawls the duplicates less often.
- Stage D2-C706 Day 2 · Deciding what goes in the index?
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
extendsD1-C065 Day 1 · How crawling worksThe scheduler is shared infrastructure that decides what to fetch and when and sends URLs to the crawler. Each team decides the scheduling parameters for its own user agents.
- Stage D2-C706 Day 2 · Deciding what goes in the index?
'Discovered – currently not indexed' in Search Console is a crawl-scheduling state: Google knows the URL exists but does not want to crawl it yet. Of the two not-indexed statuses discussed, it was called the 'kind of nastier' one.
extendsDocs D1-C097 Day 1 · How Google thinks about crawl budgetGoogle's crawl budget guide is written for sites with over 1 million unique pages that change weekly, over 10,000 pages that change daily, or many URLs reported as 'Discovered – currently not indexed'.
- Analysis D2-C708 Day 2 · Deciding what goes in the index?
The help page explains 'Discovered – currently not indexed' by expected server overload (capacity), while on stage it was explained as Google not wanting the URL yet (demand); the crawl budget guide covers both, so first rule out slow responses and server errors, then treat the status as a quality and demand problem.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Stage D2-C709 Day 2 · Deciding what goes in the index?
Site owners can influence 'Discovered – currently not indexed' by getting other URLs of the site indexed and showing Google's systems that the site's content is good and useful to users.
extendsD1-C093 Day 1 · How Google thinks about crawl budgetCrawl demand is driven by the quality of the site, the change frequency of its URLs and their popularity on the internet.
- Stage D2-C714 Day 2 · Deciding what goes in the index?
'Crawled – currently not indexed' is most of the time a quality issue rather than a technical one: the pages are usually low quality or useless for the index, for example duplicates or soft 404s.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Analysis D2-C716 Day 2 · Deciding what goes in the index?
Treat 'Crawled – currently not indexed' as a quality audit list: compare those URLs with indexed pages of the same type for thin, duplicate or soft-404-like content and for template differences before looking for technical faults.
extendsAnalysis D1-C098 Day 1 · How Google thinks about crawl budgetOn smaller sites, slow indexing is almost always a demand problem, meaning quality, not a capacity problem.
- Stage D2-C717 Day 2 · Deciding what goes in the index?
Search Console's Page indexing report is the place to check for index selection issues, and its not-indexed reasons are useful when testing changes on a site.
extendsStage D1-C368 Day 1 · How crawling errors affect SearchSearch Console's page indexing report breaks down the reasons why pages do or do not show in Search, and its categories help find patterns in how a site's content is crawled and served to Google's crawlers.
- Stage D2-C722 Day 2 · How does the index look like?
Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
extendsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C726 Day 2 · How does the index look like?
AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Stage D2-C727 Day 2 · How does the index look like?
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
extendsD1-C051 Day 1 · How Search works and where's AI?Three reasons were given: generative AI features are built directly on the core ranking systems, query fan-out expands the original query to find related information, and generative AI features highlight content indexed by Google Search.
- Stage D2-C727 Day 2 · How does the index look like?
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
extendsDocs D1-C053 Day 1 · How Search works and where's AI?Query fan-out means running several related searches at once to gather more results; a question about lawn weeds may also search herbicides and weed prevention.
- Stage D2-C727 Day 2 · How does the index look like?
Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
extendsStage D1-C172 Day 1 · Welcome and opening keynotesGoogle said the Gemini model lets Search understand the user's intent, and query fan-out then adds further queries to the first one to enrich the quality of the answer.
- Stage D2-C741 Day 2 · How does the index look like?
In embedding-based retrieval, the distance between the embeddings of documents and the embedding of the user's query decides which documents are returned.
extendsDocs D1-C129 Day 1 · How Search works and where's AI?Google's guide says creating separate content for every variation of how people might search, including fan-out queries, primarily to manipulate rankings or AI responses violates its scaled content abuse policy. It adds that its AI systems can understand a page's relevance even without an exact match to the query.
- Stage D2-C846 Day 2 · Welcome to indexing day!
Google said a URL disallowed by robots.txt might still be indexed if the URL is important, in which case the URL is indexed but not its content.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Analysis D2-C849 Day 2 · Welcome to indexing day!
Google's robots.txt introduction names links from elsewhere on the web as the reason a disallowed URL can still be indexed, while on stage Google spoke of the URL's importance; either way, well-linked important URLs are the disallowed ones most likely to appear in results, so keep such pages crawlable with noindex if they must stay out of Search.
extendsAnalysis D1-C110 Day 1 · How Google thinks about crawl budgetA URL blocked in robots.txt can still be indexed without its content if other pages link to it, and Google cannot see a noindex on a page it is not allowed to fetch.
- Stage D3-C006 Day 3 · Making sense of users' queries
Google's first step in understanding almost any query is to detect its language, which tells Google roughly what content the user wants: a query in German suggests German content, a query in English English content.
extendsStage D2-C654 Day 2 · Calculating (some) signalsIn ranking, country and language signals help Google serve users the right content for their country and language.
- Stage D3-C013 Day 3 · Making sense of users' queries
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
extendsStage D2-C737 Day 2 · How does the index look like?A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
- Stage D3-C059 Day 3 · Making sense of users' queries
Google treats fan-out queries generated by the LLM the same way as queries typed by users, so understanding how normal queries work explains fan-out queries too.
extendsStage D2-C727 Day 2 · How does the index look like?Google said that when AI Overviews or AI Mode run a query fan-out, the generated queries are sent to Google's Search index and documents come back with their snippets, which then feed the AI-generated answer (part of this passage is unclear in the recording).
- Stage D3-C075 Day 3 · Making sense of users' queries
At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.
extendsStage D2-C736 Day 2 · How does the index look like?In posting-list retrieval, the posting lists of the query's words are intersected, which yields an unranked list of candidate URLs.
- Stage D3-C075 Day 3 · Making sense of users' queries
At retrieval, Google splits the query into words, applies query understanding and expansion, and matches the words and their expansions against the posting lists.
extendsStage D2-C738 Day 2 · How does the index look like?At retrieval, Google looks up the posting lists of the query words that are actually important rather than of every word in the query.
- Stage D3-C077 Day 3 · Making sense of users' queries
The first condition for retrieving a document is that the query's words, or its concepts in the case of vectors or embeddings, are in the document or related to it.
extendsStage D2-C740 Day 2 · How does the index look like?Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.
- Stage D3-C079 Day 3 · Making sense of users' queries
To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.
extendsStage D2-C649 Day 2 · Calculating (some) signalsCountry and language are among Google's most important signals and have been used since Google's early days.
- Stage D3-C079 Day 3 · Making sense of users' queries
To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.
extendsStage D2-C722 Day 2 · How does the index look like?Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
- Stage D3-C080 Day 3 · Making sense of users' queries
At retrieval, Google tries to match results to the user's language wherever possible: someone searching in Spanish does not necessarily want results in Italian.
extendsStage D2-C654 Day 2 · Calculating (some) signalsIn ranking, country and language signals help Google serve users the right content for their country and language.
- Stage D3-C082 Day 3 · Making sense of users' queries
Country is the second retrieval signal: a user searching from Switzerland wants cheese from Switzerland, not from Germany, and a user in Spain is poorly served by results targeting a South American country.
extendsStage D2-C649 Day 2 · Calculating (some) signalsCountry and language are among Google's most important signals and have been used since Google's early days.
- Stage D3-C083 Day 3 · Making sense of users' queries
Google called quality the most important of the signals used to order candidates at retrieval: a URL of high quality is more likely to be retrieved from the index for specific queries.
extendsStage D2-C722 Day 2 · How does the index look like?Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
- Analysis D3-C109 Day 3 · Lightning session K: Facets of quality
The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.
extendsStage D2-C318 Day 2 · Understanding what's on a pageGoogle does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
- Stage D3-C205 Day 3 · How Google thinks about Quality
Google's quality talk said AI has fundamentally changed how Google builds spam updates, letting it evaluate many more candidates and drastically increasing its velocity, so it launches faster with more impact.
extendsStage D2-C668 Day 2 · Calculating (some) signalsGoogle uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is highly accurate.
- Stage D3-C206 Day 3 · How Google thinks about Quality
Google's quality talk said AI lets Google catch new types of spam and catch more of it.
extendsStage D2-C668 Day 2 · Calculating (some) signalsGoogle uses more and more AI to detect spam, and Google's testing shows that this AI-based detection is highly accurate.
- Docs D3-C287 Day 3 · What are quality updates
Google's spam updates page says its automated spam detection systems run constantly, and a notable improvement to them, such as to the AI-based SpamBrain system, is called a spam update and listed with Google's ranking updates.
extendsStage D2-C670 Day 2 · Calculating (some) signalsSpamBrain is central to Google's spam-fighting efforts and has been improved many times since its launch.
- Stage D3-C311 Day 3 · How Search results are born
Google handles an image used as a search query much like a text query interpreted as an embedding: the image is broken down into vectors (embeddings) that are then searched for in the index.
extendsStage D2-C740 Day 2 · How does the index look like?Besides posting lists, Google can retrieve documents through vector embeddings: parts of documents are associated with embeddings, which form a vector space used for retrieval.
- Stage D3-C316 Day 3 · How Search results are born
Google generates the parts of a text result, such as title link and snippet, from its understanding of the underlying web page, even when the site owner provides nothing extra.
extendsStage D2-C724 Day 2 · How does the index look like?The snippet shown for a web result is reconstructed from the tokens stored in Google's index: Google knows the position of each token in the document and rebuilds the snippet from those positions.
- Stage D3-C325 Day 3 · How Search results are born
AI Mode and AI Overviews are not rich results but standard search features: they need no structured data to function and work with the normal text results from Google's index.
extendsStage D2-C726 Day 2 · How does the index look like?AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.
- Analysis D3-C487 Day 3 · Lightning session L: Understanding SERPs and your users
Google's ranking systems guide describes 'query deserves freshness' systems that show fresher content where it would be expected, which is narrower than a general preference for fresh content; refresh pages whose queries expect current information, and judge other refreshes by quality.
extendsStage D2-C656 Day 2 · Calculating (some) signalsFreshness is a signal for queries that deserve fresh results ('query deserves freshness'): when a breaking event hits a city, such as possible closure of Barcelona's airport, users want really fresh results, not results from two weeks ago.
- D3-C612 Day 3 · How long does it take to..?
Google may never fetch a lower-quality site's sitemap again: once it figures out the site is of lower quality, it no longer wants to fetch the sitemap.
extendsStage D1-C331 Day 1 · How crawling worksGoogle's crawl scheduler very likely deprioritises a URL when the URL or its site is known to be historically spammy.
- Stage D3-C695 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
The cheaper tokens become, the more AI slop is created, and Google counts AI slop as scaled content abuse.
extendsDocs D2-C611 Day 2 · Focusing on Internationalisation and LocalisationGoogle's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.
- Stage D3-C697 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem, but how it was used and what for.
extendsDocs D2-C611 Day 2 · Focusing on Internationalisation and LocalisationGoogle's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
repeatsDocs D1-C073 Day 1 · How crawling errors affect SearchA soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
- Stage D2-C334 Day 2 · Understanding what's on a page
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
repeatsStage D1-C355 Day 1 · How crawling errors affect SearchA soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.
- Stage D2-C679 Day 2 · Calculating (some) signals
Combing Google's documentation for signals is not the best use of an SEO's time; creating content that users will like is a better one.
repeatsStage D1-C059 Day 1 · Welcome and opening keynotesThe opening keynote closed with the advice to think about UEO, user engine optimisation, next to SEO and GEO: focus on the user and the rest will follow.
- Stage D2-C680 Day 2 · Deciding what goes in the index?
Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
repeatsD1-C037 Day 1 · How Search works and where's AI?For classic Search, crawling means Googlebot, scheduling and robots.txt, with AI used in parts such as scheduling. Indexing is one big but not limitless index that calculates signals and understands more than words, with AI such as BERT. Serving uses hundreds of signals tailored to the moment, with AI such as RankBrain.
- Stage D2-C719 Day 2 · How does the index look like?
The speaker recapped Google's pipeline up to the index: Google crawls pages, processes the fetched documents and then stores them in its index.
repeatsD1-C036 Day 1 · How Search works and where's AI?Search runs as three stages, crawling, indexing and serving, and the event covered one stage per day.
- Stage D3-C074 Day 3 · Making sense of users' queries
Google's index uses posting lists: for each word, a list of the URLs associated with that word.
repeatsStage D2-C733 Day 2 · How does the index look like?For most of the tokens Google finds on the web, though not every single one, the index keeps a posting list of the URLs that contain that token.
- Stage D3-C076 Day 3 · Making sense of users' queries
For retrieval, Google uses signals attached individually to each document in the index.
repeatsStage D2-C722 Day 2 · How does the index look like?Each document in Google's index has pretty much all the signals calculated for it attached, for example quality signals plus the page's country and language, according to an illustration the speaker called an approximation of the real structure.
- Stage D3-C256 Day 3 · What are quality updates
Google does not index every URL on the web; because it cannot index everything, it has to rank results better and better to satisfy users' information needs.
repeatsStage D2-C680 Day 2 · Deciding what goes in the index?Google's index is immense but finite, so Google cannot index every URL it finds on a web with a practically infinite number of URLs.
- Stage D3-C275 Day 3 · What are quality updates
Google aims to keep more than 99% of search results free from spam and said it already achieves this, thanks to advances in AI.
repeatsStage D2-C676 Day 2 · Calculating (some) signalsGoogle's testing shows that, thanks to SpamBrain, more than 99% of visits from Search are now spam-free.
- Stage D3-C702 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
AI Overviews and AI Mode are built on the Search infrastructure Google has used for 25 to 30 years and have very few processes of their own.
repeatsStage D2-C726 Day 2 · How does the index look like?AI Overviews and AI Mode use the same index structures and token-based snippets as classic web results, a point Google called important but not obvious.