A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
Speaker Gary IllyesEvidence video, transcript
Used byrequirement DEV-HTM-01
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Day 2 · Thursday 1 October 2026 · 11:30
Speaker from the author's recording label. Opened with Slido quiz questions, then the part on main content, which only a second attendee recording captured.
A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
Speaker Gary IllyesEvidence video, transcript
Used byrequirement DEV-HTM-01
On Google's example blog page, the post title and opening sentence counted as important because they sit in the main content, in front of the user, while the site tagline, the 'Categories' sidebar and category links such as 'Hugo (7)' counted as less important supplementary text.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-HTM-02
In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.
Speaker Gary IllyesEvidence slide photo, transcript
Used byglossary term Tokenization
Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.
Speaker Gary IllyesEvidence 2 slide photos
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-ERR-03
Mistakes in Google's own systems are a further cause of soft 404s, and Gary Illyes asked site owners to report such mistakes in Google's forums.
“BONUS: Mistakes in Google's systems (that you should notify us about)”
Wording checked against the slide or recording
Speaker Gary IllyesEvidence slide photo, transcript
Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.
Speaker Gary IllyesEvidence transcript
When Google processes a page for indexing, it gives words different weights depending on the part of the page where they appear.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-02
Words in the footer of a page get a lower weight, so text placed in the footer is unlikely to contribute much to ranking the page.
“if you put something in a footer, it's more likely that it's not going to contribute much to ranking”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-02
To make a word count for ranking a page, Gary Illyes said the simplest step is to move it into the main content, because where text sits on a page already contributes quite a bit to ranking.
“where you position text on a page will already contribute quite a bit to ranking that page”
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-HTM-02
Gary Illyes said not everything on a page can or should be important: if everything were placed in the main content, nothing would stand out as main content, which he called working as intended.
Speaker Gary IllyesEvidence transcript
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
Speaker Gary IllyesEvidence slide photo, transcript
Used byglossary term Tokenization
For languages written with spaces between words, such as English and German, Search tokenization splits a sentence into its individual words.
Speaker Gary IllyesEvidence transcript
Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
Speaker Gary IllyesEvidence transcript
For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
Speaker Gary IllyesEvidence transcript
Gary Illyes said a colleague, John, would cover how Google interprets the words of a query on the morning of Day 3.
Speaker Gary IllyesEvidence transcript
When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-HTM-07
Google also stores spam metadata with the tokens of a page, for example that text was white on a white background, so ranking can use that information.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-06
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Speaker Gary IllyesEvidence slide photo, transcript
Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.
Speaker Gary IllyesEvidence slide photo, transcript
Used byglossary term Tokenization
Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real, but it matters at the level of an AI model's context window.
Speaker Gary IllyesEvidence transcript
Used bymyth M-002
Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.
Speaker Gary IllyesEvidence transcript
Used bymyth M-002
Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-AIF-03myth M-002
Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-AIF-03
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Speaker Gary IllyesEvidence 2 slide photos, transcript
Used byrequirements DEV-ERR-01, DEV-ERR-03glossary term Soft 404
Because error pages are worded in endless variations, of which 'page not found' is only the classic one, Google cannot detect soft 404s with simple error, word or keyword matching.
Speaker Gary IllyesEvidence transcript
Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.
“This is basically an LLM thing, something like BERT, that is specifically trained to understand page structure”
Speaker Gary IllyesEvidence transcript
For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.
Speaker Gary IllyesEvidence slide photo, transcript
Used byrequirement DEV-ERR-03
Google's soft 404 detection tells the page chrome, such as navigation and footer, apart from the main content by analysing the page's visual hierarchy alongside its text.
Speaker Gary IllyesEvidence transcript
Gary Illyes said soft 404 detection keeps searchers from clicking into dead ends and avoids wasting site owners' resources on visitors sent to error pages.
Speaker Gary IllyesEvidence transcript
Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
Gary Illyes defined a page's main content as any part of the page that directly helps the page achieve its purpose, what it was built for.
“Main content is any part of the page that directly helps the page achieve its purpose”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
Main content is not only text: images, videos, a tool or anything else that helps a page achieve its purpose can be main content.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
Content created by other users can be main content: on a user-generated content site, the user-generated content can be the page's main content.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
A comment section below a blog post can still be part of the page's main content and can contribute to Google's understanding of the page.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-REN-02
A page's main content includes all of its headings and its visible title.
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
Gary Illyes said the main content is what Google considers when ranking a page.
“It's the main content that we consider for ranking.”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-HTM-01
Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.
“the context window is perhaps 900,000 or even closer to a million big”
Speaker Gary IllyesEvidence transcript
Used byrequirement DEV-AIF-03
Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.
“When Google indexes a page, it determines the primary content (or centerpiece) of each page.”
Publisher Google Search Central
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
Google's Search Quality Rater Guidelines define main content as any part of the page that directly helps it achieve its purpose, including text, images, videos, page features such as calculators and content created by users, and they count the title at the top of the page as part of it.
“Main Content is any part of the page that directly helps the page achieve its purpose.”
Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)
Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)
Google's Search Quality Rater Guidelines say navigation links are a common type of supplementary content, and that content behind tabs and user reviews or comments may count as main content on some pages and as supplementary content on others, depending on the page's purpose.
Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)
Used byrequirement DEV-HTM-01
Put the terms a page should rank for in its main content (title, headings, opening paragraphs) rather than only in sidebars, tag lists or footers; footer keyword blocks are unlikely to add ranking value.
Author Ibrahim Anjro
Used byrequirement DEV-HTM-02
Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.
Author Ibrahim Anjro
Do not rewrite pages into short, self-contained chunks for AI systems; Google says Gemini reads context windows of millions of tokens, so structure content for readers, with clear headings and complete explanations.
Author Ibrahim Anjro
Used byrequirement DEV-AIF-03
Return real error status codes for error states, 404 or 410 for missing content and 503 for outages such as a failed database connection, also in single-page apps; a 200 page whose main content is only an error message is treated as a soft 404 even when header and navigation look normal.
Author Ibrahim Anjro
Used byrequirements DEV-ERR-03, DEV-SRV-03
Hidden text is recorded at token level: Google stores spam metadata such as white-on-white text with the tokens, so leftover hidden keyword blocks are a liability, not neutral clutter, and should be removed.
Author Ibrahim Anjro
Used byrequirement DEV-HTM-06
The talk gave Gemini's context window both as millions of tokens and as roughly 900,000 to a million; Google's long-context docs say Gemini models have context windows of 1 million or more tokens (about eight average novels per million), so plan with about one million tokens as the documented floor rather than several million.
Author Ibrahim Anjro
Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.
Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.
Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real, but it matters at the level of an AI model's context window.
Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.
Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.
Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.
A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.
Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.
BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.
For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.
For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.
Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.
Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden attribute; Google indexes such hidden content.
Gary Illyes said the main content is what Google considers when ranking a page.
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.
Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.
Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.
A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.
A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.
A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.
Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.
When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.
Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.
Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.
A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.
Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.