Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Day 2 · Thursday 1 October 2026 · 11:30

Understanding what's on a page

Speaker Gary Illyes, Search Relations

TalkCoverageTranscriptSlidesVideo

Speaker from the author's recording label. Opened with Slido quiz questions, then the part on main content, which only a second attendee recording captured.

Shown on screen 6

SlideConsistent with docsD2-C309

A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

Speaker Gary IllyesEvidence video, transcript

Used byrequirement DEV-HTM-01

  • Repeats D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
  • Extended by D2-C861 Day 2: Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not…
SlideConsistent with docsD2-C314

On Google's example blog page, the post title and opening sentence counted as important because they sit in the main content, in front of the user, while the site tagline, the 'Categories' sidebar and category links such as 'Hugo (7)' counted as less important supplementary text.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-HTM-02

SlideConsistent with docsD2-C326

In tokenization for AI models, common English words stay whole and each maps to a numeric token ID, so the model works with IDs rather than with the words; on Google's slide the word 'can' had the same ID, 740, both times it appeared.

Speaker Gary IllyesEvidence slide photo, transcript

Used byglossary term Tokenization

SlideNot in docsD2-C328

Google's two tokenization slides showed the difference on the same sentence: the Search tokenizer kept 'robots.txt' and 'tl;dr' as single tokens, while the AI-model tokenizer split them into pieces such as 'tl' and 'dr' or 'robots' and 'txt', with the punctuation as separate tokens.

Speaker Gary IllyesEvidence 2 slide photos

SlideConfirmed by docsD2-C336

Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.

Speaker Gary IllyesEvidence slide photo, transcript

Things

Used byrequirement DEV-ERR-03

  • Extends D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D2-C213 Day 2: A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as…
SlideNot in docsD2-C337

Mistakes in Google's own systems are a further cause of soft 404s, and Gary Illyes asked site owners to report such mistakes in Google's forums.

“BONUS: Mistakes in Google's systems (that you should notify us about)”

Wording checked against the slide or recording

Speaker Gary IllyesEvidence slide photo, transcript

Things

Said on stage 33

StageConsistent with docsD2-C311

Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C170 Day 3: Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content…
StageNot in docsD2-C313

Words in the footer of a page get a lower weight, so text placed in the footer is unlikely to contribute much to ranking the page.

“if you put something in a footer, it's more likely that it's not going to contribute much to ranking”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-02

StageConsistent with docsD2-C315

To make a word count for ranking a page, Gary Illyes said the simplest step is to move it into the main content, because where text sits on a page already contributes quite a bit to ranking.

“where you position text on a page will already contribute quite a bit to ranking that page”

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-HTM-02

StageNot in docsD2-C316

Gary Illyes said not everything on a page can or should be important: if everything were placed in the main content, nothing would stand out as main content, which he called working as intended.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C318

Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

Speaker Gary IllyesEvidence slide photo, transcript

Used byglossary term Tokenization

  • Repeated by D2-C721 Day 2: Google's Search index does not hold the full content of pages; Google said storing full pages and pulling…
  • Extended by D3-C109 Day 3: The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens…
StageNot in docsD2-C320

Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

Speaker Gary IllyesEvidence transcript

  • Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
  • Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
StageNot in docsD2-C321

For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

Speaker Gary IllyesEvidence transcript

  • Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageD2-C322

Gary Illyes said a colleague, John, would cover how Google interprets the words of a query on the morning of Day 3.

Speaker Gary IllyesEvidence transcript

StageNot in docsD2-C323

When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-HTM-07

  • Repeated by D2-C720 Day 2: Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the…
StageNot in docsD2-C325

Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

Speaker Gary IllyesEvidence slide photo, transcript

  • Contradicts D1-C039 Day 1: Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization…
StageConsistent with docsD2-C327

Tokenizers for AI models split long words into sub-word pieces that may make no sense on their own, because a token for every possible word would make the vocabulary too big, and a generative model only cares about closeness in vector space.

Speaker Gary IllyesEvidence slide photo, transcript

Used byglossary term Tokenization

StageConsistent with docsD2-C330

Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real, but it matters at the level of an AI model's context window.

Speaker Gary IllyesEvidence transcript

Used bymyth M-002

  • Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
StageConsistent with docsD2-C331

Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.

Speaker Gary IllyesEvidence transcript

Things

Used bymyth M-002

  • Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
  • Extended by D2-C869 Day 2: Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps…
StageConsistent with docsD2-C332

Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-03myth M-002

  • Extends D1-C054 Day 1: Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise…
  • Extended by D2-C825 Day 2: Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window…
StageNot in docsD2-C825

Gary Illyes added that, once chunk size is thought of in millions of tokens as Gemini's context window allows, chunking has perhaps lost its meaning anyway.

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-03

  • Extends D2-C332 Day 2: Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller…
StageConfirmed by docsD2-C334

A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

Speaker Gary IllyesEvidence 2 slide photos, transcript

Used byrequirements DEV-ERR-01, DEV-ERR-03glossary term Soft 404

  • Repeats D1-C073 Day 1: A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It…
  • Extends D1-C070 Day 1: Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the…
  • Repeats D1-C355 Day 1: A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found'…
  • Extended by D2-C903 Day 2: A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that…
  • Extended by D2-C700 Day 2: Index selection drops soft 404 pages that were not dropped earlier, for example when a document is…
StageNot in docsD2-C335

Because error pages are worded in endless variations, of which 'page not found' is only the classic one, Google cannot detect soft 404s with simple error, word or keyword matching.

Speaker Gary IllyesEvidence transcript

Things
StageNot in docsD2-C338

Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

“This is basically an LLM thing, something like BERT, that is specifically trained to understand page structure”

Speaker Gary IllyesEvidence transcript

Things
  • Extends D1-C042 Day 1: BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at…
StageNot in docsD2-C339

For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

Speaker Gary IllyesEvidence slide photo, transcript

Used byrequirement DEV-ERR-03

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
StageConsistent with docsD2-C341

Gary Illyes said soft 404 detection keeps searchers from clicking into dead ends and avoids wasting site owners' resources on visitors sent to error pages.

Speaker Gary IllyesEvidence transcript

Things
StageConsistent with docsD2-C861

Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
  • Extends D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
StageConfirmed by docsD2-C862

Gary Illyes defined a page's main content as any part of the page that directly helps the page achieve its purpose, what it was built for.

“Main content is any part of the page that directly helps the page achieve its purpose”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConfirmed by docsD2-C864

Content created by other users can be main content: on a user-generated content site, the user-generated content can be the page's main content.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConsistent with docsD2-C865

A comment section below a blog post can still be part of the page's main content and can contribute to Google's understanding of the page.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConfirmed by docsD2-C866

Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-REN-02

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
  • Extends D2-C202 Day 2: Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden…
StageConsistent with docsD2-C868

Gary Illyes said the main content is what Google considers when ranking a page.

“It's the main content that we consider for ranking.”

Speaker Gary IllyesEvidence transcript

Used byrequirement DEV-HTM-01

  • Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
StageConsistent with docsD2-C869

Right after saying Gemini's context window holds millions of tokens, Gary Illyes put its size at perhaps 900,000 or even closer to a million, without a unit that the recordings capture.

“the context window is perhaps 900,000 or even closer to a million big”

Speaker Gary IllyesEvidence transcript

Things

Used byrequirement DEV-AIF-03

  • Long context Google AI for Developers (Gemini API docs) · checked 3 October 2026
  • Extends D2-C331 Day 2: Gary Illyes said Gemini's context window, where chunking actually matters, holds millions of tokens.

What Google's documentation says 3

DocsSourceD2-C310

Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.

“When Google indexes a page, it determines the primary content (or centerpiece) of each page.”

Publisher Google Search Central

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

DocsSourceD2-C870

Google's Search Quality Rater Guidelines define main content as any part of the page that directly helps it achieve its purpose, including text, images, videos, page features such as calculators and content created by users, and they count the title at the top of the page as part of it.

“Main Content is any part of the page that directly helps the page achieve its purpose.”

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
DocsSourceD2-C871

Google's Search Quality Rater Guidelines say navigation links are a common type of supplementary content, and that content behind tabs and user reviews or comments may count as main content on some pages and as supplementary content on others, depending on the page's purpose.

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026

Analysis by the author 6

AnalysisD2-C329

Day 1's slide said Gemini shares technologies such as tokenization with Search, while on Day 2 Gary Illyes showed that the two tokenizers split the same text differently ('or mostly'); read this as a shared processing step with different outputs, so Search's word tokens and Gemini's sub-word tokens are not the same units.

Author Ibrahim Anjro

AnalysisD2-C333

Do not rewrite pages into short, self-contained chunks for AI systems; Google says Gemini reads context windows of millions of tokens, so structure content for readers, with clear headings and complete explanations.

Author Ibrahim Anjro

Things

Used byrequirement DEV-AIF-03

AnalysisD2-C872

The talk gave Gemini's context window both as millions of tokens and as roughly 900,000 to a million; Google's long-context docs say Gemini models have context windows of 1 million or more tokens (about eight average novels per million), so plan with about one million tokens as the documented floor rather than several million.

Author Ibrahim Anjro

Things
  1. Stage D2-C325 Day 2 · Understanding what's on a page

    Tokenization for AI models such as Gemini, in training and in inference, differs from tokenization for Search, although Gary Illyes qualified this with 'or mostly'.

    contradicts
    Slide D1-C039 Day 1 · How Search works and where's AI?

    Gemini is not part of Search, but it uses crawlers for data, shares some technologies such as tokenization and deduping, and grounds on the Search index.

  2. Stage D2-C330 Day 2 · Understanding what's on a page

    Gary Illyes said the common SEO advice to chunk content for AI systems is misunderstood: chunking is real, but it matters at the level of an AI model's context window.

    extends
    Slide D1-C054 Day 1 · How Search works and where's AI?

    Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.

  3. Stage D2-C332 Day 2 · Understanding what's on a page

    Gemini does not need content cut into small chunks of 100 or 200 words, Gary Illyes said, since a smaller book fits in its context window.

    extends
    Slide D1-C054 Day 1 · How Search works and where's AI?

    Myth: optimise for AI over readers. Google's answer: optimise for people, with no need to obsess over precise keywords or AI phrasing, no need to chop content, and no need for llms.txt.

  4. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    extends
    Stage D1-C070 Day 1 · How crawling errors affect Search

    Soft 404s were named as a crawl problem alongside DNS and firewall issues, and described as one of the biggest problems on the internet right now for crawling and showing up in Search.

  5. Slide D2-C336 Day 2 · Understanding what's on a page

    Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.

    extends
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  6. Slide D2-C336 Day 2 · Understanding what's on a page

    Google listed four common causes of soft 404s: pages that look like errors but are not, thin or empty content, server or CMS misconfigurations, and JavaScript-dependent content that fails to load.

    extends
    Stage D2-C213 Day 2 · Lightning session D: Rendering and JavaScript

    A page that renders empty for Google, such as a client-side product page hit by these mistakes, is seen as thin content and ends up treated as a soft 404 even though users see a full page.

  7. Stage D2-C338 Day 2 · Understanding what's on a page

    Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

    extends
    Slide D1-C042 Day 1 · How Search works and where's AI?

    BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.

  8. Stage D2-C339 Day 2 · Understanding what's on a page

    For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  9. Stage D2-C700 Day 2 · Deciding what goes in the index?

    Index selection drops soft 404 pages that were not dropped earlier, for example when a document is reprocessed.

    extends
    Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

  10. Stage D2-C866 Day 2 · Understanding what's on a page

    Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.

    extends
    Slide D2-C202 Day 2 · Lightning session D: Rendering and JavaScript

    Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden attribute; Google indexes such hidden content.

  11. Stage D2-C868 Day 2 · Understanding what's on a page

    Gary Illyes said the main content is what Google considers when ranking a page.

    extends
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  12. Stage D2-C903 Day 2 · Lightning session E: Managing Duplicates and Site Moves

    A community speaker warned that HTTP 200 responses across a new domain show only that the URLs work, not that the content users came for is still there, so after launch the redirect map becomes the test.

    extends
    Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

  13. Stage D3-C011 Day 3 · Making sense of users' queries

    Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.

    extends
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

  14. Stage D3-C013 Day 3 · Making sense of users' queries

    Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

    extends
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  15. Analysis D3-C109 Day 3 · Lightning session K: Facets of quality

    The inspector metaphor blends two steps: Googlebot fetches pages during crawling, while tokenization happens later when the page is processed for the index, as Google explained on Day 2.

    extends
    Stage D2-C318 Day 2 · Understanding what's on a page

    Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

  16. Stage D3-C170 Day 3 · How Google thinks about Quality

    Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.

    extends
    Stage D2-C311 Day 2 · Understanding what's on a page

    Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.

  17. Slide D2-C309 Day 2 · Understanding what's on a page

    A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

    repeats
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  18. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    repeats
    Docs D1-C073 Day 1 · How crawling errors affect Search

    A soft 404 is a page that returns a success code while its content looks like an error or an empty page. It is kept out of the index but continues to be crawled, wasting crawl budget.

  19. Stage D2-C334 Day 2 · Understanding what's on a page

    A soft 404 is a page that should return an error status code but returns HTTP 200; from a crawling point of view it looks indexable, but because it has no real content Google throws it out of the index.

    repeats
    Stage D1-C355 Day 1 · How crawling errors affect Search

    A soft 404 is a 404 in disguise: the page returns 200 but its content says something like 'page not found', information the site should have sent as the HTTP status.

  20. Stage D2-C720 Day 2 · How does the index look like?

    Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.

    repeats
    Stage D2-C323 Day 2 · Understanding what's on a page

    When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

  21. Stage D2-C721 Day 2 · How does the index look like?

    Google's Search index does not hold the full content of pages; Google said storing full pages and pulling them out at serving time would be a very inefficient way of doing search.

    repeats
    Stage D2-C318 Day 2 · Understanding what's on a page

    Google does not store the complete sentences or the full HTML of a page in the Search index, because large pieces of text would be unsearchable; it tokenizes the text into the smallest segments that still allow search.

  22. Stage D2-C737 Day 2 · How does the index look like?

    A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

    repeats
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  23. Slide D3-C070 Day 3 · Making sense of users' queries

    Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.

    repeats
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.