Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · International and multilingual sites

Non-Latin scripts and word segmentation

Gary Illyes said Google segments text in languages written without spaces, such as Thai and Chinese, into words using statistical models built from web content in that language, and applies exactly the same segmentation to queries so they can match the index (said at the event, not in Google's docs). Google's internationalisation talk said Google usually understands query words written with or without diacritics, consistent with a 2006 Google blog post, and usually understands non-English words typed 'in English', probably meaning romanised spellings (not in Google's docs). A community talk on Persian and Arabic showed that letters which look identical can be different Unicode characters (such as the Arabic and Persian forms of the letter ye), so a searcher may type one variant while a site uses another, and keyword data for one term splits across the variants (a community observation, undocumented). Day 3's query understanding talk filled in the details: Google generally treats spellings with and without diacritics as synonyms behind the scenes (a German ü written as ü, as ue or without the dots), which the 2006 post supports, but sometimes gets variants wrong, so search a variant to check before picking one spelling. Users expect content in the form they search with, Latin letters or the local script, and some, such as Hindi users, use both; Google's advice was to focus on what users actually search for rather than on what the search engine does (said at the event, not in Google's docs). Thai was named again as a language whose lack of spaces between words complicates query understanding. A second recording covered the rest of that talk: mixing right-to-left and left-to-right scripts can scramble the display order of titles, product names and URLs, and users may type the same Persian query in Persian script or in Latin letters with the same intent, as Google's 2023 post on multilingual searches describes for Hindi, so keyword research should ask how people actually type, especially on mobile. The presenter observed that bought links and paid editorial content still visibly influence competitive Persian, Turkish and Arabic results (an observation, stressed as not a recommendation), that Persian offers far fewer natural link opportunities, and that AI answers to Persian queries cite English sources where Persian content is thin; Google's October 2023 spam update improved its coverage per language. The talk concluded that multilingual SEO is more than translation: script, language behaviour and market maturity matter. Author’s view: the October 2023 update names neither Persian nor Arabic and covers cloaking, hacked, auto-generated and scraped spam, not link spam, so it does not show that the bought links seen working there were addressed.

Things in this topic 3

Counts are claims that name the thing. All things

What to do

  • For Persian and Arabic, research each lookalike-character spelling of a keyword separately, add up the volumes and check which variant your own content uses.
  • Write Thai, Chinese and other unspaced languages naturally; Google segments pages and queries with the same models.
  • Do not build separate pages for accent-free or Latin-letter spellings of the same words; Google usually understands both forms in queries.
  • Check in Search Console which scripts and spellings your audience types, and use those forms in headings and key text.
  • Research how your audience actually types queries, in the local script or in Latin letters, and check that titles mixing right-to-left and left-to-right scripts display in the right order.

Day 2: Indexing 22

Said on stage 19

StageNot in docsD2-C320

Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

  • Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
  • Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
StageNot in docsD2-C321

For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

  • Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
  • Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
StageNot in docsD2-C599

Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.

Speaker GoogleIn Day 2, 14:15 · Focusing on Internationalisation and LocalisationEvidence transcript

  • Extended by D2-C965 Day 2: Users of non-Latin-script languages do not always search in their own script: the same Persian query may be…
  • Extended by D3-C051 Day 3: Users expect content written the way they search: in some languages they search in Latin characters, in…
StageConsistent with docsD2-C600

Google usually understands a query word whether it is written with or without diacritics (accents).

Speaker GoogleIn Day 2, 14:15 · Focusing on Internationalisation and LocalisationEvidence transcript

  • Extended by D3-C048 Day 3: Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a…
StageConsistent with docsD2-C965

Users of non-Latin-script languages do not always search in their own script: the same Persian query may be typed in Persian script or in Latin letters, with the same intent and the same expected results.

Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript

Used byrequirement DEV-INT-11glossary term Transliterated queries

  • Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…
StageNot in docsD2-C967

The presenter of the non-Latin-script talk said that in competitive Persian, Turkish and Arabic searches, bought backlinks and paid editorial content still visibly influence rankings and are widespread (an observation; no data was shown); the presenter stressed this described the situation and was not a recommendation.

Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript

StageConfirmed by docsD2-C968

The presenter of the non-Latin-script talk said Google announced in October 2023 that it had improved its spam-detection coverage for languages: the spam policy was global, but the coverage improvement was language-specific.

Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript

StageConfirmed by docsD2-C971

Google search features often launch in some languages or countries first and expand later (the presenter's example was site names, launched in several languages and then extended to all languages in 2023), so comparisons of performance across languages and countries should not assume a feature is live everywhere at once.

Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript

StageConsistent with docsD2-C972

In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English sources; the presenter explained that Persian content on a topic is often thinner, of lower quality or less relevant, so the systems retrieve from languages with better content, which the presenter called cross-lingual retrieval.

Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript

  • Extends D2-C603 Day 2: Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is…

What Google's documentation says 1

DocsSourceD2-C981

Google's post on multilingual searches (8 September 2023) says that, because of typing difficulty on some keyboards, a person in India might search in Hindi using Latin rather than Devanagari characters and want and receive Hindi results written either way.

Publisher Search Central blog (8 September 2023)Annotates Day 2, 14:30 · Lightning session G: Internationalisation

Used byrequirement DEV-INT-11glossary term Transliterated queries

Analysis by the author 2

AnalysisD2-C979

The October 2023 spam update cited in the non-Latin-script talk names neither Persian nor Arabic (only 'other languages') and lists cloaking, hacked, auto-generated and scraped spam, not link spam, so it does not show that the bought links the presenter saw working in Persian or Arabic search were addressed.

Author Ibrahim AnjroAnnotates Day 2, 14:30 · Lightning session G: Internationalisation

Day 3: Serving: Ranking, Search Console, and Performance 10

Shown on screen 1

Said on stage 7

StageNot in docsD3-C011

Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.

Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript

  • Extends D2-C320 Day 2: Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings…
StageConfirmed by docsD3-C048

Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a German 'ü' written as 'ü', as 'ue' or left out.

Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript

Used byrequirement DEV-INT-11

  • Extends D2-C600 Day 2: Google usually understands a query word whether it is written with or without diacritics (accents).
StageNot in docsD3-C051

Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.

Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript

Used byrequirement DEV-INT-11

  • Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…

What Google's documentation says 1

DocsSourceD3-C091

A 2006 Search Central blog post says Google considers pages with and without accents for a query word (México and Mexico), and that which accented characters count as equivalent depends on the searcher's interface language.

Publisher Search Central blog (1 September 2006)Annotates Day 3, 10:25 · Making sense of users' queries

Analysis by the author 1

Across days and sessions 8

  1. Stage D2-C965 Day 2 · Lightning session G: Internationalisation

    Users of non-Latin-script languages do not always search in their own script: the same Persian query may be typed in Persian script or in Latin letters, with the same intent and the same expected results.

    extends
    Stage D2-C599 Day 2 · Focusing on Internationalisation and Localisation

    Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.

  2. Stage D2-C972 Day 2 · Lightning session G: Internationalisation

    In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English sources; the presenter explained that Persian content on a topic is often thinner, of lower quality or less relevant, so the systems retrieve from languages with better content, which the presenter called cross-lingual retrieval.

    extends
    Stage D2-C603 Day 2 · Focusing on Internationalisation and Localisation

    Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.

  3. Stage D3-C011 Day 3 · Making sense of users' queries

    Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.

    extends
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

  4. Stage D3-C013 Day 3 · Making sense of users' queries

    Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.

    extends
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  5. Stage D3-C048 Day 3 · Making sense of users' queries

    Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a German 'ü' written as 'ü', as 'ue' or left out.

    extends
    Stage D2-C600 Day 2 · Focusing on Internationalisation and Localisation

    Google usually understands a query word whether it is written with or without diacritics (accents).

  6. Stage D3-C051 Day 3 · Making sense of users' queries

    Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.

    extends
    Stage D2-C599 Day 2 · Focusing on Internationalisation and Localisation

    Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.

  7. Stage D2-C737 Day 2 · How does the index look like?

    A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.

    repeats
    Stage D2-C321 Day 2 · Understanding what's on a page

    For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.

  8. Slide D3-C070 Day 3 · Making sense of users' queries

    Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.

    repeats
    Stage D2-C320 Day 2 · Understanding what's on a page

    Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.

Built on these claims 2

Developer requirements 2

Sources 7