Day 2: Indexing 22
Said on stage 19
Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
- Extended by D3-C011 Day 3: Google named Thai as a language that makes query understanding more complex because it does not separate…
- Repeated by D3-C070 Day 3: Google's summary slide on query understanding noted that some languages do not use spaces between words…
For languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript
- Repeated by D2-C737 Day 2: A search query is broken into words with the same segmenter or tokenizer that Google used to build the index.
- Extended by D3-C013 Day 3: Google's query processing deliberately mirrors indexing: a query is transformed into something that can be…
Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.
Speaker GoogleIn Day 2, 14:15 · Focusing on Internationalisation and LocalisationEvidence transcript
- Extended by D2-C965 Day 2: Users of non-Latin-script languages do not always search in their own script: the same Persian query may be…
- Extended by D3-C051 Day 3: Users expect content written the way they search: in some languages they search in Latin characters, in…
Google usually understands a query word whether it is written with or without diacritics (accents).
Speaker GoogleIn Day 2, 14:15 · Focusing on Internationalisation and LocalisationEvidence transcript
- Extended by D3-C048 Day 3: Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a…
A community talk, presented on its author's behalf by a colleague, was about multilingual SEO for languages written in non-Latin scripts, such as Persian and Arabic.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Arabic and Persian contain letters that look identical to users but are different characters to software, with different Unicode code points; the example given was the letter ye, which has an Arabic and a Persian form.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Because Arabic and Persian lookalike letters have different code points, a searcher may type one variant of a word while a website's text uses the other.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Lookalike Arabic and Persian character variants cause problems for data analysis and keyword research: data for one term can be split across the variants, which makes keyword research less accurate.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Persian and Arabic are written right to left and English left to right, so a title that mixes the two scripts can display in a confusing, unpredictable order even when its content is correct.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Used byrequirement DEV-INT-12glossary term Bidirectional text
The display problem of mixing right-to-left and left-to-right text also affects product titles and URLs, the presenter of the non-Latin-script talk said, showing an example from a large Iranian e-commerce site.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Used byrequirement DEV-INT-12glossary term Bidirectional text
Users of non-Latin-script languages do not always search in their own script: the same Persian query may be typed in Persian script or in Latin letters, with the same intent and the same expected results.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Used byrequirement DEV-INT-11glossary term Transliterated queries
- Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…
Keyword research for non-Latin-script languages should ask how people actually type and spell in local search, not only what the right keyword is, because switching keyboards is inconvenient, especially when typing on mobile.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
The presenter of the non-Latin-script talk said that in competitive Persian, Turkish and Arabic searches, bought backlinks and paid editorial content still visibly influence rankings and are widespread (an observation; no data was shown); the presenter stressed this described the situation and was not a recommendation.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
The presenter of the non-Latin-script talk said Google announced in October 2023 that it had improved its spam-detection coverage for languages: the spam policy was global, but the coverage improvement was language-specific.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
The presenter of the non-Latin-script talk said Persian offers far fewer natural link opportunities than English, with fewer publications, niche blogs and websites.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
For SEO in languages such as Persian, the presenter of the non-Latin-script talk advised accepting different competitive dynamics and basing strategy on the actual maturity of that language's market, not only on the global policy.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
Google search features often launch in some languages or countries first and expand later (the presenter's example was site names, launched in several languages and then extended to all languages in 2023), so comparisons of performance across languages and countries should not assume a feature is live everywhere at once.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
In the presenter's observation, AI answers to Persian queries are written in Persian but cite some English sources; the presenter explained that Persian content on a topic is often thinner, of lower quality or less relevant, so the systems retrieve from languages with better content, which the presenter called cross-lingual retrieval.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
- Extends D2-C603 Day 2: Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is…
The non-Latin-script talk concluded that multilingual SEO is not just translation: script, language behaviour and market maturity also matter, and basic SEO tactics, though the same everywhere, must be fitted to each market.
Speaker a second community speakerIn Day 2, 14:30 · Lightning session G: InternationalisationEvidence transcript
What Google's documentation says 1
Google's post on multilingual searches (8 September 2023) says that, because of typing difficulty on some keyboards, a person in India might search in Hindi using Latin rather than Devanagari characters and want and receive Hindi results written either way.
Publisher Search Central blog (8 September 2023)Annotates Day 2, 14:30 · Lightning session G: Internationalisation
Used byrequirement DEV-INT-11glossary term Transliterated queries
Analysis by the author 2
For Persian and Arabic keyword research, look up each lookalike-character spelling of a term separately and add up the volumes, and check which variant the site's own content uses.
Author Ibrahim AnjroAnnotates Day 2, 14:30 · Lightning session G: Internationalisation
The October 2023 spam update cited in the non-Latin-script talk names neither Persian nor Arabic (only 'other languages') and lists cloaking, hacked, auto-generated and scraped spam, not link spam, so it does not show that the bought links the presenter saw working in Persian or Arabic search were addressed.
Author Ibrahim AnjroAnnotates Day 2, 14:30 · Lightning session G: Internationalisation
Day 3: Serving: Ranking, Search Console, and Performance 10
Shown on screen 1
Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence slide photo, transcript
- Repeats D2-C320 Day 2: Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings…
Said on stage 7
Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
- Extends D2-C320 Day 2: Text in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings…
Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a German 'ü' written as 'ü', as 'ue' or left out.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
Used byrequirement DEV-INT-11
- Extends D2-C600 Day 2: Google usually understands a query word whether it is written with or without diacritics (accents).
Google sometimes gets diacritic variants wrong, so it is worth searching to see whether Google understands a variant; if it does, pick one spelling and use it.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
Google's recommendation for diacritics, non-English words, product names and spelling variants is to focus first on what users actually search for.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
Used byrequirement DEV-INT-11
Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
Used byrequirement DEV-INT-11
- Extends D2-C599 Day 2: Google usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter…
Google said focusing on what users actually do makes more sense than trying to match what the search engine does.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
Spelling variants need no artificial focus: users are often fine finding one version of a word on one page and another version elsewhere.
Speaker John MuellerIn Day 3, 10:25 · Making sense of users' queriesEvidence transcript
What Google's documentation says 1
A 2006 Search Central blog post says Google considers pages with and without accents for a query word (México and Mexico), and that which accented characters count as equivalent depends on the searcher's interface language.
Publisher Search Central blog (1 September 2006)Annotates Day 3, 10:25 · Making sense of users' queries
Analysis by the author 1
For audiences that type the same words in two scripts or spellings (Hindi in Devanagari and in Latin letters, German with and without umlauts), check which forms appear in Search Console's queries and use those forms in headings and key text.
Author Ibrahim AnjroAnnotates Day 3, 10:25 · Making sense of users' queries
Used byrequirement DEV-INT-11