Area
International and multilingual sites
Language versions, hreflang, country targeting and localisation, including languages written in non-Latin scripts.
Topics in this area 5
53 claims · 6 sessions
hreflang annotations
Google extracts hreflang while processing a page's HTML to learn its other-language equivalents, called it useful in ranking, and said it tries to use hreflang alternates when it clusters same-language pages for different countries (German pages for Germany, Austria and Switzerland) so it can show the right country version. Google's internationalisation talk listed the common mistakes: missing return links (Google ignores one-way annotations, since an alternate URL can be on any domain and any site could otherwise claim to be another site's language version), missing self-references, several methods at once with conflicting values, and invalid codes such as se, dk or cz instead of sv, da or cs, UK instead of GB, EU, city codes or a region code with no language. Google's hreflang guide agrees and adds detail: the HTML head, HTTP header and sitemap methods are equivalent and combining them brings no benefit (so using only one is the talk's advice against conflicts, not a documented rule), alternate URLs must be fully qualified and may be on other domains, and hreflang is recommended even for small regional variations or when only the template is translated. Google does not use hreflang to determine a page's language, and its canonical documentation says the canonical of an hreflang page should be in the same language and that URLs in hreflang clusters are preferred as canonicals; a community speaker called canonical links between language versions a misuse that would leave only one language ranking. A community speaker cited a third-party study (reported in 2017, auditing 20,000 multilingual sites) that found at least one hreflang error on 75% of them, and stressed that hreflang only routes users to the right version: it does not make a site trusted in a new country, and understanding each market is the harder part. Day 3 added the query side: a brand-only query does not reveal which language the user wants, so Google sometimes shows the wrong language version for brand searches, falling back on location and browser settings, and it applies language and country already at retrieval, before ranking. Author’s view: that makes hreflang and a visible language switcher the safeguard for brand searches. Google's internationalisation talk added that incorrect language and region codes are a mistake it finds very often, with the same wrong codes coming up again and again, especially in Europe, because site owners assume the codes are easy; a second recording of the deduplication talk confirmed that its short advice for same-language, different-country pages was simply to use hreflang.
Day 2Day 3
36 claims · 7 sessions
Page language and language versions
Google determines a page's language for indexing from its content, not from hreflang or a language code in the URL, and its internationalisation talk said each page is annotated with only one language in the index (said at the event, not in Google's docs). Pages should therefore make their target language obvious and avoid mixing languages, and pages that translate only the template around untranslated main content are near-duplicates that Google may cluster. Each language version needs its own URL, because Googlebot mostly crawls from US IP addresses without an Accept-Language header and versions switched by cookie, in place or by redirect may not all be crawled; a community speaker showed that a language selector built as a button leaves the whole alternate-language cluster without crawlable links. Google called language and country among its most important signals since its early days, and said index selection weights language so the index does not turn into an English one; in under-served languages such as Basque that makes lower-quality documents more likely to be selected and gives new local content a good chance to replace spam (the index-selection details were said at the event, not in Google's docs). Author’s view: an orphaned language cluster is a case of Day 1's point that pages with no internal links depend on sitemaps and are found slowly. Day 3 showed the query side: detecting a query's language is Google's first step in understanding almost any query, and at retrieval Google matches results to the user's language wherever possible. A brand-only query does not reveal the language, so Google falls back on location and browser settings and sometimes shows the wrong language version; Google's help says the results language comes from the query, the user's language setting, the device's languages and location. Google called a missing result in a user's language an opportunity for whoever targets those words. Google said on Day 2 that working out a site's localized pages and which countries it targets is one of the difficult, complex tasks for it in indexing.
Day 2Day 3
34 claims · 6 sessions
Country targeting
Google's internationalisation talk called the ccTLD one of the strongest country-targeting signals, yet said a separate ccTLD per country is not necessarily better: ccTLDs, subdomains and subdirectories are all acceptable depending on availability, cost, local rules and commitment, URL parameters are not recommended, and sites should not move domains for the signal. The speaker added that server location is no longer reliable and little used (the word 'server' is unclear in the recording), whereas Google's documentation still lists it as a possible but not definitive signal, alongside hreflang, local addresses and phone numbers, currency, local links and Business Profile signals, without ranking them. Language subdomains such as en or de do not set a target audience, hreflang targets countries only (EU and city codes are invalid), and Google's duplication talk advised against 'clever' geo-redirecting, which Google's documentation explains: Googlebot mostly crawls from US IP addresses. Google said country has been one of its most important signals since its early days, helps serve users the right content, and is weighed in index selection so that smaller countries sharing a language, such as the UK or Switzerland, are not crowded out by the US or Germany (the index-selection part was said at the event, not in Google's docs). A community speaker warned that neither hreflang nor a country version makes a brand trusted in a market: local authority has to be earned, and a brand that no local source mentions should fix its existing markets before expanding. Day 3 placed country at retrieval: it is the second signal Google uses to order candidates, after language, so a user in Switzerland gets Swiss rather than German results, and Google's multi-regional guide says geotargeting a site to one country can improve its rankings there at the expense of other locales. Google added that working out a site's localized pages and which countries it targets is one of the difficult, complex tasks for it in indexing.
Day 2Day 3
70 claims · 8 sessions
Localisation beyond translation
Google's internationalisation talk treated localisation as more than translation: machine translation may be acceptable, at the site owner's judgement, after weighing translation quality, local conventions such as date formats and calendars (in Thailand the current year, as of 2026, is 2569) and cultural adaptation, and Google's spam policies list automated translation of scraped content as an example of scaled content abuse. Consumer data shown in the talk (source not captured) suggested that trust cues differ by market: promotions and discounts matter to shoppers in Europe and the US, reviews weigh more in the US, brand reputation in Europe, expert recommendations and certifications in Germany, and product origin in France (said at the event, not in Google's docs). The talk also warned that AI answers are still language-dependent: if an answer is synthesised from the top results, a query in another language draws on different data, so missing or weak local content becomes an invisible gap that site owners should check per language and market. John Mueller explained that the notranslate rule opts a page out of Google Search's translation features (documented) and that its form in a meta tag named google also turns off Chrome's automatic translation (said at the event; a 2008 Search Central blog post says that form stops Google Translate, and Chromium's Translate design document says it disables Chrome's translation), adding that he dislikes the rule because translation is a form of accessibility. A community speaker argued that understanding each market is harder than hreflang, citing an audit in which a correctly translated Dutch site converted far worse than its sister sites because the region had never been checked (Dutch homes are narrow, on several floors, with steep stairs), and noting that regional demand shifts over time, as a chart of search interest in ceiling fans in northern versus southern Europe showed. Day 3 added that Google does not prefer English words on pages in other languages, and that where it recognises an English term and its local equivalent as synonyms a page needs only one of them; Google called a gap in results for a language an opportunity for whoever targets those words, and repeated that some countries rely heavily on reviews as social proof (said at the event). A second recording of Day 2 added that working out a site's localized pages and which countries it targets is one of the difficult, complex tasks for Google in indexing, and the non-Latin-script talk saw AI answers to Persian queries written in Persian but citing English sources where Persian content was thin, concluding that multilingual SEO is more than translation: script, typing habits and market maturity must be fitted to each market. On Day 1 a community speaker showed that AI fan-out queries are often in another language than the prompt, citing a study heard as Peec AI's and a Spanish prompt about Barcelona restaurants that fanned out in English. Author’s view: Peec AI's published study found English fan-outs in nearly 78% of non-English prompt runs, from 66% for Spanish to 94% for Turkish, and 43% of the fan-out searches for those prompts ran in English.
Day 1Day 2Day 3
32 claims · 4 sessions
Non-Latin scripts and word segmentation
Gary Illyes said Google segments text in languages written without spaces, such as Thai and Chinese, into words using statistical models built from web content in that language, and applies exactly the same segmentation to queries so they can match the index (said at the event, not in Google's docs). Google's internationalisation talk said Google usually understands query words written with or without diacritics, consistent with a 2006 Google blog post, and usually understands non-English words typed 'in English', probably meaning romanised spellings (not in Google's docs). A community talk on Persian and Arabic showed that letters which look identical can be different Unicode characters (such as the Arabic and Persian forms of the letter ye), so a searcher may type one variant while a site uses another, and keyword data for one term splits across the variants (a community observation, undocumented). Day 3's query understanding talk filled in the details: Google generally treats spellings with and without diacritics as synonyms behind the scenes (a German ü written as ü, as ue or without the dots), which the 2006 post supports, but sometimes gets variants wrong, so search a variant to check before picking one spelling. Users expect content in the form they search with, Latin letters or the local script, and some, such as Hindi users, use both; Google's advice was to focus on what users actually search for rather than on what the search engine does (said at the event, not in Google's docs). Thai was named again as a language whose lack of spaces between words complicates query understanding. A second recording covered the rest of that talk: mixing right-to-left and left-to-right scripts can scramble the display order of titles, product names and URLs, and users may type the same Persian query in Persian script or in Latin letters with the same intent, as Google's 2023 post on multilingual searches describes for Hindi, so keyword research should ask how people actually type, especially on mobile. The presenter observed that bought links and paid editorial content still visibly influence competitive Persian, Turkish and Arabic results (an observation, stressed as not a recommendation), that Persian offers far fewer natural link opportunities, and that AI answers to Persian queries cite English sources where Persian content is thin; Google's October 2023 spam update improved its coverage per language. The talk concluded that multilingual SEO is more than translation: script, language behaviour and market maturity matter. Author’s view: the October 2023 update names neither Persian nor Arabic and covers cloaking, hacked, auto-generated and scraped spam, not link spam, so it does not show that the bought links seen working there were addressed.
Day 2Day 3
Across days 18
- D2-C187 Day 2 · Lightning session D: Rendering and JavaScript
A market or language selector built as a button works for users but leaves the whole cluster of alternate-language pages without crawlable links, so the cluster is orphaned for Google.
extendsDocs D1-C115 Day 1 · session not recordedGoogle's crawlers do not click buttons. Each page in a series needs its own URL and an <a href> link to the next page, should not use page 1 as its canonical, and rel=next and rel=prev are no longer used.
- Analysis D2-C189 Day 2 · Lightning session D: Rendering and JavaScript
Build market and language selectors as plain <a href> links to each alternate URL, not buttons or script handlers; otherwise the language versions have no internal links and depend on sitemaps to be found, which is slow.
extendsAnalysis D1-C067 Day 1 · How crawling worksA page with no internal links depends on sitemaps alone to be found, so it is discovered slowly and attracts little crawl demand.
- Stage D2-C603 Day 2 · Focusing on Internationalisation and Localisation
Users often assume AI is all-knowing and borderless, but AI is still language-dependent: if an AI answer is synthesized from the top results, a query in another language draws on a totally different set of data.
extendsD1-C038 Day 1 · How Search works and where's AI?AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
- Analysis D2-C607 Day 2 · Focusing on Internationalisation and Localisation
Check AI Overviews and AI Mode with native-language queries in each target market, not with translated English keywords, and compare with the country breakdown of Search Console's generative AI performance report; topics where competitors are cited and you are not point to missing or weak local content.
extendsDocs D1-C124 Day 1 · How Search works and where's AI?Google's launch post for the Generative AI performance reports in Search Console (3 June 2026; rolled out to all sites worldwide by 31 August 2026) lists impressions, pages, countries, devices (Search only) and dates, and no click or query metrics.
- Analysis D2-C612 Day 2 · Focusing on Internationalisation and Localisation
Treat machine translation as a first draft: have a native speaker review it and adapt dates, calendars and units before publishing, because unreviewed bulk translation that adds little value can also fall under Google's scaled content abuse policy.
extendsAnalysis D1-C048 Day 1 · How Search works and where's AI?The answer is not a claim that Google detects AI text. It says ranking favours text that reads as natural to people. The risk with AI content is scale without value, which falls under Google's scaled content abuse policy, not the tool itself.
- Stage D3-C006 Day 3 · Making sense of users' queries
Google's first step in understanding almost any query is to detect its language, which tells Google roughly what content the user wants: a query in German suggests German content, a query in English English content.
extendsStage D2-C654 Day 2 · Calculating (some) signalsIn ranking, country and language signals help Google serve users the right content for their country and language.
- Stage D3-C007 Day 3 · Making sense of users' queries
Query language detection works poorly when someone searches only for a brand name, such as Facebook or Google, because the query does not show which language the user wants results in.
extendsStage D1-C215 Day 1 · How Search works and where's AI?Serving starts with interpreting the query, which includes cleaning it up, detecting its language and expanding it.
- Stage D3-C011 Day 3 · Making sense of users' queries
Google named Thai as a language that makes query understanding more complex because it does not separate words with spaces; the speaker added, hedging with 'apparently', that Thai uses spaces to separate sentences.
extendsStage D2-C320 Day 2 · Understanding what's on a pageText in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.
- Stage D3-C013 Day 3 · Making sense of users' queries
Google's query processing deliberately mirrors indexing: a query is transformed into something that can be matched against the index, and stop word removal is part of that transformation.
extendsStage D2-C321 Day 2 · Understanding what's on a pageFor languages written without spaces, such as Thai and Chinese, Google uses exactly the same word segmentation when indexing a page as when interpreting the user's query, because otherwise the query could not be matched against the index.
- Stage D3-C048 Day 3 · Making sense of users' queries
Google generally treats spellings with and without diacritics as synonyms behind the scenes, for example a German 'ü' written as 'ü', as 'ue' or left out.
extendsStage D2-C600 Day 2 · Focusing on Internationalisation and LocalisationGoogle usually understands a query word whether it is written with or without diacritics (accents).
- Stage D3-C051 Day 3 · Making sense of users' queries
Users expect content written the way they search: in some languages they search in Latin characters, in others in the local script, and Hindi users, for example, search both in Hindi and in Latin letters.
extendsStage D2-C599 Day 2 · Focusing on Internationalisation and LocalisationGoogle usually understands non-English words typed 'in English' (probably meaning romanised, Latin-letter spellings, which the speaker did not spell out); the speaker said this belongs to query interpretation, a Day 3 (serving) subject.
- Stage D3-C079 Day 3 · Making sense of users' queries
To order candidates at retrieval, Google uses signals collected during indexing, and the first two are language and country.
extendsStage D2-C649 Day 2 · Calculating (some) signalsCountry and language are among Google's most important signals and have been used since Google's early days.
- Stage D3-C080 Day 3 · Making sense of users' queries
At retrieval, Google tries to match results to the user's language wherever possible: someone searching in Spanish does not necessarily want results in Italian.
extendsStage D2-C654 Day 2 · Calculating (some) signalsIn ranking, country and language signals help Google serve users the right content for their country and language.
- Stage D3-C082 Day 3 · Making sense of users' queries
Country is the second retrieval signal: a user searching from Switzerland wants cheese from Switzerland, not from Germany, and a user in Spain is poorly served by results targeting a South American country.
extendsStage D2-C649 Day 2 · Calculating (some) signalsCountry and language are among Google's most important signals and have been used since Google's early days.
- Stage D3-C341 Day 3 · How Search results are born
Some countries rely a lot on social proof such as reviews, Google's speaker said, recalling a point from Google's Day 2 internationalisation talk.
extendsStage D2-C614 Day 2 · Focusing on Internationalisation and LocalisationAccording to consumer data shown on a slide in the talk (source not captured), US consumers judge product quality more by user feedback and reviews, while European shoppers seem to look more at brand reputation.
- Stage D3-C695 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
The cheaper tokens become, the more AI slop is created, and Google counts AI slop as scaled content abuse.
extendsDocs D2-C611 Day 2 · Focusing on Internationalisation and LocalisationGoogle's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.
- Stage D3-C697 Day 3 · Wrapping all up: AI, Search, and making sense of everything.
Google's Search Quality Rater Guidelines point out that the tool used to create content is not the problem, but how it was used and what for.
extendsDocs D2-C611 Day 2 · Focusing on Internationalisation and LocalisationGoogle's spam policies define scaled content abuse as generating many pages mainly to manipulate rankings, with little or no value to users, no matter how they are created, and list automated translating of scraped content among the examples.
- D3-C070 Day 3 · Making sense of users' queries
Google's summary slide on query understanding noted that some languages do not use spaces between words, which complicates query understanding.
repeatsStage D2-C320 Day 2 · Understanding what's on a pageText in languages written without spaces, such as Thai and Chinese, would end up in the index as long strings that might never be searched for, so Google segments it into words with statistical models built from other web content in that language.