Skip to content
AdsChatGPT now shows ads.Get early access

Playbooks/August 21, 2026

Multilingual GEO: Why Your AI Visibility Does Not Translate

Robin Pautigny

Robin Pautigny

Co-founder, Refine

Multilingual GEO: Why Your AI Visibility Does Not Translate

Summary

AI visibility does not carry across languages. A brand that dominates English-language answers in ChatGPT or Perplexity is frequently invisible when the same question is asked in French, German or Spanish, because the model retrieves from a different pool of sources and reasons over a different set of competitors. Fixing it requires per-language prompt tracking, genuinely localised content rather than machine translation, and local third-party corroboration. Expect four to twelve weeks before the change is measurable.

Ask ChatGPT for the best project management tool and you get one shortlist. Ask it for le meilleur outil de gestion de projet and you often get a different one. Same model, same day, same underlying question - different brands named. This is the single most under-measured gap in GEO, and if you sell in more than one market it is almost certainly costing you answers you assume you already own.

The short answer

AI visibility does not transfer between languages. Language models retrieve and rank sources in the language of the prompt, so your English-language authority is largely invisible to a French or German query. Winning a second market means tracking prompts in that language, publishing content written natively rather than translated, and earning mentions on sources that market actually trusts.

Why Visibility Breaks at the Language Border

The intuition most marketing teams carry over from SEO is that a strong domain travels. In classic search that is broadly true: authority accrues to a domain, hreflang stitches your language versions together, and a well-linked site tends to rank reasonably everywhere it has content.

Generative engines do not work that way. When a model answers a question, it is doing two things: recalling what it absorbed during training, and - for anything grounded in live retrieval - fetching a handful of documents to cite. Both stages are language-filtered. A retrieval layer answering a French prompt overwhelmingly surfaces French-language documents. A model reasoning about a French prompt leans on the French-language portion of its training data, which is a fraction of the English portion and weighted towards entirely different publishers.

The practical consequence is that your competitive set changes when the language changes. In English you may be benchmarked against three global players you know well. In German you are being benchmarked against a local incumbent that has fifteen years of trade-press coverage in that language and no international presence at all. You are not losing to the competitor you prepared for.

What Actually Changes When the Prompt Language Changes

It helps to be precise about which variables move, because each one has a different fix.

  • The source pool. Retrieval returns documents in the prompt language. Your English documentation, however good, is rarely the citation for a Spanish answer.
  • The competitive set. Local players with deep native coverage appear in answers where they have no English-language footprint whatsoever.
  • The vocabulary. Categories are named differently. What English calls AI visibility tracking is discussed in French as suivi de visibilite IA or referencement dans les IA - and a query using the second term retrieves a different corpus than the first.
  • The intent mix. The same market can be at a different stage of adoption in each language. English prompts may skew towards comparison and pricing; a less mature market may still be asking definitional questions.
  • The trusted intermediaries. English answers lean on Reddit, G2 and large US publications. French answers lean on French trade media and specialist blogs; German answers on Fachpresse and local community forums. The intermediaries are not interchangeable.
  • The sentiment. Descriptions of your brand are reconstructed from local sources, so an inaccuracy that only exists in one language can persist there long after you have corrected the English version.

Notice that only one of these is a translation problem. The rest are distribution and positioning problems that translation cannot touch.

How to Audit Your Visibility Market by Market

You cannot fix a gap you have not sized. The audit is straightforward but has to be done per language, not once with the results extrapolated.

  • Build a native prompt set per market. Do not translate your English prompts. Write them the way a buyer in that market would actually type them, including the local category name and any local regulatory or workflow framing. Twenty to forty prompts per language is enough to start.
  • Run each set across the engines that matter locally. Coverage differs: Mistral carries real weight in France, Copilot in enterprise-heavy German-speaking markets, Perplexity in technical audiences. Do not assume ChatGPT alone represents the market.
  • Record mention rate, average position and citation source for each prompt, per language. The citation source column is the one that pays for the exercise - it tells you which domains the model trusts in that market.
  • Extract the competitive set per language. List every brand named more than twice. You will usually find at least one name you were not tracking.
  • Check accuracy separately. Have a native speaker read how the model describes you. Stale pricing, a wrong category or an outdated feature list is common and quietly disqualifying.
  • Score the gap. Compare mention rate in your strongest language against each other language. A brand at 60 percent in English and 8 percent in German has a distribution problem, not a product problem.

Track languages as separate markets, not one dashboard

The most common measurement mistake is averaging across languages, which hides the exact gap you are trying to find. Refine tracks prompt sets per language across ChatGPT, Gemini, Perplexity, Claude, Copilot and Mistral, so mention rate, share of voice, competitive set and cited sources are reported market by market. If your French mention rate is a third of your English one, that shows up as a number rather than a suspicion - and you can see which domains are being cited instead of yours.

Writing Content That Earns Citations in a Second Language

Machine translation of an existing English article is the default move and the weakest one. It produces text that is technically correct and contextually foreign: the examples are American, the currency is wrong, the regulatory references do not apply, and the category is named with a calque no local buyer uses. Models can retrieve it, but it competes badly against a page written by someone who understands the market.

Localisation that actually earns citations means starting from the local query rather than the English article. Take the native prompt set you built during the audit and write directly to those questions. Use the vocabulary that appears in the prompts, not the vocabulary your translator preferred. Replace examples with local companies, local price points and local regulation - GDPR framing lands differently in Germany than in the UK, and both differ from how a US-authored page treats it.

The extractability rules that govern GEO in English apply unchanged in every other language. Answer the question in the first two sentences. Keep paragraphs short enough to lift whole. Attach numbers and dates to claims. Use headings that mirror the question being asked. A model looking for a clean French sentence to quote will take one from whichever page makes it easiest.

On the technical side, keep the hygiene tight: correct lang attributes, reciprocal hreflang, one canonical per language version, and no client-side language switching that leaves crawlers on a default locale. None of this wins you citations on its own, but any of it broken will stop a correctly written page from ever being retrieved.

Local Corroboration Is the Signal Most Teams Miss

Models are conservative about naming brands. They are far more comfortable recommending a company that several independent sources agree exists, works and serves the described use case. That agreement has to be found in the prompt language, and this is where most international GEO programmes stall.

Your G2 profile, your English case studies and your coverage in a major US publication are close to useless as corroboration for a Spanish-language answer. What counts is the local equivalent: the trade publication that market reads, the review platform it uses, the community where practitioners actually argue about tools in their own language.

  • Get listed in local directories and comparison sites - they are cited far more often than their traffic numbers suggest.
  • Earn coverage in native-language trade media rather than syndicating an English press release.
  • Encourage reviews written in the local language on the platforms that market trusts, which are often not the ones you use for English.
  • Participate honestly in local communities and forums. Reddit dominates English answers; other markets have their own equivalents that carry the same weight locally.
  • Publish at least one substantive comparison page per language, naming the local competitive set you found in the audit rather than the global one.
  • Give local partners, integrators and resellers something accurate to publish about you in their own language.

What to Measure and How Long It Takes

Report the same four numbers per language and never blend them: mention rate across your prompt set, average position when mentioned, share of voice against the local competitive set, and the proportion of answers where the description of you is factually correct. A fifth column - which domains were cited instead of you - is what turns the report into a work plan.

On timing, be realistic with whoever is funding this. Retrieval-grounded engines such as Perplexity and Google AI Mode can reflect new content within days to a few weeks. Answers that lean on training recall move on a much slower cycle, often a quarter or more. Corroboration is the slowest input of all, because you are waiting on other people to publish. Four to twelve weeks before a clear trend appears is a fair expectation, and the sequence that works is audit, then content, then corroboration - measured continuously rather than checked once at the end.

The reason this is worth doing now is that most of your competitors are still measuring one language and assuming the rest follow. In a second-language market the answer is being generated whether or not you have optimised for it, and the shortlist is currently much shorter than the one you are fighting over in English.

Short on time? Have an assistant summarise this page for you.