AI translation is now good enough to open new markets on a small budget. It is also good enough to publish a confident, fluent, completely wrong version of your brand into a language nobody at your company can read.
Machine translation has got dramatically better, and the research community that runs the benchmarks still titled its 2024 findings paper “the LLM era is here but MT is not solved yet.” That tension is the whole job. This guide sorts marketing content by the cost of an undetected error, explains the real difference between translation, localisation and transcreation, gives you an honest tool comparison, and covers the multilingual SEO rules Google actually documents, including the hreflang mistake that quietly makes the whole thing pointless.
You have seen the statistic. CSA Research surveyed 8,709 consumers across 29 countries and found that 76% prefer to buy products with information in their own language, and 40% will never buy from a website in another language.[1] Every localisation vendor on earth leads with it, usually without the sample size and usually without the year.
It is a real number and it is worth acting on. But I am going to give you the parts of that same study nobody quotes, because they change how you should spend the budget.
From the identical survey: 66% of respondents said they would still choose the cheaper product even without information in their own language, and 69% said they would choose a strong global brand over a product with local-language information.[1] Language preference is real, and it loses to price and to brand strength.
So translation is not a growth strategy on its own. It is a friction remover. If people already want what you sell and cannot buy it comfortably, translating fixes something real. If nobody in that market has heard of you, translating your site gets you a beautifully localised page that nobody visits. I have watched a client spend four months and a serious budget localising into five languages before running a single paid test in any of those markets. Two of the five had no meaningful demand at all. We could have found that out in three weeks with translated ad copy and a landing page.
The other number worth having: W3Techs data shows English accounts for 49.5% of websites whose content language can be identified.[2] Slightly over half the web is already not in English. That is the opportunity framed properly, as a competitive gap rather than a moral obligation.
These get used interchangeably in briefs and then everyone is disappointed by what comes back. The difference is simply how much licence you are giving someone to change the original.
Translation converts meaning from one language to another. The source is authoritative. Fidelity is the job. Spec sheets, shipping policies, technical FAQs.
Localisation is translation plus everything around the words. Date formats, because 11/07 means July 11 in the US and 11 July in the UK. Currency and price psychology, because .99 pricing is not universally persuasive. Units, address formats, payment methods that people actually use in that market, imagery, colour associations, name order, and form fields, because plenty of countries have no concept of “state.” And text expansion: German commonly runs 20 to 35% longer than English, which breaks your buttons and your navigation. Localisation happens at the level of locale, not language. Brazilian and European Portuguese are two different jobs. So are Mexican and Castilian Spanish.
Transcreation means the source is a brief, not a text. A copywriter recreates the effect of the original with freedom to change the words entirely. Taglines, campaign concepts, headlines, anything built on a pun. It is priced by the hour or by the concept, never by the word, and it comes with a rationale document explaining the choices.
The short version: translation changes the words, localisation changes everything around the words, transcreation changes the words on purpose.
Where AI sits across those three is not what most people assume. It is strongest on translation, decent on the mechanical parts of localisation, and genuinely useful on transcreation drafts in a way dedicated translation engines are not, because you can brief a general model on tone, audience and formality the way you would brief a writer. You can tell Claude or ChatGPT “keep the joke, change the reference, we need tú not usted, this must fit 40 characters.” DeepL cannot take that brief. What AI cannot do is be the last person who reads it.
Stop asking whether the AI is good enough. Ask CSA Research’s question instead, which is what their 2025 trends report frames as treating quality as risk management: decide where to invest “based on the potential harm from errors to customers, brand, reputation, and financial or legal status.”[3]
That reframing does something useful. It means the same tool can be perfectly appropriate for one page and reckless on the next.
| Tier | Content | Process | Why |
|---|---|---|---|
| Green | Support and knowledge base articles, internal comms, user reviews and UGC, product attributes and filters, gisting inbound enquiries | Raw AI output, spot-check only | The alternative is usually nothing at all, and errors are cheap to correct |
| Amber | Blog posts, SEO landing pages, email sequences, social posts, hero product descriptions, video subtitles, case studies | AI first draft, native-speaker review before publishing | Errors are visible, embarrassing and indexed, but rarely legally dangerous |
| Red | Taglines and campaign concepts, paid ad copy, legal terms and privacy policies, regulated claims, pricing and contracts, checkout and refund flows, crisis comms | Human translator or transcreator, always | Money, regulators, or your brand for years |
A risk-based sorting of marketing content types, following CSA Research’s guidance to treat translation quality as risk management rather than a fixed quality bar. [3]
One thing to layer on top: the risk tier shifts by language. A study evaluating GPT-4 against professional translators found it performed comparably to junior translators on total errors but lagged mid-level and senior ones, and that its capability “gradually weaken[ed] from resource-rich to resource-poor directions.”[4] In plain terms, your amber-tier blog post is amber in French and closer to red in Vietnamese. Same content, same tool, different risk.
Two rules I would treat as absolute. Never machine-translate something that was already machine-translated, because the errors compound and become undetectable. And never publish AI translation into a language nobody at your company or in your network can read. If you cannot detect a failure, you have not managed the risk. You have hidden it.
The honest headline first. The Conference on Machine Translation, which is where the field’s benchmarks are actually run, titled its 2024 findings paper “The LLM Era Is Here but MT Is Not Solved Yet.”[5] That is the researchers who evaluate eight large language models and four commercial translation providers with professional human annotators. If they will not declare it solved, no vendor blog should convince you otherwise.
| Tool | Best at | Weak at | Use it when |
|---|---|---|---|
| DeepL | Highest-quality raw output for major European pairs, glossaries, document formatting | Narrower depth outside its strong languages; no workflow layer | You need one excellent translation into German, French, Dutch or Spanish |
| Google Translate | Enormous language coverage, free, instant, great for comprehension | Flat literal tone; quality drops sharply on rare pairs; free-tier data handling | Reading inbound content, or a language nothing else supports |
| ChatGPT / Claude | Tone, formality, brand voice, transcreation drafts, explaining its own choices | Non-deterministic, drifts over long documents, no translation memory | Adapting a campaign line, or sense-checking another engine’s output |
| Weglot | Getting a small site multilingual in days, hreflang handled for you | Pricing scales with words and languages; default is raw MT across everything | You have no developer and need two to five languages live this month |
| Smartling / Lokalise | Translation memory, glossaries, workflow routing, in-context review | Overkill and over-priced below real volume; implementation is a project | You are running continuous localisation across many languages |
Practical positioning of the main tools marketers use. Quality claims made by individual vendors about their own products are excluded here; see the caution below.
A word on vendor benchmarks. DeepL publishes blind-test results claiming language experts preferred its translations 1.3 times more often than Google Translate and 1.7 times more often than ChatGPT-4.[6] That may well be true, and DeepL’s European output genuinely is excellent in my experience. But it is DeepL testing DeepL, and the page does not publish the protocol, sample size or how annotators were selected. Use it as “what DeepL claims,” not as evidence.
My actual stack for a mid-sized marketing team: DeepL or Google for the raw engine pass depending on language, ChatGPT or Claude for anything where tone matters and for reviewing the other engine’s output, Weglot if the website needs to go multilingual before the next quarter, and a paid native-speaker reviewer per language. That last line is the one people cut. It is also the one that stops you shipping something humiliating.
On budget: the language services market is estimated somewhere between roughly $31.7 billion by Slator’s narrower definition and $75.7 billion by Nimdzi’s broader one for 2025.[7][8] I quote both deliberately. When you see one of those numbers alone in a pitch deck, someone has picked the definition that suits their argument.
Five steps, and the order matters more than the tooling:
Step one is on the list because most localisation budgets get committed before anyone checks whether the demand is real.
Step two is the one that gets skipped and it is the cheapest quality win available. Before any translation happens, write a one-page glossary: product names that must never be translated, your preferred term for each core concept, the formality register per language (tú or usted, du or Sie), and anything legal that must be worded exactly. DeepL, Smartling and Lokalise all accept glossaries directly. With ChatGPT or Claude you paste it into the prompt. Without it, you will get four different translations of your own product category across one website, which is exactly what happened to a client of mine who now has “smart scheduling” rendered three different ways in German across their pricing, homepage and help centre.
The prompt shape that works for the amber tier:
“Translate the following marketing copy from English into [locale, for example de-DE]. Audience: [who]. Register: [formal/informal, and the pronoun to use]. Keep these terms untranslated: [list]. Preserve the structure and any character limits noted in brackets. Where a phrase relies on English wordplay or an English cultural reference, do not translate it literally: give me your best adaptation and a one-line note explaining what you changed and why. At the end, list anything you were unsure about.”
That last instruction is the important one. A model asked simply to translate will smooth over its own uncertainty. A model asked to report uncertainty hands you a review list, which is what turns an unreviewable wall of foreign-language text into a twenty-minute check. It is the same principle behind keeping brand voice consistent at scale: constrain the output, then ask the tool to flag where it strained against the constraints.
You can do all of the above well and still get no traffic, because the technical layer is where multilingual sites quietly fail. Four things Google documents explicitly.
Google’s documentation is unambiguous: “Each language version must list itself as well as all other language versions,” and “if two pages don’t both point to each other, the tags will be ignored.”[9] The identical block of link tags goes on every language version, including a self-reference. Missing return links is the single most common implementation failure. Use absolute URLs, and remember that language code comes first: de is German, de-be is German for Belgium, and be on its own means Belarusian, not Belgium. Google names that exact trap in its own docs.
This surprises people every time. Google states plainly that it “uses the visible content of your page to determine its language” and does not use code-level signals such as lang attributes or the URL.[10] Which means a page tagged as French that is still 80% English navigation and boilerplate will not be treated as French. Google specifically warns that “translating only the boilerplate text of your pages while keeping the bulk of your content in a single language” creates a bad experience.[10]
Google advises against automatically redirecting users between language versions because “these redirections could prevent users (and search engines) from viewing all the versions of your site.”[10] There is a mechanical reason this breaks indexing: Googlebot usually crawls from the USA and sends requests without an Accept-Language header. Your language-sniffing redirect therefore sends the crawler to your English page every time. Offer a visible language switcher and a suggestion banner instead.
Google’s spam policies name translation directly within scaled content abuse: generating many pages “through automated transformations like synonymizing, translating, or other obfuscation techniques, where little value is provided to users.”[11]
Read that carefully, because it is widely misquoted. Google is not banning machine translation. The test is user value and scale, not production method. A genuinely useful translated support library is fine. Ten thousand auto-generated location-plus-language pages is not. If you want the wider context on how Google evaluates this kind of content, our guides to using AI for SEO and local SEO cover the same distinction.
Translating the website before testing the market. Covered above, and it is the most expensive mistake on this list by a wide margin. Translate ads and one landing page. Spend money. See what happens. Then commit.
Translating without localising the conversion path. Your German page is beautiful and checkout offers only credit cards, when a large share of that market expects other payment methods. Your Dutch page converts badly because the form demands a state. This is where the money leaks, and it is invisible if you only review the copy.
No glossary, so your own product name mutates. Cheap to prevent, tedious to fix afterwards.
Nobody who reads the language ever looks at it. If your entire Japanese presence has never been read by a Japanese speaker, you do not have a Japanese presence. You have a liability that is currently silent.
Letting the plugin publish everything on install. Tools like Weglot default to machine-translating your whole site the moment you switch them on. That is genuinely useful for a support centre and genuinely risky for your legal terms and pricing page. Set the tiers before you set the plugin live, not after.
Something worth knowing about where this is heading: CSA Research notes that many enterprises are already labelling AI translation explicitly as a risk-avoidance strategy, and that standards work is under way to formally distinguish unreviewed AI translation from professionally reviewed translation.[3] Disclosure is becoming normal. If you are treating “nobody will notice it was machine translated” as part of your plan, that assumption has a shelf life.
Here is what I would actually do this month if I were you. Pick your single highest-traffic support article and one product page. Translate both into the one market where you already see traffic you cannot serve. Get a native speaker to spend an hour on the product page and ten minutes on the support article. Put hreflang on properly. Then look at what happens over six weeks before you spend another pound. That is a real test, it costs very little, and it tells you more than any vendor deck will.
For some of it, yes. The honest position comes from the research community that runs the benchmarks: the 2024 Conference on Machine Translation titled its findings paper “The LLM Era Is Here but MT Is Not Solved Yet.”[5] A separate study found GPT-4 performed comparably to junior translators but lagged mid-level and senior ones, with quality weakening on lower-resource languages.[4] Use it freely for support content and internal material, use it as a first draft for blogs and product pages, and keep humans on taglines, ads, legal copy and anything regulated.
Not for being machine translated. Google’s spam policies list translation within scaled content abuse only where many pages are generated “where little value is provided to users.”[11] The test is user value and scale, not the production method. A useful translated help centre is fine. Thousands of thin auto-generated pages are not. The bigger practical risk is technical: Google determines page language from visible content, so a page with translated headings but English body copy will not be treated as that language at all.[10]
Translation converts the words. Localisation adapts everything around them: dates, currency, units, payment methods, address and form fields, imagery, colour associations and text length, since German commonly runs 20 to 35% longer than English and breaks layouts. Localisation works at the level of locale rather than language, so Brazilian and European Portuguese are separate jobs. Transcreation goes further again, recreating the intended effect with freedom to change the words entirely, which is what taglines and campaign lines need.
The ones where you already see demand you cannot serve. Check your analytics for traffic from non-English-speaking markets, look at where support enquiries arrive in other languages, and run a small paid test with translated ad copy and one landing page before committing to a full site. CSA Research found 66% of consumers would still choose the cheaper product without local-language information and 69% would choose a strong global brand over one with local-language content, so language alone does not create demand.[1]
For anything customer-facing and consequential, yes, though the role shifts from translating to reviewing and post-editing, which is faster and cheaper. Budget for one native-speaker reviewer per language you publish in. The rule I would treat as non-negotiable is never publishing into a language nobody in your company or network can read: if you cannot detect a failure, you have not managed the risk, you have concealed it. CSA Research frames this well by treating translation quality as risk management based on the potential harm from errors.[3]
This guide draws on consumer research from CSA Research, web language data from W3Techs, peer-reviewed machine translation evaluation from the Conference on Machine Translation and an academic study comparing GPT-4 with professional translators, plus Google’s own published documentation on multilingual sites and spam policies. Vendor-published quality claims are labelled as such in the text. Market size figures are quoted from two firms with different definitions, deliberately.