How to Check AI Translations Before Customers See Them

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Check AI Translations Before Customers See Them.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Check AI Translations Before Customers See Them.

Check AI translations in three layers before publishing: run mechanical checks anyone can do (numbers, dates, prices, links, placeholders, codes, length), have a second AI review the translation against the source using a fixed error list, then have a fluent reviewer read customer-facing copy for meaning and tone. Score errors so "ready to publish" is a threshold, not a feeling.

The layer people skip is the one that matters most for marketing. Back-translation and AI review catch wrong meanings. They don't catch a slogan that's accurate but sounds machine-made, or an email that switches between formal and informal "you" halfway down. Only someone who reads the language natively notices those, and customers always do.

Follow me on Instagram@sagnikteaches

Sort the content by what a mistake would cost

Not everything needs all three layers. Consider an illustrative marketing agency localising an outdoor clothing retailer's autumn campaign into Spanish, Dutch and Polish. Before anything is translated, each asset goes into a tier:

Connect on LinkedInSagnik Bhattacharya
AssetTierChecks
Returns policy, sale terms, sizing guide1: legal, money or safetyProfessional translator, or AI plus full revision by one; never AI alone
Landing page and three campaign emails2: persuasive, publicMechanical checks, second-AI review, native reviewer
Twelve ads and social posts2: short but paidSame as above; length and platform limits checked twice
Replies to social comments during the campaign3: short, reversibleMechanical checks and a weekly spot check by the reviewer
Internal briefing notes for the client's staff3: internalAI translation, read once for obvious errors

Tier 1 is where the decision about machine translation at all belongs; whether machine translation is good enough for client-facing documents covers that judgement. The rest of this tutorial is about tiers 2 and 3, where AI translation is reasonable and the checking decides whether it embarrasses anyone.

Subscribe on YouTube@codingliquids

Checks anyone can run without speaking the language

A surprising share of the errors that reach customers have nothing to do with fluency. They're broken mechanics, and a non-speaker with the source open alongside can find them in minutes. Work through this list for every asset:

  • Numbers and percentages. Every figure in the source appears in the translation, and no new ones appear.
  • Prices and number formats. Many languages use a decimal comma, so 59.99 becomes 59,99. Check the format your market expects, and that the price itself didn't change.
  • Dates. Write dates out in words in the source ("12 October") so they can't flip between day-month and month-day readings.
  • Placeholders and merge tags. Anything in curly brackets or your email platform's tag format must come through untouched.
  • Codes, product names and URLs. Discount codes, SKUs, brand names and links stay exactly as they are.
  • Length. Subject lines, buttons and ad headlines have limits; translations often run longer than the source.
  • Formality. Pick formal or informal "you" per language and check it's consistent. You can spot a switch by searching for both forms even without reading the sentence.

Here is the mechanical check on the agency's first Spanish email. The source:

Hi {first_name}, our autumn sale starts today. Get 20% off all waterproof jackets until 12 October. Use code AUTUMN20 at checkout. Free returns within 30 days.

The AI translation, illustrative:

Hola {nombre}, nuestras rebajas de otoño empiezan hoy. Consigue un 20% de descuento en todas las chaquetas impermeables hasta el 12 de octubre. Usa el código OTOÑO20 al pagar. Devoluciones gratuitas en 30 días.

Two mechanical failures, both found without reading Spanish. The merge tag {first_name} became {nombre}, so the email platform would have sent the literal text "{nombre}" to every subscriber. And the discount code AUTUMN20 was translated to OTOÑO20, a code that doesn't exist at checkout. Everything else in the mechanical list passes: 20%, 12 October and 30 days all survived. The fix goes into the translation prompt so it doesn't recur:

Translate the text below into Spanish for customers of an outdoor
clothing shop. Use informal "tú" throughout.
Never translate or change:
- anything inside {curly brackets}
- discount codes, product codes and anything in CAPITALS with digits
- the brand name, product names and URLs
Keep every number exactly as in the source. If a phrase has no
natural equivalent, translate the meaning and mark it with [ADAPTED].

If the same terms recur across campaigns, move them into a glossary the translation tool enforces. DeepL's glossary feature, for instance, fixes how specific terms translate and adjusts them grammatically rather than doing a blind find-and-replace; how many glossaries and entries you get depends on the plan. Building a product glossary AI must use in every draft covers what to put in one.

Some mechanics only look fine. Polish needs different noun forms for 1, for 2 to 4, and for 5 or more: "1 produkt", "2 produkty", "5 produktów". A basket reminder built as "Masz {n} produkty w koszyku" is correct for 2, 3 and 4 items and wrong for everything else. Unless your email platform supports plural rules per language, rephrase the template so the number sits on its own: "Liczba produktów w koszyku: {n}", which reads correctly for any number. Ask your reviewer to check every template that contains a number variable, in every language.

Back-translation: useful for meaning, blind to tone

Back-translation means translating the target text back into English, ideally with a different tool, and comparing it with the source. It's cheap and it does catch real problems: a dropped sentence, a negation that vanished, a "one size up" that became "one size down". Run it on anything with instructions or conditions in it.

What it can't do is tell you how the translation sounds. The campaign's tagline shows the gap:

  • Source: "Weatherproof. Worry-proof."
  • AI translation: "A prueba de clima. A prueba de preocupaciones."
  • Back-translation: "Weatherproof. Worry-proof."

A perfect round trip, and a tagline a native reader would find stilted: it's a word-for-word calque of an English pun that doesn't work the same way. The Spanish reviewer replaced it with a line built on the idea (rain or shine, you're covered) rather than the words. Back-translation would have signed off the calque with full marks. Slogans, puns, headlines and calls to action need a human who reads the language natively, every time.

A second AI as reviewer, with a fixed error list

A second model, ideally a different one from the model that translated, makes a useful reviewer as long as you constrain what it looks for. Free-form "is this translation good?" prompts produce vague praise or dozens of stylistic suggestions. A fixed list with severities produces something you can act on:

You are reviewing a Dutch translation of English marketing copy.
Compare the translation with the source, sentence by sentence.
Report only these error types:
ACCURACY (wrong meaning, addition, omission)
NUMBERS (any changed figure, date, price or code)
TERMINOLOGY (glossary terms below not used as specified)
REGISTER (mixing formal "u" and informal "je"; the brand uses "je")
LOCALE (number formats, punctuation conventions)
For each error: quote source and translation, give the category,
rate it minor, major or critical, and suggest a fix.
Don't report stylistic preferences.
[glossary] [source] [translation]

An illustrative review of the Dutch delivery email:

1. LOCALE, minor. "Bestel voor 15:00 uur" - "voor" can read as "for";
   write "vóór 15.00 uur" to make "before" unambiguous.
2. REGISTER, major. Line 2 uses "je bestelling", line 5 "Heeft u
   vragen?". Change line 5 to "Heb je vragen?".
3. ACCURACY, minor. "Retourneren is gratis voor 30 dagen" is a
   literal rendering; natural: "Je kunt 30 dagen gratis retourneren."
4. NUMBERS. None found.
5. LOCALE, minor. Price "59.99" should use a decimal comma: "59,99".

Items 1, 2, 3 and 5 are genuine and quick to fix. What you'd still watch for: a second model sometimes reports a preference as an error (suggesting a synonym that's no better), and it can miss a problem both models share. That's why this layer sits between the mechanical checks and the human reviewer, not in place of either. Asking the same model that translated to review its own work is the weakest version, because it tends to approve its own choices.

Briefing a native reviewer so they check the right things

A reviewer given only "please check this" will either skim or rewrite everything to their taste. A one-page brief gets you a focused review and makes the fee go further. The agency's brief for the Polish reviewer, filled in:

REVIEW BRIEF: autumn campaign, Polish          Due: Thursday 5pm
Client:     outdoor clothing retailer, online only
Readers:    existing customers, mostly 25-45, bought before
Voice:      friendly, direct, informal address throughout
Assets:     landing page (650 words), 3 emails (250 each),
            12 ads and posts (about 30 words each)
Please compare against the English (attached side by side) and mark
each issue as minor, major or critical:
  critical = could cost money, break the law, or cause offence
  major    = wrong or misleading meaning, wrong register, broken
             template, anything a customer would notice at once
  minor    = unnatural but understandable
Do not change: brand and product names, codes, {placeholders}, URLs
Glossary:   attached (14 terms)
Length:     subject lines max 45 characters; ad headlines max 40
Ignore:     personal style preferences where the text is correct
Also check: every template containing {n} reads correctly for 1, 3
            and 7 items

It helps to know what you're buying. The ISO 17100 standard for translation services distinguishes revision, a bilingual comparison of the translation against the source, from review, a read of the target text alone for fitness for purpose. For marketing you want both, or at least revision with an eye on tone. A reviewer reading only the Polish can't tell you that a condition was dropped from the English. A separate standard, ISO 18587, covers full post-editing of machine translation output, which is the service to ask for if you want a professional to take AI output to publishable quality rather than just flag problems; the machine translation post-editing workflow shows what that work involves.

Tell both the AI and the reviewer which variety of the language your customers use. Spanish, Portuguese and several other languages differ noticeably between the countries that speak them, in vocabulary, in how formal marketing usually sounds, and even in everyday words for clothing. Name the market in the translation prompt and in the brief; a reviewer from a different market may "correct" words that were right for your customers.

Bilingual colleagues can be excellent reviewers, but check two things first. Are they fluent in the variety of the language your customers use? And are they comfortable criticising copy, including copy that a client or a senior colleague approved in English?

Score errors so "good enough" is a number

Without a score, every review ends in a discussion about whether the remaining issues matter. The translation industry's MQM framework (Multidimensional Quality Metrics) offers a simple fix: weight each error by severity, with minor errors counting 1 penalty point, major 5 and critical 25, and express the total per 1,000 words. You then set a threshold in advance. For the agency's tier 2 content the rule was: no critical errors, and no more than 10 points per 1,000 words.

The Spanish landing page, 800 words, first review:

  • Major: "order one size up" rendered as "pide una talla menos", which means one size down. That would have driven returns. 5 points.
  • Major: the page switched from "tú" to "usted" in the FAQ block. 5 points.
  • Minor: six unnatural phrasings, including the tagline. 6 points.

That's 16 points in 800 words, or 20 per 1,000: over the threshold, so it went back for fixes and a second look at the changed sentences only. The second pass found two remaining minor issues, 2 points in 800 words or 2.5 per 1,000, and the page was cleared. The score also gives you a way to compare languages and tools over time: if Dutch consistently scores 3 and Polish 14, you know where to spend reviewer hours.

Check it where customers will see it

A translation can pass every text check and still fail once it's placed in the page or the email. Before launch, someone opens each language version in the real channel, on a phone and a laptop, and looks for problems no document review can show:

  • Truncation. A button that says "Shop now" in English may cut off mid-word in a longer language, and the ad preview may drop the end of a headline.
  • Special characters. Letters such as ñ, ó, ą or ę should display correctly in subject lines, preview text and page titles, not as question marks or boxes.
  • Untranslated leftovers. Text inside images, alt text, form error messages, cookie banners and the email footer are easy to miss because they weren't in the document that got translated.
  • Links. Each language version's links go to the matching language page, not back to the English one.
  • The language switcher and page titles. The right language is shown for each version, and the browser tab title is translated too.

A realistic example from the campaign: the Polish emails read perfectly in the review document, but the test send showed the subject line's "ą" as a question mark in one mail app, because the subject had been pasted into the platform through a spreadsheet that mangled the encoding. Retyping the subject directly into the platform fixed it. Ten minutes of test sends per language catches this kind of thing; skipping them means 8,000 customers see it first.

Mistakes that still reach customers

Even with checks in place, some errors recur often enough to deserve a named line in your checklist. Each of these is realistic, and each would pass a quick read:

  • The translated merge tag. "{nombre}" or a translated platform tag goes out to the whole list as literal text. Found by the mechanical check; prevented by the do-not-translate line in the prompt.
  • The translated discount code. Customers type a code that doesn't exist, then email support. Worse than no code at all.
  • The formality flip. Written by different prompts on different days, the emails in one sequence address the customer in different registers.
  • The unchecked last-minute edit. The English subject line changes after review, someone re-translates just that line, and nobody reviews it. Any change after review goes back to the reviewer, even one line.
  • The critical number. "Free returns within 30 days" becomes "within 3 days" in one language. That's a critical error in MQM terms and a customer-service problem the day it lands.

If the translated content also feeds a support inbox, the same principles apply to replies; multilingual customer support with AI translation covers the reply flow. For terms that must never vary between languages, see how translators keep terminology consistent with AI.

The autumn campaign in three languages, costed

Here's the illustrative campaign end to end. Per language: one 650-word landing page, three 250-word emails and twelve 30-word ads and posts, about 1,760 words. Three languages make about 5,280 words.

  • AI translation: covered by the agency's existing AI subscription. About 20 minutes per language including prompt set-up.
  • Mechanical checks by an account executive: about 30 minutes per language. Found in this campaign: the merge tag, the discount code, two over-length subject lines.
  • Second-AI review: about 15 minutes per language to run and triage.
  • Native reviewer: say $45 an hour and about 3 hours per language, including the re-check of fixes: $135 per language, $405 for three.
  • Agency time: roughly 65 minutes per language, 3 hours 15 minutes in all. At $50 an hour, about $163.

Total: about $570 of checking for a three-language campaign. Set that against a single miss. The Polish email list for this client has 8,000 subscribers; if the translated code had gone out, support would have spent a day answering "your code doesn't work" and the client would have extended the sale or honoured the discount manually. And the tagline, left as a calque, would have sat on every ad in the campaign, telling Spanish-speaking customers that nobody who speaks their language had looked at it.

Further reads

Sources: MQM (Multidimensional Quality Metrics) scoring model and severity definitions; ISO 17100 (translation services: revision and review) and ISO 18587 (post-editing of machine translation output) as described by ISO and standards summaries; DeepL glossary feature pages.

Launching in other languages with AI translation?

On a 1:1 call we'll sort your content by risk, set up the checks your team can run without speaking the language, and decide where a native reviewer earns their fee.

Book a 1:1 call with me