Machine Translation Post-Editing: A Translator's Workflow

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Machine Translation Post-Editing: A Translator's Workflow.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Machine Translation Post-Editing: A Translator's Workflow.

Post-edit efficiently by deciding the level before you start: light post-editing, full post-editing or fresh translation, based on the text's risk and a sample of the raw output. Load the termbase and translation memory before machine translation runs, then edit in three passes: meaning against the source, target fluency, then automated QA. Change only what is wrong.

The biggest time leak isn't bad machine output. It's over-editing: rewriting acceptable segments because they aren't how you would have phrased them. The biggest risk runs the other way. Modern engines produce such fluent text that meaning errors, a dropped "not", a missing "may contain", a number that changed, read perfectly well. Post-editing is a different reading discipline from revising a colleague's work: you read the source first, every time, and you judge the target against it rather than against your own taste.

Follow me on Instagram@sagnikteaches

Decide the level before you touch a segment

ISO 18587:2017, the standard for post-editing machine translation output, sets its requirements for full, human post-editing and distinguishes it from light post-editing, which aims for text that is understandable and accurate rather than polished. That distinction is the first decision on every job. A working triage:

Connect on LinkedInSagnik Bhattacharya
ContentLevelWhy
Internal emails, supplier correspondence, gist of a reportLightReader needs meaning, not polish
Product descriptions, web pages, help articlesFullCustomers read it; brand and fluency matter
Allergen, safety, dosage or instruction textFull, with a second reviewerAn error can hurt someone
Contracts and legal termsFull at minimum; often fresh translationPrecision and liability
Slogans, marketing headlines, wordplayFresh translation (transcreation)MT output is rarely a useful start

One job can contain several levels. An illustrative farm shop's online catalogue has 300 product descriptions (full post-editing), a delivery FAQ (full), and a set of allergen statements on the labels of its own-made pies and chutneys (full, plus a second reviewer). Splitting the file by level before starting keeps the allergen lines from getting the same quick treatment as "a hearty seasonal chutney". Whether MT output is good enough for client-facing documents at all is covered in when machine translation is good enough for client-facing documents.

Subscribe on YouTube@codingliquids

Pre-flight: make the machine output better before you start

Ten minutes of set-up saves an hour of editing. Before running machine translation:

  1. Load the client's termbase and translation memory into your CAT tool, and attach the termbase to the engine where the tool allows it. Trados Studio 2026's AI Assistant can be directed to use your termbase, and memoQ AGT builds its output from your translation memories, termbases and LiveDocs. For DeepL, glossaries depend on the plan: the Team plan allows five glossaries with up to 1,000 entries per language pair, the Business plan unlimited glossaries. The terminology side is covered in depth in keeping terminology consistent with AI.
  2. Protect what mustn't change: tags, placeholders, product codes, prices and units. Lock them or mark them as non-translatable so the engine can't reformat them.
  3. Sample before you quote or schedule. Post-edit 300 to 500 words of representative text and time yourself. That sample sets your rate and deadline, as below.
  4. Use quality estimation carefully. Some tools score each machine-translated segment. Phrase's Quality Performance Score runs from 0 to 100, and project managers can set a threshold above which segments are locked as good enough; Trados Studio 2026 includes machine translation quality estimation. Skip high-scoring segments only if the client has agreed to that way of working, and still spot-check a sample.

Picture how the scoring shortcut goes wrong. A threshold locks every segment scoring above 90, and one of the locked segments is an allergen line where "may contain traces of" came out as "contains". The output was fluent and well-formed, so it scored well. A score predicts quality; it doesn't compare meaning against the source for you. If you use locking at all, filter safety, allergen, legal and numeric segments into a separate file first so none of them can be locked.

Three passes, each with one job

Trying to fix meaning, style and formatting in one pass is slow and misses things. Split the work:

Pass 1: meaning, source first

Read each source segment, then the target. Ask only: does the target say what the source says, no more, no less? Fix omissions, additions, mistranslations, wrong numbers and negation. On light post-editing, this is most of the job. Don't touch style unless it changes meaning.

Sometimes the problem is in the source. On the farm shop's labels, an illustrative cheese and onion pie description says "suitable for vegetarians", while the ingredients two lines below list lard in the pastry. Don't fix it silently in the target, and don't translate the contradiction as if nothing were wrong. Translate what the source says, flag the segment, and send a query:

Query, segment 418 (cheese and onion pie label): the source says
"suitable for vegetarians", but the ingredients include lard. I've
translated it as written and flagged the segment. Could you confirm
which is correct before this goes to print?

Pass 2: target fluency (full post-editing only)

Read the target on its own, as the customer will, ideally in a preview rather than segment by segment. Fix awkward word order, inconsistent register (formal and informal address mixed), sentences that are correct but unnatural, and repetition across segments that the engine couldn't see.

The line between fixing and over-editing is easiest to see on one segment. Source: "Our apple chutney is slow-cooked in small batches and pairs well with a mature hard cheese." Illustrative MT output, back-translated: "Our apple chutney is cooked slowly in small batches and goes well with a mature hard cheese, you'll see." The correct full post-edit is one move: delete the added "you'll see", which isn't in the source and slips into an informal register, then confirm. An over-edit recasts the whole sentence into the translator's preferred rhythm ("Slow-cooked in small batches, our apple chutney is the perfect partner for…") and adds "perfect", a claim the client never made. The first edit takes 15 seconds, the second two minutes, and only the first is what the client paid for.

Pass 3: automated QA

Run your CAT tool's QA or a dedicated checker for numbers, tags, punctuation, double spaces, untranslated segments and terminology against the termbase. Fix, then re-run until clean. Most missed errors on post-editing jobs are ones a QA tool would have caught in seconds.

An illustrative QA report on the product descriptions, and what each line deserved:

QA flagReal or false?Action
Segment 87: tag missingRealRestore the bold tag round the product name
Segment 212: "cured bacon" is not the termbase term "dry-cured bacon"RealFix it, then search the file for inflected forms the check may have missed
Segment 140: "250 g" in source, "250g" in targetDependsFollow the client's style guide on spacing before units
Segment 301: hyphen in "4-6" changed to a dashFalse for meaningAccept if the style guide allows it; ignore for this segment only
38 double-space warnings at deliberate line breaksFalseAdjust the QA profile so they stop appearing

Deal with false positives one at a time or fix the profile. Never tick "ignore all", because the real terminology miss in segment 212 sits in the same list as 38 harmless warnings.

The errors fluent machine output hides

These are the patterns to hunt for in pass 1. The examples are illustrative, with the target shown back-translated:

Error typeSourceMT output (back-translated)Fix
Dropped qualifierMay contain traces of mustardContains mustardRestore "may contain traces of"; this changes a legal allergen statement
Negation lostDo not refreeze once thawedRefreeze once thawedRestore the negation
Number reformattedServes 4-6Serves 4.6Lock ranges and numbers in pre-flight
Term driftDry-cured bacon (termbase term)Salted bacon in one segment, cured bacon in anotherApply the termbase term everywhere
Register switchYou'll love it with cheeseFormal address here, informal in the next linePick the register in the style guide and apply it throughout
Fluent inventionHand-raised pork pieHandmade pork pie with a crisp crustRemove the added claim; "crisp crust" isn't in the source

The last row is the one that newer, more fluent engines produce more often: small additions that sound right and slip past a reader who isn't checking against the source. A butcher's product range, with cure methods, cuts and allergen lines, is full of exactly these traps, which is why food and safety text gets full post-editing and a second reviewer.

Speed habits that don't cost accuracy

  • Confirm and move. If a segment is correct and acceptable at the agreed level, confirm it. Learn the confirm-and-next shortcut in your CAT tool.
  • Time-box research. Set a limit, say three minutes, for any one term. Beyond that, flag a query for the client and move on.
  • Batch similar segments. Filter the file for a recurring phrase (a delivery line, a storage instruction) and fix every instance at once, so the translation memory propagates the fix.
  • Keep a personal error log. Note the engine's repeated errors on this client's text. After a few jobs you'll know where to look first. A filled-in example is below.
  • Don't re-translate in light post-editing. If a segment is accurate but clunky, and the level is light, leave it. That's the service the client bought.

The farm shop's error log after its first job, illustrative:

CLIENT: farm shop   ENGINE: [engine + client glossary]   UPDATED: Oct 2026
- "may contain traces of" dropped to "contains" (3 times in 500 words)
- ranges such as "4-6" reformatted as decimals unless locked
- "hand-raised" gains extras such as "with a crisp crust"
- "best before" sometimes rendered with the "use by" term
- register slips to informal after a question in the source

Read it before pass 1 on the next job for that client and check those five patterns first. Delete a line once the engine or glossary stops producing that error for two jobs running, so the log stays short enough to read.

An AI chat assistant as a second reader, not the editor

A large language model can do a useful extra check after pass 1 on high-risk text, provided the client allows the text in that tool and you use a plan that doesn't train on your content (the question is covered in whether it's safe to put client documents into AI translation tools). Ask for meaning differences only:

Compare each source segment with its translation. List ONLY segments
where the meaning differs: something omitted, added, negated, or a
number, unit or quantity changed. For each, quote both and explain the
difference in one line. Do not comment on style. If there are none,
say "no meaning differences found".

An illustrative reply on 40 label segments flagged three: one real omission ("traces" missing, as in the table above, which you had already fixed in pass 1 in a different segment but missed in this one), one false alarm where the model misread an idiom, and one genuine question about whether "best before" had been rendered as "use by", which has a different meaning on food labels. You check each flag against the source yourself; the model's flags are prompts to look, not corrections to accept.

Price the work from your sample

Your pre-flight sample gives you the numbers to price with. Keep a simple log per job: words, level, minutes, and a rough note of how much you changed. Illustratively, a translator who translates new text at about 400 words an hour might log these post-editing speeds on the farm shop's text, covering passes 1 and 2 together:

  • Delivery FAQ, clean output, full post-editing: 900 words an hour.
  • Product descriptions, decent output with term drift: 650 words an hour.
  • Allergen statements, full post-editing plus careful checking: 350 words an hour, slower than translating fresh, because every line needs source comparison and a lookup.

Those numbers justify different prices within one job. A per-word discount makes sense on the FAQ and none on the allergen lines.

A quick sum shows why. Take an illustrative standard rate of $0.12 a word, which at 400 words an hour earns $48 an hour. If the client offers a flat 60% of that, $0.072 a word, across the whole file, your effective hourly rate depends entirely on the section: $64.80 an hour on the FAQ at 900 words an hour, $46.80 on the descriptions at 650, and $25.20 on the allergen lines at 350. The flat rate pays well on the easy text and about half your normal rate on the text where a mistake could hurt someone.

When a client asks for a flat post-editing rate, send the sample results:

Thanks for sending the files. I've post-edited a 400-word sample from
each section to check the machine output. The FAQ and most product
descriptions are good enough for a reduced post-editing rate. The
allergen statements need full checking line by line, which takes as
long as new translation, so I've quoted those at my standard rate.
Breakdown attached.

The farm shop catalogue, start to finish

Putting it together for the illustrative 6,000-word job: 4,500 words of product descriptions, 1,000 words of delivery FAQ and 500 words of allergen statements.

  1. Pre-flight, 40 minutes. Termbase with 60 client terms loaded and attached to the engine; prices and codes locked; three 400-word samples post-edited and timed.
  2. Quote. Reduced rate on descriptions and FAQ, standard rate on allergen lines, with the email above.
  3. Passes 1 and 2, about 9.5 hours. Descriptions at 650 words an hour (about 7 hours), the FAQ at 900 (just over an hour) and the allergen lines at 350 (about an hour and a half). The preview read for fluency happens at the end of each section rather than at the very end, while the text is fresh.
  4. Second reader and pass 3, 45 minutes. The AI meaning check on the allergen lines, then QA until clean.
  5. Reviewer, separately. A second linguist reviews the allergen lines, as agreed with the client.

That's roughly 11 hours of the translator's time, against around 15 for translating from scratch at 400 words an hour, and the saving comes almost entirely from the descriptions and FAQ. The allergen lines saved nothing and were never going to. A translator who had accepted a flat discount on the whole file would have given away the margin on exactly the lines that carried the most risk. The error log from this job goes into the next one for the same client, so the second catalogue update starts with a list of where this engine slips on this client's text. The workflow at agency level, with several linguists and clients, is set out in how small translation agencies build an AI-assisted workflow.

Post-editing questions translators ask

Should post-editing be paid per word or per hour?

Either can be fair if it's based on evidence. Per-word rates work when you've sampled the raw output and know your speed on that kind of text. Hourly rates protect you when raw quality is unknown or uneven. What isn't fair is a fixed per-word discount applied regardless of quality, because poor output can take as long as translating from scratch.

Can I use an AI chat assistant to post-edit for me?

You can use one as a second reader that flags possible meaning differences, but not as the post-editor. It can miss errors, invent problems and smooth over mistranslations. Check the client allows the text to go into that tool, use a plan that doesn't train on your content, and make every change yourself after reading the source.

What if the client wants light post-editing for text that will be published?

Explain the difference in writing: light post-editing makes the text understandable and accurate enough for internal use, not polished for customers. Offer full post-editing for the published parts, or ask the client to confirm in writing that they accept light-edited quality for publication. Keep that confirmation with the job file.

Further reads

Sources: ISO 18587:2017 catalogue entry (requirements for full, human post-editing); Phrase support pages on Phrase QPS thresholds; RWS documentation on Trados Studio 2026 AI features and MTQE; memoQ AGT documentation; DeepL Pro plan page.

Want an MT post-editing set-up that pays properly?

On a 1:1 call we'll look at the work you post-edit, set up MT, termbase and QA in the tools you use, and build a sampling and pricing routine you can apply to every quote.

Book a 1:1 call with me