Build it in layers. Define risk-tiered services first: raw machine output, light or full post-editing, and human translation with revision. Run every client through one translation management system holding their memories and termbases, connect machine translation or LLM engines, add quality sampling, and get written client consent for each tier. Start with one client and one language pair.
The technology is the easy half. The hard half is service definition and pricing. Clients hear "AI" and expect lower prices on everything, and agencies without clear tiers end up delivering full-quality work at post-editing rates. There's a certification wrinkle too: ISO 17100, the standard many agencies hold for translation services, says that raw machine translation output plus post-editing is outside its scope. AI-assisted work needs its own description, usually against ISO 18587, the post-editing standard, rather than being sold as 17100 translation.
Define the service tiers before choosing any tools
| Tier | Process | Suits | Price basis |
|---|---|---|---|
| 1. Machine output, unedited | Engine plus automated checks; clearly labelled | Internal gist, triage of incoming documents | Per job or low per-word fee |
| 2. Light post-editing | One linguist fixes meaning errors only | Internal reports, supplier correspondence | Reduced per-word or hourly |
| 3. Full post-editing | One linguist to publishable quality, QA, sample review | Web pages, product content, help articles | Per word, set from a sample of raw quality |
| 4. Human translation with revision | Translator plus a second-person reviser | Legal, safety, marketing headlines, anything high-risk | Standard rates |
Map a real client to the tiers before you sell them. An illustrative independent wine merchant that sells abroad might send internal supplier emails (tier 1 or 2), 400 tasting notes a season for its website (tier 3), and back-label legal text and a campaign slogan (tier 4). Seen on one page, the client understands why one job is quoted at three different rates. A written policy that describes these tiers to clients is covered in whether your translation business needs an AI usage policy.
The stack: one TMS, connected engines, shared resources
A small agency needs one translation management system (TMS) that holds every client's translation memory and termbase, connects to machine translation engines, and lets freelancers work inside it. Three widely used options with AI built in:
- Phrase TMS. Its Phrase Quality Performance Score (QPS) rates machine-translated segments from 0 to 100, and project managers can set a threshold above which segments are locked as good enough. The default is 100, so nothing is auto-locked until you choose a lower score.
- memoQ with AGT. memoQ Adaptive Generative Translation combines a large language model with your translation memories, termbases and LiveDocs, running through Microsoft Azure OpenAI, and memoQ says no engines are trained on your data.
- Trados Studio 2026. RWS added built-in AI features including an AI Assistant that can be directed to use your termbase, Smart Review and machine translation quality estimation, alongside Language Weaver engines.
For engines, DeepL's paid plans state that content isn't used to train its models, with glossary limits that rise by plan (the Team plan allows five glossaries of up to 1,000 entries per language pair; Business has unlimited glossaries). Whatever you choose, check four things with each vendor: where data is processed, whether it's retained, whether it's used for training, and whether you can switch AI features off for a client who refuses them. The client's side of that question is covered in whether it's safe to put client documents into AI translation tools.
Intake and routing: AI's job before translation starts
The least glamorous use of AI in an agency pays back fastest: reading incoming files so the project manager routes them correctly. Paste or upload the file (inside the TMS's AI tools or a business AI plan the client has agreed to) and ask:
Analyse this file for a translation project manager. Report:
1. Content types present (e.g. marketing, product description, legal,
safety or allergen, UI strings, internal) with approximate word
counts for each.
2. Any lines that carry legal, safety, health or financial risk.
Quote them.
3. Text that shouldn't be translated (product codes, brand names,
prices, units).
4. Candidate terms for the termbase, with frequency.
Don't translate anything.
An illustrative result for the wine merchant's seasonal file: 11,200 words, of which 9,800 are tasting notes, 1,100 are shipping and returns text, and 300 are back-label text including an allergen statement and alcohol-content lines. The model flags the allergen and alcohol lines, lists 42 term candidates such as grape varieties and "old vines", and notes that prices appear in the tasting notes. The project manager splits the job: tasting notes and shipping text to tier 3, back-label lines to tier 4 with a reviser, prices locked. Without that split, the back-label lines would have gone through with the tasting notes, and a light-touch pass on an allergen line is exactly the risk tiers exist to prevent.
The split changes the quote too. With illustrative rates of $0.09 a word for tier 3 and $0.14 for tier 4, the 10,900 words of tasting notes and shipping text come to $981 and the 300 back-label words to $42: $1,023 in all. Quoted entirely at tier 4, the same file would be $1,568. The client saves on the flowing text, pays the full rate, reviser included, only where the risk sits, and can see why on one page.
Small, frequent jobs show the value of intake even more. Picture a craft brewery client that launches a new beer every few weeks and needs the same bundle each time: a 120-word website description, a four-line social post and the back-label text, which carries the alcohol content, allergen line and responsible-drinking statement. Each bundle is under 300 words, and in a busy week the temptation is to push the whole thing through one linguist at one tier. With intake in place, the project manager's routine takes two minutes: description and social post to tier 3, back label to tier 4 with a reviser, the beer's name and hop names locked as non-translatable. The job is tiny, but the back label is the part a regulator or an allergic customer will read.
Intake also catches words that would otherwise be billed by mistake. In an illustrative 3,200-word returns-policy file, the analysis reports that 1,400 words are already in the target language: someone at the client pasted last year's translation beneath the new source. The project manager asks whether the old translation is still approved, runs it against the translation memory, and quotes only the 1,800 new words. Image-only PDFs are the other trap. The model can't count text it can't read, so run scans through OCR (optical character recognition) first and compare a page of the converted text with the scan before trusting any word count.
Linguists: briefing, rates and feedback
Post-editing is a skill, and many good translators dislike it until they're briefed properly and paid fairly. A filled-in post-editing brief for the tasting notes:
JOB: seasonal tasting notes, 9,800 words, tier 3 (full post-editing)
ENGINE: [engine] with client termbase v6 attached; TM pre-applied
QUALITY: publishable web copy; client voice guide attached (warm,
knowledgeable, no scores or medals unless in the source)
DO: fix meaning, terminology and fluency; keep prices and codes locked
DON'T: rewrite accurate segments for style preference; add descriptors
that aren't in the source
QPS: segments scoring 95+ are locked; flag any locked segment you think
is wrong in the comments rather than skipping it
RATE: agreed per-word rate for this client's content type (see PO)
DEADLINE: 14 Oct, 17:00; queries in the TMS query log by 10 Oct
Rates should come from evidence. Ask each linguist to time a sample on the client's content, then set the per-word rate so their hourly earnings are at least what they'd make translating. Collect their error notes after each job ("the engine renders 'old vines' three different ways"), because that feedback improves the termbase and engine settings for the next one. The linguist's side of this routine is set out in a translator's post-editing workflow.
Make tool use part of every delivery, so the confidentiality question gets asked on each job rather than once a year. A short form in the TMS delivery step, as one illustrative freelancer might fill it in:
Tools used on this job: TMS editor and its QA only
Machine output quality: fair; term drift on grape varieties (12 segs)
Locked segments I think are wrong: 1 (seg 2231, negation dropped)
Queries still open: 0
Suggested termbase additions: "old vines", "barrel-aged", "late harvest"
The first line is the confidentiality check; the rest feeds threshold calibration and the termbase. If a linguist ever names anything other than the TMS on that first line, talk to them before the next job, not after a client asks.
Quality: estimation, sampling and scoring
Quality estimation scores tell you where to look; they don't tell you the work is right. Two practices keep them honest.
Calibrate thresholds on real jobs. Before locking segments above a score, have a linguist review a sample of segments at different score bands and record how many needed changes. Only lock the bands where almost none did, and repeat per client and language pair, because engines behave differently on different content.
An illustrative calibration on the wine merchant's first job, where a linguist reviewed 200 machine-translated segments spread across score bands and recorded which needed any change:
| Score band | Segments reviewed | Needed a change | Share |
|---|---|---|---|
| 95-100 | 60 | 2 | 3% |
| 85-94 | 70 | 11 | 16% |
| Below 85 | 70 | 38 | 54% |
On those numbers you'd lock 95 and above for this client and pair, and nothing below. Two changes in 60 still means locked segments get spot-checked, which is how the dropped "not" in the scorecard below came to light. The same exercise on the brewery's labels could look quite different, because short, dense label text behaves differently from flowing tasting notes.
Sample and score every tier 3 job. A reviser checks a fixed sample, say 10% or 500 words, whichever is larger, using an error typology such as MQM (Multidimensional Quality Metrics) with severity weights. A filled-in scorecard for one job, illustrative:
| Category | Minor | Major | Critical | Example |
|---|---|---|---|---|
| Accuracy | 2 | 1 | 0 | Major: "hints of" rendered as "notes of strong" |
| Terminology | 1 | 0 | 0 | Grape variety spelt inconsistently once |
| Fluency | 3 | 0 | 0 | Repetitive sentence openings |
| Locked segments | 0 | 1 | 0 | A 97-score segment with a dropped "not" |
That last row is the finding that matters. One locked segment carried a real error, so the threshold for this client goes up and the linguist's brief keeps the instruction to flag suspicious locked segments. Agree pass marks with each client for each tier, and send the scorecard with the delivery for tier 3 and 4 work; it turns "quality" from a promise into a number.
Client consent and contract wording
Put the tiers and the AI processing into your terms and your quotes, not just a conversation. Illustrative quote wording:
Tier 3 (full post-editing): your text will be pre-translated with a
machine translation engine that does not use your content for training,
using your approved translation memory and termbase, then edited to
publishable quality by a professional linguist and sample-checked by a
second linguist. Tier 4 content (back labels) will be translated and
revised by two linguists without machine translation. Tell us if any
content must not be processed by machine translation.
The last sentence matters. Some clients, often in legal or regulated work, forbid machine processing entirely, and you need that recorded before a project manager routes their file.
Expect some clients to ask for everything at the tier 2 price. A reply that holds the line without lecturing, illustrative:
Thanks, that's a fair question. Tier 2 fixes meaning errors only, so
the text is accurate but can read as machine-translated. That suits
your supplier emails, and we're happy to quote those at tier 2. The
tasting notes go on your website, so we'd recommend tier 3 for them,
and the back labels need tier 4 because they carry the allergen and
alcohol lines. If you'd still like the tasting notes at tier 2, we can
do that: we'll mark the delivery "light post-edited, not for
publication" and ask you to confirm by email that you accept it.
That final offer protects both sides. If the client publishes tier 2 text and a customer complains, the record shows what was ordered and what was delivered.
A five-person agency's first 90 days
An illustrative agency with a project manager, three in-house linguists and a pool of freelancers:
- Days 1 to 30: one client, one pair. Choose the wine merchant and its main language pair. Move its translation memory and termbase into the TMS, connect one engine, and run the next job as tier 3 with a full review, not just a sample. Record linguist time and reviser findings.
- Days 31 to 60: calibrate and price. Set QPS thresholds from the review data, fix the termbase from linguist notes, and set tier 3 rates from timed samples. Draft the client-facing tier menu and the quote wording.
- Days 61 to 90: extend. Add two more clients whose content suits tier 3, add the intake analysis to every new job, and brief two freelancers on post-editing with paid sample work.
In this illustration, by day 90 tier 3 work takes around 60 to 70% of the linguist time the same content took as human translation, with the reviser's time about the same as before. Your own figures will depend on the engine, the language pair and how clean the client's memories were to begin with. Whether that becomes margin or a lower price for the client is a commercial decision, and it's easier to make with a quarter of real data than with a vendor's projection. Track it the way measuring time saved after rolling out AI suggests.
Where agency AI workflows go wrong
- Polluted translation memories. Light-edited segments saved into the main memory come back on the next tier 3 job as 100% matches. Keep them separate or tagged.
- Thresholds set once and forgotten. A new engine version or a new content type changes score reliability. Re-sample quarterly.
- Freelancers working outside the TMS. A linguist who pastes text into a consumer chatbot to "check" it breaks the confidentiality you promised. Provide the checks inside the tools you control.
- Tier creep. A client orders tier 2 for text that's then published. Keep the tier on every delivery file name and invoice, so the record is clear when a customer complains about the wording.
- Terminology left to the engine. Engines apply glossaries imperfectly, especially with inflected forms. The routine in keeping terminology consistent with AI is worth building into every tier.
The memory problem surfaces months later, which makes it hard to trace. An illustrative case: a tier 2 job for the wine merchant's warehouse team included "Returns must be made within 30 days of delivery", post-edited to accurate but stiff wording and saved into the main memory. Four months on, the same sentence appeared in the website's returns page, a tier 3 job. The TMS offered it as a 100% match, the linguist's settings skipped confirmed 100% matches, and the stiff wording went live. The client's marketing lead asked why one line on the page read like a machine translation. The fix took two changes: tag every segment with the tier it was edited to, and apply a match penalty to any memory holding tier 1 or 2 work, so those matches drop below 100% and a linguist always sees them. Most TMSs let you set a penalty per memory; check where the setting sits in yours.
Questions agency owners ask about AI-assisted translation
Can we still sell ISO 17100 translation if we use machine translation?
ISO 17100 states that raw machine translation output plus post-editing is outside its scope, so AI-assisted jobs shouldn't be sold as ISO 17100 work. Many agencies describe those services against ISO 18587, the post-editing standard, or their own documented process. If you're certified, ask your certification body how they expect the two services to be separated in your procedures and on quotes.
Should post-edited segments go into the client's translation memory?
Only once they've been through the same review you'd apply to human translation for that tier. Unreviewed or light-edited output pollutes the memory, and bad matches then reappear on future jobs looking like approved translations. Many agencies keep a separate memory for light post-editing work, or tag segments by tier, so full-quality memories stay clean.
How do we stop freelancers pasting client text into consumer AI tools?
Put it in the linguist agreement and the job brief, give them the tools they need inside the TMS so they have no reason to go outside it, and explain why: client confidentiality and the client's own AI terms. Spot-check by asking linguists which tools they used on each job, and treat a breach as seriously as any other confidentiality breach.
Do clients expect AI-assisted translation to be cheaper?
Most do, and some will be right for their content. The answer is a clear tier menu: raw machine output and light post-editing cost less because they are a different service, while full post-editing and human translation with revision are priced on the work they involve. A client who understands the tiers usually chooses by risk rather than by the lowest price.
Further reads
- Is Machine Translation Good Enough for Client-Facing Documents? — Help clients decide when MT output is good enough.
- How to Check AI Translations Before Customers See Them — A final check before translations reach customers.
- AI Translation vs Human Translators: What Documents Really Cost — How clients compare AI and human translation costs.
- Questions to Ask Before Buying AI That Touches Client Data — Questions to ask any AI vendor that touches client text.
- SOC 2 and ISO 27001 Explained: Checking an AI Vendor's Security — Read a translation-tool vendor's security claims properly.
- How to Offer Multilingual Customer Support With AI Translation — A service line some clients will ask about next.
- How to Audit Your Website for AI Content That Needs Fixing — A content audit for sites with AI-drafted pages: inventory, leftover searches, a five-check score, keep-fix-merge decisions and a translation agency's numbers.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: ISO 17100 (scope statement on machine translation and post-editing; revision by a second person); ISO 18587:2017 catalogue entry; Phrase support pages on Phrase QPS thresholds; memoQ AGT documentation; RWS Trados Studio 2026 AI documentation; DeepL Pro plan page (glossaries, no training on content).