Translation

The AI translation workflow, from glossary to delivery

The professional translation workflow today has three parts: a glossary closed before any translating starts, a machine pass constrained by that glossary, and human post-editing, which is the only place anything is decided. The order matters more than the choice of model, because a glossary built after the translation is not a glossary, it is a list of things you have to go back and change.

  • Lesson 3 of 6
  • Intermediate
  • Free, no signup

The five stations of one document, and where the deciding happens

The third station is the only one where anything is decided. The other four are preparation and checking.

  1. 1

    Glossary and brief

    Audience, what the text has to do, register, and the do-not-translate list.

  2. 2

    The machine pass, constrained

    Paired output, not continuous paragraphs. Continuous text is not post-editable.

  3. 3

    Human post-editing

    Source sentence, target sentence, decision. The one station where the machine does not stand in for you.

  4. 4

    The consistency sweep

    Mechanical and cheap. It catches consistency and not correctness.

  5. 5

    Delivery, and a bigger glossary

    This project's glossary is the input to the next project for the same client, if you keep it somewhere.

This order is for a document that gets delivered. To understand a text quickly, the second station on its own is enough and the rest are overhead.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

How many stations the real workflow has

Five stations, and not one of them is translating in the old sense. First a glossary and a short brief about the job get closed. Then a machine pass runs, constrained by that glossary. Then a human sits down with the output and decides. Then the whole document is checked for consistency. Then delivery, with a glossary that is now one entry bigger.

The third station has a name, post-editing, and it is not a new coinage. There is an international standard for it, ISO 18587:2017, titled Translation services, post-editing of machine translation output, requirements, and its own scope covers full human post-editing of machine output plus the competences of the post-editor. Anyone who thinks this is a shortcut somebody invented last year is looking at a profession that has had written requirements since 2017.

Why does the order matter? Because four of the five stations sit outside the translating itself. And the usual shape of the job going wrong is exactly this: someone hands the text to a model, reads the output, likes it, and halfway through discovers that one term has to change everywhere. At that moment you are not building a glossary any more. You are repairing.

Three stations: a locked green glossary, a blue translating machine and a red reading lamp over the output

Close the glossary before you translate, not after

A glossary has three kinds of row and each needs separate thought. The first is a term with a settled equivalent that you only have to remember. The second is a term with two live equivalents, where you pick one and hold it to the end of the document. The third is a word that does not get translated at all.

The third kind is where the work limps, because that list is never as short as you assume. We counted it on this very learning section: in the Persian text of 150 lesson files, outside code blocks and prompt recipes, 4,192 Latin-script tokens were sitting inside Persian sentences; 1,055 of them distinct, and 561 of those appeared exactly once. More than half the list is a tail seen a single time and never again.

That shape does not make the job easier; it tells you where to look. The head of the list, the ten that keep repeating, announces itself and nobody forgets it. What gets dropped is the tail: a product name, a file format, a legal name that occurs once in the whole document and, precisely for that reason, has no memory behind it.

So the glossary gets closed by reading the source once, not by remembering. And when the client has earlier documents, you take the list out of those rather than out of your own head. This is the one place in the whole workflow where an hour spent saves several hours at the stations after it.

Light or full post-editing: which one did you sell

There are two kinds of post-editing and the difference between them is commercial, not technical. In the light kind you catch only the errors that change the meaning and you leave style alone; the output is correct, it smells of the machine, and for a text that merely has to be understood that is enough. In the full kind the output has to come out as though a human wrote it from the start: tone, register, sentence shape, all of it.

The scope of ISO 18587 is closed around the full kind. So when somebody says their work follows that standard, they are talking about the heaviest version and not the lightest. Which of the two you are selling has to be settled with the client before you start; it is the step that almost always gets skipped and comes back later as "this still reads like a machine".

And here the economics of the job change, without our needing to hand you a number. Translating from scratch, you were selling time for producing text. In post-editing, producing text has become nearly free and what you sell is judgment: whether this equivalent is right, whether this sentence carries the same meaning in this context, whether this figure came from the source or the model produced it. Anyone who does not accept that sells post-editing at a discount and takes less money for not typing, while the hardest part of the work is still exactly where it was.

Light and full post-editing, and which is for which text

Neither is better than the other. The difference is what the text has to do.

Light post-editing

  • Only errors of meaning get fixed
  • Style and sentence shape are left alone
  • For a text that only has to be understood
  • The output is correct and smells of the machine

Full post-editing

  • Tone and register get fixed as well
  • Sentences get rewritten where that is needed
  • For a text that gets published or signed
  • The scope of ISO 18587 is closed around this one

If you do not settle which one you are selling before you start, the client expects full and you priced light. That gap always opens up at the end of the job.

Post-edit segment by segment, not paragraph by paragraph

Do not read the machine output as flowing text. Read it in pairs: source sentence, target sentence, decision. The reason is that flowing text carries you along; the reader's mind guesses the next sentence from the previous one and sees what it expected to see. With the target sentence sitting next to the source sentence, that guessing is not possible.

So ask the model for paired output rather than continuous paragraphs. Continuous text cannot be post-edited; it can only be read again, which is a different job.

You also need one check that the human eye cannot do: that a term is rendered the same way everywhere in the document. That work is purely mechanical and it is the best thing to hand to a cheap fast model. In a fresh conversation, not the one where the translating happened:

You have a target text and its glossary.

Glossary, each row as "source term = chosen form":
{glossary}

Target text:
{target text}

For every glossary row, find each place in the target text where that term should appear and write which form was used. Report only the rows that have more than one form, and give the sentence number of each form. Do not choose any form and do not change the text.

This check catches consistency and not correctness; if you rendered a term the same wrong way everywhere, this pass says nothing at all. Quality control in the full sense is a separate job and the lesson on translation quality control in this same path deals with it.

When the text sits inside tags and placeholders

A real document is rarely plain text. It is HTML with its tags, or a software file with placeholders like {name} and %s, or a subtitle file with timings, or a spreadsheet where every cell is a separate string. And when a language model sees something like that it does not merely translate; it helps. It moves a tag, translates an attribute inside a tag, renders a placeholder into the target language, and writes the time format more tidily.

None of these things raise an error, which is what makes them dangerous. A missing tag skews the page and somebody sees it the same day. A translated placeholder skews nothing; it only makes the software print the word {name} where the user's name should be. And that one usually gets reported by the user, not by you.

The professional route is simple and older than language models: separate the structure from the text, give the machine the text alone, then put it back. Computer aided translation tools were built for exactly this, and that is their real value rather than the translation memory. If you have no tool and the job is small, the two column table is enough, provided the second column never sees any structure.

And if you have no choice but to hand over marked-up text directly, hold two rules. First, write every placeholder and every tag verbatim into that prompt's do-not-translate list. Second, after each pass, search the target yourself for all of them and compare the counts against the source. It takes thirty seconds and it is the only thing that finds this class of error before the user does.

Subtitles have rules of their own, from reading-speed limits to the mechanics of the file, and the lesson on subtitles deals with that technical layer.

What exactly to look at before you deliver

A language model's error is fluent and does not announce itself, so the last look should not go to the overall quality of the text. It should go to the categories where the model breaks quietly. Four of them: numbers and units, proper names, dates, and negation.

Numbers and units, because the model rewrites the figure and in rewriting it sometimes rounds it or swaps the unit. Proper names, because any name that was not on the do-not-translate list is a candidate for being translated. Dates, because date formats are not the same across languages and the conversion opens another place to be wrong. And negation, which is the dangerous one: a "must not" that becomes a "must" makes a sentence that is perfectly fluent and exactly inverted.

One more look is needed that has nothing to do with the text: what left your machine. The client's document is not yours to decide the destination of, and by delivery time this question is meaningless; it has to have been answered before the machine pass. The next part of this lesson says what each of these services does with the text you hand it.

Four things to see before delivery, and four that guarantee nothing

Before delivery sheetPost-edited document

Seen before delivery

  • Every number and unit compared against the source, one by one.
  • Every proper name was either in the glossary or deliberately translated.
  • Every negative sentence read again, because its inversion is invisible.
  • The consistency sweep was run on the glossary and the multi-form rows resolved.

These guarantee nothing

  • That the text reads smoothly, which is a property of the model and not a sign of correctness.
  • That a more expensive model was used, which changes the class of error and not its existence.
  • That there are no spelling mistakes, a thing tools have handled for years.
  • That you read it twice yourself, if both times were as continuous text.

The right hand column is not a list of worthless things; they are the things that feel like correctness, and not one of them says anything about a fluent error.

The fast path, with AI

The translation prompt below does three things an ordinary prompt does not: it forces the glossary on the model, it asks for paired output so that the result is post-editable at all, and it makes the model mark where it is unsure with a fixed tag instead of choosing silently. The third one matters most. The problem with a language model is not that it does not know; it is that when it does not know, it produces something that looks confident. Marking the uncertainty turns that invisible error into a list.

  1. Close the glossary, even if it is five rows. Three columns: source term, chosen form, and the words that are not translated at all. When the client has earlier documents, take the rows out of those.
  2. Cut the text into manageable chunks, not because the model runs out of window but because a long output cannot be set next to the source and read. A section, or a few pages, at a time.
  3. Run the prompt below with a cheap fast model of the Flash class. This pass is mechanical and an expensive model does not do it better. Which model is currently the best pick for Persian sits, with its criteria, in the translation category of <a class="text-link" href="/en/ai/translation/">this site's AI section</a>, and gets updated there.
  4. Post-edit the output pair by pair and go to the tagged rows first. They are your task list, not a list of the model's mistakes.
  5. Every decision you make that was not in the glossary goes into the glossary at that moment. Skip this step and the next project for the same client starts from zero.

Copy-ready recipe

Translate the text below from {source language} into {target language}.

Audience: {audience}
What the text has to do: {purpose}
Register: {formal / semi-formal / conversational}

Glossary, these forms are mandatory:
{source term = chosen form}

Do not translate, keep verbatim:
{list of words}

Rules:
1) Give the output as a two column table: source sentence, target sentence. Keep the sentence numbers.
2) Do not merge or split sentences. If the source has twenty sentences, the target has twenty sentences.
3) Wherever you are not sure of a term's equivalent, write your chosen form and put [?] immediately after it. Do not guess and do not explain.
4) Carry numbers, units, dates and proper names across verbatim and do not convert them.
5) Write no commentary outside the table.

Source text:
{source text}

Before you trust the output: The uncertainty tag is the model's own guess about its own not knowing, and it misses exactly the place where it was confident and wrong. So the tagged list is the priority of your work, not the boundary of it. The do-not-translate list is enforced by you and not by the model: after each pass, search the target text for those words yourself. And before any of this, do not put a document you have no right to send outside into this workflow at all.

AI in this kind of work

There are two roles for a model in this workflow and collapsing them into one is the common mistake. The first is mechanical: the translation pass and the consistency sweep, which are the work of a cheap fast model, and an expensive one does not do them better. The second is judgment: where you ask whether this equivalent is right in this context, and there a larger model genuinely makes a difference. Our position is to spend the money on the second and to take the cheapest thing that reads your language for the first.

Tools that actually help

  • Gemini A fit for the mechanical pass, because its Flash class is cheap and it reads Persian directly, so the text does not have to go through English first. Google's own page says the app works in over 230 countries and territories, and Iran is not on that list.
  • Claude For the judgment role, meaning the place where a decision about an equivalent gets made. Its own privacy page says your chats are not used for model training by default and only enter that path if you choose it or in a safety review.
  • Cohere The only model in our reference whose own vendor presents it as a translation model is Command A Translate, and it lists its 23 languages one by one, with Persian, Arabic, Turkish and English among them. Its own commercial agreement writes Iran by name into the Restricted Location definition.

Where it backfires

The main risk here is not translation, it is confidentiality. The client's document is not yours, and when you put it into a service you are deciding on the client's behalf. And the answer is not the same for every service: Google's own privacy page for the Gemini apps says human reviewers see some of the data, tells the user not to enter confidential information they would not want a reviewer to see, and says that chats reviewed by a human are not deleted when you delete your activity and are retained for up to three years. Claude's data page, by contrast, says chats are not used for training by default. So this is not an "AI" question, it is a "which service, with which setting" question, and it has to be answered before the first paragraph is sent. For a document under a confidentiality agreement, the right path is to ask the client rather than to guess.

Sources: Google: Gemini Apps Privacy Hub (human reviewers, retention) Anthropic: is my data used for model training Google: where Gemini Apps are available Cohere: models, with the Command A Translate language list

Where this advice stops

This workflow assumes you know both languages well enough to judge. If you do not read the target language well enough to see a fluent error, this method increases your speed and not your accuracy, which is the worst possible combination. It has three other boundaries. It does not apply to sworn translation; that is a permit and not a skill and no lesson hands it over. It does not apply to literary translation either, because holding a term consistent and preserving sentence numbers are precisely the constraints a literary text needs to break. And the third is said less often: post-editing is not always cheaper than translating from scratch. If the machine pass comes out badly on your kind of text, repairing it costs more than writing it again, and you find that out on the first page rather than the twentieth. When you do find it out, throw the machine pass away.

From our own work

This learning section is itself a four-language production line, and we paid for its worst failure elsewhere on this site. In the AI section, a string that should have carried four languages was sometimes a plain Persian string, and the result was that the English page printed the Persian text. The page came up fine, returned a 200, and nothing looked broken. The only person who could tell was the English-reading visitor, who never reported it. The cause is visible in the code: this plugin's language-value function falls back to Persian when it cannot find the requested language. That behaviour is right for an interface label and a disaster for body content. Which is why this section's validator demands a four-language array for every free-text field, rejects a plain string, and raises a separate error for an empty language. The general point for anyone who delivers translations: in a multilingual production line the dangerous failure is not a bad translation. It is a missing translation that looks exactly like a present one.

Real follow-up questions

Is post-editing really faster than translating from scratch?

On most practical texts yes, and on all texts no. What decides it is how well the machine pass comes out on that kind of text, and that shows up on the first page. If you find yourself rewriting most of the sentences, throw the machine pass away; carrying on repairing a bad text costs more than writing one.

Where do I start if I have no glossary at all?

From the client's own earlier documents and their site, because terms they have already chosen are not your decision and have to be respected. When none of that exists, read the source once and take only the terms that have more than one live equivalent. A glossary that lists everything is a glossary nobody maintains.

Can I hand the model the whole document at once?

Technically, on today's models, usually yes, but the reason for chunking is not the size of the window. A long output cannot be set beside the source and read pair by pair, and that reading is what post-editing is. A document translated in one go usually gets read in one go too, and that is not post-editing any more.