Translation

How machine translation works, from Google Translate to language models

Machine translation means software turning text from one language into another with no human involved at the moment of translation. Four generations have followed one another, and the important difference with the latest is not that it translates better but that its errors no longer look like machine errors.

  • Lesson 2 of 6
  • Beginner
  • Free, no signup

Four generations of machine translation, and what changed each time

The order is this. Exact dating is deliberately left out, because the generations overlap and no single day divides two of them.

  1. 1

    Rule based

    A dictionary plus grammar rules people wrote by hand. Predictable, and it broke on any sentence nobody had foreseen.

  2. 2

    Statistical

    Learned from a large volume of bilingual text which phrase pairs with which. Smoother, and still working phrase by phrase.

  3. 3

    Neural

    Turns the whole sentence into a numeric representation and rebuilds the target sentence from it. This is where public machine translation stopped being a joke.

  4. 4

    Language model

    A general model for which translation is one task among many, and which takes instructions: the reader, the tone, and the words to leave alone.

A new generation has not retired the previous one. A large share of the machine translation working inside real products today is still neural, and for bulk repetitive work that remains the right choice.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

Machine translation changed four times, and what changed each time

Machine translation started from a simple idea: give a computer a dictionary and the rules of grammar and it should be able to build a sentence. The first generation was exactly that, with people writing the rules by hand. Its output was predictable, and any sentence the rule writer had not foreseen broke it.

The second generation dropped the rules and took shelter in statistics. Instead of being told what the grammar is, the system was given a large volume of bilingual text and left to work out which phrase usually pairs with which. Sentences got smoother and still smelled of machine, because the system worked phrase by phrase and never saw the whole sentence.

The third generation is where public machine translation stopped being a joke. Neural models turn the whole sentence into a numeric representation and regenerate the target sentence from that. The result was sentences with correct structure that a person could read without laughing.

The fourth generation is the language model, and its difference from the three before it is one of kind rather than degree. The first three knew one job only. A language model is a general model for which translation is one task among many, and more importantly it takes instructions: you can tell it who the reader is, what the tone should be, and which words to leave alone. None of the three earlier generations understood any of that.

One common misconception is worth correcting here: these four generations have not replaced one another. A large share of the machine translation working inside real products today is still neural rather than a language model, and for plenty of jobs that is the right choice.

Four machines in a row, gears, an abacus, a neural web and a glowing sphere, lit red, green, blue and white

Where a language model genuinely beats classic machine translation

Three places, and all three come back to one thing: a language model sees outside the sentence.

First, context. An ambiguous pronoun or a word with two meanings is left to luck in classic machine translation, because the system has only that sentence. A language model can read the rest of the paragraph, and if you tell it what the text is about, its guess stops being a guess.

Second, tone. You cannot tell a classic translation engine that this text is for a customer and should be formal but warm. You can tell a language model, and the effect shows. The register question raised in the first lesson of this path has become, for the first time, something you can instruct.

Third, consistency of terminology across a document. Give the model a term list and you can require that chapter seven uses the equivalent chapter one used. This is what multilingual sites fail at more than anything else.

One important condition applies to all three: these are capabilities, not default behaviour. Without context, tone and a term list, a language model behaves exactly like a classic engine, only with smoother sentences. Anyone who just pastes the text and presses the button gets nothing from this generation except more fluency.

Two kinds of error, and which one you can see

Both make mistakes. The difference is which one's mistakes announce themselves.

Classic machine translation

  • Sentences fall apart and verbs go missing
  • The error is visible at a glance
  • Takes no instruction and understands no tone
  • Cheap and predictable at volume

Language model

  • The sentence is correct and the term may be invented
  • The error is fluent and hides itself
  • Takes context, tone and a term list
  • Without those instructions it behaves like the previous generation

This table does not say which translates better. It says which is more expensive to review. For a text that only has to be understood, the right column is the best we have; for a text that gets signed, the right column needs more review rather than less.

Where it is more dangerous than the generation before

When old machine translation broke, you could tell. The sentence fell apart, a verb went missing, and a reader saw at a glance that a machine had written it and that it needed review. Those errors were ugly, and the ugliness was itself a warning.

The language model has removed that warning. Its errors are fluent. A correct sentence, a suitable tone, and in the middle of it a term that nobody in the field actually uses. Someone who does not know the field has no way of catching it, because nothing in the text looks out of place.

Its most dangerous form is confident invention of terminology. Rather than saying it does not know, the model produces an equivalent that is structurally sound, sounds professional, and exists in no source. In a medical or legal text, that is exactly where financial and human damage comes from. The behaviour is explained more fully in the lesson on AI hallucination, and it takes a worse shape in translation, because the reader of the target text does not have the source to compare against.

Our position is plain: for a text that only has to be understood, this generation is the best thing we have ever had. For a text that will be published or signed, this generation is riskier than the one before it rather than safer, precisely because it puts the reviewer to sleep.

Which model actually names Persian

"Multilingual" is a marketing claim until the vendor prints the list of languages. In the translation category of this site's AI section we took that as the criterion, and the result is two rows, neither of them complete.

On one side is Cohere. Command A Translate is the only model in that catalogue whose maker calls it its machine translation model outright and lists its 23 languages one by one, with all four languages of this site on that list. The other side of it is that its context window is 8,000 tokens, the shortest window in the whole catalogue, and Cohere's own commercial agreement writes Iran by name into its Restricted Location definition. So the model that knows Persian has no official purchase route from Iran.

On the other side is Meta. Llama 4 ships with open weights and a licence that permits commercial use, meaning you can take it and run it on your own server. Its language list has twelve entries and is closed: Arabic is on it, Persian and Turkish are not. And its Scout variant has a ten million token window.

Hold on to that inversion, because it decides the shape of any website translation project. The model that names your language has room for a few pages, and the model that does not has room for a book. The practical consequence is that you translate a site page by page rather than in one go, and if you want a long document to go through at once, you are knowingly choosing a model that was not built for translation and has not declared your language.

There is a licensing point people grasp late as well: open weights and free use are not the same thing. A model like Aya Expanse has published weights and does name the languages, but its licence is non commercial; translate a client site with it and take money, and you have broken the licence.

The model that names your language and the model you can run

Two pans that never fill together in this category. Both columns come from the sourced files of this site's AI section.

Names Persian

  • Its maker calls it a translation model outright
  • All 23 of its languages are listed one by one
  • Its window is 8,000 tokens, the shortest in that catalogue
  • The vendor contract writes Iran by name

You can run it

  • Open weights, and it comes up on your own server
  • Its language list has twelve entries and is closed
  • Persian and Turkish are not on that list
  • Its Scout variant has a ten million token window

Neither column says anything about translation quality. They say what the vendor declared and what the licence permits; reading the output is a different job which has not been done here.

What all this means for a multilingual site

Three practical consequences, and none of them has anything to do with picking the best model.

One: your unit of work is the page, not the site. The short window of dedicated translation models forces this, and as it happens it is better for review too, because an error on one page is found on that page. Two: the term list has to be written before the first page and travel with the text in every call, or the consistency described in the second section simply does not happen. Three: human review concentrates on the pages a decision or a payment passes through, rather than spreading evenly across every page.

And one honest statement that rarely gets written down. Not one of the vendors we examined in that category publishes a citable number for the translation quality of its own model. Cohere writes that Command A Translate is its advanced translation model and prints no score; the others print none either. So any table in Persian claiming to rank the best translation AI is either using an external benchmark, in which case it should say which, or making its numbers up.

We say the same about ourselves: that category tells you which vendor listed your language for that specific model, not how natural its translation reads. Knowing the second means reading the output, and that is a different job which we have not claimed to have done.

The fast path, with AI

The assumption everyone makes is that a language model always beats a classic engine, and for some texts that assumption is wrong. On simple repetitive text the difference between the two generations goes to nearly zero and you are paying for nothing; on text full of back references and specialist terms, that same difference is the whole game. It takes twenty minutes, once per kind of text, to find out which case you are in, and after that you no longer have to guess.

  1. Take one representative paragraph of that kind of text, and do not pick the easiest. Take a paragraph with at least one pronoun whose referent sits further back and one specialist term, because the difference between the generations only shows up there.
  2. Run the first pass with nothing attached: just the paragraph and one line saying translate this into the target language. That is the classic generation, simulated with a modern model.
  3. Run the second pass with the same model and the same paragraph, but this time put the four things from the first lesson on top of it: who the reader is, what the text has to do, where the register sits, and which words are not translated.
  4. Give the two outputs to the model in a fresh call and run the recipe below. Open a new conversation for it; asking it to compare in the same place it produced them means it is judging its own work.
  5. Now read only the categories. If every difference landed in style and register, this kind of text does not need a context brief and can be translated cheaply and in bulk. If even one landed in terminology or reference, this kind of text is context sensitive and every batch of it has to travel with that brief.

Copy-ready recipe

You have two translations of one single source text.

Source text:
{source text}

Translation A:
{output of the first pass}

Translation B:
{output of the second pass}

List only the differences and put each difference into one of these four categories:

1) Style: a word or a structure changed and the meaning stayed the same.
2) Register: the level of formality changed.
3) Terminology: the equivalent of a specialist term changed.
4) Reference: a pronoun or a pointer now refers to something else.

For each item, quote both forms verbatim. Do not call either one better or worse and do not propose a third form. If a category is empty, leave it empty.

Before you trust the output: This is a diagnosis, not a verdict. Categories three and four are exactly where the model does its worst work, so check every item it put in those two against a real source in that field yourself, and do not trust its classification. And one paragraph is one sample: if the two passes come out identical on your hardest paragraph, that is a statement about that kind of text and not a licence covering everything you translate.

AI in this kind of work

On this topic the useful question is not which one is better. The useful question is whether the vendor named your language for that specific model, and whether you can reach it at all. In this site's AI section we applied exactly those two criteria to the candidates in this category and got two rows, neither of which satisfies both conditions. Each tool below is introduced on those two questions, not on a quality claim.

Tools that actually help

  • Cohere Its own documentation presents Command A Translate as its machine translation model and lists its 23 languages, Persian, Arabic, Turkish and English among them. Its window is 8,000 tokens, and its own commercial agreement writes Iran by name into the Restricted Location definition.
  • Llama 4 Scout Open weights and a licence permitting commercial use, with a ceiling of 700 million monthly users and a requirement to display the Llama name. Its language list is closed and has twelve entries, without Persian or Turkish on it, while its window is ten million tokens.
  • Mistral Large 3 It carries the freest licence in this category, Apache 2.0 with open weights and full commercial use. Even so, Mistral publishes neither a complete language list nor a country list for this model, which is why it took no rank in our table. Absence of evidence is itself an answer.
  • Gemini A fit for the diagnostic pass above, because it reads Persian itself and the text does not have to go through English first. Google's own page says the app works in over 230 countries and territories and more than 70 languages, and Iran is not on that list.

Where it backfires

Three risks, and the first is what this whole lesson is about: the fluent error. When a language model does not know the correct equivalent, instead of saying so it produces one that is structurally sound and sounds professional. Nothing marks it. The only defence is having the output read by someone who knows the field, and that is what makes machine translation cheap rather than free. The second risk is language lists: "multilingual" is a marketing claim until the vendor prints the list, and the list for a model family is not the list for a specific version. The third is for site owners: publishing machine translated text in bulk without human review is exactly the pattern Google's spam policies describe as scaled content abuse. The issue is not the language or the tool; it is publishing something nobody has read.

Sources: Cohere: models, with the Command A Translate language list and context window Cohere: commercial SaaS agreement (Restricted Locations) Meta: Llama 4 Scout model card Meta: Llama 4 licence Google: where Gemini Apps are available Google Search Central: spam policies for Google web search

Where this advice stops

This lesson says nothing about the translation quality of these models and should not. We have read the vendors' published lists and cited them; how natural each one's Persian reads requires reading the output, and that has not been done here. Two other boundaries. The generations are deliberately written without years, because they overlap and any precise dating is a claim we cannot stand behind. And language lists, licences and access are all things a vendor can change tomorrow; the verification date at the top of this page is the boundary of validity for every number on it. The practical workflow, meaning how you actually translate and post edit a document with these tools, is the subject of the next lesson in this path; here it was only diagnosed.

From our own work

When we were building the translation category of this site's AI section, the first column of the table was going to be a translation quality score. That column was never built, because not one of the seven candidates had published a number for the translation quality of its own model; Cohere writes that Command A Translate is its advanced translation model and prints no score, and the others print none either. A column with no figure for any candidate is not a column. Another column was tried and withdrawn as well: the context window. The ranking engine normalises every numeric column between its lowest and highest candidate, and here the lowest was 8,000 tokens and the highest ten million, three orders of magnitude apart. The effect was that a real thirty two fold difference appeared in the table as 0.4 against 0.2. The number left the table and its point stayed in the prose. Those two decisions carry a lesson beyond translation: a table whose column has no evidence behind it only has the shape of a table.

Real follow-up questions

Is Google Translate better, or translating with a language model?

It depends on what you are doing. For understanding a text quickly, or for high repetitive volume, a dedicated translation engine is cheaper and more predictable. For a text with tone, context and terminology to keep consistent, a language model wins, but only if you give it that context; without a brief, the difference between the two shrinks to the smoothness of the sentences.

How do I find out whether a model really supports my language?

Find the language list for that specific version in the vendor's own documentation, not in an article and not for the model family. If the vendor gives no list and only says multilingual, the answer is unknown and it should be written that way. There is a trap too: a list left open with an "including" is not the same thing as a closed list, and only the second one can be cited.

Does publishing machine translated text cause an SEO problem?

The problem is not machine translation itself but publishing in bulk something nobody has read. Google's spam policies describe that pattern as scaled content abuse, and the same page shows that the criterion is unhelpfulness and scale rather than the tool that produced it. So one translated page that a human reviewed and settled in for the reader of that language is fine; two hundred pages that went straight from the engine to the site are not.