Translation

What specialized translation is, and why domain knowledge beats language skill

Specialized translation means a text where one wrong term of art ruins the whole thing, no matter how fluent the sentence is. In these texts a fluent translation with the wrong term is worse than a stiff one with the right term, and that single sentence is the entire difference between specialized work and general work.

  • Lesson 4 of 6
  • Intermediate
  • Free, no signup

Six fields where the term, not the sentence, decides

Medical and pharmaceuticalLegal and contractual
Scientific and academicDomain knowledgenot language skillTechnical and engineering
Software and technologyFinancial and accounting

What bounds these six is not the name of a field. Any text where a professional community has fixed the precise meaning of a word falls under the same rule.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

What makes a text specialized

It is not the name of the field. A general article about medicine is not a specialized text, and one clause of a sales contract, in which there is no difficult word at all, is. The test is simple: is there a word in this text that has a precise, agreed meaning inside a professional community, and is there anybody who would notice it being wrong?

That word is called a term of art, and what separates it from an ordinary word is that it does not take its meaning from a dictionary, it takes it from use inside that community. A translator who knows both languages well and does not know the community produces an equivalent that is correct as vocabulary and that nobody in practice uses. The text stays fluent and loses its authority.

This is what reorders the skills in specialized work. In a general text, command of the language decides the outcome. In a specialized text, command of the language is the entry condition and what decides the outcome is domain knowledge: knowing what this thing is called in this field and, more than that, knowing where you do not know.

A honeycomb of term cards where one red wrong cell cracks the whole comb, the rest green and blue

Why a wrong term is worse than a clumsy sentence

A clumsy sentence gives itself away. The expert reader sees that it is a translation, is mildly irritated, and takes the meaning correctly. A wrong term inside a fluent sentence carries no signal at all; the reader reads it, believes it, and acts on it.

The consequence takes a different shape in each field. In a medical text, dosage, route of administration and negation are the three places where an error reaches a person directly. In a legal text, a defined term is not a word at all, it is a pointer: "the party" or "the product" in a contract means whatever the definitions clause said, and a translator who renders it as two different words has created two different things. In a technical text, a wrong unit or part name goes into a build, and at that point it is no longer a text, it is an object. In a financial text, a wrong accounting term produces a report whose numbers are right and which says something else.

Our position on these texts is plain: in specialized translation, when you are not sure, do not translate. Keep the source form and flag it so that somebody who knows can decide. A professional translator does not lose credibility by flagging. They lose it by inventing a convincing equivalent.

One wrong term in four fields, four different consequences

  • Medical

    Dosage, route of administration, negation. The error leaves the text and reaches a person.

  • Legal

    A defined term is not a word, it is a pointer. Two forms means two different things.

  • Technical

    A unit or a part name goes into a build. At that point it is not a text, it is an object.

  • Financial

    It produces a report whose numbers are right and which says something else.

what a wrong term breaks

In all four the sentence is fluent and nothing marks the error. The only defence is having the text read by somebody who knows that field.

The personal glossary, an asset that grows with every project

Domain knowledge is not acquired overnight, but the thing that holds it is a plain file. The personal glossary is the only thing in this work that compounds: every project adds a few rows and the next project starts where the last one stopped. A translator who has worked for ten years and has no glossary has worked one year ten times.

A good glossary is not a list of all the terms in a field, though. That is called a dictionary and it is not your job. A glossary holds only the rows where a real choice exists; the places where the language has more than one live form and, if you do not decide, you will write it two ways on two different pages yourself.

We measured this on the texts of this site and the result is usable by anyone building a glossary. We took nine concepts that have more than one common Persian form and counted their forms across 150 lesson files. Eight of the nine had settled on a single form by themselves and needed no glossary row at all. One had not: the concept of a backup, written three different ways inside this one corpus, 89 times as one form across 13 files, 38 times as a second form across 12 files, and 15 times as a third across 6 files.

What these figures say is simple: inconsistency is not spread evenly, it clusters. It lands where the language genuinely offers a choice, and you find those points by counting rather than by guessing. A glossary that lists everything gets abandoned after a fortnight; a glossary that holds only the choice points stays short and gets used.

Five steps to a glossary that does not get abandoned

  1. Gather your past deliveries

    Your glossary already exists; it is just scattered and it disagrees with itself.

    1
  2. Find the terms with more than one form

    Only where you yourself wrote it two ways. The rest need no glossary row.

    2
  3. Choose one form and write down why

    The reason matters more than the choice; six months later it is the reason that stops an arbitrary change.

    3
  4. One file, not memory

    The format does not matter, it only has to be copyable into a prompt or a tool.

    4
  5. Re-run it periodically on new work

    New inconsistency gets made in new work, not in old work that was resolved once.

    5

These steps build a glossary and they do not build domain knowledge. A glossary only holds your decisions; whether those decisions are right is a different job.

Where a Persian term can actually be checked

The honest answer is that Persian has fewer central references available than the large languages do, and it is better to say that plainly than to assemble a list of resources that do not work.

One checkable example: the World Health Organization's ICD-11 browser, which is the reference for naming diseases. Today, 9 September 2026, we read the language links on that page in the 2025-01 release: fourteen languages, including Arabic, Turkish, English, Spanish, French, Russian, Chinese and several others. Persian is not on that list. So an Arabic-reading and a Turkish-reading medical translator can look a disease name up in its own official reference, and a Persian-reading one cannot.

There is also IATE, the multilingual terminology database of the EU institutions, which for the languages it covers is a real place to check a term in fields such as law, medicine and transport. We have not checked its language list, so we make no claim about Persian coverage in it.

So in practice the Persian translator's reference is three things, in this order: the client's own previously approved documents, which whatever they contain are not your decision; Persian books and papers in that field, which show what practitioners actually say; and asking somebody who works in it. None of these is fast, and that is exactly what makes specialized translation expensive.

When not to take the job yourself

There are three situations where the right answer is no, and saying it is not a sign of weakness.

First, sworn translation. That is a permit and not a skill; a sworn translator is licensed by the judiciary and no lesson and no tool hands it over. If the text has to carry a seal, the route starts somewhere else entirely.

Second, a text where a terminology error has a medical or legal consequence and you have no expert reviewer. A glossary, care, and even several years of experience do not stand in for somebody who works in that field. Here the right move is to find the reviewer before you accept the job and to put their cost into the price, rather than hoping it will not be needed.

Third, when the volume and the deadline leave you room for none of the above. A specialized text done in a hurry is exactly the text that comes out fluent and wrong. If you do not have the time or the field is not yours, the sensible thing is to hand the job to a translator inside that field and to spend your own hours where your expertise actually is.

The fast path, with AI

Your glossary already exists, and its problem is that it is scattered across everything you have delivered so far and it disagrees with itself. What the model does well here is exactly the part that blinds a human: reading a large amount of old text and finding the places where you wrote one term two ways. What you must not hand it is the choosing. The audit below holds that line: the model finds, you choose.

  1. Take ten to twenty of your own past deliveries in one field. One field at a time, because a term can have two correct equivalents in two fields and mixing them makes the audit meaningless.
  2. If you still have the source texts, give them as pairs. If you do not, the target texts alone are enough: the aim is to find your own internal disagreement, not to judge whether the translation was right.
  3. Run the audit with a cheap fast model of the Flash class. This is mechanical, high volume work and an expensive model has no advantage in it. The stronger model belongs to the next step, when you are deciding about one specific row.
  4. Read the output row by row and choose the house form for each yourself, with one sentence of reason. Where you do not know either, leave the row open and flag it to ask somebody in that field.
  5. Put the result in the same file you keep at hand on every project, and re-run the audit on new work every three months. A file kept somewhere else is forgotten by the next project.

Copy-ready recipe

You have several texts from my past work in the field of {field name}.

Texts:
{the texts, each with a number}

Your job is only to find terminology inconsistency.

1) Isolate the terms of art belonging to that field. Do not count ordinary words.
2) Report only the terms that appear in these texts with more than one form.
3) Give the output as a table: concept, forms found, the number of the text each form appeared in, how many times each form occurs.
4) Do not call any form right or wrong and do not propose a new form.
5) If you grouped two forms as one concept and you are not sure, write the row and put [?] after it.
6) Do not report a term that has only one form at all.

Before you trust the output: Two things this audit does not catch, and both matter. First, it only sees what you gave it; a term that appears once in your entire body of work is not in the output, and that tail is exactly what causes trouble on the next project. Second, the model will merge two forms that are genuinely two different concepts into one row, and it will do it confidently; so open every row yourself before deciding. And a row attached to a medical or legal consequence still needs an expert reviewer, even after you have chosen.

AI in this kind of work

In specialized work one simple split makes all the difference and is rarely stated: a language model is good at recognising that a word is a term of art and bad at choosing that term's equivalent. The first is pattern work and is what it was built for; the second is the knowledge of a professional community, and there is no guarantee that knowledge was in its data. That is our position: put the model on finding and flagging and do not take the decision from it. Each tool below is introduced on that criterion.

Tools that actually help

  • Gemini A fit for the high volume audit, because its Flash class is cheap and it reads Persian directly. Google's own page says the app works in over 230 countries and territories, and Iran is not on that list.
  • Claude For arguing about one specific row, where you ask whether these two forms really are one concept. Its own privacy page says chats are not used for model training by default, which is a criterion that matters for a client's text.
  • Mistral Large 3 It carries the freest licence in this category, Apache 2.0 with open weights and full commercial use, which means you can run it on your own hardware. For a confidential medical or legal document that is the only architecture which does not take the text out of the building. In exchange, Mistral publishes no complete language list for this model.
  • Cohere The only model in our reference whose own vendor presents it as a translation model is Command A Translate, and it lists its 23 languages one by one, with Persian, Arabic and Turkish among them. Its own commercial agreement writes Iran by name into the Restricted Location definition.

Where it backfires

The risk specific to this topic is a wrong term delivered with high confidence. When a language model does not know the right equivalent for a term of art, instead of saying so it produces one that is structurally sound, sounds professional, and carries no mark at all. In a general text that is a stylistic flaw; in a medical or legal text it is the thing that has consequences. So the fast path in this lesson works with an expert review and not instead of one. The second risk is confidentiality, and it is sharper in exactly these fields: a medical record and a contract are texts where sending them to an outside service can itself be a breach, regardless of how good the translation comes out. And the answer is not the same for every service; Google's own privacy page for the Gemini apps says human reviewers see some of the data and reviewed chats are retained for up to three years, while Claude's data page says chats are not used for training by default. For a genuinely confidential document, a model running on your own hardware is the only architecture that removes the question.

Sources: Google: Gemini Apps Privacy Hub (human reviewers, retention) Anthropic: is my data used for model training Google: where Gemini Apps are available Cohere: models, with the Command A Translate language list

Where this advice stops

This lesson does not make you a specialized translator and does not intend to. A glossary is not domain knowledge; it only holds your decisions, and whether those decisions are right is judged somewhere else. Three more boundaries, plainly. Sworn translation is not here: it is a permit and it comes through a judiciary licence. The ICD-11 finding in the fourth part is one measurement, on one day, on one release, and not an eternal rule; the WHO could add a language tomorrow and void that paragraph, and the verification date at the top of this page is the boundary of its validity. And the measurement in the third part is nine hand-picked concepts across texts that were written in four languages side by side rather than translated from one another; it is not a survey of Persian and should not be generalised into one.

From our own work

The measurement in the third part has another side that did not fit there and matters more for building a glossary. Eight of the nine concepts, with no glossary and no coordination between the sessions that wrote these files, had settled on a single form by themselves, their disagreement no more than noise: server, 765 times across 110 files and always the same form; browser, 379 times across 78 files and always the same form; download, 112 times across 37 files; plugin, 273 times out of 274. Writing a glossary row for any of those would have been wasted time. The inconsistency landed only where the language genuinely offered a choice. The method, so it can be re-run or disputed: the Persian text of every lesson file was extracted from the data, code blocks and prompt recipes were set aside, and each form was counted with word boundaries rather than as a substring, because without word boundaries a count of the Persian word for cache also lands inside the Persian word for country and inflates the figure fourfold. The practical result for any translator: before you write a glossary, count your own past work. The list that comes out of counting is shorter than the list written out of fear and, unlike it, gets used.

Real follow-up questions

Do I need a degree in the field to do specialized translation?

No, but you need a way to check and you need to know where you do not know. Many good translators in a field got there by working in it continuously rather than through a qualification. The boundary is clear, though: for a text where a terminology error has a medical or legal consequence, an expert reviewer takes the place of the qualification, and nothing fills the gap when there is none.

Can AI translate a term of art correctly?

Sometimes yes, and it never tells you which time. The issue is not average accuracy but the distribution of the errors: the rare term that matters in your text is exactly where the model has seen the least data and shows the most confidence. Which is why in these fields you put the model on flagging and do the choosing yourself or with an expert reviewer.

What format should I keep the glossary in?

The simplest format you can copy and paste straight into a prompt or a tool; a three column table in a spreadsheet or a text file is enough. The format does not decide anything. What decides it is that there is one file rather than several and that it sits where you actually open it while working. A glossary you have to go looking for does not exist by the next project.