Translation

Subtitle translation: what has to fit inside the speaker's time

Subtitle translation means a translation that has to be read in the same time the speaker took, and that single constraint is what separates it from translating text. A sentence that is correct but cannot be read in that time is, in a subtitle, the wrong sentence.

  • Lesson 5 of 6
  • Intermediate
  • Free, no signup

The five decisions taken for every single event

The order matters: the time is settled before the text, not after it.

  1. 1

    What was said, not what was written

    The basis is the sound of the scene; a script is another thing and does not always match what was performed.

  2. 2

    The event's clock

    The in and out points come from the picture and not from your sentence. The sentence has to fit inside them.

  3. 3

    The reading ceiling

    How much can be read in those seconds while still watching the picture. The platform gives you the number.

  4. 4

    The condensing decision

    If there is time, nothing is cut. If there is not, negation, numbers and names stay and the rest goes.

  5. 5

    Breaking the line and reading it at real speed

    Break at punctuation, then read it once against the picture. If you read it twice, the line is long.

This cycle repeats for every event, and that is what makes subtitle translation slow. Somebody who has only translated the text has not yet started the third stage.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

What separates subtitle translation from translating text

When you translate text, the reader has time. In a subtitle the speaker has already spent the time and you are only filling the gap they left. The unit of work is no longer the sentence either: it is the event, a block that arrives on one exact second and leaves on another, and every decision about it has to make sense inside that interval.

Three constraints press on that one block at once. The clock says how many seconds this text stays on the picture. The box says how many characters fit on a line and how many lines are allowed. And the reader, whose actual job is watching rather than reading; if your line takes all of their attention, they have lost the film and you have lost the job.

There is one more difference, and it is rarely stated. A subtitle is the only translation whose source is still audible. The viewer hears the original, sees your line, and compares them without meaning to. Somebody who knows ten words of the source language will look for exactly those ten words, and if they cannot find them they feel something went missing, even when the translation is entirely right. That is why a free but correct rendering smells of error in subtitling more than anywhere else.

A subtitle box on a cinema screen with a red bar overflowing it and a blue bar fitting inside, green timer marks

The numbers that actually tie your hands

Two published guides carry numbers you can actually read off, and both were written for delivery into European languages. The BBC subtitle guidelines take a reading rate of 160 to 180 words per minute and conclude from it that you should allow around 0.3 seconds per word, so a four word subtitle stays on screen for at least 1.2 seconds. Netflix gives a floor and a ceiling for the event instead of a rate: a minimum of five sixths of a second and a maximum of 7 seconds.

ConstraintBBC guidelinesNetflix guide
Maximum linestwo for landscape video, three for verticaltwo
Line length37 characters on Teletext broadcast; online, 68 percent of the width of a 16:9 videoset in each language guide
Event durationaround 0.3 seconds per wordfrom five sixths of a second to 7 seconds
Where to breakat punctuation, and never between article and noun or inside a complex verbafter punctuation, before conjunctions and prepositions

One thing the BBC says plainly and almost nobody quotes: in a proportional font, counting characters is not the right measure at all, because letters differ in width and what limits you is the width of the region. To keep the number usable it gives an equivalence: 37 characters in a 75 percent wide region of a landscape video is about 25 characters in a 90 percent wide region of a 9:16 vertical video. Same sentence, same font, a third less room, purely because the video turned vertical.

Now the honest boundary: neither number is Persian. Netflix keeps a separate guide per language and its Farsi one sits behind a partner login; we opened the public address today and it redirected to a sign-in page, so we quote no number from it. What a professional does instead is take the number from the platform that will receive the file, and where that platform publishes none, measure once and make that the house rule. The technical layer of the file itself, the difference between SRT and VTT and between burned and soft subtitles, is explained elsewhere: the subtitles and captions lesson.

Persian fits the box; the clock is another matter

The practical question for every subtitle translator is whether the text fits the line. That is measurable, and we measured it on our own text. This section of the site holds 152 lesson files, each written in four languages side by side, so every Persian passage has an English counterpart. Across the 2,827 long passages in that corpus the Persian came out shorter than the English in 98.4 percent of cases, and took 87 percent of the English characters overall. Arabic took 75 percent and Turkish 95 percent.

Now count the same passages in words and the direction reverses: Persian holds 104 percent of the English word count and uses more words in 69.4 percent of the passages. The reason sits in one more number, the average characters per word: 4.61 in Persian and 5.5 in English. Persian pours the same content into shorter and more numerous words.

For subtitling those two numbers say two different things, and this is exactly where translators get it wrong. The box counts characters, so a Persian line usually fits. The clock, though, is measured with something the guides wrote in words, and 160 to 180 words per minute means every extra word costs extra time. Time your Persian with an English rule of thumb and you produce a line that fits the box and not the interval: it does not overflow and it raises no warning, the reader simply never reaches the end of it.

So for Persian the binding constraint comes from time rather than from width. There is a boundary worth stating plainly: these texts are instructional prose rather than dialogue, and they are not translations either, since the four languages are written separately. The numbers show a direction and not the real size of speech.

The same content, by character and by word

against English, over the 2,827 passages of this section, 9 September 2026
  • Persian 87 percent104 percent

    fits the box and costs more time

  • Arabic 75 percent74 percent

    both measures come down together

  • Turkish 95 percent74 percent

    fewer and longer words

characterswords

Only in Persian do the two dots pull apart, and that gap is the difference between the box and the clock. One method warning: this site's writing rule renders the half space as an ordinary space, so the Persian word count here sits a little above the conventional one; the character count is untouched.

Condensing: what goes and what never does

Condensing has a bad name because it is usually explained badly. Condensing is not a style, it is what the clock forces. The BBC guidelines put it from the other side: if there is time for verbatim speech, do not edit unnecessarily, and the aim is to give the viewer as much access to the soundtrack as you can. Do not automatically cut small words like "but" or "too", because they carry meaning. And it rejects simplification for deaf viewers outright: it is condescending, and it frustrates anyone who lip-reads.

So the working rule is this: keeping is the default, and cutting is imposed on you by the clock rather than chosen by taste. When it is imposed, the order is not random. First to go is whatever the picture is already saying or the ear has just heard: repetition, hesitation, addressing somebody who is in frame. Then the markers of ordinary speech, carefully, because those are what carry tone. Then a second explanatory sentence, if the first one did its work.

And the things that are never sacrificed to condensing: negation, numbers, names, and the single word the joke or the threat or the plot turn is sitting on. The BBC is explicit about names too, saying not to edit out names used to address people, because they are easy targets and following the plot depends on them. Cutting a negation is the worst of the set, because it produces an error the reader cannot even feel: the sentence stays sound and its meaning turns over.

In practice the test is simple. Read the line at real speed while watching the picture. If you had to read it twice, the line is long; if you lost the picture, the line is long; if you got to the end and something was missing, you cut too much.

The two pans that are always at war in a subtitle

The pressure of the clock

  • Repetition and the speaker's hesitation come out
  • What the picture already says is not written again
  • The second explanatory sentence, if the first did its work
  • Addressing somebody who is in frame at that moment

What is not touched

  • Negation, whose removal turns the sentence over
  • Numbers, dates and amounts, exactly as they are
  • Names, even when they are used to address people
  • The one word the joke or the plot turn rides on

The pans are not level and should not be: the default is the right-hand side, and the left only weighs more when the clock refuses. Any cut the clock did not require is your decision, and you should be able to defend it.

Where to break the line, and a trap that only happens in Persian

Both guides share one logic: break the line where the language already breaks. The best place is at punctuation, and after that before a conjunction or a preposition. The list of things not to separate is nearly the same in both: an article from its noun, an adjective from what it describes, a first name from a surname, a pronoun from its verb, and the parts of a complex verb.

In Persian that same logic has two cases the English lists do not carry. The first is the ezafe, the linking vowel: "project manager" or "the original sound of the film" is one unit in Persian, and breaking the line inside it is exactly what those guides forbid between an article and a noun. The second is the Persian compound verb, whose nominal part and light verb are one verb together; putting "decision" at the end of one line and "took" at the start of the next leaves the reader hanging inside a verb.

Then a trap that only visits right-to-left text. A subtitle file is plain text and has no attribute for declaring direction; a web page has dir and an SRT has nothing. So when a Persian line ends in a Latin name or a number, the Unicode bidirectional algorithm decides where the final full stop sits, and in some players it appears at the other end of the line. The common fix is one of the invisible direction control characters such as U+200F, from the same family this site bans in its own copy, because on the web the dir attribute does that job and an invisible character is only a machine's fingerprint. In a subtitle file that attribute does not exist, so this is the exception and it should be used knowingly.

There is a simpler route, and we prefer it: arrange the line so that it does not end on a Latin name or a number. One Persian word after it is enough and the problem never arises. If you cannot avoid it, open the file in the very player your audience uses and look, because this is the class of thing that looks right in the editor and breaks in someone's hands.

The things a subtitle does not translate at all

The Netflix general guide carries several rules that are rarely quoted in Persian, and each of them takes one decision out of the translator's hands, to the benefit of the work.

Currency is not converted. Any amount of money spoken in the dialogue stays in the original currency. A hundred dollars remains a hundred dollars and does not become the local equivalent, not only because the rate moves, but because that line belongs to a scene and the character said dollars.

A brand name has three routes and only these three: the same English name where it is widely known in that territory, the name the brand is known by in that territory, or a generic term for the product. And one explicit prohibition: do not swap one company's brand for another company's trademarked item.

Quotations are the delicate case. The guide says it is best practice to originate a new translation for any quoted text, because a new translation is free of rights issues, and that an existing translation may be used only if it is in the public domain or documented permission has been granted and payment made. A translator who lifts a line of poetry or a sentence from a published novel off the internet is making a legal decision, not a linguistic one.

And one rule that comes back to the translator: the translator credit is written as the last event of the file, in the target language, over the copyright card at the end of the programme, with a duration of up to 5 seconds. Your professional credit has a defined place in the file, and it stays empty only if you leave it empty.

The fast path, with AI

What a language model does well in subtitling is squeezing one sentence to a stated ceiling; what it does badly is touching the structure of the file. So the whole craft of this pass sits in the constraints: you state the ceiling, you make it report the count, and you take away its permission to merge events. Without that third constraint the model joins two short subtitles into one because that makes a smoother sentence, and quietly ruins your timing.

  1. Take your own platform's number rather than a general rule. If the platform publishes none, measure once over ten real events and keep that fixed. A constraint without a number is not a constraint.
  2. Paste the SRT blocks with their numbers and timecodes, not the text alone. The model has to see each event's duration to know which line will not fit; give it only the sentences and it cannot see the problem, so it makes you a smooth translation instead.
  3. Keep the batch small. Context windows are not the same across models, and the one whose vendor names Persian in its language list has the shortest window in this category; a feature film's file does not fit in one go. A small batch means the consistency problem returns, and the answer to that is the same glossary the AI translation workflow lesson builds.
  4. Run the recipe below. The three constraints that do the work are these: merging and splitting events is forbidden, the character count of the longest line goes in its own column, and a line that overruns gets tagged rather than silently shortened.
  5. A human pass over three things: names and numbers, the tagged lines, and then one read at real speed against the picture. That third one cannot be handed over, because it is the only place where you find out whether the line is read or merely fits.

Copy-ready recipe

Role: subtitle translator. Source {source language}, target {target language}.

Constraints:
- At most {number of lines} lines per event, at most {number of characters} characters per line.
- Keep the event numbers and timecodes exactly as they are. Do not merge or split any event.
- Do not convert currency, units or numbers. Do not drop a negation.
- These are not translated: {do-not-translate list}

Output, a table only, with these columns:
event number | duration (seconds) | source text | target text | longest line (characters)

Output rules:
- If a line exceeds the ceiling, tag it [LONG] and offer a shorter alternative. Do not shorten it yourself.
- Wherever you are unsure of the meaning, tag it [?] and leave your rendering in place.
- Write nothing outside the table.

Text:
{SRT blocks}

Before you trust the output: The character-count column is filled in by the same model that wrote the line, so that number is a signal and not a measurement; the real check happens in your own subtitle editor and belongs there. The [LONG] and [?] lines are your work list, not the model's. And before pasting any unreleased dialogue into any service, read that service's data retention policy; unreleased material is the one thing that cannot be repaired after it leaks once.

AI in this kind of work

In subtitling a language model is good at one specific job: shortening a sentence without losing its meaning, dozens of times in a row, without getting tired. It is bad at two others: deciding anything about timing, and knowing which word the scene turns on. Our position is to put the model on the sentence and keep the clock out of its reach.

Tools that actually help

  • Gemini For the high volume condensing pass, because its Flash class is cheap and a subtitle file usually holds hundreds of events. Google's own page says the app is available in over 230 countries and territories, and Iran is not on that list.
  • Claude For the few lines that refuse to shrink and have to be argued about, where you ask which half of this sentence can go without killing the joke. Its own privacy page says chats are not used for model training by default, which is a criterion that matters for unreleased dialogue.
  • Cohere Command A Translate In Cohere's own documentation it is the only model its vendor presents as a translation model, listing its 23 languages one by one with Persian, Arabic and Turkish among them. But its context window is 8,000 tokens, so a film's file has to go in batches and consistency falls back on your glossary.

Where it backfires

The risk specific to this job is damage that raises no error. When a model sees a subtitle file it sees text rather than structure; it merges two short events, moves a sentence across an event boundary, and its output looks perfectly sound because the numbers still run in order. Whoever opens the file is the one who discovers that the lines have shifted against the picture. That is why the "do not merge" constraint in the recipe above is not optional. The second risk is confident condensing: the model shortens the sentence and drops exactly the word the scene turned on, and the result reads more smoothly than your version. The third, and it is serious in this profession, is confidentiality; the dialogue of an unreleased film and a series script are precisely the material that should not be pasted into a service whose policy you have not read. Google's privacy page for the Gemini apps says human reviewers see some of the data and reviewed chats are retained for up to three years, while Claude's data page says chats are not used for training by default. These two policies are not the same, and choosing between them should be a decision.

Sources: Google: where Gemini Apps are available Google: Gemini Apps Privacy Hub (human reviewers, retention) Anthropic: is my data used for model training Cohere: models, Command A Translate languages and context window

Where this advice stops

This lesson carries no Persian number for reading speed and cannot: Netflix keeps its Farsi guide behind a partner login, and we have run no reading-speed measurement on a Persian-speaking audience. The BBC and Netflix numbers are two houses' delivery specs, not a universal law. Our own measurement has boundaries too: it is instructional prose rather than dialogue, the texts are written in four languages rather than translated, and a character count is only an approximation of width, since in a proportional font what really constrains you is the width of the region. Three things are deliberately absent: the technical layer of the file, which the subtitles lesson owns; captioning conventions for deaf viewers, which are their own craft; and dubbing, whose translation carries a different constraint and lives in the AI voice and dubbing lesson.

From our own work

The question of whether a Persian line fits could have been guessed at; we counted instead. Today, over the 152 lesson files of this section, every string that carried all four languages was extracted, its HTML tags and code blocks stripped, giving 11,439 four-language units of which 2,827 were long passages. The result we did not expect was in the gap between two measures rather than in the ratio itself: Persian takes 87 percent of the English characters and 104 percent of its words, because the average Persian word runs 4.61 characters against 5.5 in English. For an interface designer that means the Persian line fits; for a subtitle translator it means the same line needs more time to be read, because reading speed in the guides is written in words. We also carry a method warning we ran into ourselves: this site's writing rule renders the half space as an ordinary space, so our Persian word count sits a little above conventional orthography and that 104 percent should be read as a ceiling rather than an exact figure. The character count is unaffected by the rule, because either way it is one character.

Real follow-up questions

How many characters per line is right for a Persian subtitle?

There is no published Persian number we could point you to, and inventing one just to have an answer is worse than having none. Take the number from the platform you are delivering to; if it publishes none, start from the international broadcast figure for Latin-script languages and test it over ten real events of your own to see whether it is read.

Can you hand a whole SRT file to a model in one go?

It depends on that model's context window, and the windows differ widely; the model whose vendor names Persian in its language list has an 8,000 token window and a feature film's file does not fit in it. Batching works, provided the glossary and the list of names are repeated in every batch, or your character changes name halfway through the film.

Is translating subtitles different from translating a screenplay?

Yes, and the difference is one thing: a screenplay has no clock. In a screenplay you translate the sentence in full and the reader takes as long as they need; in a subtitle that same sentence has to be read inside the interval the speaker occupied. This is why a well translated screenplay does not directly make a good subtitle file and has to be rewritten for the clock.