Motion and video

Video editing basics: from raw footage to a locked cut

Video editing means deciding what does not stay in the video, then arranging what is left so the viewer never gets lost. The work moves in six passes: organising the files, a rough assembly, the first cut, the sound pass, picture lock, and finally colour and export.

  • Lesson 5 of 10
  • Beginner
  • Free, no signup

The six passes an edit goes through

  1. 1

    Organise and name

    Every file in one structure, under names you can search. The dullest pass, and the one that gives hours back later.

  2. 2

    Rough assembly

    Every usable shot in a row, with no decision about length. All this pass tells you is what you actually have.

  3. 3

    First cut

    The story is complete but long. If it does not work here, no effect saves it later.

  4. 4

    Sound pass

    Dialogue is levelled, music ducks under speech, intrusive noises come out.

  5. 5

    Picture lock

    From here shot lengths stop changing, because every change breaks the colour and sound work again.

  6. 6

    Colour and export

    The last pass and the shortest. Done before picture lock, it is work you will do twice.

This order is for a short project with one editor. On a team, sound and colour run in parallel and picture lock becomes a stricter, more formal moment.

Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.

What is editing, exactly?

Editing is choosing what does not stay in the video. A one-minute promo usually comes out of tens of minutes of footage, and the editor's job, before anything else, is deciding what happens to those tens of minutes.

That sounds simple and it is the hardest part of the work. A beginner looks for something to add: a transition, an effect, louder music. Someone who has delivered a few projects looks for something to remove without breaking the meaning.

One exercise that pays off from day one: take your first cut and tell yourself it has to get twenty percent shorter. It almost always can. And it almost always gets better.

There is a second point that gets said less often. A cut is not a technical decision, it is a narrative one. You place it where the viewer has already asked the next question; cut earlier and they are confused, cut later and they are bored. No software tells you where that moment is.

How do you lay out a timeline so you do not get lost halfway?

Every editing program, free or expensive, shares one shape: horizontal tracks stacked on each other. The upper tracks are picture, and the higher one sits, the more it covers what is under it. The lower tracks are audio, and they are all heard together.

That single difference causes most of the day-one confusion. In picture, the layer above hides the layer below. In audio nothing hides anything, everything sums; so if the music is sitting on top of the dialogue, the problem is not track order, it is level.

A rule that buys time: give every kind of sound its own track. Dialogue on one, music on another, effects on a third. Then when the client says the music is too loud, you pull one fader and you are done; if everything is on one track you have to touch every clip.

And a habit no tutorial takes seriously: naming. A file still carrying the name the camera gave it will not be findable three days later. Five minutes spent naming folders by day and scene is the difference between a project you can reopen next week and one you have to rebuild.

The timeline from top to bottom: what goes on which track

  • Top video track: titles and graphicsThe higher it sits, the more it covers.
  • Main video track: the shotsThe edit itself lives here; every other track serves it.
  • Dialogue trackThe level reference for the whole mix. Everything else is set against it.
  • Music trackIts own fader, so it can duck under speech and come back.
  • Effects and ambience trackLooks unimportant and is what makes the scene feel real.

Track names differ between programs and the order is not mandatory. What is mandatory is that each kind of sound gets its own track.

What are J and L cuts, and how do they differ from a straight cut?

In a straight cut, picture and sound change together. That sounds logical and it is tiring in dialogue, because your ear does not work that way in the real world: you hear someone before you turn your head towards them.

A J cut imitates exactly that: the sound of the next shot arrives before its picture. An L cut is the reverse: the picture changes while the sound of the previous shot keeps running. The names come from the shape the two offset pieces draw on the timeline.

CutWhat happensWhere it earns its place
Straight cutPicture and sound change on the same frameScene change, topic change, anywhere a pause is wanted
Cut on actionCut in the middle of a movement that continues in the next shotHiding the cut in action scenes and product demos
J cutThe next shot's sound arrives firstDialogue, interviews, entering a scene not yet seen
L cutThe previous shot's sound keeps runningLeaving a scene softly, carrying a feeling into the next shot
Jump cutA chunk removed from the middle of a static shotShortening a straight-to-camera piece, if you accept the jump

The part tutorials mention less: do not put J and L cuts everywhere. If every edit in a three-second interview rhythm overlaps its audio, the result is one continuous mush in which no sentence stands apart. These two cuts join two scenes; they are not a treatment for every cut.

The jump cut needs a position too. In straight-to-camera video it is accepted and it shortens the work; in a brand promo it usually reads as sloppy. If you hesitate, lay another shot over the same audio and hide the cut behind it.

Why does a professional edit start from the sound?

A short test: listen to your first cut once with the picture off. If it is still clear what is happening and where it ends, the edit is standing up. If you understand nothing without the picture, the problem is in the cuts, and colour and effects will not fix it.

In narrative and promo work, what really decides the length of each shot is the sound, not the picture. A sentence finishes, a breath is taken, a musical beat lands; the cut sits on those points. That is why many editors build a clean audio track first and lay picture over it, rather than the other way round.

And a fact that is hard to accept: a viewer tolerates average picture with good sound, but switches off great picture with bad sound. Our own motion graphics and video production page says it in one line, and that is exactly why voice and mix are a stage of their own in our process, not a layer added at the end.

The minimum the sound pass has to do is three things: level the dialogue so it sounds even from shot to shot, duck the music wherever someone speaks, and remove the intrusive noises that were in the recording. None of them needs professional mixing knowledge, and all three make the biggest difference.

How do you set the pacing so it does not get boring?

There is no correct number for how long a shot should be, and anyone handing you one has really turned the taste of a single genre into a number. Pacing is decided by what the video has to do.

Three pressures land on top of each other: the viewer needs time to read and see, the message must not be stretched, and shots of the same shape in a row put people to sleep. Setting the pace means finding the point where all three are bearable, not reaching a particular speed.

Two tests that genuinely work and need no tool. First, watch it muted: wherever your eye jumps ahead, that shot is long. Second, watch it at double speed: anything still comprehensible at double speed is probably slow at normal speed.

And a rule that is almost always right for promo video: find the longest shot in the piece and ask what it shows that the previous shot did not. If there is no clear answer, that is where it gets halved.

Which editor should you start on?

Our position is simple: the software is the third thing you learn, not the first. Cutting, sound and pacing are shared across all of these programs, and someone who knows them finds their way around any of them inside a week.

Even so the first choice is not neutral, because it decides how soon you hit a ceiling. There are three real rungs.

CapCut is built for short vertical clips and social platforms, and it is the fastest route to something publishable. Its own site leans on being free and needing no credit card. Its ceiling is in the same place: a multi-scene project with serious media management is not this tool's job.

DaVinci Resolve is the only option whose free version is a complete program rather than a crippled one. Blackmagic's product page says the free version works with virtually all 8-bit formats at up to 60fps and resolutions as high as 3840 by 2160. The paid version, Resolve Studio, is $295 and adds things like the neural engine, noise reduction and text-based editing. If you are serious, start here.

Premiere Pro is subscription only; Adobe's own page says it is available only as a membership and offers a seven-day trial. The reason to pick it is usually not the software but the team and the client: when a project has to be handed to someone else, a shared project format is worth the cost.

One thing this lesson does not say, deliberately: which of the three opens, activates or can be paid for on your connection in Iran. We do not test that from this server, so we make no claim about it.

Three rungs of editing software, and where each one ends

  1. 1

    CapCut: short vertical clips

    The fastest route to something publishable; its own site says free and no credit card. Its ceiling is a multi-scene project with media management.

  2. 2

    DaVinci Resolve free: the serious starting point

    The free version is a complete program: up to 60fps and up to 3840 by 2160, per Blackmagic's own page.

  3. 3

    Resolve Studio or Premiere: when the team and client decide

    Studio is $295 once, Premiere is subscription only with a seven-day trial. The reason is usually project handover, not a feature.

This ladder is about what the software can do, not about access from Iran; we do not test that from this server.

The fast path, with AI

The fast path in editing is not letting AI edit; it is building the first cut from text. For anything with talking in it (an interview, a tutorial, a straight-to-camera video) the whole file gets a timecoded transcript, the keep-and-cut decisions get made on that text, and only those decisions are then executed on the timeline. An afternoon of work becomes an hour, and more importantly no good take gets forgotten.

  1. Pull the audio out of the file. One channel, 16 kHz, uncompressed. That is exactly what speech-to-text models expect, and it is a fraction of the size of the video.
  2. Run the transcription on your own machine. Whisper is released under the MIT licence and runs locally, so a client's unreleased footage never leaves your computer. The medium model is a reasonable starting point, and its own README says ffmpeg has to be on the system.
  3. Hand the timecoded text to a strong model with the recipe below. What comes back is a decision table, not a summary: each chunk with its in and out times, and whether it stays and why.
  4. Execute the keep column on the timeline. The rough assembly and the first cut now arrive together in one sitting, and you go straight to pacing and sound.

Copy-ready recipe

# 1. Pull the audio (run and checked on this server with ffmpeg 6.1.1)
ffmpeg -i interview.mov -vn -ac 1 -ar 16000 -c:a pcm_s16le interview.wav

# 2. Transcribe locally. Full option list with whisper --help
whisper interview.wav --model medium --language Persian

# 3. The decision prompt, for a strong model (not a cheap fast one)

Role: assistant editor. The text below is the timecoded transcript of a raw file.
Goal of the video: {one sentence}
Target length: {e.g. 90 seconds}
Audience: {who they are and what they already know}

Transcript:
{paste the timecoded text here}

Build a table, one row per meaningful chunk, with these columns:
1. In and out times, exactly as they appear in the transcript
2. One line of what is being said
3. Keep / cut / maybe
4. Reason, in one sentence, based on the goal of the video and not on how
   well the sentence is phrased
5. If it stays, its suggested place in the final order

Rules:
- Never write a time that is not in the transcript. Do not guess a timecode.
- Do not rewrite the speaker. Column two is a quotation.
- List repeated takes separately, with the best one marked.
- If the total of the keep rows exceeds the target length, say by how much
  right there, and which rows should go first.
- Do not propose anything that is not in the transcript; if the story has a
  hole, put it in a separate list headed "shot needed".

Before you trust the output: This path builds the first cut, not the video. Three things stay with you: pacing, because the model does not know where a breath belongs; delivery, because the best sentence on paper is sometimes the worst take on camera; and accuracy, because automatic transcription gets names, numbers and technical terms wrong and that error then walks into a cut decision. Every row the model marked "maybe" you watch, rather than read.

AI in this kind of work

In editing, AI works well where the work is mechanical and reviewable: transcription, finding silences, generating captions, separating voice from noise. It works badly where the decision is narrative: where to cut, how long to hold, which take was better performed. One question separates the two: if you can look at the output for thirty seconds and accept or reject it, handing it to a model makes sense.

Tools that actually help

  • Whisper A speech-to-text model under the MIT licence that runs on your own machine. For client work that is the property that matters: nothing gets uploaded anywhere. Its README says ffmpeg has to be on the system, and it ships six model sizes.
  • Claude A good fit for the decision step over a transcript: it holds long text without rewriting the speaker, so the quotation column stays untouched. Iran is not on Anthropic's supported-country list; the payment layer is covered on this site's buying-AI page.
  • Gemini When the transcript spans several files and several hours, it is the cheaper choice for the first mechanical pass: grouping chunks and finding repeats. Iran is not on Google's supported-country list.
  • DaVinci Resolve Studio If you want it inside the editor itself, the Studio version, per Blackmagic's page, has the neural engine and text-based editing. The free version does not, and that difference is more of a reason to buy than the effects are.

Where it backfires

The main risk on this topic is automatic silence removal. Auto-cut tools strip silences and filler words, and on a tutorial the result is excellent; on narrative work the same feature destroys the pacing, because a pause in a story is not silence, it is meaning. A video with every breath removed sounds fast and lifeless and nobody can say why. The second risk is data, and it belongs to client work: unreleased footage, a product not yet launched, the voices of real people. Uploading that to a cloud service is a decision, not a technical step, and whether your content is used to train a model depends on that service's plan and settings and is written on their own page; that is why the fast path above uses local transcription. And a third risk gets less thought: automatic captions get names and numbers wrong, and unreviewed, that error goes out burned into the picture.

Sources: OpenAI Whisper: README, licence and model sizes Blackmagic Design: DaVinci Resolve, free version and Studio Anthropic: is my data used for model training

Where this advice stops

Everything in this lesson is written for short narrative and promotional video: a promo, a product introduction, an interview, a tutorial. For live broadcast, for long documentary and for multi-camera event editing, the order of work and the tools are different and this lesson does not cover them. One more honest limit: none of these rules replaces watching good work. Study the cuts in ten professional videos and you will learn more than from any list.

From our own work

The five-stage process written on our own motion graphics page shows this same order from the production side: script and storyboard before any camera; and at stage four, editing, voice and mix are set together rather than one after the other. Stage five is the part beginners learn late: the final export is not one file but three, for Instagram, the website and television. Which means framing and text-safety decisions have to be made in that first timeline, or the vertical version is a re-edit and not another export.

Real follow-up questions

What computer do you need to start editing?

It depends on the footage, not the software. Phone clips and ordinary camera files edit fine on any laptop from the last few years; what brings a machine down is heavier formats and higher resolutions. If it slows, look for the proxy or optimised-media option in the program before you buy hardware, because it does the same work on a lighter file.

Where do you get free music for a video?

The rule matters more than the source: for every track you use, you should have its licence saved somewhere, with the date and the terms. Free music libraries exist and their terms differ, especially about commercial use and required attribution. This lesson recommends no particular service, because those terms change and checking the current one is your job.

Is editing on a phone enough, or does it have to be on a computer?

For a short vertical clip published the same day, a phone is enough and is sometimes faster. Where it falls short is a multi-scene project with several audio tracks: media management, precise levelling and multi-format delivery do not fit on a small screen. The practical boundary is the first time you go looking for a third audio track.