Learning motion graphics
Motion graphics means moving graphic elements to deliver a message: text, shapes, icons and logos, not characters and not footage. Learning it does not start in the software, it starts by separating it from animation and editing, because those are three different jobs with different skills and different tools.
- Lesson 1 of 10
- Beginner
- Free, no signup
The five stages of a project, in the order they belong
Stage three is where everybody starts, and it is exactly where they should not.
-
1
Script
One message, not three. Any sentence whose removal does not break the message is removed.
-
2
Style frames
Two or three complete stills. The argument about how it looks ends before anything moves.
-
3
Animation
This is where keyframes go, onto shots whose lengths are already known.
-
4
Sound
Mix, music and effects. The narration itself was recorded far earlier.
-
5
Export
One version per place it will play. The default preset is usually right for none of them.
Sound is stage four on this strip because the mix comes last; the raw narration take has to exist before stage three. The lesson body opens that distinction up.
Last checked: Facts and tool names in this lesson are re-checked against their sources on this date.
How motion graphics differs from animation and editing
Motion graphics is moving graphic elements in order to deliver information. Its subject is text, shapes, charts and logos, and its main instrument is time: what gets seen first, what comes next, and how much space sits between them. The simple test is this: if that same message arrived complete on a static poster, the movement added nothing and only raised the cost of making it.
Character animation is a different job. There a creature has to seem alive, and the core skill is performance rather than layout: weight, balance, gaze, attitude. Video editing is a third, and it starts somewhere else entirely, from footage that already exists; your work is choosing, ordering and pacing, not making the picture. Visual effects is a fourth: bolting invented layers onto filmed images so that the seam does not show.
In the Iranian job market these four are mixed together constantly. One ad for a motion designer may actually want somebody to cut Instagram promos, and the next wants somebody to open a logo in three seconds. Those are two separate skills, and a person who has one does not necessarily have the other. Without that distinction from the start, you spend months finding out you learned the right tool for the wrong job.
The way to choose is not the one usually recommended, either. Do not look at the work that is pleasant to watch, look at the work whose fifth hour is not unbearable. Editing means hours of watching repetitive footage. Motion graphics means hours of nudging things whose difference only you can see. Both look appealing, and neither is what the behind-the-scenes videos show.
The four jobs in the moving-picture family
The cells carry no axis labels because there is no axis: each cell is a job with its own core skill.
-
Motion graphics
Text, shapes, logos. Core skill: ordering and timing information.
-
Character animation
A creature that has to seem alive. Core skill: performance and weight.
-
Video editing
The footage already exists. Core skill: choosing, ordering, pacing.
-
Visual effects
Invented layers over filmed images. Core skill: hiding the seam.
In a real project these four combine, and one person may do two of them. This grid separates the jobs, not the people.
What stages a project really passes through
Five stages, in the order shown in the figure above: script, style frames, animation, sound, export. Beginners usually start straight at stage three, and that is exactly where they lose the two stages that would have saved them the most time.
The thing whose order you most need to change is sound. Record the narration before you animate, even if it is your own phone and a temporary take. The reason is that the audio, not your taste, sets the length of every shot: a sentence that runs four seconds ruins a two second shot. If you animate first and record afterwards, almost every shot comes out the wrong length and the whole thing has to be re-timed. It is the one decision in this lesson I can call almost always right.
The second skipped stage is the style frame: two or three complete still frames showing the colour, the type, the space and the texture of the piece, before anything moves. They take a few hours and replace several days of revisions, because they move the argument about how the piece feels from the middle of the project to its start. Any change that costs five minutes on a still frame costs two days once it has been animated.
Our own motion graphics service page lists this process in five steps, with sound sitting in step four beside the edit. That order is right for delivery, because the mix and the master genuinely do come last; what this lesson adds is that the raw narration take has to exist much earlier, before the first keyframe. They are two different statements about two different things: record first, mix last.
Which application should I open
The answer depends on the kind of work, not on which name is louder. If the thing you are making is built from text, shapes and logos, you need a compositor, meaning an application that moves layers on a timeline with keyframes. If the thing you are making is built from footage and your work is choosing and cutting, you need a linear editor. These are two separate tool families and each does the other job badly.
Three practical options and what their makers publish about system requirements. After Effects is the working standard for graphics compositing, and Adobe puts its Windows minimum at 16 GB of RAM and 4 GB of GPU memory, with 32 GB recommended for 4K and above. Blender is free and open source, and its minimum is 8 GB of RAM and 2 GB of GPU memory, exactly half. DaVinci Resolve has a free version and carries both an edit page and a Fusion page for node compositing in one application; its Studio version is listed on their own site as a one-off $295.
I put those numbers in deliberately, because in Iran the real bottleneck is usually the hardware rather than the tutorials. Somebody with eight gigabytes of RAM can learn serious work in Blender or Resolve, but in After Effects they will spend most of their time waiting for a preview. Payment is a separate matter: we sell no accounts and recommend no way around the restriction, and we say so plainly in our access and payment layer.
And a position you may not like: pick one tool and stay with it for a year. Swapping applications every two months feels like progress and is not; what builds skill is repetition on the same shortcuts and the same menus. Skills transfer between these programs, habits do not.
Two tool families, and which one is yours
What is the thing you are making built from?
A layer compositor
- After Effects: the working standard, published Windows minimum 16 GB RAM and 4 GB GPU
- Blender: free and open source, published minimum 8 GB RAM and 2 GB GPU
- The work happens on keyframes and easing curves, not on cuts
A linear editor
- DaVinci Resolve: has a free version, Studio listed at a one-off $295
- Its Fusion page gives you node compositing in the same application
- The work is choosing, ordering and pacing cuts, not making the image
The hardware figures were read on the vendors own pages on 2026-09-06 and the links sit at the foot of the lesson. None of these three guarantees your work is any good.
What your first three projects should be
Project one: a logo opening in five seconds, with at most three moving elements. The three element limit is the most important part of the exercise, because it forces you to decide what moves and what stays put. The finish criterion is not taste either: you must be able to say which frame each element starts on, which frame it stops on, and why.
Project two: a thirty second explainer cut to a real narration you recorded yourself. Here you learn the thing project one did not contain at all, which is cutting to audio. Finish criterion: no on-screen text stays up for less time than it takes to read. The simple way to measure that is to read it aloud twice yourself; if you cannot finish, neither can the viewer.
Project three: a fifteen second vertical version of that same piece for social, with subtitles. Here reality arrives: video autoplays muted in a feed, the aspect ratio has changed, and things that were margin in the horizontal version now fall outside the frame. Finish criterion: the message arrives complete with the sound off.
What your first project should not be is a showreel. A showreel displays a skill you do not have yet and imposes no constraint on you, so it teaches nothing. Each of the three above carries one specific constraint, and it is the constraints that teach, not the software.
The first three projects, each with one constraint
The order matters: each rung adds something the one below it did not contain at all.
-
1
A logo opening, five seconds
Constraint: at most three moving elements. Done means you can name the start and stop frame of each one.
-
2
An explainer, thirty seconds
Constraint: cut to a real narration you recorded. Done means no text stays up for less than its reading time.
-
3
A vertical cut, fifteen seconds
Constraint: subtitled and playing muted. Done means the message still arrives complete with the sound off.
These three rungs do not get you to sellable work. What they do is drill three basic skills separately, so that they do not gang up on you in project four.
What gives a beginner piece away
Five tells, and none of them is about creativity. The first is linear motion: an element that travels from one point to another at a constant speed looks artificial, because nothing in the physical world moves that way. It is the only item on this list that a single setting fixes, and the next lesson in this path explains all of it with numbers.
The second is everything moving at once. When five elements enter together the eye does not know where to look, and the result is clutter rather than energy. The third is text: a line that stays on screen for less time than it takes to read was better off absent. The fourth is the absence of sound, and I do not mean music; a silent video looks unfinished even when the picture is clean.
The fifth is the export, and it is the cruellest, because it happens after the work is done. A file that comes out on default settings may be tens of megabytes and stutter on playback, or its colour may differ from what you were seeing in the application. This is where the formats and codecs lesson of this path belongs, and until you have read it one simple rule carries you: watch the export where it is going to be played, not in the application preview window.
Where motion graphics is the wrong answer
Movement takes attention, and that is exactly why it gets used where it should not be. If the viewer has to read something carefully or compare values, movement is interference rather than help: a static table or chart the reader can dwell on almost always works better than an animated version. Every figure you see in these lessons is motionless, and that is a decision rather than a limitation.
The second case is accessibility. Some viewers are made unwell by heavy movement, which is why browsers carry the prefers-reduced-motion setting. If your piece plays on a website, it needs a state in which that setting stops it or simplifies it. The web accessibility lesson covers this more fully.
The third case is commercial: a video that eats the viewer bandwidth in the first seconds may be the reason they never see the page at all. The decision here is between how much impact and how much weight, and the answer does not always favour the video. On this very site we did not decide in favour of it, and the next block says why.
The fast path, with AI
The fast path here is not making the video with AI. It compresses the two slowest parts of the job: turning a recorded narration into a shot sheet with real timings, and cutting the on-screen text that does not fit. What separates this from "ask a model for a video" is that every timing number comes out of your own audio file and the model is forbidden to invent one. This wants a frontier model, because it is judgement work, and our current pick is listed in <a class="text-link" href="/en/ai/">the AI section of this site</a>.
- Finish the narration script and record it, even on a phone. That file is the timing spine, and without it the rest of this path does not work.
- Run the audio through a transcription model and take a transcript with a timecode on every segment. Note the total length too, because the bottom of the table needs it.
- Feed the timecoded transcript to the model with the recipe below. The output is a shot table whose in and out points are real numbers from your own audio, not guesses.
- From that table, turn only the two or three shots the model named into style frames, and get those approved. The rest do not need designing before animation.
- The human pass: check the table against the waveform, because machine timecodes drift on Persian. Then watch the finished piece once with no picture and once with no sound. The model does none of these three.
Copy-ready recipe
Role: a motion designer experienced in explainer videos.
My input is a narration transcript with a timecode on each segment. The timecodes are real: do not change them and do not invent any number.
Give the output as one table, one row per shot:
1. Shot number
2. In and out, taken exactly from the input timecodes
3. Shot duration in seconds, the difference of those two numbers
4. The narration sentence spoken over this shot
5. What is seen on screen, in one sentence
6. On-screen text, six words maximum, or "no text"
7. What moves, one element per shot only
Rules:
- No shot shorter than two seconds. If one is, merge it with the next.
- The shots must fill the whole file with no gaps. Print the sum of the durations at the foot of the table, and if it does not equal the total length, say where it fails to match.
- Wherever on-screen text runs past six words, cut it yourself and say what you removed.
After the table, name two or three shots worth building style frames for before animation, each with a one sentence reason.
Total length of the audio file: {duration}
Transcript with timecodes:
{transcript}
Before you trust the output: The model does not know how long each shot takes to build, so the table it gives you is not your schedule. The timecodes came from a transcription model too, and on Persian they slip, especially where the speaker pauses or says a proper noun; check every in and out against the waveform before you place the first keyframe. And the last thing the model cannot tell you is whether the movement you built reads. That is only learned by watching it.
AI in this kind of work
The honest division is this: AI is strong on everything around the picture and weak on the picture itself when the picture has to be exactly your brand shape. Script, timing, transcription, subtitles and per-platform versions genuinely get faster. So does a background shot or a mood image. What it cannot handle is anything that has to be revised afterwards, and the reason is technical rather than aesthetic: a generator output has no keyframes.
Tools that actually help
- Claude For turning a timecoded transcript into a shot table and for cutting on-screen text, because it writes out its reasoning and reasoning can be refused. Iran is not on Anthropic supported-countries list, so there is no official signup or payment from Iran.
- Gemini For the same text work, and for transcription when the audio file is long. Google own page says the Gemini web app works in more than 230 countries and territories, and Iran is not on that list.
- Seedance 2.5 For background shots and mood imagery. Our own reference, which read the vendor documentation on 2026-08-12, records 30 seconds in one generation and up to 50 reference files, with a 720p resolution ceiling. On Iran we have read no document and we do not guess.
- Kling VIDEO 3.0 When the output goes to a client, because it is the only model in that table with native 4K, at fifteen seconds per generation. One constraint they wrote down themselves: its reference-driven mode caps at 1080p, so 4K and heavy referencing do not combine.
- RGB The access and payment layer for Iran, separate from the tools themselves.
Where it backfires
Three risks, all three specific to this job. First, Persian: none of the four video models we assess in our own reference lists Persian among its speech languages. Kling, the only one that publishes a list at all, counts five and Persian is not among them. So for a Persian video the voice is a separate step whichever route you take. Second, the missing keyframes: when a client says "this, but that movement slightly slower", on a motion graphics project you change one number, and on a generated clip you must regenerate, which also changes the faces, the colour and the background. That is why generator output is B-roll in professional work rather than a brand asset. Third, the ceilings do not add up: the longest clip in that table caps at 720p, and the model with native 4K gives fifteen seconds. If your plan needs long and 4K together, none of these does that in one generation today.
Sources: Anthropic: supported countries Google: where Gemini apps are available ByteDance Seed: Seedance 2.5 Kling: VIDEO 3.0 model user guide
Where this advice stops
This lesson is about short pieces made by one person: a promo, an explainer, a social cut. For a long film, a project with a team, or a delivery to broadcast specification, the order of the stages still holds but the scale and the management tooling differ, and these ten lessons do not build that. The second limit matters more: nowhere in this lesson do we say how long the first three projects take, because we do not know. What anyone spends depends on their hardware, on how much design they already have, and on how many times they are willing to start over, and any number we printed here would be one we made up.
From our own work
On mrgym.ir, which is our own client project, every exercise page carries two moving things built in two completely different ways, and that difference is this lesson argument. We counted today: 1,069 published exercise pages, and all 1,069 carry a demonstration video. We read one of those files with ffprobe: h264, 1800 by 960 pixels, 59.94 frames per second, 16.5 seconds, 1,980,182 bytes. But the thing that actually teaches, meaning the display of which muscle is working, is not video at all: it is an inline SVG whose paths are written by hand in the site template, with the target muscle pulsing through a CSS keyframe from scale 1 to 1.08 and back over 1.9 seconds and the supporting muscle over 2.4, both inside a prefers-reduced-motion guard. Not one extra byte is downloaded for that part. And an honest observation about our own file: for a silent loop that shows a single movement, 59.94 frames per second is more than the job needs, and we have not fixed it.
Real follow-up questions
Can I learn motion graphics without knowing graphic design?
To start, yes. To reach work somebody pays for, no, and the reason is simple: every frame of your piece is a graphic design, and if its layout and contrast are wrong, movement does not fix them. The practical route is to advance in both at once and to drop into the design principles lesson whenever a frame is not working.
Which application is free to start with?
Blender is free and open source and works for both 2D and 3D motion. DaVinci Resolve has a free version and carries the edit and the compositing in one application. Blender own published minimum is 8 GB of RAM and 2 GB of GPU memory, half of what Adobe publishes for After Effects, and at the start that is a large difference.
Can motion graphics be made on a phone?
For short social content, yes, and mobile editing apps are genuinely enough for cutting to audio and adding subtitles. For work where you have to stop a specific element on a specific frame, no, because precise keyframe control is exactly what those apps do not give. There is a further practical limit that gets mentioned less: Persian fonts. On a phone your brand typeface is usually unavailable and the text comes out in a default one.