Minting the First Frame Before You Animate Anything
The founding charter
Two ways into a video model, and most of you have only used one. Text-to-video invents the entire look of a shot from words every single time you press render — the face, the wardrobe, the room, the light, all re-rolled from scratch. Image-to-video starts from a still you've already approved, so the model's only job is making that picture move the way you asked. A start frame is that exact first picture; a keyframe is any frame you hand the model on purpose instead of one it guesses. Stills cost a fraction of what video does and land in seconds, so the fussing over faces, wardrobe, background and light happens cheaply before a single second of motion gets billed.
The method: take a filled-out prompt, sort it into what holds still (subject, lens, lighting, tone) and what moves (action, camera move, pacing). Build the still first — mint several, judge them hard on face, wardrobe, background and light, and pick a winner before spending anything on motion. Then feed that winner into the video tool with a prompt containing nothing but the movement. No re-describing the woman, the apron, the light — that's already in the picture. One prompt sheet split this way went from six wasted video renders down to two, with the second only needing a fix to a verb.
There's a trap waiting inside a "good" still, though: a perfectly resolved, centred, finished-looking image often animates badly, because the model has nowhere to move into. Frames that work leave headroom, real depth behind the subject, and a pose caught mid-action rather than settled — a moment with somewhere left to go, not a finished photograph.
Consequential AI-video news and practical production
Runway retired both Gen-3 Alpha models over the summer and moved image-to-video onto Gen-4.5, folding the old first-and-last-frame trick into a separate workflow — so the button has moved if you learned it on Gen-3. Currently only a small set of tools take both a start and end frame in their hosted interfaces: Luma's Dream Machine, Runway, Kling, and MiniMax's Hailuo, whose newest model usefully adopts the input image's own aspect ratio instead of forcing a crop. On the image-to-video leaderboard with audio, a post-trained MiniMax H3 currently sits at the top, with ByteDance's Seedance 2.0 close behind.
Last week you filled in seven slots and rendered a shot, then six more with one slot changed each time. If you kept the notes, look at them now. Count how many of those slots describe something that moves.
Two. Action and pacing. Camera move is halfway there, since it describes how the frame travels but not what is in it. The other four — subject, lens, lighting, tone — describe things that hold perfectly still for the entire clip. The woman's face does not change. The lens does not change. The light does not swing across the room. The mood does not turn from warm to cold in five seconds.
And yet every time you pressed render, you paid for the model to invent all four of them again from scratch. New face. New wardrobe. New window. New light. You were not directing a shot. You were pulling a lever and hoping.
This chapter is about the move that stops that, and it is the most useful single thing in this whole opening stretch of the show. You settle the still picture first, cheaply, and only then pay for motion. Everything that holds still gets locked before anything moves. That is where art direction actually starts, because you cannot art-direct something you are re-rolling from zero on every attempt.
The two doors into a video model
Your generator has two ways in. You have probably only used one.
The first is text-to-video. Words go in, a clip comes out. The model reads your description and invents the entire look of the shot — who the person is, what they are wearing, what is behind them, how the light falls — and then animates whatever it invented. Every single render, it invents all of that again. Change one comma and the face may change. That is not the model misbehaving. That is the job you gave it.
The second door is image-to-video. A still picture goes in, plus a short instruction about movement, and the clip starts from a picture you have already looked at and approved. The model is no longer inventing the look. The look arrived with the image. Its job shrinks down to one thing: make this picture move the way I asked.
Two terms, because this show glosses its jargon the first time it shows up. A start frame is the exact first picture of the clip — literally frame one, the thing on screen at zero seconds. Feed a still into the image input and that still becomes the start frame. A keyframe, more generally, is a frame you hand the model on purpose rather than one it guesses. The start frame is the first and simplest keyframe you will ever use. Later in the course you will hand over more than one, and that will unlock things you cannot do yet. For now, one is plenty and one changes everything.
Here is why this is not just a nicer workflow but a cheaper one, in plain arithmetic. Video costs by the second, and a five-second clip is five seconds of billing plus a wait long enough to make coffee. A still image costs a small fraction of that and lands in seconds rather than minutes. So twenty attempts at getting the look right is a rounding error when those attempts are stills. Twenty attempts at getting the look right when each attempt is a five-second clip is a real bill and most of an afternoon, and at the end of it you still have not started working on the movement. The still is where you can afford to be fussy. Be fussy there, and you arrive at the video stage with one job left instead of six.
That is the whole idea. The rest is doing it.
Before you can do it, you need one thing, and it is worth checking now rather than discovering halfway through. You need a source of still frames, and you need your generator to accept one.
Open the tool you picked in the first episode and look at the generation panel — the box where you normally type your prompt. Somewhere near it there will be a place to attach a picture. On some tools it is a tab labelled image-to-video sitting next to text-to-video. On others it is a small upload square, or a plus button, or a slot labelled start frame. If your tool has a mode switch, set it to the image path and watch what appears; on several platforms a second empty slot shows up beside the first, offering to take a final frame as well. Ignore that second slot for today. It exists, it is useful, and it is a later lesson. All you need to establish right now is that the first slot is there and that it will take a file from your machine.
If your generator makes stills as well as clips, you are done — you can mint your frames in the same place you animate them. If it only does video, you need an image generator sitting alongside it. Any of the hosted image tools will do. This show does not care which one you use, because the workflow survives a tool swap and the tool of the month does not. What matters is that you end up with a still picture you chose, in a file you can upload, in the aspect ratio you intend to render.
One more practical note while you are in there. Look at what the tool does to an uploaded image that does not match the canvas you selected. Several platforms will quietly crop it, or pad the edges, to force it into the shape you picked. That means a frame you carefully composed can arrive at the model with the top of someone's head sliced off. Make your stills in the same shape you plan to render — sixteen by nine wide, nine by sixteen tall, whichever the job needs — and that problem never appears.
Splitting the prompt sheet
Now the actual work. Take one of last week's filled prompt sheets. Not a new idea, not a better idea. One you already wrote and already rendered, so you can see the difference the method makes rather than taking my word for it.
I will use one of mine so you have something concrete to follow. Seven slots, filled in the order the sheet runs them.
Subject: a woman in her fifties in a paint-stained grey apron, short silver hair, standing at a workbench in a cluttered ceramics studio. Action: she lifts a finished bowl off the bench and turns it slowly in both hands. Camera move: slow push in. Lens: long, so the background goes soft. Lighting: hard afternoon sun through a dusty side window, coming from frame left. Tone: quiet, worn-in. Pacing: one unhurried movement, nothing else.
Read that back and sort the slots into two piles: things that hold still, and things that move.
The first pile is subject, lens, lighting, tone. The woman, her apron, her hair, the bench, the studio, the soft background, the hard side light, the worn-in mood. All of that is true at frame one and still true at frame one hundred and twenty. That pile is a description of a photograph.
The second pile is action, camera move, pacing. She lifts and turns the bowl. The camera pushes in. It happens once, slowly. That pile is a description of movement, and it only makes sense once the photograph exists.
So: build the photograph.
Take the first pile and write it as an image prompt. The wording changes a little, because you are now describing a moment rather than an event. Mine came out as a paint-stained grey apron, short silver hair, a woman in her fifties at a cluttered ceramics workbench, holding a finished bowl, long lens with a soft background, hard afternoon sun from frame left through a dusty window, quiet and worn-in. Notice that the action turned into a pose. She is holding the bowl. She is not lifting it, because a still cannot lift anything. Notice also that camera move and pacing are simply absent. There is nothing for them to describe.
Then generate stills. Not one. Several, and keep going until one is actually right.
Right means something specific here, and this is the part where people rush and pay for it later. Right means the face is a face you would cast. The wardrobe is the wardrobe you pictured, not a near miss. The background has the clutter you wanted and not a tidy showroom. The light comes from the side you asked for. Frame one is the frame that will be on screen when your clip starts, so any compromise you accept here is a compromise you have baked into the shot before it moves.
The reason to judge on a still rather than a clip is that a still is honest. Mistakes sit there and let you look at them. In a moving clip your eye is busy following the motion and a mediocre face slides right past you, and you tell yourself the shot is fine. Freeze it and you see the wrong hair, the wrong decade, the plastic sheen, the hand with four fingers. Cheap to see, cheap to fix, and you have not paid for a second of video yet.
I went through eleven stills on mine. Two were close. The first close one gave me a woman about fifteen years too young. The second had the light coming from behind her, which is a different shot with a different mood, and I noticed that on the still instead of after a render. The eleventh had the age, the apron, the silver hair, the dusty window light from the left, and a bench with genuine mess on it. That one got saved.
Now the video prompt, and this is where the discipline pays. Upload the chosen still into the image input. Then write a prompt containing nothing but the second pile.
Mine: she lifts the bowl off the bench and turns it slowly in both hands, camera pushes in slowly, one unhurried movement.
That is the whole prompt. No woman in her fifties. No silver hair. No apron. No long lens, no hard afternoon light, no worn-in tone. All of that is already in the picture, and repeating it in words does nothing useful — at best the model ignores you, and at worst your description drifts a little from the image and you have handed it two slightly different instructions to reconcile.
The clip came back opening on exactly the frame I approved. Same woman, same apron, same light from the left, same mess on the bench, because that frame was not generated, it was supplied. And the only thing I could still be unhappy about was the movement. Which I was — the first pass turned the bowl too fast and the push in was more of a lunge. So I rewrote the motion and rendered again, from the same still. Second pass held. Two video renders total, and the second one only had to fix a verb.
Compare that with what the same shot cost me the week before, when every render was a fresh roll of the dice on all seven slots. Six clips in, I had six different women in six different studios and I still had not seen the lighting I asked for twice in a row.
Here is a variation to run yourself, and I want you to actually run it, because it makes the point in a way reading cannot. Take two stills instead of one — the winner and the runner-up. Feed them to the generator one at a time with the identical motion prompt, word for word, not a comma changed. Then watch both clips.
You will see two completely different shots. Different person, different room, different mood, different film the shot appears to belong to. And the words were the same words. Which tells you where the shot actually lives: not in the sentence, in the frame. All the sentence did was choose the movement.
That is the same working move this show has been building for two episodes now — change exactly one thing, re-render, read the difference. First we did it with the render loop, then with one prompt slot at a time. This week the one thing you change is the start frame. Same tool, same habit, aimed somewhere new.
There is a pitfall waiting for you here, and it is a sneaky one because it does not look like failure. It looks like success. You mint a still so good you want to print it, you animate it, and the clip barely moves. Or it moves and something goes wrong with the subject.
The cause is that a beautiful still can be a dead end. The model has to find the movement somewhere inside the picture you gave it, and some pictures leave nowhere to go. If the subject is cropped tight, there is no space for them to move into. If there is nothing behind them — a flat backdrop, a blank wall, a blur — there is nothing that can shift as the camera travels, so a push in has no depth to reveal. And if the pose is already finished, the model has a problem. A woman standing with the bowl held still at chest height, arms settled, is a person at the end of a movement. Ask her to lift the bowl and the lift already happened. There is no runway.
Three signs you have hit this. First, the clip looks like a slow zoom on a photograph rather than a moving shot — the frame drifts, but nothing inside it lives. Second, the subject smears, drifts or bends, usually near the edges of frame or at the hands, because the model has been asked for movement it cannot find and starts inventing it out of the pixels it has. Third and most common, the first second is perfect and then the rest comes apart, which is the model coasting on your frame and then running out of information.
The fix is to mint frames for what they can become, not for how they look on their own. Give the subject headroom and space on the side they are heading toward. Put something real behind them — depth, objects, a receding room — so a camera move has something to move through. And catch the pose mid-action rather than settled. My winning still has the bowl already off the bench and in her hands, but her wrists are turned in and her weight is slightly forward, so there is an obvious next instant. That is the difference between a picture and a start frame. A picture is complete. A start frame is a moment with somewhere left to go.
Which means the still you would hang on a wall is often the wrong one. A perfectly balanced, symmetrical, centred, finished composition is a photograph that has resolved itself, and the model animates it by nudging the camera and hoping. Slightly off-centre, slightly unresolved, slightly leaning into a movement — that is what animates well. Judge your candidates on that, alongside face and wardrobe and light.
Two habits to carry forward before the homework. First, keep every approved frame filed next to the prompt that made it. A frame with no prompt is a dead end you cannot repeat, and in two months when a client asks for another shot of that woman in that studio, the frame plus its prompt is the difference between five minutes and half a day. This is the start of something the show comes back to properly much later, when the frames and prompts you have been filing turn into an asset library rather than a folder.
Second, keep pricing the work per finished clip rather than per press of the button. Splitting the prompt sheet increases the number of generations you run — eleven stills instead of none. It also collapsed my video renders from six to two. Count the whole path to a shot you would deliver, and the split wins comfortably, because you moved the expensive uncertainty onto the cheap medium.
Homework. Take one filled prompt sheet from last week. Split it into an image prompt and a motion prompt, sorting each slot by whether it moves. Mint three candidate stills and pick a winner on face, wardrobe, background and light, plus room to move. Animate the winner with a motion-only prompt. Then animate your second-choice still with the exact same motion prompt, word for word, and write down what changed. That last render is the one that teaches you the lesson, so do not skip it because you already like the first clip.
What shipped, and what it does to your start frames
Which brings us to this month, because the start-frame path is exactly where the platforms have been moving, and one change lands directly on the workflow you just learned.
Runway retired both of its Gen-3 Alpha models over the summer — the standard model at the start of July, the Turbo version at the end of it — and moved image-to-video onto Gen-4.5. If you learned the old first-and-last-frame trick on Gen-3, the button has moved: start-frame animation lives in the main image-to-video path, and the two-frame work has been folded into a separate Animate Frames workflow. Practical consequence, if you are on Runway: check which panel you are in before you upload, because the old muscle memory now lands you somewhere else.
The tools that will currently take both a start and an end frame in their hosted web interface are a small club — Luma's Dream Machine, Runway, Kling and MiniMax's Hailuo. That matters even though the end frame is a later lesson, because it tells you which platforms have built out this path and which have not. Several of the biggest video models, Google's Veo and ByteDance's Seedance and OpenAI's Sora among them, will happily take a start frame but do not expose an end-frame option on the web at all.
What each caps on the start-frame path is worth knowing before you commit a project to one. Luma's keyframe route generates five seconds at twenty-four frames a second, up to ten-eighty, and then extends from there — up to nine seconds in one go, and further in increments if you keep pushing. Kling gives you five or ten seconds by default with the newest version reaching fifteen, and it renders at thirty frames a second, which is a different feel and something to notice if you are cutting it against twenty-four. Hailuo is fixed at six or ten seconds at twenty-five frames, and its newest model does something quietly useful: it adopts your input image's own aspect ratio instead of forcing you to crop to a standard shape. If you have ever had a start frame cropped against your will, that is the fix. Runway's turbo path locked first-and-last generations to five-second steps at seven-twenty; on Gen-4.5 you will want to check the durations yourself rather than assume the old numbers carried over.
And a snapshot from the leaderboard, framed as a snapshot because it will have moved by the time you hear this. On the image-to-video board with audio, as of August, a post-trained version of MiniMax H3 sits at the top on Elo, with ByteDance's Seedance 2.0 close enough behind that the gap is not information, the base H3 just behind that, then Google's Gemini Omni Flash and Alibaba's Wan 3.0. Runway's Gen-4.5 and Kling 3.0 Omni, both near the front not long ago, have drifted toward the middle of that top bracket. Read that the way this show always reads standings: a tight cluster means take your pick, and a real gap between tiers means something.
Your smallest next action is the same one from earlier, and it is a two-minute job. Open your generator, find the image input on the generation panel, and confirm it takes a file. Everything this week depends on that one slot.
