3. Starting From a Still
Why the courier kept changing
Three takes of the same action produced three different women: a face that shifted, a collar that stood up in one clip and lay flat in another, a step with three stone treads in one and two in the next. The cause is not a badly written prompt. A sentence names a category, not a person — lengthening it narrows the cloud without ever reaching one specific face.
What the image owns and what the prompt owns
Supplying a still as the first frame fixes what a prompt can only describe. The image anchors facial structure, hair, wardrobe cut and fabric, composition at the start of the clip, and the direction of key, fill and rim light. The prompt owns action, contact, timing, camera path and duration — none of which exist in a still. Once the picture is fixed, stop describing it: Runway's image-to-video prompting guide advises complementing the still rather than duplicating it, using plain nouns, and phrasing everything affirmatively because negative instructions work poorly.
Setting up the shot
Gen-4.5 in Image to Video mode requires a paid plan, takes JPEG or PNG stills, outputs 720-line silent video, runs two to ten seconds, and bills twelve credits a second. Bring the still in at 16:9 or the interface will demand a crop — or widen it first with Expand Image, remembering that the new edges are invented. Motion Brush belongs to the deprecated Gen-2; camera moves are named in words. First-and-last-frame work lives in a separate keyframes app.
Two failures and five repairs
An illustrative take holds the face and wardrobe but pushes a hand through the parcel for four frames, and drifts the shadow direction once the dolly reveals unrendered doorway. The repairs, ranked: change the reference, subtract from the prompt, shorten the action, trim in the edit, or regenerate and hope. Regenerating proves nothing about cause — one clip against one clip is a different seed. Fix the seed and change one variable, or run five to ten under identical settings and count how often the failure appears. Failed generations are not refunded unless the system itself broke.
Where the labour moves
Text-to-video still wins when you do not yet know what the thing looks like. The image-led path costs you the still: ten to thirty candidates, retouching, sometimes a small reference sheet of angles.
Conforming frame rates upstream
A frame-rate enhancement endpoint added in mid-September conforms video up to 300 seconds to targets from 24 up to 120, including the fractional broadcast variants, at one credit per two seconds. The next step needs no code: note your clip's actual rate and your destination's.
Look back at the three takes of the courier shot from the first chapter. The action was the same every time: Mara walks up to a bakery step and leaves a parcel on it. But Mara was not the same person. Her face changed between takes. In one the jacket had a collar that stood up; in another it lay flat. The step itself moved — three stone treads in one clip, two in another, with different lettering on the shutter behind her.
That happened for a reason, and the reason is not that the prompt was badly written. A sentence cannot specify a face. When you type "a young courier in a dark jacket", you are naming a category, not a person. The model has seen an enormous number of young couriers in dark jackets, and each time you generate, it draws a fresh one out of that cloud. Make the sentence longer and you narrow the cloud a little — "a dark green canvas jacket with a high collar" rules some things out — but you never get down to one specific human being, because there is no number of words that adds up to a face.
So the question this chapter answers is: how do you get that shot on purpose? The answer is to stop asking the model to invent the courier every time, and instead hand it a picture of her.
That is what image-to-video means. You supply one still image as the first frame of the clip, and the model animates outward from it. The difference between this and what you did before is worth stating carefully, because it is the whole idea: a prompt describes an intention, and a reference frame fixes a fact. Anything you can settle in the still before you generate is one less thing the model gets to reinvent.
Which means the work splits in two, and it is worth knowing which half owns which decision.
The image owns who and where. The face is the clearest case — the source image fixes the facial structure, the eye shape, the bone lines, the way the hair falls. Wardrobe is next: the cut of the jacket, the collar, the fabric texture, any logo or buckle or strap. Those are anchored from frame zero, because they are already drawn. The image also owns composition, and it owns it absolutely at the start of the clip: where Mara sits in the frame, how much empty space is above her head, how sharp the background is behind her, what overlaps what. And it largely owns light. The key light, the fill, the rim on her shoulder, the direction the shadows fall — all of that is inherited from the still rather than guessed.
The prompt owns what happens and how the camera behaves. Movement, contact, timing, the path the camera travels, and the length of the clip. None of that is in a still image, because a still image has no time in it. A photograph of a woman mid-stride does not say whether she is about to kneel or about to keep walking.
That split has a practical consequence. Once the picture is fixed, you should stop describing the picture. Runway's own guidance for image-to-video is to complement the image rather than duplicate it: do not re-describe the wardrobe, the facial features, the lighting, the colour of the wall or the artistic style. Those details are already decided, and repeating them in words does not reinforce them. It gives the model a second, vaguer version of the same information to reconcile with the first. If your still has warm late-afternoon light and your prompt says "golden hour lighting", you have invited the model to reinterpret light it could simply have copied. The recommended habit is to refer to subjects with plain nouns — "the woman", "the parcel" — and confine the sentence to physical action and camera path.
There is one more piece of good news buried in that. Negative instructions do not work well here, so the advice is to phrase everything affirmatively. Instead of "she does not drop the parcel", you write what she does do: "the woman crouches and sets the parcel down on the step". Write the shot you want, not the shot you fear.
Setting up the generation
Now the actual operation, in the tool you have already been paying for. In the Runway web app, in Tools mode with the Video tab active, you select Gen-4.5 and use it in Image to Video mode. The only difference from the first chapter, at the level of the screen, is that the image input you deliberately left empty last time is now the important field. Everything else about the surface is familiar.
Access first, because there is no point preparing a still for a mode you cannot run. Gen-4.5 needs a paid plan — Standard, Pro, the unlimited tier or Enterprise. A free account does not unlock it. That is the same gate as the text-only path, but it is worth confirming rather than assuming, because requirements do not automatically carry across modes. Two models sitting next to each other in the same menu can want different things.
Then the still. Runway accepts JPEG and PNG files for the image input. Gen-4.5 generates at seven hundred and twenty lines of resolution, which for a widescreen sixteen-by-nine frame means twelve hundred and eighty pixels across by seven hundred and twenty down. Your source image does not have to be exactly that, but its shape matters, and here is the behaviour to know. By default, the interface takes the shape of the image you give it. Hand it a square picture and it will plan a square video. That sounds convenient until you remember that this shot has to cut against a widescreen clip you have already accepted and a widescreen still of the step. The editing project took its geometry from the first thing imported into it, which was that generated clip: twelve hundred and eighty by seven hundred and twenty, sixteen by nine. Anything of a different shape dropped into that timeline gets black bars and a notice in the properties panel.
So you want the still to be sixteen by nine before it goes in. If you override the ratio to sixteen by nine with a mismatched still — a square one, or a tall phone-shaped one — the interface will require you to apply a crop to fit the frame, which means throwing away the top and bottom or the sides of a picture you carefully composed. There is a way out that is not cropping: expand the still first, using Runway's Expand Image tool or the Edit Studio surface, which paints new material outward into the wider frame rather than cutting the existing frame down. Expanding means the model invents the new edges, so look at what it invented — a doorway that was not there, a second step, a stray sign — before you commit to it.
Now the prompt. For this shot, something like: the woman walks the last two steps to the bakery step, crouches, and sets the parcel down on the stone, then straightens and turns away; the camera holds a slow steady dolly in. Read that again and notice what is missing. No jacket. No hair. No face. No light. No mention of the bakery's colour or the time of day. Just movement, contact, and one camera behaviour, phrased as things that happen rather than things that must not.
Camera direction is entirely a language job in Gen-4.5. The old Motion Brush, where you painted a region of the image to move, belongs to the legacy Gen-2 model and is not the control here. You get camera movement by naming it the way a crew would: dolly in, slow pan left, tilt down, tracking shot alongside her, or a locked-off static camera if you want the frame to sit still. Naming a real camera move also tends to help with a specific failure. Vague speed words — "hyper-fast", "dynamic" — encourage the model to thrash, and thrashing shows up as flickering and stutter. A mechanical description like "slow steady dolly push" gives it something stable to aim at.
Duration is a slider, and it runs from two seconds to ten. Set five, to match the clip already in the project. Frame rate lives in the advanced settings, where you choose twenty-four or twenty-five frames a second — pick the one your project is already running so the editor is not forced to duplicate or drop frames — and where you will also find the fixed seed option, which we will use in a moment for a reason worth explaining properly.
Cost is shown before you commit. Gen-4.5 bills twelve credits for every second of video it makes, so a five-second clip is sixty credits, and the total appears on or beside the Generate button as you change the duration. Credits come off when the generation completes. If the system itself fails, the charge is refunded automatically; a clip that comes back technically fine but artistically useless is not a system failure, and you pay for it. That is the honest arithmetic of this craft, and it is why we keep counting rejected takes.
One last thing to have straight before generating: the file that comes back is silent. Gen-4.5 outputs video only, with no dialogue and no ambient sound. Sound arrives separately, either from Runway's own audio generation or as tracks you place in the editor — which is exactly what you already did with the ambience file sitting in the project folder.
While we are near the edge of what this mode does: the base Image to Video prompt bar takes a single starting image, the first frame. Giving the model both a first and a last frame and asking it to travel between them is a real capability, but it is not this button. It lives in a separate Animate with Keyframes app inside Runway, with its own inputs. It is a different operation with different handling, and it deserves to be set up properly rather than mentioned in passing.
Here is an illustrative pass, invented to show the shape of the work rather than to report anything you have done.
The still: Mara mid-stride, three-quarters on to camera, about a metre from the bakery step, the parcel already held in both hands against her chest, the shutter behind her plain and unlettered. Sixteen by nine. Late light coming from frame left, so the step has a hard shadow line across it.
Two of those choices are deliberate and worth naming. The parcel is in both hands, clearly gripped, with her fingers visible against it — not hovering near it, not half behind it. And the shutter is blank. Both of those are there because of failures the model is prone to, and we will see one of them anyway.
The prompt is the one above. Duration five seconds. Twenty-four frames a second, to match the project. Sixty credits on the button. Generate.
Then the part that is not glamorous and is the whole job: watch it properly. Not once at full speed while feeling pleased. Play it, then scrub it slowly, then step through the moment of the action frame by frame. This is the same review habit from the first chapter, and it is the habit that separates having generated something from having a shot.
The first take, illustratively, comes back mostly right. Mara's face is Mara's face. The jacket is the jacket, collar and all. The framing at the start is the framing you composed, because it could not be anything else. The dolly moves in slowly and does not lurch. Two things are wrong. At around the two-and-a-half-second mark, as she crouches, her right hand passes through the corner of the parcel — for four frames the cardboard edge and her fingers occupy the same space, and then the hand re-forms on the far side of it. And in the last second, as the camera has moved close enough to see past her shoulder into the doorway, the shadow on the step shifts direction and the whole frame warms up.
Both of those have names and both have causes worth understanding, because the cause tells you the repair.
The hand through the parcel is a contact failure, and it comes from what these models fundamentally are. A diffusion video model is a predictor of what pixels plausibly follow other pixels. It is not running a physics simulation. Nothing in it knows that cardboard is solid. It knows what hands holding boxes usually look like, and when the hand and the box move relative to each other in a complicated way, "usually" runs out. Fingers warp, duplicate, fuse to the prop, or slide through it.
The light shift is a different failure with a similar root. The still told the model where the light came from for the part of the scene the still contained. The dolly in pushed the camera close enough to reveal the doorway, which the still never showed. The model had to invent that space, and inventing space means inventing the light in it — so the shadow geometry it made up for the new area did not agree with the shadow geometry it inherited, and the mismatch resolved itself by drifting the whole frame.
Now the repair decision, which is the real skill. You have five moves available, and they are not interchangeable.
You can change the reference. For the hand, that means recomposing the still so the difficult contact is easier: hands already flat and clearly separated from the parcel, or the parcel already resting on the step with her hands empty, or the frame cropped from the chest up so the manipulation happens out of shot and is implied rather than shown. Changing the reference is the strongest move available in this mode, because the reference is the part the model is not allowed to guess at.
You can change the prompt. For the light shift, the fix is subtractive: strip any environmental or lighting words out entirely and restrict the sentence to movement, then restrict the movement itself. A slow pan or a small tilt keeps the camera inside territory the still already described. A big push or an orbit forces the model into unrendered space, which is where invented light comes from.
You can shorten the action. Identity holds best early and drifts as the clip runs on, and the same is true of light and physics — the further from frame zero, the less the reference is constraining and the more the model is extrapolating. Asking for three seconds instead of five, and asking for one movement instead of a walk plus a crouch plus a turn, removes the hardest stretch rather than trying to fix it.
You can edit the result. The light shift here happens in the final second. The project already has a timeline, and trimming a second off the tail is faster and cheaper than any regeneration, provided the shorter clip still carries the readable action and still leaves handles at both ends for the cut. Not every flaw needs to be generated away.
Or you can regenerate and hope for a better draw. This is the weakest move, and it is where the trap is.
Suppose you shorten the prompt, generate again, and this time the hand grips the parcel cleanly. It is tempting to conclude that the shorter prompt fixed the hands. You do not know that. Generation is stochastic: every run starts from a different random seed and noise pattern, and a hand that resolves cleanly on a new seed may have resolved cleanly on the old prompt with that same seed. One clip against one clip proves nothing about cause. If you actually want to know whether a change helped, there are two honest ways. Hold the seed fixed in the advanced settings and change exactly one thing, so the two clips are comparable. Or run a batch — five to ten generations under identical settings — and look at how often the failure appears, rather than whether it appeared once. Everything short of that is a guess, and the practical answer most of the time is to stop trying to prove causation and just keep the take that works.
For the record, the illustrative outcome: the second attempt used a revised still with the parcel already gripped low and clear, a prompt trimmed to the crouch and the set-down with a locked-off camera, and a three-second duration. That one is clean, and it is the accepted clip. Thirty-six credits for the second take on top of sixty for the first, and the first one is not footage. It is cost. It goes in the notes beside the prompts and settings that are already there, along with the reference still — because the still is now an input to the shot, as much as the prompt is, and a clip whose reference has been lost cannot be revised, only redone.
So which method wins? Neither, and the comparison is the useful part.
Text-to-video remains better when you do not yet know what the thing looks like. You cannot supply a reference frame for a look you have not found. The first chapter's three takes were not wasted work even though they disagreed with each other; disagreeing is how you discover that the high collar reads better than the flat one, and that the step wants two treads rather than three. Text is for searching. It gives you range, cheaply, in exchange for control.
Image-to-video is better when you know what the thing looks like and need it to stay that way. It buys you a face that holds, wardrobe that holds, a composition that is exactly yours at the moment the clip begins, and light that is inherited rather than reinvented. Those are the four things that break continuity between shots, and fixing all four in a picture is the reason this path exists.
What it costs is the still. That is not one click. Getting a reference frame that is compositionally sound often means generating ten to thirty candidates and choosing one. It may mean retouching the chosen one in an image editor — cleaning up a bad hand, removing stray lettering, evening out a hard lighting edge — because any flaw in the still gets animated, and a small oddity in a frame becomes a moving artefact that draws the eye. And if the scene needs Mara from more than one angle, the honest version of this work is a small reference sheet: front, three-quarter, profile, so the shots you cut together agree with each other.
Which is to say the labour does not disappear when you move to the image-led path. It moves earlier, out of the generation and into the preparation, where it is cheaper to inspect and where a mistake costs nothing but your own time to fix.
The frame-rate endpoint Runway added
One small change worth knowing about, from the middle of September, and it touches the finishing end of this work rather than the generating end.
Runway added an endpoint called enhance frame rate to its developer interface — the programmatic side of the service, not the web app you have been clicking in. What it does is conform a video to a target frame rate, or smooth its motion to a higher one, on Runway's machines. The production task it affects is the one that bit you in the editor: frame rates that do not match. Your clips come out of Gen-4.5 at twenty-four or twenty-five frames a second, and if a destination or a client wants thirty, or a broadcast-flavoured twenty-nine point nine seven, you have historically had to do that conversion on your own computer with optical flow interpolation in a desktop tool, which is slow and is one of the easier places to introduce smeared frames without noticing.
The conditions are specific. It accepts video up to three hundred seconds long — five minutes, which comfortably covers a short, though not a feature. The target rates it will write include twenty-four, twenty-five, thirty, forty-eight, fifty, sixty and a hundred and twenty, plus the fractional broadcast variants of twenty-four, thirty and sixty. Billing is one credit for every two seconds of video, which puts it in a different price bracket entirely from generation: the same five seconds that cost sixty credits to generate costs two and a half credits to conform.
The smallest useful next action does not require touching the developer interface at all. Open your accepted clip's properties in the editor and write down its actual frame rate, then write down the frame rate your intended destination wants. If those two numbers agree, you have nothing to do and now you know it. If they disagree, you have found a real conversion step in your own project, and you know what it would cost to have it done upstream rather than fought with locally.
