Your First Usable Shot
The founding charter
The core of the chapter is a single, deliberately small production exercise: turning a text description into five seconds of AI-generated video and learning to see what actually came out. It sets up a fictional courier named Mara, leaving a parcel on a bakery step, and uses that plain, one-action brief to expose exactly where generation succeeds and fails — hand contact, facial identity drift, invented lettering, and inconsistent light.
Before generating, it walks through what a new Runway account actually gives you: a one-time deposit of credits and storage, watermarked clips, and a credit cost per second that differs sharply between Turbo and base models — arithmetic that shapes every decision that follows. It covers the mechanics of writing a prompt within a real, practice-tested length limit rather than the documented maximum, using the text-to-video prompting guide, the Gen-4 Video prompting guide, and camera terms and prompt examples for describing shot type and camera movement. It explains fixed duration and aspect ratio options via changing aspect ratio and resolution, and the plain fact that the standard video output has no sound — audio is a separate step. Three generations are run, each judged in a structured, repeated viewing rather than a single glance, with one targeted change made between takes and a written record kept of what was rejected and why.
Explain AI-video changes and production shortcuts
A shorter turn looks at how other hosted tools have started folding audio directly into the first generation pass, rather than treating it as a later step — pointing to Hailuo's MiniMax H3 and Kling's newer synchronized-audio model as places where that line has already moved, alongside Runway's own shift toward stronger physical motion in its current default models.
You type a sentence. A few minutes later a computer hands you five seconds of moving picture that nobody filmed. There is a woman, a street, a grey morning, and she sets a parcel down on a step. No camera was carried anywhere. No actor was hired. And when you watch it closely, her hand goes half an inch into the parcel at the moment she puts it down.
That gap — between the shot you asked for and the frames you actually got — is what this chapter is about.
The thing making those frames is a video generation model. It is a program trained on an enormous amount of existing film and footage, which has learned to produce new footage from a description. You give it words. It gives you a short clip. Some of these models run on a computer you own, which needs a powerful graphics card and some patience with software. Others run on somebody else's computers and you reach them through a web browser, the same way you reach your email. Those are the hosted services, and that is where we are starting, because you can be generating today without installing anything.
Hosted generation sits between two older ways of getting moving pictures. On one side is filming: cameras, locations, people, permission, time. It gives you total control over what is in the frame and it costs real money and real days. On the other side is animation: total control again, and enormous labour per second. Generation is fast and cheap by comparison, and in exchange you give up a lot of control. You cannot tell the model to move the actor's hand two inches left. You can only describe, look at what came back, and describe again. That trade is why it is useful to a person working alone, and why the working habit in this chapter matters more than any clever wording.
So here is what we will do. We will take the smallest brief that still counts as a shot. We will open one hosted service, and I will tell you exactly how you get in and what it will and will not do for you. We will build a prompt out of named parts, so you know why each word is there. We will generate. Then we will spend real time looking at what came out. We will make one targeted change, generate again, and compare. And we will finish with a clip saved where you can find it, next to a written record of how it was made.
Let me give you the brief first, because everything else serves it.
There is a courier named Mara. She rides a bicycle around a city delivering packages. Early one morning she leaves a small paper parcel on the stone step of a bakery that has not opened yet. That is the whole shot: one person, one plain physical action, one place, about five seconds long.
Mara is invented. The bakery is invented. The story she belongs to is invented, and I made it up for teaching, not because it exists somewhere. I will keep coming back to her, because a small story with one character and one clear action makes problems visible. When her jacket changes colour between shots, you can see it. That is the only reason she exists.
Notice how little is in that brief. One action. No dialogue. No second person. Nothing that requires the model to understand a relationship or a joke. When you are learning, a small brief is not a compromise. It is the instrument you are measuring with.
Now let us get you into a tool.
The service we will use is Runway. You reach it in a web browser at its web app — that is, you log into a site rather than installing a program — and its documentation recommends Google Chrome. There is no local installation, no graphics card requirement, and no code. You register with an email address, or by signing in with a Google or Apple account, and no credit card is required to start.
What a new account gets you is worth knowing precisely, because it is the budget your practice runs on. A new account receives a one-time deposit of one hundred and twenty-five credits and five gigabytes of cloud storage. One-time means it does not refill next month. Clips you generate on that free deposit carry an on-screen watermark. And accounts that look like repeat signups, or that come in from certain networks, may be given no free credits at all — so if you open the credit balance and see nothing there, that is a known outcome and not something you did wrong.
Credits are spent per generated second, and the rate depends on which model you choose. The Turbo models cost five credits for each second generated. The base models, Gen-4 and Gen-4.5, cost twelve credits per second. Run that against your one hundred and twenty-five. A five-second clip on a base model is sixty credits, so the free deposit buys you two of them and then you are done. A five-second clip on a Turbo model is twenty-five credits, so the same deposit buys you five. Five attempts is a lesson. Two attempts is barely a coin toss. So for this chapter we will generate on Gen-4 Turbo at five credits a second, and I want you to feel that arithmetic in your hands, because from here on every second you generate has a price, whether or not you look at it. When the free credits are gone, generating means a paid plan; the Standard plan is fifteen dollars a month for six hundred and twenty-five credits.
Now, what does the service actually ask you for?
The main input is a text prompt: a plain-language box where you describe what should be in the frame and how the camera behaves. The documented limit is one thousand characters, and the official guidance is to be concise and says the order of your terms does not matter. Creators using it report something more cautious, which is that prompts running past roughly sixty to eighty words start to leak — the model quietly ignores some of your instructions. So treat sixty to eighty words as your real ceiling, not one thousand characters.
You may also, if you want, drag in an image as a reference or first frame, in JPEG, PNG or WEBP format. Doing that changes the job: the picture then fixes the framing, the style and who the subject is, and your text only directs how things move. Our brief does not need it, so we are not using it today. I mention it so that when you see the upload box, you know what it is for and can leave it alone.
Two more things about how you talk to it. There is no separate box for saying what you do not want. The documentation tells you to state your requirements as positives instead — you write static camera rather than no camera shake — because these models often read the thing you named in a negative as a thing you asked for. Say what you want to see. And there is a seed setting, tucked under the generation or advanced settings. A seed is just a number that fixes the model's starting randomness. By default it is randomised, so every run is a fresh roll. You can switch on a fixed seed and lock a number, and then the composition and motion path stay roughly put across runs. We will use that in a moment.
Camera movement you can either describe in words — a slow push-in, a macro close-up, an orbiting camera — or pick from the interface's presets, which include pan, tilt, zoom, truck, pedestal, roll and static.
Then there is what comes out, and this part you need before you plan anything. Duration comes in two sizes: five seconds or ten seconds. Nothing in between. Creators generally report that motion and physics hold together better across five seconds than across ten, with the trouble showing up toward the end of the longer clips. Generation renders at seven-twenty — that is, twelve hundred and eighty by seven hundred and twenty pixels for a widescreen clip — and there is a separate upscale to four times that resolution available from the clip's own menu afterwards, which costs credits on entry tiers. Aspect ratio, meaning the shape of the frame, you choose from a fixed set: sixteen-by-nine widescreen, nine-by-sixteen for phones held upright, square, four-by-three and three-by-four, and a very wide twenty-one-by-nine. Downloads come as an MP4 video file, which is the format everything plays.
And the point I want to nail down before you press anything: the file that comes out of the standard text-to-video pass is silent. Runway's own video models render picture with no audio track, and audio is not synchronised into that first pass. Sound effects, speech and lip-sync are made afterwards, as their own step, using the service's separate audio tools or an outside model. So the sound is available to you inside the same service — it is simply not part of the generation you are about to run. This surprises people, because you watch someone set down a parcel and your ear supplies the paper crunch and the birds, and then you play the file and there is nothing. Plan for a silent picture. The advice for people on the free deposit is to keep generations silent precisely so you do not burn credits on audio you are not ready to use.
So: browser open, logged in, and in the Tools area you open Video and select Gen-4 Turbo. Set duration to five seconds. Set aspect ratio to sixteen-by-nine.
Now the prompt. I want to build it out of named parts, because from here on those names are how you and I will talk about a shot.
A shot is one continuous run of camera, from the moment it starts recording to the moment it stops. Cut, and you have a new shot. A scene is bigger: it is the action in one place and one stretch of time, and it is usually made of several shots. Mara leaving the parcel at the bakery is a scene. The five seconds of her hands putting it down is a shot. Everything a generation model gives you is a single shot. That is worth sitting with, because it means a scene is something you assemble, not something you ask for.
The parts of the shot, then. First the subject: who or what the frame is about. Mara, a young bike courier in a grey rain jacket. Second the action: the one physical thing that happens. She sets a small paper parcel down on a stone step. Third the framing: how much of the subject the frame contains. A medium shot takes a person from roughly the waist up. A close shot is head and shoulders, or a pair of hands. A wide shot puts the whole figure in its surroundings. Framing is a decision about what the audience is allowed to notice.
Fourth, camera movement: whether the camera holds still or travels. Still is called static, and for a first shot it is the right choice, because a moving camera gives the model one more thing to get wrong. Fifth, light: where it comes from and what it feels like. Soft light from the left, early morning. Naming a direction gives the model something to be consistent about. Sixth, duration, which we have set at five seconds. Seventh, aspect ratio, which we have set at widescreen.
Here is the prompt I typed, at forty-four words:
Medium shot of a young bike courier in a grey rain jacket. She crouches and sets a small brown paper parcel down on a stone bakery step, then straightens up. Empty city street at dawn, wet pavement, soft light from the left. Static camera.
Forty-four words sits comfortably under that sixty-to-eighty ceiling, which means every term in it has a fair chance of being read.
Before we generate, one thing about that action, because it is the difference between a nice clip and a useful one. A shot can look handsome and still be no use to whoever cuts the film together, and often that is because the action does not finish. If Mara begins to crouch and the clip ends while she is still bending, an editor has nothing. They cannot cut away from a half-finished movement without it looking like a mistake. They need to see the parcel land. That is why the action is simple: crouch, place, straighten. A complete action in five seconds means one small movement, not three. Simplicity here is not modesty. It is what makes the footage cuttable.
Press Generate. That is twenty-five credits, leaving a hundred.
Now the part that decides whether any of this was worth the credits: look at what came back.
The temptation is to watch it once, feel pleased or disappointed, and move on. Do not. Watch it at least three times, and give each pass a job. This is a habit, and it is the habit that separates directing from gambling.
First pass, watch the action only. Does it start and does it finish inside the five seconds? In my clip, Mara crouches at about half a second, the parcel is on the step by three and a half, and she is upright again by four and a half. There is half a second of nothing at the end. That is fine. A little air at the end is useful.
Second pass, watch the contact points. Where a body touches a thing, generation models are weakest. Do the fingers close around the parcel, or does the parcel simply travel with the hand? Do her feet meet the pavement, or float a few pixels above it? In my clip her right hand passes into the corner of the parcel as she lets go. Not badly. But if you froze the frame and pointed, someone would see it. That is a real flaw and I am going to name it: broken contact at the moment of release.
Third pass, watch the face. Generated faces drift. Does she look like the same person at second one and second four? Mine holds until she straightens, and then her face changes — the jaw narrows and the eyes sit differently. That is called identity drift, and it matters more than it looks, because a story needs the same person to keep being the same person.
Then check the light. Does it keep coming from one direction? Mine does, mostly, though the puddles brighten oddly halfway through as if a second sun arrived. And check for text. Models love to invent lettering, especially on signs and shopfronts. There is a smear of not-quite-letters across the bakery window in mine — not words, just the shape of words. In a finished film that is a distraction. And check the sound: play it with the volume up and confirm what we already know, which is that the file has no audio track. Confirming it yourself is worth ten seconds.
So one generation, five flaws noticed, none of them fatal. Here is the attitude I want you to have about that. Those seconds you decide not to use are not failure. They are rejected work, and rejected work is a normal cost of producing one accepted shot. A working production counts credits and seconds spent per usable shot, not per generated shot. If you throw away four clips to keep one, your cost for that shot was five clips. That is a number, not a shame. People who hide their rejected work cannot estimate anything, and they end up believing every generated second is footage.
Now, one change. Exactly one.
The flaw I am going after is the broken contact, because that is what an audience notices without knowing why. And I am changing one element of the brief to address it: the framing. I will go from a medium shot to a close shot on the hands and the parcel. The reasoning is that a close frame gives the model far more pixels for the thing that has to be right, and takes the face — the other flawed thing — out of the frame entirely.
There is a decision to make about the seed. If the composition from the first run had been what I wanted, I would open the asset in the session, read off its locked seed number, and reuse it, so the layout stayed put while I changed one phrase. But I am deliberately changing the framing, which means I want a new layout. So I leave the seed randomised. That is the rule: lock the seed when you want everything except your one change to hold still, and release it when your change is the layout.
The revised prompt, at thirty-nine words:
Close shot on the hands of a young bike courier in a grey rain jacket. Her fingers set a small brown paper parcel down on a wet stone step and release it. Soft light from the left. Static camera.
Generate. Another twenty-five credits, seventy-five left.
The second clip fixes the thing I aimed at. The fingers open, the parcel sits, the shadow under it lands where the light says it should. The face problem is gone because there is no face. And the revision brought in a new problem: in this version the parcel is already resting on the step when the clip begins, and her hands only let go. The placing is not there. As a piece of storytelling it is weaker, because a hand releasing something already at rest does not read as delivery.
I want to be straight with you: a targeted revision can come back worse, and often the way it comes back worse is exactly this — you fix what you named and lose something you were not watching. When that happens you have three honest options, and which one you pick is a budget decision as much as a craft one. You can accept the earlier clip, flaws and all, if its flaws are small enough to survive on screen. You can run the revision once more with the same single change plus one sentence protecting what you lost — in my case, adding that the hands lower the parcel onto the step. Or you can accept the new clip and adjust the brief around it, deciding that this shot is about release and something else will carry the placing.
I ran the third version: close shot, hands lowering the parcel onto the step and releasing it, same light, static camera. Twenty-five credits, fifty left. That one gave me the full movement, clean contact, and consistent light, with a faint wobble in the paper's edge in the last few frames. That is the one I am accepting. Three generations, fifteen seconds produced, five seconds kept.
One more thing about repeating a result, because people assume more is saved than is. In Runway, generations live inside sessions and your workspace assets, and when you select an asset it shows you the prompt text, the model version, the locked seed number, the duration and the aspect ratio. Re-running the same prompt text with the same seed number reproduces the framing, the subject layout and the motion path. That is a lot. But it is what the service chose to keep, and it is not everything you did. It does not hold the two versions you rejected, or why you rejected them, or the fact that you moved to a close shot because of a hand going into a parcel. Whatever the tool saves, you can repeat. Whatever it does not save, your own notes have to cover. That is not a workaround. That is what a production record is for.
So here is the finish line for this chapter, and it has two pieces sitting next to each other.
The first is the clip. Download it with the download arrow on the preview pane; it comes down as an MP4. Put it in a folder you will find again, with a name that says what it is rather than what the service called it — something like Mara, bakery step, close on hands, take three.
The second is a short written record beside it, in a plain text file or a note. Write down the service and the model version. Write down the prompt, exactly as you typed it. Write down the seed number, the duration and the aspect ratio, and how many credits it cost. Then one line on what you rejected and why: two earlier takes, first for a hand passing through the parcel and a face that changed, second because the parcel started already on the step.
Six lines. Two minutes of typing. And that record is the first item in a library you will keep adding to, because the moment you want this shot again, or a shot that rhymes with it, the note is the only thing that lets you start from where you finished instead of from a blank prompt box.
Where generated sound is arriving
Now something from the wider world of these tools, because the silent-picture rule I just gave you is not the same in every service, and it is worth knowing which part differs.
For the model we used, the picture arrives with no audio track and sound is a later step. But several hosted services have moved audio into the first pass, and if you are shopping around, that is the difference to watch for.
Hailuo AI released a model called MiniMax H3 at the end of July twenty twenty-six. Its headline is what they call an omni-pass: stereo audio generated in the same text pass as the picture, including room tone, footstep-and-object sound, and dialogue. It also runs single clips up to fifteen seconds at two-K resolution, across six frame shapes from very wide to upright phone. The production task that changes is delivery of a standalone shot — no separate sound step, no separate upscale. The smallest useful thing you could do with it today: pick MiniMax H3 in the Hailuo web console, set fifteen seconds at two-K, and add sound words to the end of your prompt the way you added light words — wind howling, gravel crunching — then listen to what comes back and ask whether the crunch lands on the frame where the foot lands.
Kling AI has done the same thing on a different schedule. Its VIDEO three-point-zero Turbo, updated through the middle of June twenty twenty-six, has native synchronised audio, clips up to fifteen seconds at ten-eighty or two-K, and it reads camera paths directly out of your wording. It has also scheduled its older models, versions one through two-point-one, for full retirement on the fifteenth of September twenty twenty-six, which matters if you have prompts you rely on there. The smallest next action: switch the model selector to three-point-zero, turn the audio toggle on, and write your camera move as an explicit direction — a slow push-in, a low-angle tracking shot — rather than hoping.
And Runway itself moved its default web generation to Gen-4.5 and Gen-4 Turbo in the middle of twenty twenty-six, bringing in ElevenLabs audio capability alongside the video models — still as its own step rather than inside the text-to-video pass — and, they say, better adherence to how things move and carry weight. That is also where the twelve credits a second on base models comes from. If motion is what keeps breaking for you, the small test is to check your credit balance, pick Gen-4.5, and prompt something with real physics in it — running water, footsteps on gravel — and watch whether the weight holds.
Every one of those tests is the same loop you just ran. Small brief, one change, and look hard at what actually came out.
