Writing a Video Prompt in Parts Instead of Wishes
The seven slots of a video prompt
A video prompt isn't a sentence, it's a form with seven fixed slots, and if you leave one blank the model fills it in with whatever's most average, which is where flat, stock-footage sludge comes from. The chapter walks through each slot in turn: subject (concrete, ruling-out detail rather than a category like "a woman"), action (one verb, one event, not a plot), camera move (static, push in, pull back, pan, orbit, handheld follow — defined plainly, with warnings about what happens when the slot is left blank), lens (wide vs. long, and what each does to background and space), lighting (direction, hardness, time of day), tone (mood and look kept separate, and kept to one or two words instead of a pile of adjectives), and pacing (how much can actually fit into a five-to-ten-second clip before motion starts to smear).
It lays out a practical routine: write the seven labels down, fill each with a short phrase, render once cheaply, then change exactly one slot at a time and re-render to learn what your specific tool actually obeys. It also names the failure mode this same lesson tends to produce — an overstuffed prompt with multiple ideas crammed into every slot, which renders as generic porridge rather than as an error — and gives three signs to spot it (genericness, vague motion, small edits stop changing the output), plus the fix: strip back to three slots and rebuild. A related trap gets flagged too: a camera move that physically fights the action, like orbiting around a subject who's also running at the lens, which produces smearing and warped geometry rather than what was asked for.
What moved this month
A quick pass through the current state of video generation: a blind-comparison arena, read job by job rather than as one ranking, shows new entrants from Alibaba and a MiniMax variant pushing into the top of text-to-video, image-to-video and editing charts, while incumbents like Runway's Gen-4.5 have slid down the same table since spring. On silent (no-dialogue) boards, a model called HappyHorse leads outright. Release-wise, ByteDance's Seedance 2.5 now does thirty seconds in one pass with synchronized audio, dialogue and effects included at no extra charge. The overall pattern: audio in the same pass is becoming standard, clip lengths are stretching well past the old five-second ceiling, and price per second keeps dropping.
You can render now. That is the whole of what the last episode bought you: one tool, paid into, opened in a browser tab, and one clip that came back and got looked at properly instead of being posted or deleted in a panic. Good. That is the floor.
Here is the thing that is still broken, and it is almost certainly still broken, because everybody's is. What you are typing into that box is a wish. It looks something like this: cool shot of a woman walking in the rain, cinematic. Then you press the button, you wait, and you take whatever comes out. Maybe it is lovely. Maybe her hands do something upsetting. Maybe the camera drifts sideways for no reason you asked for. And when it is wrong, you do not actually know what to change, so you rewrite the whole sentence, add the word "beautiful," and roll again.
That is gambling. It works often enough to feel like skill and fails often enough to eat your afternoon.
So let us fix the sentence. Not by making it longer or more poetic, but by understanding what it actually is. A video prompt is not a sentence. It is a form. It has slots in it, and each slot controls one specific thing about the clip. When you leave a slot blank, the model does not skip it. It cannot skip it. Every clip that comes out has a camera position, a lens, a light source, a length of motion. Those things have to exist for a picture to exist. So if you did not decide them, the model decided them for you, using whatever is most average in everything it was ever trained on. That is where the vague, sludgy, stock-footage feeling comes from. It is not the model being uncreative. It is the model filling in your blanks with the middle of the road.
Most bad shots are not caused by a bad prompt. They are caused by an incomplete one.
This show works from seven slots. Subject, action, camera move, lens, lighting, tone, pacing. That is the whole form. I want to walk each one slowly, because half of them use words from film crews that nobody bothers to explain, and a word you cannot define is a word you cannot use on purpose.
Subject is who or what is in the shot. This is the one everybody already fills in, and it is still usually the weakest, because people write a category instead of a thing. "A woman" is a category. There are billions of them and the model will average them into a face you have seen in a thousand adverts. What you want is enough concrete detail that someone could draw it without asking you a follow-up question. Age range, build, hair, what she is wearing, what condition she is in. "A woman in her fifties, short grey hair, soaked through, wearing a heavy navy raincoat that is too big for her." Now there is a person. Note what I did not do: I did not add "beautiful," "striking," or "8k hyperdetailed." Those are not description. They are hope. A useful test for this slot is to read your subject aloud and ask whether it rules anything out. "A woman" rules out nothing. "Short grey hair, soaked through" rules out almost everything, which is exactly the job.
Action is the second slot, and it is where most people quietly wreck the shot before the render even starts. Action is one verb. One thing happening. Not a plot, not a sequence, not a story. You are describing a few seconds of the world, and a few seconds of the world contains one event. "She walks toward the camera and stops, then turns and looks back over her shoulder, then smiles" is three events, and I will come back later to what happens when you ask for three. For now: pick the verb. She walks. Not "she walks, arrives, and reacts." If the blank version of this slot is "cool shot of a woman in the rain," the model has to invent motion from nothing, and what it invents is usually a slow drift with no intent behind it, which is the visual equivalent of someone shrugging. Before: a woman in the rain. After: she walks slowly toward the camera, shoulders hunched against the rain, not looking up.
Camera move is the third slot, and here I owe you vocabulary. This is the one where film people rattle off six words as though everyone was born knowing them, so let us define them in plain terms, because once you have them you will use them for the rest of the show.
A static shot means the camera does not move at all. It sits there like a tripod, and everything that moves in the frame is the subject. A push in means the camera travels forward toward the subject, so the subject gets larger and the world closes in around them. A pull back is the same thing reversed: the camera travels away, the subject shrinks, and you reveal more of where they are. A pan across means the camera stays in one place but rotates, like a person standing still and turning their head to watch something go past. An orbit means the camera circles around the subject while staying pointed at them, so you see them from changing angles and the background sweeps behind them. And a handheld follow means the camera is held by a person walking with the subject, so it moves with them and it wobbles slightly, which reads as urgency or documentary realism depending on the rest of the shot.
Six words. That is genuinely most of what you need for the next several months of work, and they cover most of what you have ever seen in an advert.
What happens when you leave camera move blank? The model picks, and it usually picks the same thing: a slow, unmotivated drift, often forward, because that is the most common camera behaviour in the footage of the world. It will look almost fine and mean nothing. Before: a woman walking in the rain. After: static shot, camera at chest height, she walks into frame and toward the lens. Suddenly the shot has a point of view. Somebody is standing there watching her come.
The fourth slot is lens, and this is the one where people either skip it entirely or type a number they do not understand. Let me give you the plain version. A lens setting controls two things at once: how wide or how tight the view is, and how much of the background stays sharp.
A wide lens, used close to your subject, takes in a lot of the room. It makes the space feel big and slightly stretched, it puts the subject in their environment, and it tends to keep the background fairly readable. Stand close to someone with a wide lens and their nose gets a bit large and the walls curve away behind them. It feels immediate, a little raw, a little uncomfortable.
A long lens, used from further away, takes in a narrow slice. It compresses everything, so the background looks like it is stacked right up behind the subject rather than receding away, and it throws that background soft and blurry while the subject stays sharp. This is what people are usually reaching for when they type "cinematic." They want the long lens look: subject crisp, world melted into shape and colour behind them.
Those are the two ends. You can say them in exactly those words in a prompt and the model will understand you, because those descriptions appear all over the way people write about pictures. If you leave the slot blank you get the default middle: a view roughly like human eyesight, with a background that is neither sharp nor soft, which is the least expressive option available. Before: cinematic shot. After: long lens from across the street, she is sharp and the rain and traffic behind her fall into soft blur.
Fifth slot, lighting. Three questions and you have filled it. Where is the light coming from, how hard or soft is it, and what time is it.
Where it comes from is direction. Light from behind the subject makes a rim of brightness around their edges and pushes their face into shadow. Light from the front flattens them out and shows everything. Light from one side carves out the shape of a face. Light from above and slightly to the side is what most rooms and most portraits look like.
Hard or soft is about the edges of shadows. Hard light comes from something small and bright — the sun on a clear day, a bare bulb, a single street lamp — and it makes shadows with sharp edges and strong contrast. Soft light comes from something large and diffuse — an overcast sky, light bouncing off a white wall, a window on a grey afternoon — and its shadows fade off gently. Hard is drama and tension. Soft is calm, warmth, kindness, and also blandness if you are not careful.
Time of day carries colour and angle for free. Late afternoon is low and warm. Midday is high and harsh. Blue hour after sunset is dim and cold and even. Night is whatever people have switched on.
Leave this blank and you get flat, even, sourceless light — the lighting of a stock photo, which is nobody's lighting. Before: rainy street. After: night, wet street, the only light is a single sodium street lamp behind her, so she is rimmed in orange and her face is mostly in shadow.
That is one line of writing and it has done more for the shot than any adjective could.
Sixth slot is tone, and I am going to split it into two things that people mash together, because mashing them is how prompts go bad. Tone is the mood: what the shot should make somebody feel. Lonely. Tense. Tender. Triumphant. Separately, there is the look: the visual treatment. Grainy. High contrast. Desaturated, meaning the colours are pulled toward grey. Warm. Clean and glossy.
Both are legitimate. Both belong here. And this is the slot where everything goes wrong, because this is the bucket that adjectives fall into, and adjectives are cheap. It is very easy to type: cinematic, moody, epic, dramatic, award-winning, atmospheric, beautiful, professional. Eight words, zero information. Every one of those either means nothing specific or means the average of everything, and stacking them does not narrow the shot down, it widens it out. Two is plenty here. One mood, one look. Before: cinematic, moody, epic, dramatic, atmospheric. After: lonely and quiet, desaturated with visible grain. That is a decision. A machine can act on a decision.
The seventh slot is pacing, and I have saved it for last because it is the single most common reason a clip comes back mushy, and almost nobody thinks of it as a slot at all.
Pacing is how much should happen inside the clip's few seconds. That is it. And the reason it matters so much is arithmetic. Your tool is probably giving you something in the range of five to ten seconds. Take five. Five seconds is not very long. Say "one" out loud five times at a normal pace and watch how little of anything you could do in that window. You could take four or five steps. You could turn your head and settle. You could raise a cup to your mouth. That is genuinely the budget.
Now think about what people type. "She walks toward the camera, stops, turns, looks back over her shoulder, and smiles sadly as the camera pushes in." Five actions and a camera move, in five seconds. The model does not refuse. It does something worse: it tries. It attempts to fit all of that into the window, so nothing gets its proper duration, everything happens at once and at speed, limbs pass through each other, the face changes mid-turn because there was no time to be one face, and the whole thing has a smeared, rushed, dreamlike quality that you will recognise instantly once you have seen it named.
So fill the slot. Say how much happens. Before: she walks toward the camera, stops, turns, looks back and smiles. After: one continuous action across the whole clip, she simply keeps walking, no cuts, nothing else happens. That sounds boring written down. It will look ten times better than the ambitious version, because it is achievable in the time available. And you get the turn and the look back by making them a second shot later, which is what a real crew would do anyway.
Now put it together, because the point of naming seven slots is to stop improvising and start filling in a form.
Open a plain text file, or a note on your phone, or whatever you already use. Write the seven labels down the left, one per line: subject, action, camera move, lens, lighting, tone, pacing. That is your prompt sheet. It takes twenty seconds to make and you are going to use it for months.
Then fill each one in with a short phrase. Not a paragraph. A phrase, the way I have been doing. Then assemble them into a prompt in that order — subject, then action, then camera, then lens, then lighting, then tone, then pacing — and render it once, at the shortest duration your tool offers and at whatever its default resolution is. Short and cheap, because this is a draft pass, and drafts are for information rather than for beauty.
Our worked example, assembled, reads roughly like this: a woman in her fifties with short grey hair, soaked through, in a navy raincoat too big for her, walking slowly toward the camera with her shoulders hunched, static camera at chest height, long lens from across the street so the traffic behind her falls into soft blur, night, lit only by a single sodium street lamp behind her, lonely and quiet, desaturated with grain, one continuous action for the whole clip and nothing else happens.
That is not a wish. Every part of it is a decision you made, which means every part of it is a decision you can change.
Which brings me to the revision rule, and it is one line long. Change exactly one slot. Re-render at the same settings. Keep both clips and write one line about what moved.
You already know this move. Last episode we did the render, look, adjust, render again cycle, and the discipline was to actually watch what came back instead of immediately rerolling. Same discipline, aimed better. Before, you were changing the entire sentence and comparing two clips that differed in every possible way, which taught you nothing you could reuse. Now you change camera move from static to a slow push in, leave the other six alone, and re-render. Whatever is different between those two clips was caused by that one word, because nothing else moved. That is not a vibe. That is a fact about your tool, and it is yours to keep.
Do that six or seven times over an afternoon and you will know things about your generator that no leaderboard can tell you. You will know whether it obeys "static" or drifts anyway. You will know whether it understands "long lens" or ignores it. You will know how much action it can actually hold in five seconds before it starts to smear — and that number is specific to your tool and is genuinely valuable.
Keep the file. I mean this literally, and it is going to sound like housekeeping right now and stop sounding like housekeeping around the time we start building scenes with more than one shot in them. The prompt sheet is not a beginner's crutch you outgrow. It is the substrate. When we get to holding a character steady across five shots, what you will be holding steady is a subject slot. When we get to making a whole sequence sit in the same world, what you will be locking is a lighting slot and a tone slot. The sheet grows into the thing you actually work from. So keep it somewhere you will find it, and start a new one per shot rather than overwriting the last.
Now the pitfall, and it is a very particular one, because it is caused by this episode.
Here is what happens. A listener learns there are seven slots. Seven slots feels like seven opportunities. So they fill each one generously. The subject gets three ideas in it, because she could be tired and also defiant and also elegant. The action gets three verbs, because more is happening. The camera does a push in and an orbit. And then, at the end, for safety, they add five style adjectives: cinematic, hyperreal, moody, dramatic, film-grain masterpiece. It is a magnificent prompt. It reads well. It renders as porridge.
The clip comes back and it appears to ignore everything you asked for, and this is the confusing part, because you asked for so much. So let me tell you how to recognise it, because it does not look like an error.
The first sign is that it is generic rather than wrong. A prompt that is missing a slot fails in a specific direction — the light is flat, the camera drifts. An overstuffed prompt fails toward the middle. It looks averaged. Competent, glossy, and about nothing. Second sign: the motion is vague. Not broken, just uncommitted, as though the subject is doing several things at ten percent each. Third sign, and this is the reliable one: small changes to the prompt stop changing the output. You swap a word, you re-render, and you get essentially the same clip. That is the tell. It means your instructions are now competing with each other so heavily that no single one has enough weight to steer anything. You have written a prompt with no handles on it.
The fix is subtraction, and it should feel almost insultingly aggressive. Strip back to one subject, one action, one camera move. Three slots, one idea each. Delete every adjective. Re-render. That clip will probably be plainer than you want and it will be responsive, and responsive is what you need, because now you add back one slot at a time using the revision rule and watch each one actually land. You will usually find you get where you were going with four or five filled slots and no adjective pile at all.
There is a quieter version of the same problem that is worth knowing by sight, because it looks like the model is broken rather than like you overpacked it. It happens when the camera move fights the action. The classic is an orbit around a subject who is also running toward the lens. Think about what you have asked for: the camera is circling to keep her in view from changing angles, and she is charging out of the space the circle was drawn around. Those two instructions cannot both be honoured, so the model splits the difference, and what you get is smearing, or geometry that seems to slide and warp, or a background that bends in a way no real place bends. Legs sometimes acquire an extra one.
The cure is not a better prompt. It is picking a side. Either she runs and the camera is static or follows handheld, or she stands and holds still and the camera orbits. Before you render anything with a camera move in it, do a five-second sanity check: could a real person with a real camera physically do this, in this time, around this action? If a human operator would have to be in two places at once, the model is being asked for the same impossibility and will hand you a compromise.
Two things I am deliberately not opening today. Telling the model what to keep out of the shot, and the fact that different models want to be spoken to in noticeably different ways — both of those come later, once you can fill the form. And there is a bigger lever than any of this, which is making a still frame first and animating that instead of describing everything in words; that arrives in a few episodes and it changes the game, so let this one settle first.
Your homework is the sheet. Seven labels, one shot, one render, then six renders where you change one slot each. Keep them all. Keep the notes.
What moved this month
Standings first, and the usual health warning, which is not a formality: the public blind-comparison arena that this show sends you to is a snapshot, it is rebuilt from fresh human votes constantly, and the order genuinely reshuffles month to month. Read it by job — text to video, image to video, editing, with audio and without — rather than looking for an overall winner, because there is no such thing.
The big movement is that Alibaba's Wan 3.0 and a MiniMax H3 variant post-trained by fal have both walked straight into the top three across text to video, image to video and editing. Wan 3.0 currently leads text to video with audio and leads editing with audio outright. Google's Gemini Omni Flash is sitting essentially level with it on text to video, well within the error bars, which is the point in the table where ranking stops meaning anything and you should be testing on your own shots. Image to video with audio is led by that MiniMax variant, with ByteDance's Seedance just behind. The casualties are the previous mainstays: Runway's Gen-4.5 has slid down the table since spring, and Google's Veo 3.1 family has drifted into the middle of the with-audio rankings despite being extremely capable, which tells you those boards measure preference rather than quality. On the silent boards the order is different again — a model called HappyHorse leads both silent categories — and if your work has no dialogue, the silent board is the one that describes your job.
On releases, freshest and biggest first. ByteDance launched Seedance 2.5 at the end of July and opened the developer interface a week later, and the headline is thirty seconds in a single pass, with multi-round extensions on top, at 480p, 720p or 1080p, and with synchronized audio, effects and dialogue generated in the same pass at no extra audio charge. Thirty seconds in one go changes what a shot is; you can stop thinking exclusively in five-second beats. Published rates run from around ten cents a second at the low resolution to about twenty-three cents at 720p. Next action: run one prompt at the shortest length on a hosted playground and check whether the dialogue actually locks to the mouth.
Alibaba's Wan 3.0 landed on third-party workflow platforms in late August, does 480p through 1080p with joint video and audio and lip sync, and the open-weight versions are under a permissive license, which matters if you ever want this running on your own terms. Lightricks put out LTX-2.3, an open model that reaches 4K at fifty frames a second with audio in a single pass, free to run locally for anyone under ten million in revenue and paid above that. Runway's Gen-4.5 arrived alongside a high-dynamic-range grading model, reaching 4K at up to sixty frames and clips of ten to sixteen seconds, with multi-cut storyboarding — but audio stays a separate step there rather than coming out of the same pass. MiniMax's H3 does fifteen seconds at 2K with stereo audio and dialogue at thirteen cents a second. xAI's Grok Imagine 1.5 Video does fifteen seconds with native audio at five to eight cents a second, which is the cheapest thing on this list by some distance and worth a draft pass. And Google's Gemini Omni Flash does ten-second clips with synchronized audio and conversational editing, meaning you can ask it for changes in plain language, at around ten cents a second.
The pattern across all of that: audio in the same pass is becoming standard, clip length is stretching well past the old five-second wall, and the price per second is falling. Your smallest next action this week is to open the arena page, look only at the row for the job you actually do, and note the top three. Then go fill in a form.
