GnothiGnothi
SeriesFieldsCommunityPublishing
Sign inGet started free

5. Planning the Scene Before You Generate It

Summary

From an action to a scene

The two clips on the Shotcut timeline, Mara crouching to set a parcel on the bakery step and rising to turn toward the street, show an action but not yet a scene. Generation is the expensive step and story is the cheap one, so the work comes first on paper. A scene needs a want, an obstacle and a changed situation, and each beat must leave Mara somewhere different. In the illustrative version, a parcel that needs a signature meets a bakery closed for the week, and a lit window next door gives her a new route. The existing crouch and rise clips stay, now with a reason behind them.

A general chat assistant can draft the scene. Give it the fixed facts, then ask for less when the first draft gets too busy. Judging the draft is your job: with the sound off, can a viewer see what Mara wants, what blocks her and what changed?

A shot list with eyelines

The scene becomes a numbered list. Each shot has frame size, one action, screen direction, eyeline, length, and either an existing clip or a generation method. Only two of eight shots already exist. Writing eyelines down catches mismatches before they are generated: down to the phone, up and left to the sign, right toward the street.

Testing it as an animatic in Shotcut

Rough stills held at their planned lengths show timing before any credits are spent. Set a default still duration in the Properties panel with Set Default, then adjust each clip by trimming its edges or by splitting with S and ripple deleting with X. Forum threads cover changing the default duration of images and why durations can't be typed once a clip is on the timeline. Add street ambience on an audio track, because silence makes held frames feel long.

Watching the animatic exposes two gaps that the paper version hid. Mara never visibly reacts to the closed sign, and the lit window means nothing. The fixes are a new reaction close-up and a name plate on the door, which brings the list to nine shots.

Counting every take

Pricing assumes two takes per new shot and three for the hard ones. At the course's recorded rates, the new shots come to roughly 460 credits, or about 66 per accepted shot. Budgeting only the takes you keep would cover about half of that.

Keeping Mara the same

Five new shots must match Mara's face, hair, jacket and parcel. The shortcut is a flat-lit turnaround of front, three-quarter and profile views, generated from the accepted reference still. Feed two to four crops into the reference slots, along with an identical text description in every prompt, an approach also discussed in this guide to fixing character inconsistency. Pre-production suites such as Studiovity only pay off once there are several scenes to plan. Check each angle against frames from the crouch and rise clips, and regenerate any mismatch while it still costs one image.


A scene worth filming

Two clips now sit on the Shotcut timeline. In the wide shot, Mara crouches and sets the parcel on the bakery step. In the medium shot, she rises and turns toward the street, and the cut lands on her movement. It works. But watch it again and ask what it is about. A woman puts a box down and stands up. That is an action. It is not yet a scene.

This chapter is about what to do before you generate anything else. So far each new shot has cost real credits. The two takes of the rise cost a hundred. Every rejected take was money spent learning something you could have learned on paper. Generation is the expensive step. Story is the cheap one. A missing beat found in a notebook costs a minute of thinking. The same missing beat found after generation costs a reshoot, and you pay again for each take that misses.

So the order of work here is simple. First turn the action into a small scene. Then break the scene into a numbered list of shots. Then lay quick stills on a timeline at their planned lengths and watch them. Only then estimate what the real generation will cost.

A scene needs three things a viewer can follow. The character wants something. Something gets in the way. And by the end, the situation has changed. Mara setting down a parcel has no want and no obstacle. Nothing changes except the position of a box.

Here is an illustrative version, invented for teaching and not drawn from any real brief. Mara has to hand this parcel to a person. It needs a signature, so it can't just be left. She knocks at the bakery. Nobody comes. She sets the parcel down to check the address on her phone, and the address is right. Then she notices the sign in the window says closed for the week. Now she has a problem. She can't leave the parcel, and she can't hand it over. She picks it back up, looks down the street, and sees a light on in the flat above the shop next door. She walks toward it.

Count the changes. She arrives expecting to deliver. The closed door makes that uncertain. The sign makes it impossible. The light upstairs gives her a new route. Each beat leaves her somewhere different from where it found her. That is the test of a beat. If a moment could be cut and the situation would be the same afterward, it is decoration, not a beat.

Notice also that this version keeps the clips you already have. The crouch is now Mara putting the parcel down so her hands are free for the phone. The rise and turn toward the street become her looking for another way. You don't throw the footage away. You give it a reason.

An AI assistant can help you get here faster. Any general chat assistant will do for this stage, such as ChatGPT, Claude or Gemini. You are asking for writing, not video, so no special tool is needed. The useful way to ask is to give it the facts that can't change and the shape you need. You might write: a courier named Mara, one location, a bakery doorstep on a quiet street, a parcel that needs a signature. Under forty seconds of screen time. No dialogue. Give her a goal, an obstacle, and four to six beats, each one changing her situation. Keep the action she already does, which is crouching to set the parcel down, then rising and turning toward the street.

The first draft will often be too busy. Assistants like drama, so you may get a dog stealing the parcel, a storm, a phone call with a crying relative. Each of those adds things you would have to generate and keep consistent, and most can't be shown without dialogue. Revise by asking for less. Ask for fewer props, no new characters except someone glimpsed at the end, and every beat shown through something Mara does or looks at. A second pass usually gets much closer to the version above.

Then stop letting the assistant judge. Read the draft as a viewer would, with the sound off and no notes. Ask three questions. Can I tell what Mara wants without being told? Can I see what blocks her? Can I see what changed? Run the illustrative draft through them. The want is clear only if we see her knock and wait. A shot of her just standing at a door reads as patience, not intent. The obstacle is clear only if the viewer can read the closed sign, or at least see Mara read it and react. The change is clear only if the light upstairs is something we see her see. Three of those answers depend on a shot that shows Mara looking and then shows what she looks at. Keep that in mind, because the shot list has to supply it.

The assistant drafted the words. You decide whether they tell a story. A draft can sound complete and still hide a beat that never reaches the screen.

Now the scene needs to become something you can schedule. A shot list breaks the scene into numbered pieces, each one a single thing to generate. Each row says how big the frame is, what happens, which way things move, where Mara looks, and how long it runs. One more column matters for this course: which clip you already have, or which method a new shot would need.

Keep each shot to one clear action. The models you have used produce short clips, and one action per clip is what they handle best. Two actions in one generation give the model two chances to go wrong.

Here is the illustrative list, read row by row. The screen direction is fixed by the footage you already own. In the accepted rise, Mara turns toward the street, which we'll say is screen right. Everything after that has to respect it, by the 180-degree rule you already know.

Shot one is a wide shot. Mara walks into frame from the left, carrying the parcel, and stops at the bakery step. She moves left to right and looks at the door. Four seconds. Nothing existing covers it. It needs image-to-video, starting from the bakery-step still you already have with Mara added, so the location matches.

Shot two is a medium shot. Mara knocks and waits. She faces the door and looks slightly up, at about the height of a face behind glass. Three seconds. New, and image-to-video, starting from a still of her at the door.

Shot three is the wide shot of her crouching to set the parcel down. Three seconds. This is your accepted crouch clip. It is covered.

Shot four is a close-up. She checks her phone and her eyes drop to the screen. Two seconds. New. Text-to-video would struggle to keep her face. Image-to-video from a close reference frame of Mara is the safer method.

Shot five is a medium close-up through the window glass, or on the glass. The closed-for-the-week sign. This is what she sees. Two seconds. New. It has no person in it, so a plain still or text-to-video can cover it. A still with a slow push-in in the edit may be enough.

Shot six is the medium shot of Mara rising and turning three-quarters toward the street. She faces screen right. Five seconds. This is your accepted Animate Frames clip. It is covered.

Shot seven is a wide shot of the flat above the shop next door, a lit window. Her point of view. Two seconds. New, and the first-and-last-frame method in Animate Frames could carry a light switching on if you want that change on screen. Otherwise image-to-video from a still.

Shot eight is a wide shot. Mara picks up the parcel and walks out of frame to the right. Four seconds. New. Animate Frames, with a start frame matched to the end of shot six so the cut on action holds.

Two shots out of eight already exist. That is the first useful number the list gives you. Most of the scene is still ahead.

Look at the eyeline column again. Shots four, five and seven work only if Mara's look and the thing she looks at line up. She looks down at the phone, then up and slightly left at the window sign, then right toward the street. If the window sign shot is framed as if seen from the street side, her eyeline won't match it. Writing the eyeline down now prevents a generation that looks right alone and wrong in the cut.

Now test the list on a timeline, before spending any credits on it. An animatic is a rough version of the film made from still pictures held for their planned durations. It shows timing and order. It does not need to look good. A sketch, a phone photo of a doll on a step, or a frame exported from your existing clips will all do.

Make one still per shot. For shots three and six, export a frame from the clips you have, the same way you exported frames in the last chapter. For the rest, rough images are fine. Name each still with its shot number.

In Shotcut, set a default length for stills first, so each one doesn't land at whatever length Shotcut picks. Open one still with File, Open File, so it appears in the Source player. In the Properties panel, type a duration into the Duration field. Shotcut writes time as hours, minutes, seconds and frames, so three seconds is zero zero, zero zero, zero three, zero zero. Then click Set Default. Every still you bring to the timeline after that starts at three seconds.

Drag the stills onto the first video track in shot order. Shots three and six can go in as the real clips if you like, which gives the animatic a mix of moving and still frames. That is fine.

Now fit each still to its planned length. Once a clip is on the timeline, you can't type its duration there. Instead, hover over the edge of a clip until the red trim bracket appears and drag it. Inward shortens, outward lengthens. Or put the playhead where the shot should end, press S to split the clip there, select the leftover piece and press X. That is a ripple delete. It removes the piece and closes the gap, so the shots after it slide left.

Put the ambience underneath. If your project has no audio track yet, add one with Control plus U, or Command plus U on a Mac, or from the Timeline menu under track operations. Drop the ambience file on it. Street sound changes how long a held frame feels. Silence makes three seconds seem long.

Then watch it, from the start, at full speed, more than once. Watch for three things. Is any still held so long that nothing seems to be happening? Is any cut so quick you can't take in what you saw? And is there a moment where you, the viewer, don't know something you need?

In this illustrative pass, two problems show up. First, the cut from the closed sign straight to Mara rising feels like a jump. We see the sign, then she is already moving. We never see her take it in. The scene needs a reaction, a close shot of her face as she understands. Without it, the obstacle exists on screen but never reaches Mara. That was the risk the story check pointed at. The paper version hid it, because a reader fills it in. A timeline with nothing between two pictures does not.

Second, the lit window in shot seven reads as just another window. Nothing tells us it belongs to someone who might sign. One fix is to see a figure pass behind the curtain. Another is to give that window a sign or a name plate. The cheaper fix is usually the one that adds no new character, so the illustrative choice is a name plate on the door beside it, seen in the same wide shot.

Revise the list. Add a new shot between five and six, a close-up of Mara's face as she reads the sign, eyeline up and a little left, two seconds, image-to-video from the close reference frame. Change shot seven's action to include the door and name plate. The list now has nine shots, and two are covered. Rebuild the animatic with the new still and watch again. Repeat until a viewer who has never heard the story can tell you what happened.

The shot list now tells you what generation will cost before you spend anything. The rule from earlier chapters still applies. You pay for every take, not just the ones you keep.

Use the costs already in your records. Animate Frames costs about fifty credits for five seconds on its standard setting, and half that on Turbo. Gen-4.5 Image to Video costs twelve credits a second. Text-to-video for the empty sign shot can be skipped if a still with a push-in works.

Now an illustrative estimate. The rise took two takes. Assume new shots also need about two takes each, and the hard ones three. Shots one, two and four are image-to-video at four, three and two seconds, so a take costs forty-eight, thirty-six and twenty-four credits. At two takes each, that's two hundred and sixteen credits for the three. The added reaction close-up, at two seconds and maybe three takes because faces are hard, adds seventy-two. Shot seven, the window and door, as a three-second image-to-video shot at two takes, is another seventy-two. Shot eight, the exit in Animate Frames at fifty a take and two takes, is a hundred. Shot five, if a still with a push-in works, costs nothing to generate. Add those and the scene's new shots come to roughly four hundred and sixty credits. Spread over seven new accepted shots, that is about sixty-six credits per accepted shot, even though no single take costs that much.

That gap is the whole reason to count rejected takes. If you priced only the takes you keep, you would budget about half as much and run out halfway through.

The estimate also shows where the risk is. Five of the new shots show Mara's face or body. Each one has to match the woman in the shots you already have, in face, hair, jacket and parcel. And the street has to stay the same street in every wide shot. If one generation gives her a different collar, you pay for a reshoot. That is why the next step isn't generation. It is a set of reference sheets for Mara, the parcel and the street, made so every shot on this list starts from the same pictures.

A practical shortcut for keeping Mara the same

No release this week changed how reference images or planning work in the tools this course uses. So here is a shortcut that applies to the problem you just found.

Before generating any of the new shots, make a turnaround. That is a single set of images of Mara from the front, from three-quarters and in profile, all under the same flat, even light. Generate it in an image model, starting from your existing reference still of Mara so her face and jacket begin from what you already accepted. Flat light matters. If one angle has strong side light and another doesn't, the video model may read the shadow as part of her face.

Then use those angles as references. Runway's image reference inputs can take a cropped angle or two alongside the prompt. Kling's Elements offers similar slots. Two to four crops are typical. Pair them with the same short description of Mara, written the same way in every prompt, word for word. If shot two calls her a courier in a navy jacket, shot eight shouldn't call her a woman in a dark coat. A changed phrase invites a changed person.

Pre-production suites such as LTX Studio or Studiovity let you place camera positions and framing for a whole scene before rendering. For one short scene, your Shotcut animatic already does that job, so they are worth a look only when you have several scenes to plan.

The smallest useful next action is this. Take your revised Mara reference still. Make three flat-lit angles from it. Then hold each one next to a frame from the accepted crouch and rise clips. If the jacket, hair and face match all three, save them in the project folder as Mara's reference pack. If one doesn't match, regenerate that angle now, while it costs one image and not a video take.