OCDevel AI Video Generation

Picking Your Tool and Landing Your First Clip

Learning to read the board before you touch the text box

The habit that actually fixes those endless near-miss afternoons isn't a better prompt — it's picking a method for choosing a tool rather than a favorite tool, since rankings reshuffle every few weeks. The Artificial Analysis Video Arena works by blind pairwise preference, folded into an Elo-style score with a margin of error — a margin that often means "first" and "second" are really a tie. The board splits by job (text-to-video, image-to-video, editing) and again by audio versus silent output, because a talking clip and a silent one aren't competing for the same task. The chapter walks through picking two candidates off the right pool — one near the top, one better value — dating the note, and letting your own footage break the tie.

From there it covers what the dials around the text box actually do: aspect ratio, duration (five seconds is one beat, not a scene), resolution, and frame rate — and why resolution is the cheapest lever, worth dropping to draft tier while you're still deciding what you're asking for. Pricing gets walked through concretely: Runway's starting credits and per-second rates across its models, Luma's Dream Machine daily free credits and resolution-based cost jumps, Kling's daily refresh, and Pika's capped free tier. The core discipline that ties it together: change exactly one variable per generation, so a better result actually teaches you something, and watch for the "blind re-roll" — repeatedly regenerating the same wording and calling it iteration.

What's moved on the boards

A quick pass through recent releases: Kuaishou's Kling now produces synchronized audio and dialogue in the same pass, ByteDance's Seedance and Google's Veo have added native sound and first-and-last-frame control, and Runway's newer model has slipped out of the top spot on general text-to-video as faster, open-weight engines overtook it. OpenAI's original Sora dropped off the active boards after its consumer app closed. The takeaway offered isn't the gossip — it's that every name here is a snapshot with a date attached, and the board should be checked fresh before committing a project to any single tool.


Here is how the afternoon usually goes. You find a website that makes video. You type a sentence — a woman walking through a night market, neon signs, rain on the pavement — and you wait. Something comes back. It is almost right. The rain is there. The neon is there. But she is drifting rather than walking, and her hand does something no hand does. So you type the sentence again, maybe with an extra word or two, and you wait again. Something else comes back. Also almost right, wrong in a new way. You do that eleven more times. Then you look up and it is dark outside, you have a folder of near-misses, your free credits are gone, and you have learned nothing you could repeat tomorrow.

That is not a talent problem. Nobody told you three things, and all three are learnable in an evening. Nobody told you which tool to sit in, and why that choice depends on the job rather than on which one is famous. Nobody told you what the handful of settings around the text box actually do, so you left them on whatever they came set to. And nobody told you how to look at a bad clip and work out whether the model failed you or your request was never clear enough to answer. Without that third one especially, every clip is a coin flip, and flipping a coin faster is not a craft.

So let me start with the claim this whole thing rests on, and then show you it works. You are not going to choose a model. You are going to choose a method for choosing, because the standings in this field reshuffle every few weeks, and the method does not. The tool that is best at human bodies in motion this month may be third by the time you finish this chapter, and the sentence "use the one that wins on the job you are doing today" will still be true in a year. Learn the sentence. The names are rented.

Which means the first real skill is not prompting. It is reading a leaderboard.

There is a public leaderboard called the Artificial Analysis Video Arena, and the reason to send you there rather than to my opinion is that it is not an opinion. It works by blind pairwise preference. Two systems get the same request, a human is shown both clips without being told which came from where, and the human picks one. Do that hundreds of thousands of times and you can turn all those little votes into a score — the same kind of rating system chess uses, where beating a strong opponent moves you more than beating a weak one. Each entry on the board carries its rank, the lab that made it, the specific model version, that rating score, and a small plus-or-minus range around it, which matters more than people notice. If two models sit five points apart and each carries a range of give-or-take eight points, they are not really first and second. They are tied, and you should pick between them on price or speed or which interface you can stand.

Now the part that changes how you use it. The arena does not publish one list. It splits by job, because "good at video" is not a thing any more than "good at cooking" is a thing. There is a board for text-to-video, which means you type words and get moving pictures with nothing else supplied. There is a board for image-to-video, which means you hand the model a still picture as the first frame and it animates from there. And there is a board for video editing, where the input is already a clip and the model changes it. Those are kept in separate pools on purpose, because comparing a from-scratch generation against an edit of existing footage tells you nothing — it is a different task with different inputs.

Then it splits again, and this second split is the one beginners walk past. Each of those pools is divided into models judged with audio and models judged without. Some systems now produce sound in the same pass as the picture — footsteps, room tone, actual spoken dialogue with lips roughly matching. Others produce silence and always will. If those two were rated in the same pool, the ones with sound would win on novelty regardless of whether the picture was any good, so they are separated. Practically, that means you must know before you look which pool you belong in, and the answer comes from what you are making. If you are cutting a piece where you will lay in your own voiceover and music later, silent output is fine and you should read the no-audio side, where you will often find cheaper and faster options. If you want a character to say a line, you want the audio side.

You can also filter the board a few ways that are worth knowing about: current models only versus everything ever tested, and open-weights models — ones whose files you could in principle download and run yourself — versus proprietary ones you can only reach through somebody's website. And alongside the plain ranked list, the arena plots quality against two other things: how many seconds a system takes to produce a standard short clip, and what a minute of finished video costs through its programming interface. That second chart is the one that tells you where the bargains are, because quality and price are only loosely related, and there is usually something sitting well below the leaders on price while sitting very close to them on preference votes.

Let me do one pass through this out loud, so you have seen the motion.

Say the job is a five-second product shot for social media: a bottle on a wet stone surface, light moving across it, no talking. First question: which board? There is no dialogue and no sound I care about, so it is the no-audio pool. Second question: text-to-video or image-to-video? I already have a decent still of the bottle from a photoshoot, so image-to-video, which also happens to give me far more control over what the thing actually looks like, since I am handing over the look instead of describing it. So I open the image-to-video board, no-audio pool, current models only. I read down the top five. I ignore the exact ordering, because the ranges overlap. I write down two candidates rather than one — one that sits near the top, one that sits lower but much cheaper per minute on the price chart. And I write the date next to them, because in four weeks that list will not look like this, and a note that says which day I checked is the difference between a record and a rumour.

Two candidates, not one, for a reason worth spelling out. Crowd votes are an average over other people's prompts, not over yours. A model can win the board overall and be mediocre at your particular thing — glass, or hair, or a dolly move, or hands. The board narrows the field from forty options to two. Your own footage decides between the two. That is the whole division of labour, and it means you are never at the mercy of somebody else's benchmark.

One more piece of vocabulary before we move, because it governs everything after. A generation is one paid unit of work: one press of the button, one clip out, one deduction from your balance. Almost every hosted tool sells you credits rather than time, and then charges credits per second of video, with the rate depending on which model you picked and how good you asked it to look. On Runway, a browser-based suite with a family of models, a new account gets a one-time pot of a hundred and twenty-five credits that does not refill, and paid plans start at fifteen dollars a month for six hundred and twenty-five credits a month. The rate runs from about five credits per second of video on its fast turbo models up to ten on the older flagship and twenty-five on the newest, which means the same five-second clip can cost you twenty-five credits or a hundred and twenty-five depending only on which model you had selected — a five-to-one swing you can make by accident. Luma's Dream Machine, also browser-based, gives roughly eighty credits a day free, which in practice is about one short clip every twenty-four hours, and charges around a hundred credits for a five-second clip at the middle resolution against four hundred for the same five seconds at full high definition. Kling, from Kuaishou, refreshes about sixty-six credits daily and charges roughly ten credits for a five-second clip in its standard mode. Pika's free plan gives eighty credits a month, caps you at the lowest resolution, and charges twelve credits for five seconds there against forty for the same five seconds at full HD. All four watermark their free output and forbid commercial use of it, which is fine for learning and not fine for a client.

Take the shape of that rather than the numbers. Free tiers exist and are enough to learn on. The cost of one clip is not a fixed thing — it is the model times the resolution times the length, and you control all three.

Which brings us to the settings.

There are four dials that decide most of what comes out, they all sit near the text box, they are all set before you generate rather than fixed afterward, and most people never touch any of them. The first is aspect ratio, which is just the shape of the frame — wide like a cinema screen, tall like a phone held upright, or square. Choose it by where the clip will be watched, and choose it first, because it changes what you should ask for. A tall frame has no room for a person standing next to a car; it has room for a person, or a car. If you generate wide and crop to tall later, you throw away the sides of every shot you paid for, and the composition you liked is usually in the part you threw away.

The second is duration, and every one of these tools caps it in about the same place: five seconds by default, ten at a stretch, sometimes fifteen, with the option to extend a clip in further segments after the fact. That cap is not laziness. Coherence decays over time in these systems, and the longer the run, the more chances the model has to forget what a hand is. So the cap is protecting you, and it should reshape your request. Five seconds is one action, one camera move, one beat. Not a woman entering a market, buying fruit, and turning to leave. That is three shots, and three shots is three generations you will cut together. Learning that early saves you months, because "my clip fell apart halfway" is very often "I asked for a scene and bought a shot".

The third and fourth travel together: resolution and frame rate. Resolution is how much detail is in each frame, and the practical ladder on these platforms runs from a rough draft tier around four-eighty or five-forty, through seven-twenty, up to full high definition at ten-eighty, with four-K available at the top by generating it or by upscaling afterward. Frame rate is how many pictures go by each second, and everything here lands at the ordinary web rates, roughly twenty-four to thirty. That number is why smoothness is not usually a dial you fight: at twenty-four, motion reads as motion. What makes AI video look like a slideshow is almost never the frame rate — it is a model straining at motion it cannot hold together, and the fix is a different request or a different model, not a bigger number.

The reason I have put resolution last is that it is your cheapest lever and you should abuse it. Look again at those figures: on one of these platforms, five seconds costs twelve credits at the draft tier and forty at full HD. That is more than three drafts for the price of one finished clip. Nothing about your framing, your action, or your camera move needs full resolution to be judged. So you do your thinking at the bottom of the ladder and you spend at the top exactly once, on the version you already know is right.

That is the workflow, and here it is in the order you would actually do it tonight.

  1. Decide the job in one sentence, including whether you need sound in the clip itself and whether you are starting from words or from a still image.
  2. Open the arena board that matches that job, in the right audio pool, and take the two candidates you talked about — one near the top, one that looks like better value — with today's date beside them.
  3. Open one of them in its browser playground and set the frame shape and the length on purpose, before you type anything, based on where the clip gets watched and how many beats it holds.
  4. Drop the resolution to the cheapest tier it offers and generate one draft.
  5. Watch that draft once against what you asked for, change exactly one thing, and generate again.

That last step is the spine of everything, so let me make it a rule and then defend it. Change one variable per generation. One. Either the wording, or the length, or the starting image, or the model, or the aspect ratio — never two, and certainly never four.

Here is the mechanism, and why it is not fussiness. These systems are not deterministic in the way a calculator is. The same request run twice gives you two different clips. So the only way to learn anything from a pair of generations is if exactly one thing differed between them, because then the difference in the output has one candidate explanation. If you rewrote the prompt and switched models and pushed the length to ten seconds all at once, and the result is better, you have discovered nothing you can use again. You cannot tell which change earned it. Next time you want that quality you will have to guess all over again. Whereas if you changed only the phrase describing the camera and the drift stopped, you have learned something about this model's dialect that will keep paying out for as long as you use it.

The obvious objection is that this feels slow. It is slower per generation and dramatically faster to the clip you keep, because you are moving in a straight line instead of milling around. And there is a case where the rule genuinely bends: at the very start, when you know nothing about a model, throw two or three loose, quite different requests at it at draft resolution just to see the shape of its output. That is not iteration, it is reconnaissance, and it is fine. The rule applies the moment you have something almost right, which is the moment people abandon it.

Because the failure has a name, and you should be able to catch yourself doing it. Call it the blind re-roll. You dislike the clip, so you press generate again with the same words, hoping for a better draw. And you do get something different, because of the randomness, so it feels like progress. Twenty minutes later you have four wrong clips, a smaller balance, and no more understanding of the tool than you had at the start. The tell is one question, and you can ask it out loud: which variable did I just change? If you cannot name it in a few words, you did not iterate. You gambled. Sometimes gambling wins, which is exactly what makes it sticky.

The other half of that pitfall is a number people compute wrongly. When you compare tools you naturally look at what one generation costs, and one generation looks cheap — twelve credits, ten credits, a rounding error. But you do not deliver a generation. You deliver a clip you actually kept. So the number that governs your real budget is the total spent divided by the keepers, and if it took you fifteen tries to land one, a cheap model just got fifteen times less cheap. Two tools with the same posted rate can differ enormously on that measure, because the one that listens more closely to your wording gets you there in three attempts instead of fifteen. This is real ground and the show comes back to it properly, with the arithmetic worked out, once you have the habits that make the arithmetic worth doing. For tonight, just carry the shape of it: attempts are the expensive part, and every attempt you can explain is worth several you cannot.

So where does that leave you. You have a way of picking a tool that survives the tool being replaced. You have four settings you now set deliberately instead of inheriting. You have one clip you made on purpose, at draft quality, from a request that was one beat long. And you have the one habit that turns spending into learning.

What you cannot do yet — and I want to be honest about the size of the gap — is reliably land the shot that is already in your head. Right now you can describe a thing and get a thing, and steer it by trial. Getting the specific framing, the specific move, the specific light, the first time or near it, is a different skill: it is knowing what a prompt is made of and in what order a model wants to hear it. That is next.

Which brings us to the perishable part, the standings and the version numbers, so the rest of this stays true after they change.

The clearest pattern across the recent releases is that the labs have stopped competing on picture alone and started competing on control and on sound. Kuaishou's Kling arrived in a third generation, including an Omni variant, and what matters in it is not resolution — it does the usual seven-twenty and ten-eighty tiers — but that it now produces synchronized audio and dialogue in the same pass, anchors both a first and a last frame, takes fine motion paths, and accepts several reference images of a character at once. Through its programming interface it runs somewhere around seven to eight cents a second, which is roughly eleven dollars a minute on its professional mode. ByteDance shipped Seedance in two point-oh and two point-five, defaulting to seven-twenty with a four-K tier above it, ten seconds a take, native sound effects and dialogue, interpolation from a first frame to a last, direct motion paths, and reference binding for both style and character. Its rate lands nearer five and a half dollars a minute, which puts it among the better-value options at that quality. Google's Veo three point one carries native four-K and always-on sound design, plus first-and-last-frame prompting and explicit camera controls, and it is joined by a fast Omni Flash variant that renders quickly with sound at about six dollars a minute. MiniMax's Hailuo H3 added native audio alignment and is unusually obedient when animating a supplied still, at roughly seven dollars eighty a minute. Runway's Gen-4 and Gen-4.5 run up to ten to sixteen seconds at full HD, with motion brush, camera director modes and reference-sheet inputs, at twelve credits a second, which works out near sixteen dollars eighty a minute through the interface — the expensive end. Alibaba's Wan two point six and two point seven, plus a new architecture alongside them, do cinematic ten-eighty with camera direction, and sit at the premium end too, around seventeen dollars a minute and higher.

For a producer, the line through all of that is short: first-and-last-frame support and native audio have gone from rare to expected, and the price spread between the top of the board and the top of the value chart is now three-fold or more.

On the standings themselves, by job, as of this check. For prompt adherence in text-to-video, Google's Omni Flash, MiniMax's H3 and ByteDance's Seedance two point oh sit within a handful of rating points of each other at the summit, with Kling's third generation just behind. For native audio and dialogue, Google holds both top positions. For image-to-video, MiniMax's H3 leads with Seedance two point five and two point oh close underneath. For human motion and physics — bodies that stay one body, objects that stay one object — the Kling models keep the top spots. And on value per second, a small entry called P-Video from PrunaAI is the surprise, scoring respectably at about two cents a generation, with Seedance and Omni Flash forming the rest of that efficiency frontier.

What moved: OpenAI's original Sora came off the active boards entirely after its consumer app closed and its interface was slated to be retired, and Runway's Gen-4.5 slid from the summit down toward the twentieth-place class on general text-to-video, overtaken by the fast and open-weight engines from MiniMax, ByteDance and Google. That last one is the useful lesson, not the gossip. A model that was the obvious default a short while ago is now mid-table, unchanged, still perfectly capable — the field simply moved. Which is why the instruction stays the same as it was: treat every one of these names as a snapshot with a date on it, open the board yourself before you commit a project to anything, and let your own two-candidate test settle the rest.

Your next action is one of two things, depending on where you are. If you have no clip yet, open a free tier tonight, set the frame shape and the length before you type, generate one draft at the cheapest resolution, then change one word and go again. If you already have clips, open the image-to-video board in the no-audio pool, find the entry that scores near the leaders at a third of their price, and run your own last failed prompt through it. Then write down the date.