GnothiGnothi
SeriesFieldsCommunity
Sign inGet started free

Picking One Tool and Landing Your First Clip on Purpose

Picking a tool without drowning in options

An AI video generator turns a text description, a still image, or both into a short clip, usually five to ten seconds long, and almost all of that work happens through hosted, pay-by-the-second websites rather than anything you install. These tools aren't editors and they aren't cameras — they're closer to shot factories, producing raw seconds that still need everything else filmmaking usually involves built around them. The people who actually benefit are working creators with deadlines and invoices, not hobbyists chasing a pretty feed clip.

The core discipline here is resisting tool-shopping. Every model has quirks — how it handles camera speed, hands, or the word "slowly" — and that feel only comes from rendering dozens of clips through the same door. Comparing tools honestly means reading a leaderboard by job rather than by overall score: the Artificial Analysis Video Arena, which also runs as a public space on Hugging Face, pits two blind clips from the same prompt against each other and builds ratings from thousands of human votes, split by text-to-video, image-to-video, editing, and audio versus silent output. First place inside a tight cluster is noise; a real gap between tiers is information.

Weighed against cost, resolution, commercial rights, and watermarking, the recommended starting point for someone with nothing yet is Kling's entry tier, which unlocks 1080p, commercial use, no watermark, image input, and audio in the same pass at once — something none of Runway, Luma, or Pika manage on their own cheapest paid plans (Pika, for instance, requires the tier above its cheapest just to remove the watermark and get sellable output). Anyone already paying for Runway is told to stay put and simply switch to its cheapest model rather than migrate.

From there the workflow is concrete: one short prompt, one subject and action, a deliberately chosen aspect ratio and the shortest, cheapest render setting available. After each clip, name exactly one difference from what you pictured, change exactly one thing, and render again — the antidote to the "reroll spiral" of pressing generate repeatedly on an unchanged prompt. Keeping a two-line log per render and tracking cost per finished clip, not per generation, is what turns a hobbyist into someone who can quote a client a real number.

What moved this month

Standings this cycle show Alibaba's Wan 3 and Gemini Omni Flash effectively tied at the top of text-to-video-with-audio, with MiniMax H3 and Seedance 2 close behind and a real gap opening below that to Kling 3 and Veo 3.1. On the release side, Runway's Gen 4.5 now handles four-to-twenty-second 1080p clips with reference images for character consistency, Kling has added native multi-shot generation and five-language lip-synced dialogue, and Luma's Ray2 brings composable, stackable camera moves — orbits, cranes, push-ins — layered directly onto image-to-video prompts. MiniMax's newer Hailuo models push resolution and speed but still leave audio as a separate pass.


Somewhere in the next few minutes you are going to type a sentence into a website, wait about a minute, and get back a piece of moving footage that never existed. That part is easy. Almost anyone can do it. The hard part, and the reason this show exists, is doing it again tomorrow on purpose, for a client, at a price you can live with.

So let's be precise about what we're dealing with. An AI video generator is a model that takes a text description, or a still image, or both, and produces a short piece of video — usually five to ten seconds. You reach most of them through a hosted generator: a website where somebody else runs the expensive hardware and you pay by the second of output. You don't install anything. You don't own a graphics card. You open a browser tab, type, and pay. That's the layer we're working in, and it's the layer where the money is being made right now.

That puts these tools in a strange spot among their neighbours. They are not editors — they do not cut footage together or fix your sound. They are not cameras, because there is no set and no crew. They sit further upstream than either: they are shot factories. Each one gives you raw seconds, and everything you'd normally think of as filmmaking still has to happen around them. What makes them worth your attention is the price and the speed. A shot that would have cost you a location, a permit and a day now costs you a few cents and a minute of waiting.

Which is why the people this matters to are working creators. A marketer who needs six social cuts by Thursday. An indie filmmaker who cannot afford a helicopter. A small agency that wants to pitch three concepts instead of one. Not hobbyists making pretty clips for their own feed — people with deadlines, invoices and clients who notice when the hands look wrong.

Here is the horizon this whole show is walking toward, and then I want to shrink it right back down. The end state is one person taking in a brief and putting out a finished, on-brand cut, without a crew. That is real, and it is achievable, and it is a long way from where most people are standing. So we are not going there today. Today has one rung on it, and only one: you are going to pick a single generator, deliberately, and get one clip out of it on purpose, knowing what that clip cost you before you pressed the button. That's it. If you finish this and you have chosen one tool, rendered one shot, and written down one number, you have done the whole thing.

The reason I'm being so stubborn about the word "one" is that the most common mistake in this field is not a bad prompt. It's tool-shopping. Every week there is a new model with a better demo reel, and there is a version of this hobby where you spend a year signing up for free trials and never deliver anything. The problem with switching platforms every week is not the money. It's that switching resets the only thing that actually compounds, which is your feel for how one particular model answers you.

That feel is real and it is specific. Every model has habits. One of them will consistently make your camera moves too fast. One will ignore the word "slowly" unless you also say what is moving slowly. One will give you gorgeous light and mangle anything with human hands in it. None of that is written down anywhere, because it isn't a feature — it's a personality, and you only learn it by rendering forty clips through the same door and noticing the pattern. A person who has done that with a mediocre model will out-produce a person who has done it with nothing across five excellent ones. So: one tool. For at least a few weeks. Long enough to learn its accent.

But you should choose it with your eyes open rather than because you saw an ad, and that means learning to read a leaderboard.

A leaderboard here is a public ranking of video models, updated continuously, built out of ordinary people voting. The main one worth your time is the Artificial Analysis Video Arena, which also runs as a public space on Hugging Face, the site where a lot of the machine-learning world keeps its models and demos. The mechanic is simple and it is the reason the numbers mean anything. You are shown two clips generated from the exact same prompt, side by side, with the model names hidden. You pick the better one. You are not told what you just voted for until after. Do that tens of thousands of times across thousands of people and you get a rating for each model, the way chess players get ratings: win against a stronger opponent and your number climbs more than beating a weaker one. Those ratings come with a margin of error attached, which matters more than most people notice, and I'll come back to that.

There's a second kind of benchmark alongside it, an academic one called VBench, which doesn't ask humans anything. It scores clips automatically on things like frame quality, whether motion is smooth over time, whether the subject stays the same subject, and whether the video actually matches what the prompt asked for. It's useful for a different reason: it tells you where a model breaks down rather than whether people liked it.

Now the part that actually changes how you work. Do not read a leaderboard for the overall winner. There isn't one. The arena is split by job, and those splits are the whole value of the thing. There is text-to-video, where you describe a shot from nothing. There is image-to-video, where you hand it a still picture and it animates that picture. There is video editing, where you feed in footage you already have and ask for a change. And crucially, text-to-video is split again by audio: one board for models that generate synchronized sound in the same pass, one for models that hand you a silent clip. There are also filters for whether a model is closed and API-only or has openly published weights you could in principle run yourself, and filters for how long a generation takes and what it costs.

Those splits exist because the models genuinely specialise. One wins on doing what it was told. Another wins on human bodies moving in a way that doesn't make your skin crawl. Another wins because it makes usable sound in the same breath as the picture. Another wins on nothing except being cheap enough that you can afford to be wrong twenty times.

So here's the move, and it transfers no matter what the standings look like when you read this. Look at the shot you actually need this week. Not a shot in general — the one on your list. If it's a product still that needs to drift and catch the light, you are in image-to-video, and the text-to-video board is irrelevant to you. If it's a talking person, you want the audio board, because dubbing lips afterward is a whole separate fight. Find the one column that matches that job, take the top two or three names off it, and stop reading. Do not scroll. Do not open a second tab.

And take the numbers loosely, which is where that margin of error earns its keep. Right now the top of the text-to-video board with audio is a knot of four models within about twenty rating points of each other — Alibaba's Wan three, Google's Gemini Omni Flash, MiniMax's H3, and ByteDance's Seedance two — and a gap that small is not a gap. Those four are tied, whatever order they're printed in. Meanwhile the board underneath them holds names like Kuaishou's Kling three and Google's Veo three point one clustered a hundred points lower, and that gap is real enough to feel. The lesson is to read for tiers, not for rank. First place is noise. First tier is information.

One more thing about all of this, said plainly so you don't get attached: everything I just told you about who is on top will be wrong soon. These standings reshuffle weekly to monthly, because labs ship new checkpoints and small fast variants constantly. Treat any specific ranking, including any I give you, as a photograph of one afternoon. The skill of reading the board by job survives. The names on it do not.

Which brings us to the actual choice, and I am going to make it rather than hand you five options and wish you luck.

The things that bite a paying producer are boring and there are only about six of them. What does a second of output cost you. How long can one clip run. What resolution and how many frames per second. Can you feed it a still image to animate, or is it text only. Does it make sound. And, the one people forget until it hurts, are you allowed to sell the result, and does it come out with a watermark stamped across it.

Let me weigh the real hosted options against exactly those.

Runway is the most established of the browser-based platforms, and it's the deepest — it has motion controls where you paint on the frame to say what moves, it has audio and lip-sync built in, and its clips can be extended. Its entry plan is fifteen dollars a month, or twelve if you pay for the year, and gives you a monthly credit allowance. Credits burn at different rates depending on which model you pick: its fast turbo model spends five credits per second, its standard model about twelve, its flagship anywhere from twelve up to twenty-five depending on configuration. Translated into money on that entry plan, that's roughly twelve cents per second of output on turbo and somewhere between about twenty-nine and sixty cents on the flagship. Clips run five to ten seconds, at twenty-four frames per second, in seven-twenty or ten-eighty with an upscale to four-K available. There's a free tier, but it's a one-time allotment of a hundred and twenty-five credits with a watermark, which is a demo, not a workflow.

Luma's Dream Machine is the elegant one, and its camera controls are genuinely nice — you can ask for an orbit or a crane move in plain language. It also has a draft mode that is startlingly cheap, around one cent per second, which is exactly the right shape for iterating. But its entry plan starts at thirty dollars a month, three times Runway's, and the core generator does not produce synchronized audio. Your clip comes out silent and you go find sound elsewhere.

Pika is cheap and fast and pleasant, ten dollars a month, and low-resolution drafts cost about three cents a second. The catch is licensing: on current terms, watermark-free downloads and commercial rights only start at the plan above that, in the twenty-eight to thirty-five dollar range. Which means the cheap tier is for playing, not for delivering, and if you're reading this show you intend to deliver.

MiniMax's Hailuo is the value play. Its entry tier lands around eight to ten dollars, draft generations run about three cents a second, and it can push clips out to fifteen seconds, longer than most. But its base models leave you silent too, so audio is again a second job.

Kling, from Kuaishou, is what I'd put a beginner on, and the reason is not that it's the best model. It's that its cheapest paid tier is the only one that doesn't cripple you somewhere. Ten dollars a month at list price, less on annual, gets you six hundred and sixty credits. Silent footage at seven-twenty costs six credits a second, ten-eighty costs eight, and if you want native synchronized audio with dialogue and effects it's nine credits a second at seven-twenty or twelve at ten-eighty. In money, that's roughly nine cents a second for silent seven-twenty up to about eighteen cents a second for ten-eighty with sound. Six hundred and sixty credits therefore buys you somewhere between about fifty-five and a hundred and ten seconds of finished video a month, depending which of those you choose.

That is not much video. Sit with that for a second, because it's the most useful thing in this episode. A hundred seconds. Twenty clips of five seconds, if you never make a mistake. And nobody never makes a mistake. That constraint is not a problem with the plan, it's the actual shape of the job, and pretending otherwise is how people burn a month's allowance in an evening.

What Kling's entry tier gives you that the others don't is everything unlocked at once: ten-eighty output, commercial use, no watermark, still-image input, and audio in the same pass. Pika makes you pay triple for the watermark to go away. Luma starts at three times the price and hands you silence. Hailuo is cheaper but silent. So: if you are starting from nothing, start there.

The one exception. If you already pay for Runway — a lot of people do, because it got there first — do not switch. Stay, and just change which model you're calling. Use the turbo variant for everything at first, at five credits a second. It is not the prettiest thing Runway offers and that is entirely the point: it's cheap enough to be wrong on, and being wrong cheaply is the skill we're building. Your existing subscription plus discipline beats a new subscription plus curiosity.

Now open it, because the rest of this only works if you're doing it.

Every one of these tools has a playground: the plain web page with a text box, a few dropdowns, and a generate button. That's where you live. Ignore the templates and the gallery and the featured community clips. You want the box.

Write one short prompt for one simple shot. Short and simple are both load-bearing. A prompt is just your instruction to the model, and a beginner's instinct is to write a paragraph of adjectives, which gives the model forty things to obey and guarantees it drops half of them. Instead: one subject, one action, one setting. A red enamel coffee cup on a windowsill, steam rising, morning light. That's a real shot. It has something to do. It has nothing in it that AI models famously mangle — no hands, no crowds, no text on signs, nobody talking.

Then set the aspect ratio yourself, deliberately, instead of accepting whatever the tool defaults to. Aspect ratio is the shape of the frame: wide, like a television, or tall, like a phone held upright. This is a decision about where the clip is going, not a taste question. If it's for a phone feed, choose tall now, because cropping wide footage down to tall later throws away most of your picture and usually decapitates your subject.

Set the duration yourself too. These tools mostly offer five or ten seconds. Choose five. Ten seconds is twice the price and, at this stage, twice the wrongness — models tend to hold together for a few seconds and then start drifting, melting, or inventing a second camera move you never asked for.

Then render at the cheapest quality available. If there's a draft mode, use draft mode. If there's a turbo model, use turbo. If the only lever is resolution, choose seven-twenty over ten-eighty. Resolution is how many pixels tall the picture is, and frames per second is how many still images make up each second of motion — twenty-four is the standard almost everything here outputs, and it's what film has always used, so you rarely need to touch it. Neither number will fix a shot that's wrong. A beautifully rendered wrong shot is still a wrong shot, and it cost you double.

Then the step people skip. When the clip comes back, look at it against what you actually pictured before you pressed the button. Not "is this good" — that question has no answer and leads nowhere. The question is "what specifically is different from what I wanted." Maybe the steam looks like smoke. Maybe the camera drifts left for no reason. Maybe the light is evening light, not morning. Name one difference, out loud if it helps.

Then change exactly one thing, and render again.

One thing. That's the whole discipline, and it is the difference between a hobbyist and a producer at this stage. Not a better model. Not a secret prompt. Deciding on purpose instead of rolling dice.

Which brings us to the pitfall, and I want you to recognise it in yourself, because you will do it. It's the reroll spiral. You get a clip that's nearly right. You press generate again on the same prompt, unchanged, hoping the dice land better. It comes back differently, because these models have randomness built in, so you press it again. And again. Forty minutes later you have eleven clips, none of them right, no idea which change would have helped, and a meaningfully lighter credit balance.

Here's how to catch it. If you are three renders deep and you have not changed a single character of your prompt or a single setting between them, you are in the spiral. That's the test. Not how you feel — whether anything changed.

The fix is two habits. First, one variable per render, as we just said. Second, keep a note — literally two lines per render, in whatever text file is nearest. What you changed, and what happened. "Added 'slow push in' — camera now moves, too fast." That note is worth more than the clips. In a week it becomes your private map of how this particular model hears you, and that map is the compounding thing I mentioned at the top.

And then the number, which is the third habit and the one that makes you a professional. Stop measuring your spend as cost per generation and start measuring it as cost per finished clip you would actually hand to a client.

Do the arithmetic with me on Kling's entry tier, since we chose it. Silent seven-twenty at six credits per second, on a ten-dollar plan, is about nine cents per second. A five-second clip is therefore about forty-five cents. That sounds like nothing. But if it takes you ten attempts to land that shot, the finished clip cost four dollars and fifty cents, and your six-hundred-and-sixty-credit month bought you about twenty-two deliverable clips, not a hundred and ten seconds of anything. Now suppose the note-keeping and the one-change-at-a-time discipline get you there in four attempts instead of ten. Same model, same prompt skill, same subscription. Your cost per finished clip drops to about a dollar eighty, and your month suddenly holds fifty-odd deliverable shots.

Nothing about the technology changed. The only thing that changed is that you stopped rolling dice. That is what this show is going to teach you over and over, at bigger and bigger scales, and it starts at this exact size.

So the smallest next action, and then I'll get to what moved this month. Choose one tool tonight, using the one leaderboard column that matches the shot you need. Render one shot at the cheapest setting you have. Write down one number: what that finished clip cost you, all attempts included. One tool, one shot, one number. That's the rung.

What moved this month

Standings first, and read them as a snapshot. Text-to-video with audio has Wan three on top at about twelve forty-one Elo on roughly fifty-eight hundred samples, effectively tied with Gemini Omni Flash at twelve thirty-seven on a much larger sixteen-thousand-six-hundred-vote base — treat that as a coin flip, not a lead. MiniMax H3 sits at twelve twenty-five, Dreamina Seedance two at seven-twenty at twelve twenty. The second tier is Wan two point seven around eleven fifty-six, the HappyHorse variants between eleven twenty-one and eleven forty-five, then Kling three's tiers between about ten eighty-eight and eleven-oh-six, SkyReels V4 at ten ninety-nine, and the Veo three point one family at ten eighty-eight to ten ninety. Image-to-video has now crossed one point seven million votes across forty-five models, and its top is a similar contest between Kling, Wan and Seedance image-conditioned checkpoints. New arrivals on the boards: Wan three, MiniMax H3, LTX two point five, and Gemini Omni Flash. Smallest next action: open the arena and filter it down to image-to-video, then vote thirty pairs blind before you look at any names — it recalibrates your eye faster than reading standings does.

Runway's Gen four point five is the headline release. Durations now run four to twenty seconds at ten-eighty, built on an autoregressive-to-diffusion approach developed with NVIDIA, and it takes reference images to hold a character's appearance stable across separate clips. Twenty seconds in one pass matters if you have been stitching five-second fragments. Pricing is twelve to twenty-five credits per second depending on pipeline configuration, against twelve for Gen four and five for Gen four Turbo, with legacy Gen three Alpha around ten. Act-Two, the performance-capture path, is five credits per second with a three-second minimum. Aleph, the video-to-video editing series, runs fifteen to twenty-eight credits per second — expensive enough that you want it as a surgical last step, not an iteration tool. Next action: run one existing shot through Turbo and the same prompt through Gen four point five and compare the delta against the price gap before you migrate anything.

Kling has moved from one point six and Video O1 up through Kling three and Video V3. The interesting item is native multi-shot: up to six camera cuts in a single generation. Clip lengths three to fifteen seconds, native ten-eighty up to native four-K without an external upscaler, first-and-last-frame generation across three-to-ten-second intervals, and an Elements feature binding up to four reference images to hold traits and clothing through heavy camera movement. Native audio covers synchronized dialogue, ambience and lip-sync in five languages in one pipeline. Camera control is selectable directional moves with numeric displacement parameters. Next action: test first-and-last-frame at the three-second end, where interpolation is tightest.

Luma has Dream Machine on Ray2, with Camera Motion Concepts — over fifteen composable natural-language moves including orbits, cranes, push-ins and pans — layerable directly onto image-to-video prompts. Frame controls cover first-and-last keyframing between two distinct images, plus a Loop toggle that matches the final frame to the first, and an Extend tool continuing a scene open-ended or toward an end keyframe. Composable camera language is the part to steal. Next action: stack two camera concepts on one image-to-video prompt and see where it breaks.

MiniMax has shipped Hailuo 02 and 2.3, seven-sixty-eight to ten-eighty with standard and Fast drafting variants, leaning on physical dynamics and throughput at a low credit cost. Native synchronized audio is omitted from the base 2.3 models, so budget a separate audio pass. Elsewhere on the aggregators, Veo three point one with native dialogue and Wan two point six have both been added to standard API matrices — worth checking your provider's model list rather than assuming.