Cutting AI footage together leaves every clip carrying its own invented room: a different background hum, a different reverb tail, a different volume, sometimes a scrap of orchestral music that starts fresh in every shot and gets cut off mid-phrase. None of it matches, because none of it was recorded in a shared space. The fix is to stop comparing shots to each other and instead run one continuous layer underneath all of them — an ambience bed and a music bed — so the ear judges every cut against that steady floor instead of against the shot next to it.
Before laying beds, each clip's inherited audio gets sorted into one of three fates: keep it if there's synced dialogue, mute and rebuild it if it's just atmosphere, or keep the words and pull the room down around them. Generated music is the one thing almost never worth keeping, since it can't survive a cut. New atmosphere and effects can be built from scratch with a plain description using ElevenLabs' sound-effect generator, while a score can come from text-to-music tools like Suno — mind the licensing quirk where commercial rights only cover tracks made while subscribed — or from the genre-and-mood arrangement approach of SOUNDRAW, which produces instrumental beds that don't fight dialogue and offers separated stems on its higher tiers.
The rest is levels, done in DaVinci Resolve's Fairlight room: dialogue always sits on top, music and ambience tuck ten to fifteen decibels under it, ducking pulls the bed down whenever someone speaks, and loudness gets normalized to platform targets before export. The one check that catches most mix failures: play the finished piece on a laptop speaker or phone, at low volume, and listen only for whether every word survives.
ByteDance's Seedance 2.5 now generates continuous thirty-second clips in 4K from a single pass, accepting up to fifty reference inputs at once. MiniMax released H3 with open weights, and Alibaba's Wan 3.0 has moved to the top of several rankings tracked on Artificial Analysis, though the image-to-video and text-to-video leaderboards no longer agree on a single winner. OpenAI's Sora API is set to sunset on September 24th, and mandatory watermarking on synthetic video is tightening across providers and platforms.
Last time we left the timeline in a very particular state. The shots were trimmed, the joins were hidden, the whole sequence played through without a hitch. And every scrap of audio those clips brought with them was still sitting there, untouched, parked on its own tracks. That was deliberate. Sound is not a garnish you sprinkle on a finished cut. It is the layer that decides whether a viewer believes the cut is one place or ten.
So put your headphones on and play the sequence from the top. Not watching. Listening.
Here is what you will hear, and I want to name it precisely, because the names are the whole diagnosis.
At the first cut, the background noise changes. Not the traffic, not the birds — the noise underneath everything. Every real room has a quiet sound of its own: air moving, a fridge somewhere, the hum a microphone picks up when nobody is speaking. Sound people call that room tone. Recorded footage has one room tone per location. Your AI shots have a different one per shot, because each shot was generated on its own, by a model that invented a room from scratch and then invented the sound of that room too. Cut two of them together and the floor of the mix jumps. A viewer will not say "the room tone changed." They will say something feels cheap.
At the second cut, the reverb changes. Reverb is the tail of a sound bouncing off walls before it dies. A big tiled hall has a long one. A carpeted living room has almost none. Your generator decided, per shot, how big the space was. So a line of dialogue in shot three rings like a church, and the same character two seconds later in shot four sounds like she is speaking into a duvet. The picture says these are one continuous scene. The reverb says otherwise, and the ear trusts the reverb.
At the third cut, the level changes. One shot is simply louder than the next. And at the fourth, if you were unlucky with your prompts, music appears — some little orchestral swell the model added on its own, which starts at the head of the shot, gets three seconds in, and is guillotined by your edit. Then the next shot's music starts, from its own beginning, in a different key.
That is the mess. It has one cause and one shape. Because every clip was generated in isolation, nothing matches. We have hit this exact wall before, from the picture side, and the answer was always the same: pick one reference, lock it, and bend everything else toward it. When we were making mismatched generators sit in the same world, the reference was a hero shot and the tool was a grade. Sound works the same way, and the reference has a name of its own.
The single most useful thing you can do to a sequence of AI shots is lay one continuous layer of sound underneath the entire thing and let it ignore your cuts completely.
That layer is called a bed. Usually two beds, actually. An ambience bed is a long, unchanging recording of a place — street, office, forest, kitchen — running under every shot from the first frame to the last. A music bed is the score, doing the same. Neither one respects a cut. Both run straight through them.
Why does this work so well? Because your ear uses continuous sound to decide where it is. If the same faint traffic hum, the same distant air conditioner, the same low hiss is present across a cut, the ear concludes that the space did not change, only the camera did. It will make that judgement even when the two shots came from two different models, at two different times of day, with two different grades. The continuous layer overrides the discontinuous one. That is not a trick. It is how every film you have ever seen handles the fact that a scene is shot over three days and cut from forty takes.
And here is the part that matters for a deadline: the bed does more for continuity than any per-shot repair. You can spend an hour matching the room tone of shot three to the room tone of shot four, and get a modest result. Or you can spend six minutes laying one ambience across all nine shots and get a much better one, because now the shots are not being compared to each other at all. They are all being compared to the bed, and they all match it, because the bed is on top of them.
So the bed is the reference. Everything else in this chapter is bending the inherited audio toward it.
Which brings us to the awkward question of what to do with all that native audio you have been carrying.
Some of the shots in the cut came out of models that make picture and sound in one pass — the frontier tiers do this now as standard, and the arena even keeps separate standings for with-audio and without. That was a gift when you were generating. It is a liability now, because it is the source of every seam I just described. You have three choices per shot, and you make the choice per shot, not once for the project.
You can keep the generated audio. This is right when the shot has dialogue in it. A generated line, with lips already matched to it, is genuinely hard to replace — you would be undoing the sync work the dialogue episode bought you. Keep it, and fix its faults instead: level it, and if the reverb is wildly wrong, we will cover it rather than remove it.
You can mute it and rebuild. This is right for a shot whose audio is nothing but atmosphere and one or two obvious noises. A car going past. A door. Wind. You lose nothing by throwing that away, and you gain a shot that sits perfectly on the bed because it has no competing floor of its own.
Or you can keep only part of it. Keep the dialogue and mute everything around it, which in practice means keeping the clip audio but pulling its level well down so the words survive and the room does not. This is the messy middle case and it is also the most common.
The one thing that is almost never worth keeping is generated music. Not because it is bad — some of it is fine — but because of the restart problem. Music the model made lives inside one shot. It begins where the shot begins. It cannot cross a cut, because the next shot's music does not know it exists. So every join becomes an audible reset, and no amount of fading will disguise a key change every four seconds. Music is the layer that most needs to be continuous, and native audio is the layer least able to be continuous. Mute it and score the sequence yourself.
Now let us build it, in order, in the free version of DaVinci Resolve, which is where we have been finishing.
First, get the inherited audio where you can see it. When you dropped those clips in, picture and sound came in locked together as one object. Right-click a clip and untick Link Clips, or switch off the linked-selection toggle in the timeline toolbar to unlink everything at once. Now the audio is a separate thing you can select on its own, drag down onto its own track, mute, or delete. Put all the inherited clip audio on one track and leave the tracks below it empty. Those empty tracks are your build space.
Then go through shot by shot and make the three-way choice. Muting is the M key on the track header if you want the whole track gone, or you can drop a single clip's gain in the Inspector — the panel that shows you the settings for whatever is selected. Do this pass in silence, with nothing else playing. You are deciding what survives, not how loud it is.
Now switch to the Fairlight page. That is Resolve's audio room: the tracks run across the middle, a mixer strip for every track stands on the right, and the meters live out on the far edge. Everything from here on is easier there, because you can see all the levels at once.
Lay the ambience bed. This is one long file, dropped onto an empty track below everything, stretched across the whole sequence. Where does the file come from? For discrete sounds and background textures, the sound-effect model at ElevenLabs takes a plain description and gives you audio — "quiet office, distant keyboards, faint air handling" — up to about thirty seconds per generation, with a seamless-loop toggle that matters enormously here. Switch looping on and you can repeat one clean thirty-second bed under a three-minute film without an audible bump every half minute. Watch the licence, though, and this is not a small footnote: the free tier is non-commercial and asks for credit back to them, so a client job needs a paid plan. The cheapest of those runs about six dollars a month and carries a full commercial licence with no attribution, and it exports broadcast-standard forty-eight kilohertz WAV, which is the uncompressed format you want for anything you are going to mix rather than just listen to.
Place your effects next, on the actions the picture actually shows. Footstep, door, glass down on a table, the whoosh of a camera whip. The rule here is unglamorous: only sync what the eye can see making a noise. A viewer who sees a hand set a cup down and hears nothing feels the film is fake. A viewer who hears three sounds with no visible source feels the film is cluttered. Two or three placed effects per shot is plenty, and one is often better.
Music comes last. Not first — last. This is the part people get backwards, so let me be exact about why. If you bring music in early, you start cutting the picture to fit the music, and you will end up with a sequence whose shots are the wrong length because a drum came in. You already spent an episode getting those shot lengths right. Bring the music in after the picture is locked, and cut the music to the sequence instead: find the moment in the track where it lifts, and drag the track along the timeline until that lift lands on the cut you want emphasised. Trim the front. Fade the tail out under your last shot.
For generating the track, the tools split along a line worth understanding. Suno and Udio are text-to-music: describe the thing and they compose it, vocals and all, in blocks you can extend out to several minutes. Both run about ten dollars a month for their entry paid tier, a bit less annually, and both give you commercial rights and uncompressed forty-eight kilohertz WAV on those paid plans. One trap in the small print, and it bites people: with Suno, commercial rights attach to tracks made while you were subscribed. Upgrading later does not retroactively licence the song you generated on the free tier. Generate on the paid plan or do not use the result. Both companies are also in unresolved litigation with major labels over training data, which is worth knowing before you put one under a client's brand campaign, even though Warner has settled and licensed its catalogue to Suno.
SOUNDRAW is the different animal, and often the better fit for exactly this job. You do not write it a prompt; you pick genre, mood, tempo, and arrange the sections. It makes instrumental background music, which is what a bed almost always wants — a vocal fights your dialogue for the same space in the ear. Its creator plan is around twenty dollars a month with unlimited downloads for use in video, and it trains only on its own in-house catalogue of musical phrases, so the label-lawsuit question does not apply. That plan exports high-bitrate MP3 rather than WAV, which is fine for a bed. Its higher tiers add uncompressed WAV and separated instrument stems — a stem being one isolated layer of a mix, drums on their own, bass on their own — which is genuinely useful when you want to drop the drums out under a quiet line and bring them back after.
Now set the relative levels, and understand that this is where the mix is actually made. Everything before was gathering material.
There is one rule and everything else is detail: dialogue wins. Always, in every shot, at every moment. If a viewer cannot make out a word, nothing else you did matters.
In practice, for short-form and web video, that means speech sitting up front — peaks somewhere between about minus twelve and minus six on the meter, which for a spoken track puts its average loudness around minus sixteen to minus fourteen. Your key sync effects, the ones tied to visible actions, can be louder in their peaks than the dialogue, because they are brief; somewhere around minus ten to minus four. Background ambience goes much lower — minus twenty-four to minus eighteen. And the music bed tucks in around minus twenty-four to minus eighteen as well, roughly ten to fifteen decibels under the dialogue. That gap is not a style choice. It is the amount of room speech needs to stay intelligible.
Ten to fifteen decibels down sounds like too much when you first do it. The music will feel almost inaudible. Do it anyway. Music in a mix is felt more than heard, and the moment you can comfortably enjoy the track is the moment the dialogue is in trouble.
When dialogue and music overlap, you can do better than one fixed level. The technique is called ducking: the music automatically drops when someone speaks and comes back up when they stop. Resolve has a Ducker built in on the Fairlight page, and you can also do it with the track dynamics — double-click the Dynamics block on the mixer strip for your music track, set the dialogue track as the sidechain source, and switch the compressor to listen to that sidechain. Now every time the voice crosses the threshold, the music pulls itself back by however much ratio you set, then eases back in on the release. If you would rather do it by hand, arm automation on the track and record your fader moves live, or draw the volume curve directly onto the clip. For a sixty-second spot, drawing it by hand takes two minutes and gives you exact control.
Then hide the seams that are left, and there will be some. Three moves cover almost all of them.
Carry the ambience across a cut, which the bed already does for you, but check it specifically at the joins where the picture change is biggest — a wide to a close-up, or two shots from different generators. If the bed is running there, the join has already softened.
Let sound arrive slightly off the picture change. A cut where every layer switches on the same frame reads as a cut. Let a sound start a few frames before the picture arrives, or hold a sound a few frames past the change, and the ear stops noticing the edit at all. This is the same instinct as the split edits we used on the picture side, applied to a layer that does not have to obey the frame at all.
And when a join is simply too ugly for the picture to hide, put a sound on it. A door slam, a passing car, a hard musical accent, a whoosh. A loud transient masks whatever else is happening at that instant, and the viewer's attention goes to the sound rather than to the mismatch. This is not cheating. It is one of the oldest tools in editing.
Which leaves delivery, and one number you cannot guess at.
Platforms do not play your file at the loudness you exported it. They measure it and adjust it. The measure is called loudness — integrated loudness, specifically, meaning the average level across the whole piece rather than the height of its tallest peak — and it is written in LUFS, loudness units relative to full scale. It is always a negative number, and closer to zero means louder.
YouTube's target is minus fourteen LUFS, with true peaks kept under minus one. Short-form feeds like TikTok and Reels normalise into roughly the same neighbourhood, around minus fourteen to minus twelve, under the same one-decibel peak ceiling, and mobile-first masters are commonly pushed a little hotter than that.
Here is what happens if you ignore it, in both directions, because they are not symmetrical. Deliver louder than the target and the platform simply turns your whole file down. You gain no loudness at all, and you keep every bit of squashing you applied trying to get there — so you end up with a flat, lifeless mix played quietly. Deliver quieter than the target and most video platforms do not raise it. Your film just sounds smaller than everything either side of it in the feed.
Resolve will handle this for you and you should let it. Right-click any clip for Normalize Audio Levels, which offers the broadcast standards and platform presets directly. Watch the real numbers on the Fairlight meters as you play — integrated, short-term and momentary loudness, loudness range, and true peak all read out there, and integrated is the one to trust for the whole piece. Then on the Deliver page, under the audio settings, tick Audio Normalization and choose your target. Resolve measures the render, scales it to hit the number, and limits the peaks on the way out.
And now the pitfall, which I have watched land on more people than any other in this chapter.
The mix sounds perfect in your headphones at your desk. You send it. It comes back described as muddy, or the client asks why they cannot hear the voiceover. Here is what happened. You were monitoring loud, on good headphones, in a quiet room. In those conditions your ear is generous: it can pick a quiet voice out from under a bed without effort, so you kept nudging the ambience and the music up by feel, because they sounded nice, and you let the dialogue sit lower than it should have. Then someone watched it on a phone, held at arm's length, on a single tiny speaker that reproduces almost nothing below the middle of the range and has no dynamic range to spare. The bed and the music, being broad and continuous, survive that speaker fine. The voice, being detailed, does not. It vanishes into the mush.
The check that catches it takes ninety seconds. Before you export, play the whole thing out of your laptop's own speakers at a normal, quiet, background volume — not headphones, not monitors — and if you can, put it on a phone as well. Then listen only for words. Not for whether the mix sounds good; it will not, and that is not the test. Can you make out every line? If any word is a struggle, do not raise the dialogue. Lower the bed and the music by three decibels and listen again. Repeat until every word is clean. Then go back to your headphones and you will find the mix still works there, only now with more air around the voice than you would have chosen. That version is the correct one. The small speaker is the honest judge, because it is the one your audience is actually using.
Which is the whole discipline, really. The bed makes the shots one place, and the small speaker makes sure the words survive it.
Now, this month's news, and one item at the top of it will change how much of that inherited audio you end up muting.
The biggest release is ByteDance's Seedance two point five, which generates continuous clips of up to around thirty seconds natively, at four-K, in a single pass, and accepts as many as fifty multimodal reference inputs. Both halves matter. Thirty seconds in one generation means whole beats of a scene without a join, which means fewer seams to hide with a sound effect — and it means the native audio it produces has to hold up across a much longer stretch, so listen critically to it before you commit. Fifty references is a different order of thing from what we were doing with character sheets; that is a whole cast and a style anchor in one call. Next action: take a shot you built out of three chained clips and try to get it in one.
MiniMax released H3, in the Hailuo line, in late July, with public open weights on Hugging Face and an API tier at about seven dollars eighty a minute. Open weights on a model this capable is the news, because it is the door to running it inside your own pipeline rather than someone else's website. Alibaba shipped Wan three point oh in August, on top of the earlier two point seven updates, and is testing a HappyHorse one point one iteration alongside it. Lightricks and partner platforms brought LTX two point five online in fast and pro tiers, and Sand.ai surfaced MAGI-2 in preview, also open weights. If you have been meaning to bench an open-weights model against your paid default, this is an unusually good month to do it.
On the standings, and treat these strictly as this month's snapshot, because they reshuffle every time we speak. For text-to-video with audio, Wan three point oh has arrived straight at the top, a whisker above Google's Gemini Omni Flash, with MiniMax H3 third and Seedance two point oh's Dreamina build fourth — the older Wan two point seven and HappyHorse sit noticeably further back, and Veo three point one's flagship is now well down that particular board alongside Kling three point oh Omni. For image-to-video with audio, the order flips: Seedance two point oh leads, with H3 second by a hair and Gemini Omni Flash third, then Grok Imagine one point five, HappyHorse, and MAGI-2. On video editing with audio, Wan three point oh leads again, ahead of H3 and Gemini Omni Flash. The useful read there is not who is first. It is that the with-audio boards and the picture-only boards no longer agree, and that image-to-video and text-to-video rank different models at the top — so pick per job, and bench on your own shots.
Two housekeeping items. If any of your workflows still touch OpenAI's Sora, the standalone web and mobile apps went away in late April, and the API is scheduled for final sunset on the twenty-fourth of September. Migrate anything that depends on it now rather than in the week it dies.
And the disclosure layer keeps tightening. Providers and social platforms have been accelerating mandatory watermarking and synthetic-media labelling to line up with the EU AI Act enforcement phases and industry transparency requirements — that means embedded provenance metadata travelling with the file, and visible automated labels on photorealistic AI video once it is posted. Smallest next action: export one finished piece and inspect what metadata your tools actually attached to it, so you find out what you are shipping before a client does.