How to Add Ducking Effects So Music Fades When Someone Speaks in Video

Learn how to use audio ducking to lower music during speech, with practical settings, workflow tips, and common mistakes to avoid.

*No credit card required
Audio editing timeline on a monitor with waveform tracks and keyframe points
CapCut
CapCut
Aug 11, 2026

Audio ducking lowers background music only when dialogue or voiceover is present; for most creator workflows, a starting duck of about 6-12 dB, an attack around 30-80 ms, and a release around 250-700 ms will keep speech clear without making the music pump.

Ever finish an edit and realize the music feels great until someone starts talking? That is one of the fastest ways to make a video feel harder to follow, especially on a cell phone speaker or in a noisy room. The fix is simple once you know what to control: when to automate the fade, how deep to duck the music, and which settings keep the change smooth instead of obvious.

What Audio Ducking Actually Does

Blue audio waveform display on a dark screen with a dip in the middle

Audio ducking means the music track drops in level when a voice track crosses a trigger point, then rises back up when the person stops speaking. In a sidechain workflow, the dialogue track acts as the trigger and the music track is the signal being reduced. In a keyframe workflow, you draw those level changes manually around each spoken phrase.

That matters because speech intelligibility usually fails in the consonants, not the vowels. If the music is too dense in the same frequency range as the voice, words blur even when the voice meter looks loud enough. Ducking solves that by creating temporary space under speech instead of forcing you to keep the music quiet for the entire video.

Captions still matter, but they are not the same thing as fixing the mix. A platform explains that captions are time-coded text equivalents of audio, while subtitles assume the viewer can already hear the soundtrack . High-quality captions also need correct dialogue, speaker identification when needed, and meaningful non-speech audio cues, so they support accessibility rather than replace clear spoken audio .

Choose Automatic or Manual Ducking

Person using a stylus on a tablet to edit a video timeline with audio waveform and keyframe markers

Automatic Ducking

Automatic ducking is the faster option when you have a clean dialogue track, steady voiceover pacing, and background music that stays fairly consistent. Many editors and AI-assisted video tools now offer speech detection, dialogue-aware mixing, or sidechain compression presets. In practical terms, this works best for explainers, tutorials, talking-head posts, course clips, and product demos with one main narrator.

The tradeoff is that automatic ducking follows level, not taste. If the voice track has breaths, chair noise, or long pauses between short phrases, the music can dip too often or recover too quickly. That is why automatic ducking works best after you clean the voice track first. If you need a neutral editor for that cleanup-and-balance pass, CapCut Audio Editing Tools can handle level adjustments before you decide what still needs manual refinement. A platform notes that speech enhancement tools are designed to reduce noise, reverberation, and distortion before you mix, which is a better order of operations than trying to solve muddy dialogue with deeper ducking later .

Manual Keyframes

Manual ducking is slower, but it gives you scene-by-scene judgment. That makes it the better choice for short-form social edits, brand videos with intentional musical swells, montage-heavy product spots, or any piece where the voice starts and stops irregularly. In CapCut Desktop and similar editors, this usually means putting dialogue and music on separate tracks and adding volume keyframes around each line.

Manual keyframes are also safer when the edit depends on emotion or rhythm. You can let the music stay higher during pauses, dip harder for one critical line, or bring the soundtrack back exactly on a cut, gesture, or beat. If your edit is packaging social clips, voiceover ads, or ecommerce product explainers, this extra control often produces a more polished result than one global auto-duck preset.

Set the Parameters So Speech Stays Clear

Hand adjusting an audio mixer with level lights beside a computer monitor

The five settings that matter most are threshold, duck amount or gain reduction, attack, release, and the final music bed level. Threshold is the trigger level, usually in dBFS, where speech starts the duck. Attack is how fast the music drops after speech is detected. Release is how fast it rises back up. Duck amount is how far the music drops, and the bed level is where the music sits while someone is talking.

If your dialogue peaks land around -12 dBFS to -6 dBFS, a practical starting threshold is often around -28 dBFS to -20 dBFS. Start conservative: if the duck never triggers, raise sensitivity or lower the threshold; if breaths and room tone trigger it, clean the dialogue or raise the threshold. A release that is too short causes pumping. An attack that is too slow leaves the first word masked.

Starting Ranges by Video Type

These are practical starting points, not fixed rules:

Table showing ducking settings by video type, including duck amount, attack, release, and best method

A good listening check is simple: on laptop speakers or a cell phone, every line should remain easy to understand without the music feeling like it disappears completely. If the result sounds jumpy, lengthen release by 100-200 ms. If the first syllable still gets buried, shorten attack or start the fade slightly before the line with manual keyframes.

Build a Voice-First Workflow

Start with dialogue, not music. Clean the voice track first, remove obvious background noise, trim empty gaps, and even out the loudest and softest phrases with light compression. For spoken-word video, a mic placed about 6-12 inches from the mouth and slightly off-axis is a reliable recording baseline, and keeping recording peaks around -12 dB to -6 dB gives you headroom for mixing later. If the source recording is weak, ducking will not magically create clarity.

Then add the music as a support layer, not as the foundation. Put music on its own track, set a rough starting level, enable auto ducking or create keyframes, and listen back in full context with captions turned on. Auto-generated captions can speed up the first pass, but they still need review because automated systems often miss accuracy, context, and readability . For creator workflows, this is where CapCut AI-style tools can help: generate the caption draft, clean the voiceover, then manually review the mix and text instead of assuming automation got both right.

Action Checklist

    1
  1. Put dialogue, voiceover, and music on separate tracks.
  2. 2
  3. Clean the dialogue track before adding ducking.
  4. 3
  5. Set a starting duck of 6-12 dB and adjust by content type.
  6. 4
  7. Use 30-80 ms attack and 250-700 ms release as first-pass ranges.
  8. 5
  9. Check the mix on phone speakers, laptop speakers, and headphones.
  10. 6
  11. Add captions after the spoken track is final, then review both captions and mix them together.

For branded, client, or ecommerce work, treat music rights as part of the workflow too. Use music you are licensed to use for commercial distribution, rather than assuming a platform song library covers every marketing use case. That is a production step, not an afterthought.

Avoid the Mistakes That Make Ducking Sound Amateur

The most common mistake is over-ducking. If the music drops 15-20 dB on every line, the audience hears the automation before they hear the story. That may be acceptable for a dense training clip, but it usually feels clumsy in short-form, lifestyle edits, and product promos. In those cases, it is often better to combine a lighter duck with a music track that has less vocal-range competition.

The second big mistake is using ducking to hide a bad recording. If the dialogue has room echo, distortion, or inconsistent volume, the music will still feel wrong after you duck it. Better speech cleanup matters because aggressive processing can make dialogue sound unnatural, hollow, or robotic. Cleaner speech gives ducking less work to do.

A third mistake is relying on captions as the backup plan for unclear audio. Captions are essential for accessibility and mute viewing, but viewers who do listen still need a mix where the voice wins. If the soundtrack is lyric-heavy, bright, or rhythmically busy, you may need manual keyframes, a lower music bed, or a different track entirely.

FAQ

Q: How Do I Make Background Music Automatically Lower When Someone Speaks?

A: Put the voice and music on separate tracks, enable your editor's ducking, dialogue detect, or sidechain compression feature, and set the voice track as the trigger. Start with a 6-10 dB reduction, 30-60 ms attack, and 300-500 ms release, then refine by ear.

Q: What Settings Stop Ducking From Sounding Choppy or Pumped?

A: Release time is usually the first fix. If the music jumps back too fast between phrases, lengthen release into the 400-700 ms range. If the first word gets buried, shorten attack or move manual keyframes slightly earlier. If breaths trigger the duck, clean the dialogue or raise the threshold.

Q: Should I Use Automatic Ducking or Manual Keyframes in CapCut-Style Social Editing?

A: Use automatic ducking for fast voiceover-first edits with steady narration, and use manual keyframes when timing, punchlines, beat drops, or product shots need precise control. For short-form packaging, manual correction after automation is often the best compromise.

Practical Next Steps

If you want a reliable default, mix in this order: clean the voice, place music on a separate track, apply a 6-12 dB duck, test 30-80 ms attack and 250-700 ms release, and then fix the lines that still feel too crowded with manual keyframes. That workflow is fast enough for social clips, controlled enough for marketing videos, and clear enough for education content.

The best ducking is usually the ducking nobody notices. If viewers can follow every word, the music still carries energy, and the transitions do not call attention to themselves, your settings are close.

Hot and trending