Product6 min read•Published Updated

Music Down, Automatically: Broadcast-Style Ducking in the 5am CLI

Mix voice and music locally with the 5am CLI. Learn the ducking presets, EQ-pocket fallback, loudness targets, and checks to make before publishing.

5

5AM Team

Product guides and release notes from 5AM Software Labs, the publisher of 5AM. Contact us with questions or corrections.

Music Down, Automatically: Broadcast-Style Ducking in the 5am CLI

You've recorded the narration. You've picked the perfect music bed. Now comes the part nobody warns you about: making them sit together.

Set the music too loud and it swallows your words. Set it too quiet and the track loses all its energy. Split the difference and you get both problems at once — because the right music level during speech and the right level between sentences are two different levels. A real mixing engineer solves this by riding the fader: music up in the gaps, down under the voice, all session long. That's a skill, a DAW, and an afternoon.

Today it's one command:

5am media mix --voice episode.wav --music bed.mp3 --output mix.wav

5am media mix takes a dialogue track and a music track and applies a measured mixing chain locally. It provides a starting mix that you can tune and audition.

Three moves, like a mixer would make them

It stages the music at broadcast level. Before anything else, the CLI measures both files — integrated loudness, true peak, duration — and gain-stages them the way a broadcast chain expects: your voice anchored at the target loudness, the bed sitting a fixed distance underneath (18 dB under for the podcast preset). No guessing, no "sounds about right."

It ducks in response to your voice. The music runs through a compressor that's keyed by your voice — the moment you speak, the bed dips; the moment you pause, it swells back. Attack and release are tuned like a broadcast limiter's: a short attack reduces masking at the start of speech, while the release controls recovery. Breaths, noisy recordings, and a dense music bed can require different settings.

It carves an EQ pocket right where the voice lives. This is the part a plain volume dip can't do. The music is split into three frequency bands, and the band where speech intelligibility lives — roughly 250 Hz to 4 kHz — is ducked several dB deeper than the lows and highs. Under your words the bed doesn't just get quieter; it gets thinner, leaving the bass and the air of the track intact while the midrange steps aside for the voice. When you stop talking, the pocket closes and the music is whole again.

The three-band path needs FFmpeg’s acrossover filter. If it is missing, the CLI reports a fallback to full-band ducking with a static EQ pocket. Check stderr so you know which path ran.

Pick a style, press enter

Different content wants different riding. Four presets ship today:

5am media mix --voice ep.wav --music bed.mp3 --style podcast --output mix.wav
StyleThe idea
podcast (default)Voice is king — bed 18 dB under, strong duck, tight broadcast moves, −16 LUFS target
documentaryA more present bed, gentler duck, slow cinematic recovery
promoMusic-forward and punchy — ducks just enough, louder −14 LUFS target
ambientA barely-there bed with very slow, unobtrusive moves

And because presets never fit everyone, every parameter is a flag: --bed-level, --duck-depth, --pocket-depth, --ratio, --attack, --release, --target-lufs, --true-peak, --fade-out. The style sets the defaults; your flags win.

# Deeper bed, harder duck, slower recovery
5am media mix --voice ep.wav --music bed.mp3 \
  --bed-level -24 --duck-depth 16 --release 400 --output deep.mp3

Check the final loudness and listen to the transitions

The podcast preset uses a −16 LUFS target and peak limiting. That is a preset choice, not a universal platform requirement or a guarantee of the final measurement. Check your destination’s current delivery specification and measure the encoded file. Listen to the quietest sentence, a loud passage, a pause, and the ending; lower the bed or increase duck depth if words are masked.

A few practical details:

  • The mix follows the voice duration. Short music beds loop, which can make a seam audible with some tracks. The unauthenticated spoken tag adds time after the mix. Check both the loop point and the ending.
  • Any audio in, four formats out. Inputs are anything FFmpeg can decode; output format follows your --output extension — .wav, .mp3, .m4a, or .flac.
  • JSON on stdout for this command: measured loudness, applied gains, every resolved parameter. Pipe it to jq, wire it into a script.
  • Requires a local FFmpeg (4.4+) — brew install ffmpeg, apt install ffmpeg, or winget install ffmpeg.

It completes the podcast pipeline

media mix slots into the terminal podcast workflow the CLI has been building toward:

# 1. Generate a music bed from a prompt (Lyria)
5am media generate music --prompt "warm lo-fi piano, gentle, unobtrusive" --output bed.mp3

# 2. Mix it under your episode
5am media mix --voice episode.wav --music bed.mp3 --style podcast --output episode-mixed.wav

# 3. Wrap it in video for YouTube / Reels
5am media visualize episode-mixed.wav --cover cover.jpg --subtitles episode.srt --output episode.mp4

Music generation is a separate AI request that uses credits or provider billing; mixing existing local files does not use AI credits. Use an existing licensed music track if you do not need generation. Review the exported episode before publishing. The Podcast Studio post covers the video half of that pipeline in depth.

Free to use, free to de-tag

media mix works without an account. Unauthenticated runs append a short spoken "Powered by 5AM" tag at the very end of the mix — after your audio, never over it. Running 5am login (which is free) removes it. Your mixes render fully offline either way — nothing waits on a network call.

Get it

If you have the CLI, you already have it — 5am update and go. Otherwise:

curl -fsSL https://cli.5am.app/cli/latest/install.sh | sh
5am media mix --voice your-voice.wav --music your-music.mp3 --output mix.wav

Full flags and examples are in the CLI docs. Put a voice over some music and listen to the bed step aside for your first sentence — it's the kind of thing you hear once and stop wanting to mix by hand.

Get the CLI →

Tags

#cli#audio#podcast#ducking#ffmpeg#workflow

Related posts

Creativity never sleeps.

Turn the 5 a.m. idea into shipped work. Store it, make it, sell it — in one place.

Open Podcast Studio

5 GB free · No card required