Free Auto Ducking — Voice Over Music, No Upload — Voice Over Music, Browser Sidechain
Quick answer: Free auto ducking for voice over music. Drop your voice on one lane, your music on the other, and the music ducks when you speak. Shorts and 20–90 min podcasts, MP3 or WAV, 100% in your browser. No upload, no watermark, no account.
Auto ducking is the process of automatically lowering the volume of background music whenever a voice starts speaking, then smoothly bringing it back up when the voice stops. Audio Forge Pro does it entirely in the browser with a short-window RMS sidechain: drop your voice on the top lane, your music on the bottom lane, and the tool ducks the music in real time. No upload, no watermark, and the rendered MP3 or WAV downloads straight from your device.
Audio Forge Pro is a free browser-based auto ducking studio built by M Zeshan (ShanX). It runs a short-window RMS sidechain in your browser with Web Audio API and OfflineAudioContext, so voice and music never leave the device. The mix exports as MP3 320 kbps or lossless WAV with a true-peak ceiling around -1 dBTP.
How to Auto Duck Voice Over Music (3 Steps)
- Drop voice on the top lane, music on the bottom: Two dashed lanes, both decoded locally with the Web Audio API. Nothing is uploaded — voice and music stay on this device.
- Pick a preset and tweak duck / attack / release: Shorts, YouTube, and Podcast presets ship with sensible duck depth, attack, and release. Sliders for duck amount, threshold, attack, and release let you fine-tune.
- Press Duck & mix and download MP3 or WAV: The whole timeline is rendered offline in your browser with a true-peak ceiling of about -1 dBTP so the result never clips. Download starts straight to your Downloads folder.
What Auto Ducking Supports
| File | Notes |
|---|---|
| Voice — MP3 | Decoded with Web Audio API · typical for podcast exports |
| Voice — WAV | Lossless voice master · sample-accurate start |
| Voice — M4A / AAC | iPhone Voice Memos, WhatsApp audio, mobile captures |
| Music — MP3 | Most royalty-free music beds · 320 / 192 / 128 kbps |
| Music — WAV | Lossless music bed · recommended when available |
| Output — MP3 320 kbps | Default · transparent for voice and music |
| Output — WAV | Lossless · for further editing in a DAW |
Who Uses Auto Ducking
- Shorts voiceover over a music bed — 15–60 second vertical videos.
- YouTube explainer with a quiet theme tune — 5–20 minute videos.
- Podcast intro / outro music — 30-second bed over a 5-minute cold open.
- Audiobook narration with subtle ambient music — long-form voice.
- Tutorial screen-record voice — music in the intro, voice in the body.
- Gaming commentary with BGM — menu music ducks during play-by-play.
Honest about what a browser duck can do
- Sidechain ducking with a short-window RMS envelope, soft-knee, attack / hold / release.
- Shorts, YouTube, and Podcast presets — sliders expose the same parameters for fine-tuning.
- Loop the music, trim it to the voice, or fit to the shorter track — your call.
- Negative music pre-roll so the bed is already moving when the voice starts.
- True-peak ceiling of about -1 dBTP so the rendered file never clips.
- No upload, no watermark, no account, no daily credits.
References
- MDN — Web Audio API — decode the voice and music locally.
- MDN — OfflineAudioContext — render the ducked mix faster than real time.
- MDN — GainNode — the audio-param automation that drives the duck.
Auto Ducking — Frequently Asked Questions
What is auto ducking?
Auto ducking is the process of automatically lowering the volume of background music whenever a voice starts speaking, then bringing the music back up when the voice stops. Audio Forge Pro runs a sidechain envelope in the browser: when the voice RMS crosses the threshold, the music gain drops by your chosen duck amount, holds for the configured hold time, then releases back to full level with a smooth ramp.
How do I duck background music under a voiceover online?
Drop your voice file on the top lane of the auto ducking tool, drop your music on the bottom lane, pick a preset (Shorts, YouTube, or Podcast), and click Duck & mix. The mix is rendered offline in your browser and the MP3 or WAV downloads straight from your device. No upload, no account, no watermark.
Is this a sidechain compressor?
Yes — under the hood it is a short-window RMS sidechain with attack, hold, release, and a soft-knee depth control. The duck depth (in dB), threshold, attack (10–80 ms), hold, and release (80–400 ms) are all adjustable from the sliders; the Shorts / YouTube / Podcast presets are the same parameters tuned for the most common use cases.
Do you upload my voice or music to a server?
No. Decoding, envelope analysis, and the ducked mix are all rendered in your browser with the Web Audio API and an OfflineAudioContext. Files never leave the device. The only network activity is downloading the static page and the MP3 you choose to save.
Will the mix click or pop when the voice starts and stops?
No. The sidechain has a fast attack (default 30 ms), a configurable hold (so the music does not pump between syllables), and a slower release (default 220 ms) that brings the music back up smoothly. The Shorts preset uses a 20 ms attack and a 120 ms release for tighter pacing; the Podcast preset uses 50 ms attack and 400 ms release for natural dialogue.
What happens if the music is shorter than the voice?
By default the music loops with a 5 ms crossfade at the boundary, so a 30-second music bed can carry a 60-second voiceover. The Studio UI also lets you trim the music to the voice length or fit the output to the shorter of the two. Loop works best for Shorts, trim is the safer default for podcasts.
What if the music is longer than the voice?
The default is to trim the music where the voice ends. If the music has a strong outro you want to keep, switch to “Fit to the shorter track” — both inputs stop at the shorter duration, so nothing trails off awkwardly.
Can I start the music before the voice comes in?
Yes. The music start offset slider lets you set a negative pre-roll (down to a few hundred milliseconds) so the music is already in motion when the voice begins. The Podcast preset ships with -250 ms of pre-roll for a natural intro bed.
Does it work for YouTube Shorts and TikTok?
Yes. The Shorts preset uses a fast 20 ms attack and an 18 dB duck, which is the most common pattern for voice-over-music Shorts: the music is loud at the start, gets out of the way the moment the voice starts, and eases back in at the end. For 15–60 second vertical videos, the same browser tab on a phone works fine; for longer content render on desktop.
Does it work for 20–90 minute podcasts and YouTube videos?
Yes, on desktop. The tool accepts mixes up to 90 minutes; the Podcast preset is tuned for long-form voice with a soft 10 dB duck, a 50 ms attack, and a 400 ms release. Very long mixes take a minute or two to render because the whole file is built in memory — keep the tab open until the download appears.
Can I use this with my own music and voice, or do I need royalty-free files?
Use whatever files you own or are licensed to use. The tool does not add music or generate a voice; it just mixes two files you already have. You stay responsible for the rights to both the voice and the music in the final output.
How does this compare to a desktop DAW sidechain?
A desktop DAW (Adobe Audition, Reaper, Logic, Pro Tools) gives you more surgical control — multiband sidechain, parallel compression, sidechain EQ — and ships in a paid app. Audio Forge Pro focuses on the one workflow creators actually need (drop voice + music, render duck, download) and runs it locally for free. If you already have a DAW, use it; if you do not, this is a quick, honest alternative.
AI Silence Remover | -14 LUFS Normalizer | Audio Trimmer | Audio Joiner | Extract Audio from Video | Full Audio Studio