5.1 to Stereo Downmix

Upload a 6-channel (5.1) WAV file and fold it down to stereo using ITU-R BS.775 downmix coefficients. Adjust center and surround levels, optionally include LFE, normalize the output, and export as 16-bit or 24-bit WAV. Your audio never leaves your browser.

Share

Drop your 5.1 WAV file here

Must be a 6-channel WAV in SMPTE/ITU channel order (L, R, C, LFE, Ls, Rs)

How 5.1 to Stereo Downmix Works: The ITU-R BS.775 Standard Explained

When a film or music mix is delivered in 5.1 surround and needs to play through headphones, laptop speakers, or a car stereo, the six channels have to be folded into two. That process is not arbitrary. The International Telecommunication Union defines the exact math in Recommendation ITU-R BS.775-3, which specifies the coefficients that broadcasters, streaming services, and hardware decoders use to collapse 5.1 audio into stereo without discarding the center channel dialog or washing out the surround information.

The standard downmix equations are two lines of arithmetic. For each sample position:

  • Left output: L + 0.707 × C + 0.707 × Ls
  • Right output: R + 0.707 × C + 0.707 × Rs

The value 0.707 is 1/√2, which is exactly -3 dB. The center channel (C) and both surround channels (Ls and Rs) are attenuated by 3 dB before they are mixed into the stereo output. The front left and right channels fold in at unity gain (1.0). The LFE (subwoofer) channel is omitted by default, because most stereo playback systems cannot reproduce deep bass below 80 Hz and adding unattenuated LFE content would cause clipping in the downmix.

What the Six 5.1 Channels Actually Contain

In any standard 5.1 WAV file, the six channels follow the SMPTE/ITU channel order: L (front left), R (front right), C (center), LFE (low-frequency effects, the ".1"), Ls (surround left), Rs (surround right). This is the order used by virtually every DAW when exporting multichannel WAV files, including Pro Tools, Logic Pro, Reaper, and Nuendo.

The center channel almost always carries dialog in film, and lead vocals in theatrical music mixes. It exists as a separate channel because dialog intelligibility degrades when a voice is panned hard to the left or right on a wide screen. Placing dialog in a dedicated center speaker (or in the center of a phantom stereo image) keeps it locked to the visual center regardless of how far from the screen a listener sits. When you fold down to stereo, the 0.707 coefficient ensures the dialog still appears in both ears at an appropriate level.

The surround channels (Ls and Rs) carry ambient sounds, reverb tails, and effects. In a film mix, they might hold rain, crowd noise, or off-screen action. In a 5.1 music mix, they often hold room reflections or delayed signals. The ITU-R BS.775 0.707 coefficient for surround brings them into the stereo bus at a level that reads as present but not dominant over the front channels.

Why Center Mix Level Matters for Your Downmix

The 0.707 default works well for film dialog mixes, but music engineers sometimes prefer a different center treatment. If the center channel carries a lead vocal at full level, the 0.707 coefficient blends it into both output channels at -3 dB. That gives you the vocal in stereo without it doubling to an unnaturally loud level compared to the front left and right content.

A center gain of 0.5 (-6 dB) is appropriate when the center channel is already mixed louder than the fronts in the original surround session, or when you want to pull the dialog back in a broadcast downmix. A center gain of 0 effectively drops the center channel entirely, which is occasionally used for karaoke-style outputs where the vocal needs to be removed.

This tool lets you choose 0.707 (-3 dB), 0.5 (-6 dB), or off for both center and surround independently. After processing, check your output level meters. If the left and right output RMS values are significantly lower than the input front channels, you can re-run the downmix with the normalize option enabled to bring the output to full scale without manual gain staging.

When to Include the LFE Channel

The ITU-R BS.775 standard specifies a coefficient of 0 for the LFE channel. That means the subwoofer track is excluded from the stereo downmix by default. The reason is acoustic: the LFE channel is recorded and mixed for playback through a dedicated subwoofer that rolls off steeply above 120 Hz. A stereo speaker or headphone driver reproducing that content without a crossover will distort or clip, and on small speakers the energy is inaudible anyway.

There are two exceptions where including LFE in a downmix makes sense. First, if you are exporting a downmix for a headphone mix where the listener is using large over-ear headphones or a studio monitor setup capable of flat response below 80 Hz. Second, if your LFE channel carries musical content (some electronic music productions use the LFE track for kick drum sub frequencies rather than film-style effects), you may want to blend it in at 0.5 or 0.707 to preserve the sub energy in the stereo mix.

The 0.5 (-6 dB) LFE option is a practical compromise. It brings in enough LFE energy to preserve the low-end impact of an explosion or a kick drum without overwhelming a stereo bus that was not designed to carry full-range subwoofer content.

Normalization After Fold-Down

After the downmix calculation, the output samples can exceed 0 dBFS (full scale). This happens because the fold-down process is additive. If the front left, center, and surround left channels all peak at high levels simultaneously, the summed output at any sample can exceed 1.0 in floating-point, which translates to clipping when the file is encoded to 16-bit or 24-bit PCM.

Normalization scans the entire output for the highest absolute sample value in either the left or right channel, then applies a single linear scale factor to bring that peak to exactly 0 dBFS (1.0 in floating point). The scale factor applies equally to both channels, so the stereo balance is preserved. This is not compression or limiting. It is a clean gain change that removes the headroom problem without introducing any dynamic artifacts.

If you are doing this downmix as part of a delivery workflow and need the output level at a specific LUFS target for streaming, run the normalized WAV through the LUFS Loudness Meter after downloading to verify the integrated loudness. Spotify targets -14 LUFS integrated, Apple Music targets -16 LUFS, and YouTube Music targets -14 LUFS. A downmix from a cinema-level film mix will often land somewhere between -20 and -27 LUFS, which streaming platforms will boost at playback time.

Choosing 16-bit or 24-bit Output

For most delivery scenarios, 16-bit PCM is sufficient. CD audio is 16-bit, and most streaming distributors transcode their masters to lossy formats (AAC or Ogg Vorbis at 128 to 320 kbps) where the difference between 16-bit and 24-bit source material is not audible.

Choose 24-bit output if you are delivering to a broadcaster, a film post-production pipeline, or any workflow where the stereo downmix will be processed further. The additional bit depth (24 bits provides 144 dB of theoretical dynamic range versus 96 dB for 16-bit) gives downstream processing more room to work without accumulating quantization noise. It also preserves the precision of the downmix calculation before any dithering stage in the delivery encoder.

Use the Audio Format Converter if you need to convert the exported WAV to MP3, OGG, or FLAC after downloading. And if you need to verify the stereo image quality of your downmix, the Stereo Width Meter shows the goniometer and phase correlation of the output, which tells you whether the downmix is mono-compatible and how much stereo separation the fold-down preserved.

Common Use Cases for Stereo Downmix

Film editors use stereo downmixes to create a monitoring reference for offline edit suites that do not have 5.1 speaker setups. The downmix is also the standard deliverable for streaming platforms and broadcast networks that ingest a stereo-only audio track alongside the video.

Game audio designers test surround mixes through a stereo fold-down to verify that dialog and important sound cues are audible on laptop speakers and phone speakers, where most casual players encounter the game. If an important enemy audio cue is placed entirely in the surround channels at low gain, it will disappear in the stereo downmix. Checking the downmix early in the audio design process catches those balance issues before final delivery.

Music producers working on 5.1 Dolby Atmos or Sony 360 Reality Audio releases use stereo downmixes to create the Apple Music and Spotify versions of their mixes, which are streamed in stereo to non-Atmos listeners. The ITU-R BS.775 downmix is the reference algorithm that Apple and Dolby use in their own transcoding pipelines, so running this tool gives you an accurate preview of what listeners on stereo devices will hear.

From the Blog

View All