Vision & Audio Intermediate

Source Separation

Splitting sound that's been mixed together back into its parts

Key points
  • Source separation is the job of splitting sound recorded all mixed together into separate parts and capturing each one on its own.
  • Mixed sound has no boundary. The wave is already merged into one, so there's no seam left to cut along.
  • So it's done by spreading out a sound picture and assigning each cell a share of who it belongs to.
  • Each part has a different pitch pattern, texture, and way of moving over time, and those become the clues.
  • It never comes apart perfectly. A trace of the other sounds always lingers in whatever gets pulled out.
Contents

1The analogy

One pond at a fish farm holds several species swimming together. They can still be sorted out. Different sizes, different depths in the water, different movement habits — vary the net mesh and where you scoop, and each species lands in a different tank. It's not perfectly clean, though. A few strays end up in every tank, and a species never seen before has no tank to go in at all.

Source separation is this same job done with sound. A vocal and an instrumental, or several voices in a meeting room, sit mixed together inside one file. Each part's distinct traits become the clue for sorting it into its own bin, and once that's done you can listen to the vocal alone, the drums alone, or one person's voice alone.

2In detail

Once mixed, there's no boundary left

In the pond, each species still swims separately. Sound is different. When several sounds share the same space, the air's vibrations simply add together into one wave. Inside a recorded file, the vocal and the instrumental aren't sitting side by side — only their sum survives.

In numbers, it's like being told the sum of several numbers and asked to guess the originals. There's no single correct answer. That makes source separation a job of guessing, not measuring.

Guessing is possible because each sound has a shape of its own. It's not random numbers added together — knowing that what was added looks like a voice, or looks like drums, opens room to pull them apart.

Splitting the shares on a sound picture

The actual work usually happens on a sound picture rather than cutting the waveform directly. Spread out a picture split by time and pitch, and for every cell, assign a ratio saying which part's share this sound is.

Apply this ratio map and one part survives while the rest is suppressed. Make one ratio map per part and each part's sound comes out separately. Adding all the parts back together is usually tuned to reconstruct the original sound.

It matters that no cell gets assigned to just one owner. It's common for two sounds to sit in the same cell together. Splitting it something like half and half keeps the separated sound from cutting in and out.

What gets used as a clue

The biggest clue is the pitch pattern. A voice draws regular stripes and drums draw a short, broad smear, so they look different in the picture. How long a sound sustains differs too — a string instrument rings on, a percussion hit vanishes in an instant.

How something moves over time is a clue too. A vocal rises and falls with the lyrics and breath, while an instrumental keeps the beat. This difference in flow becomes a hint for splitting apart stretches where they overlap.

If a recording is split left and right, direction is a clue as well — a sound louder on the left can be told apart from one centered in the middle. Even so, today's methods do a decent job splitting things apart even from a recording already mixed down to one channel.

It never comes apart cleanly

Listen to a separated vocal alone and a faint trace of the instrumental lingers behind it. The other side keeps a trace of the vocal's reverb. Sometimes the sound warbles as if underwater, or the edges rustle.

Some cases are especially hard: two very similar voices speaking at once, a recording made in a heavily reverberant room, or an instrument or sound the system never trained on. Sound that was already smeared together and stored that way in the original recording is hard to restore too — information that's already lost doesn't come back just because you split things apart.

Where it's used

The most familiar use is keeping just the instrumental of a song or pulling out just the vocal. It's also used in restoration work — splitting an old recording apart by instrument and remixing it — and in film, boosting dialogue while pulling background sound down.

On the speech side, it's used to keep just one person's voice in a meeting recording, or to split a voice from noise on a call made somewhere loud. Running transcription after separating noticeably improves accuracy, which is why it's often attached as a step right before it.

3More precisely

Recovering the original sounds from a signal where several are mixed together is also studied under the name blind source separation. Assigning a ratio to every cell on a sound picture is called masking, and it's now also common to split directly from the waveform without going through a picture at all. How good a result is gets measured by how much of the original sound survives in the separated track and how much of the other sounds leaked in. Music work usually splits into a preset set of parts — vocal, drums, bass, and everything else.

The analogy breaks down in one place. In the pond, each species genuinely exists separately, so scooping them out is enough. In mixed sound, there's no separate substance left to scoop. What comes out of separation isn't a recovered fragment of the original recording — it's closer to sound newly generated to match what the original probably sounded like. That's why a separated file can pick up rustling that was never in the original, and why the same file fed into a different tool comes out slightly different.

4Try it yourself

5Common misconceptions

  • It's easy to think erasing one part from a recording leaves the other exactly as it was, but actually the overlapping portion gets suppressed together, so what remains gets a little damaged too.

  • It's easy to think a separated track is the same as recording that part alone in the first place, but actually it's a plausibly reconstructed sound, and traces show up on close listening.

  • It's easy to think a high-quality file always separates cleanly, but actually similar sounds overlapping or heavy reverb keeps things from separating well even from a good file.

7One-line summary

In shortSource separation is like sorting a mixed pond of fish into separate tanks, and since sound has no boundary to cut, it's done by assigning a share to every cell and reconstructing each part from that.

Spotted an error or have a better analogy? Suggest an edit · Last updated2026-09-02