Mixing Vocals

In most songs the vocal is the most important sound. Listeners forgive a dull hi-hat, but not a vocal they can't understand. This lesson walks through a full vocal chain, step by step, with settings you can use today.

The Typical Vocal Chain

Every engineer has their own chain, but most follow the same logic: clean the sound first, control it next, then make it shine.

Edit & clean up→Pitch correction→Subtractive EQ→Compressor 1→Compressor 2→De-esser→Tone EQ→Saturation→Sends: reverb & delay
A common vocal chain. Fix problems before you add character, and add space last.

The order isn't a law. Some engineers put the de-esser first, or EQ after compression. Start with this order, then experiment once you know what each step does.

Step 1: Editing & Clean-up

Editing is boring, but it makes everything after it easier. Do it before you add any plugins.

Use short fades (about 5–10 ms) on every edit to avoid clicks.

Step 2: Pitch Correction

Pitch correction moves notes toward the correct pitch. There are two main styles of tool:

GoalRetune speedResult
Natural correctionSlow (about 20–50 ms or more), with humanize onNotes are guided gently. Slides and vibrato survive.
Hard-tune effectFastest (0 ms)Notes snap instantly. The famous robotic "T-Pain" or modern trap sound.
Set the key correctly

If the key or scale in the tuner is wrong, it will pull notes to the wrong pitches. Check the song's key first (see Notes & Scales). For natural results, tune only the notes that sound off rather than every note.

Step 3: EQ & Compression

Subtractive EQ

First, remove what you don't want:

Two-stage compression

Vocals are very dynamic. One compressor working hard often sounds pumpy. Two compressors each doing a little sound smoother.

Two-stage vocal compression

Compressor 1 "Catch the peaks" (FET style, e.g. 1176-type) Ratio 4:1 to 8:1 Attack 1–5 ms Release 50–100 ms Gain reduction: 3–6 dB on the loudest words Compressor 2 "Smooth the level" (optical style, e.g. LA-2A-type) Ratio 2:1 to 3:1 Attack 10–30 ms Release slow/auto Gain reduction: 2–3 dB, almost constant

The first compressor grabs sudden shouts quickly. The second one gently evens out the whole performance. Together the vocal stays in front without sounding squashed.

The demos below use a full beat instead of a vocal, but the controls work the same way. Use the EQ to practice the moves you would make on a voice: cut mud around 300–400 Hz, then boost presence at 3–5 kHz. Listen for how clarity changes.

Step 4: De-essing, Tone & Saturation

De-essing

Compression makes "s", "sh" and "t" sounds louder relative to the rest of the voice. These harsh sounds are called sibilance. A de-esser is a compressor that only reacts to a narrow high-frequency band.

Tone EQ

Now add what you want, with broad, gentle boosts:

Saturation

Light saturation (gentle distortion from a tape, tube or console emulation) adds harmonics. These help a vocal stay audible on phone speakers and feel more "finished". Keep it subtle. If you can clearly hear distortion, it's probably too much, unless that's the style.

Sibilance usually sits between and kHz.

A typical vocal high-pass filter sits around Hz.

Step 5: Reverb & Delay Sends

Put vocal reverb and delay on return channels, not on the vocal track itself (see Buses & Parallel FX). That way you can EQ and automate the effects separately.

Classic vocal effects setup

Plate reverb decay 1.2–2 s pre-delay 20–60 ms high-pass ~300 Hz, low-pass ~8 kHz on the return Slap delay 80–120 ms, 0–10% feedback (rock, rockabilly, rap) Tempo delay 1/4 or dotted 1/8, 20–30% feedback, filtered Throw delay automated send on the last word of a line

Pre-delay keeps the start of each word dry, so the vocal stays clear even with a big reverb. Filtering the returns stops the effects from adding mud or harshness. Learn more in Reverb & Delay.

Producer tip

Try sidechain-ducking the delay return from the dry vocal. The echoes drop while the singer is singing and bloom in the gaps. It's a modern pop staple, explained in Sidechain & Ducking.

Doubles, Harmonies & Ad-libs

Send all of these to a vocal bus with a little glue compression so the stack moves as one.

Making the Vocal Sit

A vocal "sits" in the mix when every word is clear but it still sounds like part of the song, not pasted on top.

Carve space in the music

Instead of only boosting the vocal, make room for it. Cut 1–3 dB around 2–5 kHz on busy instruments like guitars, synth pads and keys. Or use a dynamic EQ on the music bus, keyed from the vocal, so the cut happens only when the vocal plays.

Ride the vocal with automation

Even after compression, some words drop out. Vocal riding means automating the vocal level word by word, usually by 1–3 dB. Raise quiet words and the ends of phrases; lower loud ones. Professionals often ride every line. You'll learn how in Automation.

Note

Check the vocal level at low volume and in mono. If you can't understand the words quietly, the vocal is too low or masked.

Try it in your DAW
  1. Load a vocal over a beat. Lower loud breaths with clip gain and add short fades to every edit.
  2. Add an EQ: high-pass at 100 Hz, cut 3 dB at about 300 Hz if muddy.
  3. Add a fast compressor (4:1, 3 ms attack) for 3–6 dB of reduction, then a slow one (2:1, 20 ms attack) for 2–3 dB.
  4. Add a de-esser and find the "s" between 5–10 kHz. Aim for 3–6 dB of reduction on sibilant words only.
  5. Add a tone EQ: +2 dB at 3 kHz and a high shelf at 10 kHz.
  6. Create a plate reverb return and a 1/4-note delay return [Ableton: Return Tracks; FL Studio: send mixer tracks; Logic: Aux channels]. Send the vocal to both.
  7. Automate the vocal volume on any word you can't hear clearly.

1. Why do many engineers use two compressors on a vocal?

  1. One for the left channel, one for the right
  2. It's required by streaming services
  3. Each does a little work, which sounds smoother than one compressor working hard
  4. To add reverb

A fast compressor catches peaks and a slower one evens the level. Splitting the work avoids audible pumping.

2. What does a de-esser reduce?

  1. Harsh "s" and "sh" sounds (sibilance)
  2. Low rumble
  3. Breaths
  4. Pitch errors

A de-esser compresses only a narrow high band, usually 5–10 kHz, where sibilance lives.

3. Which retune speed gives the robotic hard-tune effect?

  1. Slow, with humanize on
  2. It depends on the reverb
  3. Medium, around 100 ms
  4. The fastest setting (0 ms)

A 0 ms retune speed snaps each note instantly to pitch, removing natural slides. That's the hard-tune sound.

4. Where should the lead vocal usually be panned?

  1. Hard left
  2. Center
  3. Hard right
  4. It should move constantly

The lead vocal is the focus of the song, so it sits in the center. Doubles and harmonies go out to the sides.

5. What is pre-delay on a vocal reverb used for?

  1. To tune the vocal
  2. To remove breaths
  3. To keep the start of each word dry and clear before the reverb begins
  4. To make the vocal mono

A gap of 20–60 ms before the reverb starts keeps consonants crisp while still giving the voice space.