Mixing Vocals
In most songs the vocal is the most important sound. Listeners forgive a dull hi-hat, but not a vocal they can't understand. This lesson walks through a full vocal chain, step by step, with settings you can use today.
The Typical Vocal Chain
Every engineer has their own chain, but most follow the same logic: clean the sound first, control it next, then make it shine.
The order isn't a law. Some engineers put the de-esser first, or EQ after compression. Start with this order, then experiment once you know what each step does.
Step 1: Editing & Clean-up
Editing is boring, but it makes everything after it easier. Do it before you add any plugins.
- Comping: pick the best parts from several takes and join them into one performance. (Covered in Recording Vocals.)
- Remove noise: cut or fade out silence between phrases, so headphone bleed, chair creaks and room noise disappear.
- Breaths: don't delete them all, or the vocal sounds robotic. Lower loud breaths by
6–12 dBwith clip gain, and remove only the distracting ones. - Plosives: a "p" or "b" can cause a low thump. Fix it with a short fade, clip gain, or a high-pass filter on just that word.
- Timing: nudge late or early words to the beat. Tighten doubles and harmonies so they start and end with the lead, especially hard consonants like "t" and "s".
- Clip gain: even out very loud and very quiet words before compression. The compressor then works less and sounds more natural.
Use short fades (about 5–10 ms) on every edit to avoid clicks.
Step 2: Pitch Correction
Pitch correction moves notes toward the correct pitch. There are two main styles of tool:
- Real-time, automatic tools such as Antares Auto-Tune. You set the key and scale, and the plugin pulls each note to the nearest scale note.
- Graphical, manual tools such as Celemony Melodyne, Logic's Flex Pitch or FL Studio's NewTone. You see each note as a blob and adjust pitch, drift and vibrato by hand.
| Goal | Retune speed | Result |
|---|---|---|
| Natural correction | Slow (about 20–50 ms or more), with humanize on | Notes are guided gently. Slides and vibrato survive. |
| Hard-tune effect | Fastest (0 ms) | Notes snap instantly. The famous robotic "T-Pain" or modern trap sound. |
If the key or scale in the tuner is wrong, it will pull notes to the wrong pitches. Check the song's key first (see Notes & Scales). For natural results, tune only the notes that sound off rather than every note.
Step 3: EQ & Compression
Subtractive EQ
First, remove what you don't want:
- High-pass filter at about
80–120 Hzfor most voices (lower for deep male voices, higher for thin female voices). This removes rumble and mic-stand noise. - Mud: a gentle cut of
2–4 dBaround200–500 Hzif the voice sounds boxy or muffled. - Harsh resonances: narrow cuts (Q of 4–8) somewhere in
2–5 kHzif a tone "pokes" your ears.
Two-stage compression
Vocals are very dynamic. One compressor working hard often sounds pumpy. Two compressors each doing a little sound smoother.
Two-stage vocal compression
The first compressor grabs sudden shouts quickly. The second one gently evens out the whole performance. Together the vocal stays in front without sounding squashed.
The demos below use a full beat instead of a vocal, but the controls work the same way. Use the EQ to practice the moves you would make on a voice: cut mud around 300–400 Hz, then boost presence at 3–5 kHz. Listen for how clarity changes.
Step 4: De-essing, Tone & Saturation
De-essing
Compression makes "s", "sh" and "t" sounds louder relative to the rest of the voice. These harsh sounds are called sibilance. A de-esser is a compressor that only reacts to a narrow high-frequency band.
- Find the sibilance: it usually sits between
5 and 10 kHz. Most de-essers have a "listen" button to hear only the detected band. - Aim for
3–6 dBof reduction on the "s" sounds only. - Too much makes the singer sound like they have a lisp. Back off until the "s" is smooth, not gone.
Tone EQ
Now add what you want, with broad, gentle boosts:
- Presence:
+1 to +3 dBaround2–5 kHzhelps the words cut through. - Air: a high shelf of
+2 to +4 dBfrom about10–12 kHzadds sparkle (check the de-esser still works after this). - Warmth: a small boost around
150–250 Hzfor thin voices.
Saturation
Light saturation (gentle distortion from a tape, tube or console emulation) adds harmonics. These help a vocal stay audible on phone speakers and feel more "finished". Keep it subtle. If you can clearly hear distortion, it's probably too much, unless that's the style.
Sibilance usually sits between and kHz.
A typical vocal high-pass filter sits around Hz.
Step 5: Reverb & Delay Sends
Put vocal reverb and delay on return channels, not on the vocal track itself (see Buses & Parallel FX). That way you can EQ and automate the effects separately.
Classic vocal effects setup
Pre-delay keeps the start of each word dry, so the vocal stays clear even with a big reverb. Filtering the returns stops the effects from adding mud or harshness. Learn more in Reverb & Delay.
Try sidechain-ducking the delay return from the dry vocal. The echoes drop while the singer is singing and bloom in the gaps. It's a modern pop staple, explained in Sidechain & Ducking.
Doubles, Harmonies & Ad-libs
- Lead vocal: centered, always.
- Doubles (the singer recording the same part again): pan a pair hard or wide left and right, about
6–10 dBunder the lead. They thicken the chorus without hiding the lead. - Harmonies: pan in pairs, for example 40% left and right, or spread them wider than the doubles. Cut more low end on them (high-pass at
150–250 Hz) and de-ess harder, because stacked "s" sounds pile up fast. - Ad-libs: short, extra phrases between lines. Pan them off-center (for example 30–60%), give them a different effect such as a telephone-style band-pass EQ or more delay, and keep them out of the way of the lead's words.
Send all of these to a vocal bus with a little glue compression so the stack moves as one.
Making the Vocal Sit
A vocal "sits" in the mix when every word is clear but it still sounds like part of the song, not pasted on top.
Carve space in the music
Instead of only boosting the vocal, make room for it. Cut 1–3 dB around 2–5 kHz on busy instruments like guitars, synth pads and keys. Or use a dynamic EQ on the music bus, keyed from the vocal, so the cut happens only when the vocal plays.
Ride the vocal with automation
Even after compression, some words drop out. Vocal riding means automating the vocal level word by word, usually by 1–3 dB. Raise quiet words and the ends of phrases; lower loud ones. Professionals often ride every line. You'll learn how in Automation.
Check the vocal level at low volume and in mono. If you can't understand the words quietly, the vocal is too low or masked.
- Load a vocal over a beat. Lower loud breaths with clip gain and add short fades to every edit.
- Add an EQ: high-pass at
100 Hz, cut3 dBat about300 Hzif muddy. - Add a fast compressor (4:1, 3 ms attack) for 3–6 dB of reduction, then a slow one (2:1, 20 ms attack) for 2–3 dB.
- Add a de-esser and find the "s" between 5–10 kHz. Aim for 3–6 dB of reduction on sibilant words only.
- Add a tone EQ:
+2 dBat 3 kHz and a high shelf at 10 kHz. - Create a plate reverb return and a 1/4-note delay return [Ableton: Return Tracks; FL Studio: send mixer tracks; Logic: Aux channels]. Send the vocal to both.
- Automate the vocal volume on any word you can't hear clearly.
1. Why do many engineers use two compressors on a vocal?
- One for the left channel, one for the right
- It's required by streaming services
- Each does a little work, which sounds smoother than one compressor working hard
- To add reverb
A fast compressor catches peaks and a slower one evens the level. Splitting the work avoids audible pumping.
2. What does a de-esser reduce?
- Harsh "s" and "sh" sounds (sibilance)
- Low rumble
- Breaths
- Pitch errors
A de-esser compresses only a narrow high band, usually 5–10 kHz, where sibilance lives.
3. Which retune speed gives the robotic hard-tune effect?
- Slow, with humanize on
- It depends on the reverb
- Medium, around 100 ms
- The fastest setting (0 ms)
A 0 ms retune speed snaps each note instantly to pitch, removing natural slides. That's the hard-tune sound.
4. Where should the lead vocal usually be panned?
- Hard left
- Center
- Hard right
- It should move constantly
The lead vocal is the focus of the song, so it sits in the center. Doubles and harmonies go out to the sides.
5. What is pre-delay on a vocal reverb used for?
- To tune the vocal
- To remove breaths
- To keep the start of each word dry and clear before the reverb begins
- To make the vocal mono
A gap of 20–60 ms before the reverb starts keeps consonants crisp while still giving the voice space.