Digital Audio
Your DAW asks you for a sample rate, a bit depth, a buffer size and a file format. These settings decide how faithfully your computer captures sound, how much room you have before distortion, and how big your files are. This lesson explains each one in plain words and tells you exactly what to pick.
Analog vs Digital
Sound in the air is a smooth, continuous wave of pressure (see What is Sound?). A microphone turns that pressure into a smooth, continuously changing voltage. That's an analog signal: it can take any value at any moment.
A computer can only store numbers. To record sound, it measures the voltage many thousands of times per second and writes each measurement down as a number. That's a digital signal.
- The part that measures is the ADC (analog-to-digital converter).
- The part that turns numbers back into voltage for your speakers is the DAC (digital-to-analog converter).
Both live inside your audio interface (and inside every phone and laptop).
Two settings decide how accurate that conversion is: sample rate (how often you measure) and bit depth (how precisely you measure).
Sample Rate
Each measurement is called a sample. The sample rate is how many samples are taken per second, measured in Hz or kHz.
At 44.1 kHz, the computer takes 44,100 measurements every second for each channel.
Drag the "Samples per cycle" slider down. With only a few samples per wave, the stored copy (coral) becomes a rough staircase that no longer follows the original (grey). With many samples, the copy follows the curve closely.
The Nyquist rule
The Nyquist theorem says a sample rate can capture frequencies up to half of itself. That limit is called the Nyquist frequency.
Sample rate → highest frequency it can store
Human hearing tops out around 20 kHz, so 44.1 kHz already covers everything we can hear, with a little room to spare.
If a frequency above the Nyquist limit sneaks in, it doesn't just disappear. It "folds back" and shows up as a wrong, lower frequency. This is called aliasing, and it sounds harsh and out of tune. Converters use filters to block it, and some plugins offer oversampling to reduce it inside the DAW.
Common sample rates
| Rate | Where it's used |
|---|---|
| 44.1 kHz | The CD standard; still the most common rate for music releases |
| 48 kHz | Video, film and broadcast; a popular default for music too |
| 88.2 / 96 kHz | High-resolution recording; larger files and more CPU load |
Pick 44.1 kHz or 48 kHz and use it for the whole project. Higher rates double the file size and CPU load, and most listeners can't hear a difference in the final song. If you work with video, use 48 kHz.
Bit Depth
Bit depth is how many possible values each sample can have. More bits means finer steps between the quietest and loudest levels.
- 16-bit: 65,536 possible levels.
- 24-bit: about 16.8 million possible levels.
When a measurement falls between two levels, it gets rounded to the nearest one. That rounding error is called quantization error, and you hear it as low-level noise. Go back to the demo, press Play and lower the bit depth: at 2–4 bits the tone becomes gritty and distorted.
Dynamic range: about 6 dB per bit
Dynamic range is the gap between the loudest sound a format can hold and its noise floor. Each bit adds roughly 6 dB.
| Bit depth | Dynamic range (approx.) | Use it for |
|---|---|---|
| 16-bit | 96 dB | Final release files (CD standard) |
| 24-bit | 144 dB (in theory; real converters manage around 110–120 dB) | Recording and exporting mixes |
| 32-bit float | Enormous; practically impossible to clip inside the DAW | Internal DAW processing, exporting stems and premasters |
What is 32-bit float?
32-bit floating point stores numbers in a way that can go far above 0 dBFS and far below the noise floor. Most DAWs mix internally in 32-bit float (or 64-bit). That's why a channel meter can go into the red without the file being ruined, as long as you turn it down before it reaches the master output or a fixed-point file.
When you reduce bit depth (for example, exporting a 24-bit mix to a 16-bit file), turn on dither. Dither adds a tiny amount of noise that hides the rounding errors and sounds smoother than harsh quantization distortion. Only dither once, at the very last step.
File Formats
Audio files come in two families: uncompressed / lossless, which keep every sample, and lossy, which throw away detail to make files smaller.
| Format | Type | Notes |
|---|---|---|
| WAV | Uncompressed | The everyday standard for samples, stems and masters |
| AIFF | Uncompressed | Apple's equivalent of WAV; same quality |
| FLAC | Lossless compressed | Roughly half the size of WAV with no quality loss |
| MP3 | Lossy | Small files; 320 kbps is the highest common quality |
| AAC | Lossy | Used by Apple Music and YouTube; better quality than MP3 at the same bitrate |
Lossy formats use psychoacoustics to remove sounds you're unlikely to notice. At high bitrates they sound very close to the original, but each time you re-encode, more is lost. Learn more in Psychoacoustics.
How big is one minute of stereo audio?
This is why streaming and phone storage use compressed formats, while studios work in WAV.
Never use MP3s as your working files or send them as your master to a distributor. Always keep an uncompressed WAV (or AIFF) of every final mix. Converting an MP3 back to WAV does not restore the lost detail.
Mono vs Stereo
A mono file has one channel. Played on two speakers, the same signal comes out of both, so it sounds like it sits in the middle. A stereo file has two channels, left and right, which can be different. That difference creates width.
- Record single sources (one voice, one guitar, one mic) in mono. You can pan and widen them later.
- Synths, pads, drum loops and room recordings are often stereo.
- Kick, bass and lead vocals are usually kept mono and centered in a mix.
Stereo files are twice the size of mono files. See Panning & Stereo for how to use width in a mix.
Latency and the Buffer
Your computer can't process audio one sample at a time; it works in chunks called the buffer. Collecting a chunk takes time, which creates latency: a delay between playing something and hearing it.
Calculating latency
That's one direction. Input plus output, plus the converters, gives the full round-trip latency, often 2 to 3 times larger.
Under about 10 ms round trip, most people can play and sing comfortably. Above roughly 20 ms, the delay starts to feel noticeably off. Use a small buffer (64–128) when recording and a large one (512–1024) when mixing. Many interfaces also offer direct monitoring, which sends the input straight to your headphones with no computer delay at all.
Clipping and Headroom
In digital audio, 0 dBFS is the absolute ceiling. Any sample that tries to go higher gets chopped flat. That's clipping, and it causes harsh, crackly distortion. Unlike analog gear, which distorts gradually, digital clipping is sudden and ugly.
Headroom is the space between your loudest peak and 0 dBFS. Leaving headroom gives you a safety margin for loud moments and room for processing later.
| Stage | Target peak level |
|---|---|
| Recording (24-bit) | Peaks around -12 to -6 dBFS; average around -18 dBFS |
| Individual tracks while mixing | Peaks well below 0; leave room on the master |
| Mix sent for mastering | Peaks around -6 to -3 dBFS, no limiter on the master |
| Final master | True peak no higher than about -1 dBTP |
With 24-bit recording there's no benefit to recording "hot". The noise floor is so low that peaks at -12 dBFS lose nothing, and you're safe from a sudden loud note clipping. You'll go deeper in Mixing & Gain Staging and Loudness & LUFS.
- Open your audio settings and set the sample rate to
48 kHz[Ableton: Settings › Audio · FL Studio: Options › Audio settings · Logic: File › Project Settings › Audio]. - Set the recording bit depth to
24-bitif your DAW has the option. - Note the buffer size and the latency your DAW reports. Change the buffer from 128 to 1024 and see how the number changes.
- Load any loud drum loop and turn its fader up until the master meter turns red. Listen for the crackle, then pull it back until peaks sit around
-6 dBFS. - Export 10 seconds as a 24-bit WAV and as an MP3. Compare the file sizes.
A 44.1 kHz sample rate can store frequencies up to about kHz.
Each bit of depth adds about dB of dynamic range.
The digital ceiling that causes clipping is dBFS.
1. What does the sample rate measure?
- How loud the audio can be
- How many measurements are taken per second
- How many tracks you can record
- How compressed the file is
The sample rate is the number of samples per second, for example 44,100 at 44.1 kHz.
2. According to Nyquist, what is the highest frequency a 48 kHz recording can hold?
- 48 kHz
- 96 kHz
- 20 kHz
- 24 kHz
The Nyquist frequency is half the sample rate: 48 ÷ 2 = 24 kHz.
3. Roughly how much dynamic range does 16-bit audio have?
- 96 dB
- 16 dB
- 48 dB
- 144 dB
About 6 dB per bit × 16 bits ≈ 96 dB. 24-bit gives about 144 dB in theory.
4. Which format should you keep as your final master?
- MP3 at 128 kbps
- MP3 at 320 kbps
- WAV
- Whatever is smallest
WAV (or AIFF) is uncompressed and keeps every sample. Lossy formats throw detail away permanently.
5. Why record at around -12 dBFS peaks instead of close to 0 dBFS?
- It makes the recording brighter
- It reduces latency
- It turns the file into stereo
- It leaves headroom so loud moments don't clip, and 24-bit noise is low anyway
Headroom protects you from clipping. With 24-bit audio the noise floor is so low that recording a bit quieter costs nothing.