Digital Audio

Your DAW asks you for a sample rate, a bit depth, a buffer size and a file format. These settings decide how faithfully your computer captures sound, how much room you have before distortion, and how big your files are. This lesson explains each one in plain words and tells you exactly what to pick.

Analog vs Digital

Sound in the air is a smooth, continuous wave of pressure (see What is Sound?). A microphone turns that pressure into a smooth, continuously changing voltage. That's an analog signal: it can take any value at any moment.

A computer can only store numbers. To record sound, it measures the voltage many thousands of times per second and writes each measurement down as a number. That's a digital signal.

Both live inside your audio interface (and inside every phone and laptop).

Sound wave→Microphone→ADC→Numbers in your DAW→DAC→Speakers
Recording converts sound to numbers; playback converts numbers back to sound.

Two settings decide how accurate that conversion is: sample rate (how often you measure) and bit depth (how precisely you measure).

Sample Rate

Each measurement is called a sample. The sample rate is how many samples are taken per second, measured in Hz or kHz.

At 44.1 kHz, the computer takes 44,100 measurements every second for each channel.

Drag the "Samples per cycle" slider down. With only a few samples per wave, the stored copy (coral) becomes a rough staircase that no longer follows the original (grey). With many samples, the copy follows the curve closely.

The Nyquist rule

The Nyquist theorem says a sample rate can capture frequencies up to half of itself. That limit is called the Nyquist frequency.

Sample rate → highest frequency it can store

44.1 kHz → 22.05 kHz 48 kHz → 24 kHz 96 kHz → 48 kHz

Human hearing tops out around 20 kHz, so 44.1 kHz already covers everything we can hear, with a little room to spare.

If a frequency above the Nyquist limit sneaks in, it doesn't just disappear. It "folds back" and shows up as a wrong, lower frequency. This is called aliasing, and it sounds harsh and out of tune. Converters use filters to block it, and some plugins offer oversampling to reduce it inside the DAW.

Common sample rates

RateWhere it's used
44.1 kHzThe CD standard; still the most common rate for music releases
48 kHzVideo, film and broadcast; a popular default for music too
88.2 / 96 kHzHigh-resolution recording; larger files and more CPU load
Producer tip

Pick 44.1 kHz or 48 kHz and use it for the whole project. Higher rates double the file size and CPU load, and most listeners can't hear a difference in the final song. If you work with video, use 48 kHz.

Bit Depth

Bit depth is how many possible values each sample can have. More bits means finer steps between the quietest and loudest levels.

When a measurement falls between two levels, it gets rounded to the nearest one. That rounding error is called quantization error, and you hear it as low-level noise. Go back to the demo, press Play and lower the bit depth: at 2–4 bits the tone becomes gritty and distorted.

Dynamic range: about 6 dB per bit

Dynamic range is the gap between the loudest sound a format can hold and its noise floor. Each bit adds roughly 6 dB.

Bit depthDynamic range (approx.)Use it for
16-bit96 dBFinal release files (CD standard)
24-bit144 dB (in theory; real converters manage around 110–120 dB)Recording and exporting mixes
32-bit floatEnormous; practically impossible to clip inside the DAWInternal DAW processing, exporting stems and premasters

What is 32-bit float?

32-bit floating point stores numbers in a way that can go far above 0 dBFS and far below the noise floor. Most DAWs mix internally in 32-bit float (or 64-bit). That's why a channel meter can go into the red without the file being ruined, as long as you turn it down before it reaches the master output or a fixed-point file.

Note

When you reduce bit depth (for example, exporting a 24-bit mix to a 16-bit file), turn on dither. Dither adds a tiny amount of noise that hides the rounding errors and sounds smoother than harsh quantization distortion. Only dither once, at the very last step.

File Formats

Audio files come in two families: uncompressed / lossless, which keep every sample, and lossy, which throw away detail to make files smaller.

FormatTypeNotes
WAVUncompressedThe everyday standard for samples, stems and masters
AIFFUncompressedApple's equivalent of WAV; same quality
FLACLossless compressedRoughly half the size of WAV with no quality loss
MP3LossySmall files; 320 kbps is the highest common quality
AACLossyUsed by Apple Music and YouTube; better quality than MP3 at the same bitrate

Lossy formats use psychoacoustics to remove sounds you're unlikely to notice. At high bitrates they sound very close to the original, but each time you re-encode, more is lost. Learn more in Psychoacoustics.

How big is one minute of stereo audio?

WAV 44.1 kHz / 16-bit ≈ 10 MB (44,100 × 2 bytes × 2 channels × 60 s) WAV 48 kHz / 24-bit ≈ 17 MB FLAC (same as above) ≈ roughly half of WAV, varies with the music MP3 320 kbps ≈ 2.4 MB

This is why streaming and phone storage use compressed formats, while studios work in WAV.

Warning

Never use MP3s as your working files or send them as your master to a distributor. Always keep an uncompressed WAV (or AIFF) of every final mix. Converting an MP3 back to WAV does not restore the lost detail.

Mono vs Stereo

A mono file has one channel. Played on two speakers, the same signal comes out of both, so it sounds like it sits in the middle. A stereo file has two channels, left and right, which can be different. That difference creates width.

Stereo files are twice the size of mono files. See Panning & Stereo for how to use width in a mix.

Latency and the Buffer

Your computer can't process audio one sample at a time; it works in chunks called the buffer. Collecting a chunk takes time, which creates latency: a delay between playing something and hearing it.

Calculating latency

latency (ms) = buffer size ÷ sample rate × 1000 256 samples at 48,000 Hz → 256 ÷ 48000 × 1000 ≈ 5.3 ms 1024 samples at 44,100 Hz → 1024 ÷ 44100 × 1000 ≈ 23.2 ms

That's one direction. Input plus output, plus the converters, gives the full round-trip latency, often 2 to 3 times larger.

Under about 10 ms round trip, most people can play and sing comfortably. Above roughly 20 ms, the delay starts to feel noticeably off. Use a small buffer (64–128) when recording and a large one (512–1024) when mixing. Many interfaces also offer direct monitoring, which sends the input straight to your headphones with no computer delay at all.

Clipping and Headroom

In digital audio, 0 dBFS is the absolute ceiling. Any sample that tries to go higher gets chopped flat. That's clipping, and it causes harsh, crackly distortion. Unlike analog gear, which distorts gradually, digital clipping is sudden and ugly.

Headroom is the space between your loudest peak and 0 dBFS. Leaving headroom gives you a safety margin for loud moments and room for processing later.

StageTarget peak level
Recording (24-bit)Peaks around -12 to -6 dBFS; average around -18 dBFS
Individual tracks while mixingPeaks well below 0; leave room on the master
Mix sent for masteringPeaks around -6 to -3 dBFS, no limiter on the master
Final masterTrue peak no higher than about -1 dBTP

With 24-bit recording there's no benefit to recording "hot". The noise floor is so low that peaks at -12 dBFS lose nothing, and you're safe from a sudden loud note clipping. You'll go deeper in Mixing & Gain Staging and Loudness & LUFS.

Try it in your DAW
  1. Open your audio settings and set the sample rate to 48 kHz [Ableton: Settings › Audio · FL Studio: Options › Audio settings · Logic: File › Project Settings › Audio].
  2. Set the recording bit depth to 24-bit if your DAW has the option.
  3. Note the buffer size and the latency your DAW reports. Change the buffer from 128 to 1024 and see how the number changes.
  4. Load any loud drum loop and turn its fader up until the master meter turns red. Listen for the crackle, then pull it back until peaks sit around -6 dBFS.
  5. Export 10 seconds as a 24-bit WAV and as an MP3. Compare the file sizes.

A 44.1 kHz sample rate can store frequencies up to about kHz.

Each bit of depth adds about dB of dynamic range.

The digital ceiling that causes clipping is dBFS.

1. What does the sample rate measure?

  1. How loud the audio can be
  2. How many measurements are taken per second
  3. How many tracks you can record
  4. How compressed the file is

The sample rate is the number of samples per second, for example 44,100 at 44.1 kHz.

2. According to Nyquist, what is the highest frequency a 48 kHz recording can hold?

  1. 48 kHz
  2. 96 kHz
  3. 20 kHz
  4. 24 kHz

The Nyquist frequency is half the sample rate: 48 ÷ 2 = 24 kHz.

3. Roughly how much dynamic range does 16-bit audio have?

  1. 96 dB
  2. 16 dB
  3. 48 dB
  4. 144 dB

About 6 dB per bit × 16 bits ≈ 96 dB. 24-bit gives about 144 dB in theory.

4. Which format should you keep as your final master?

  1. MP3 at 128 kbps
  2. MP3 at 320 kbps
  3. WAV
  4. Whatever is smallest

WAV (or AIFF) is uncompressed and keeps every sample. Lossy formats throw detail away permanently.

5. Why record at around -12 dBFS peaks instead of close to 0 dBFS?

  1. It makes the recording brighter
  2. It reduces latency
  3. It turns the file into stereo
  4. It leaves headroom so loud moments don't clip, and 24-bit noise is low anyway

Headroom protects you from clipping. With 24-bit audio the noise floor is so low that recording a bit quieter costs nothing.