prerequisite

Digital audio and sampling

How a sound wave becomes a list of numbers: sample rate, bit depth, channels, aliasing, and how big an audio file is.

Before this

This page assumes you are comfortable with:

Why you need this

An audio model never touches air. It reads and writes lists of numbers, and every file it saves, every editor you open it in, and every game engine that plays it agrees on how those numbers stand for a sound. Knowing the rules tells you how big a file will be, why a sound can turn into a different pitch when converted carelessly, and what the settings in an export dialog mean.

The idea

Sampling is measuring the wave at regular times. A microphone turns air pressure into a voltage that rises and falls with the sound wave. A converter measures that voltage again and again at a fixed rhythm and writes each measurement down as a number. Each measurement is a sample. Playback runs the process backwards: the numbers set a voltage, and a speaker turns it back into pressure.

Sample rate is how many samples are taken per second, written fsf_s and given in hertz or kHz. 48 kHz means 48,000 samples per second, one every 1/480001/48000 s, or about 20.8 microseconds. The two common rates are 44.1 kHz (the CD standard) and 48 kHz (video and most game engines).

The Nyquist limit. To capture a wave, you need at least two samples per cycle, one near the top and one near the bottom. So the highest frequency a sample rate can hold is half the rate:

fNyquist=fs2f_{\text{Nyquist}} = \frac{f_s}{2}

At 48 kHz that is 24,000 Hz, comfortably above the roughly 20 kHz limit of human hearing. That margin is why 44.1 and 48 kHz were chosen.

Aliasing. A frequency above the Nyquist limit does not vanish. The samples land on the wave in a pattern that matches a lower frequency, and on playback you hear that lower frequency instead. This false tone is an alias. For a tone of frequency ff sampled at fsf_s, the alias is the distance from ff to the nearest whole multiple of fsf_s:

falias=∣f−kfs∣,k=the whole number nearest f/fsf_{\text{alias}} = \left| f - k f_s \right|, \quad k = \text{the whole number nearest } f / f_s

If ff is already below the Nyquist limit, k=0k = 0 and the formula returns ff itself. Converters prevent aliasing by filtering out everything above the Nyquist limit before sampling. Software that changes sample rate must do the same; one that does not produces faint whistles that were never in the original.

Bit depth and quantization. Each sample is stored with a fixed number of bits. With bb bits there are 2b2^b possible values, so the measurement is rounded to the nearest one. That rounding is quantization, and the small error it adds sounds like a faint hiss. 16-bit integer audio has 216=65,5362^{16} = 65{,}536 levels, enough that the hiss sits far below the music. 24-bit integer gives 16,777,216 levels and leaves room for recording and editing. 32-bit float stores each sample as a floating-point number, which editors use internally because it barely loses detail and can hold values above the usual ceiling without clipping until export. Bytes per sample are the bits divided by 8: 2 bytes for 16-bit, 3 for 24-bit, 4 for 32-bit float.

Channels. Mono is one list of samples. Stereo is two, left and right, played at the same time. Most music is stereo. Most game sound effects are mono, because the game places them left or right itself.

File size. Uncompressed audio (a WAV file, for example) is just the numbers, so its size is a multiplication:

bytes=fs×bytes per sample×channels×seconds\text{bytes} = f_s \times \text{bytes per sample} \times \text{channels} \times \text{seconds}

On this page 1 MB means 1,000,000 bytes. Compressed formats such as OGG Vorbis are much smaller; Exporting game audio covers them.

The demo draws a 1000 Hz tone over 5 ms. Drag the sample rate down from 8000 Hz and watch the dots get sparser. Below 2000 Hz the dots stop following the blue tone and trace the orange alias instead, and the readout names the frequency you would hear. Drag the bit depth down to 2 or 3 and each dot snaps to one of a few levels, drawn as steps. The readout also shows the size of one uncompressed minute for the current settings.

Worked example

One minute of 48 kHz, 16-bit stereo. The rate is 48,000 samples per second, each sample is 2 bytes, there are 2 channels, and a minute is 60 s.

48,000×2×2×60=11,520,000 bytes=11.52 MB48{,}000 \times 2 \times 2 \times 60 = 11{,}520{,}000 \text{ bytes} = 11.52 \text{ MB}

Format for one minute Calculation Size
48 kHz, 16-bit, stereo 48,000 x 2 x 2 x 60 11.52 MB
48 kHz, 16-bit, mono 48,000 x 2 x 1 x 60 5.76 MB
48 kHz, 24-bit, stereo 48,000 x 3 x 2 x 60 17.28 MB
44.1 kHz, 16-bit, stereo 44,100 x 2 x 2 x 60 10.584 MB

A two-minute Lumen Clash lobby loop kept as a 48 kHz, 16-bit stereo master is about 23 MB. That is fine on your disk and far too large to ship to a web browser, which is why the shipped copy is compressed.

A 1000 Hz tone sampled at 1500 Hz. The Nyquist limit is 1500/2=7501500 / 2 = 750 Hz, so the 1000 Hz tone is above it and will alias. The nearest whole multiple of 1500 to 1000 is 1500 itself (k=1k = 1, since 1000/1500≈0.671000/1500 \approx 0.67 rounds to 1):

falias=∣1000−1×1500∣=500 Hzf_{\text{alias}} = |1000 - 1 \times 1500| = 500 \text{ Hz}

You can check this sample by sample. Samples fall every 1/15001/1500 s, about 0.667 ms. The table shows the value of the 1000 Hz sine at each sample time next to a 500 Hz sine turned upside down:

Sample Time 1000 Hz tone 500 Hz tone, flipped
0 0.000 ms 0.000 0.000
1 0.667 ms -0.866 -0.866
2 1.333 ms 0.866 0.866
3 2.000 ms 0.000 0.000
4 2.667 ms -0.866 -0.866

The two columns agree at every sample. Once the samples are written, nothing in the file says which wave they came from, and playback rebuilds the only one below the Nyquist limit: a 500 Hz tone, exactly one octave lower than the original. Set the demo's sample rate to 1500 Hz to see the same thing drawn.

In a game's audio pipeline

In setting up the studio, models generate at a fixed sample rate stated in their documentation, and you decide the one rate your whole game will use. In shaping it for the game, loop points and crossfades are measured in samples, so the rate turns milliseconds into sample counts. In shipping it, the size formula tells you what a set of lossless masters costs on disk, and why the Lumen Clash card and button sounds can be mono while the music loops stay stereo.

Common mistakes

  • Mixing sample rates. A 44.1 kHz file played as if it were 48 kHz runs about 9 percent fast and sounds sharp. The symptom is one music track that is slightly higher and quicker than the rest.
  • Resampling without a filter. Converting down with a tool that skips the low-pass filter folds high content back as aliases. You hear thin whistles or a metallic edge that was not in the original.
  • Exporting a 32-bit float edit straight to 16-bit with levels above the ceiling. Float files can hold peaks over 0 dBFS, the loudest level an integer file can store; integer files cannot. The export clips, and you hear crackle on the loudest hits.
  • Stereo sound effects by habit. A stereo click doubles the file size for no audible gain once the game pans it. Across dozens of UI sounds, the build grows noticeably.
  • Very low bit depth by accident. Saving at 8-bit to "save space" adds audible hiss under quiet passages. A compressed format at 16-bit sounds far better at a smaller size.

Cost

Uncompressed size grows linearly with every factor in the formula: double the rate, the bit depth, the channels, or the length, and the file doubles. The lossless masters for a modest game soundtrack run to hundreds of MB, which costs nothing on a modern disk but matters for anything sent to players. Sample-rate conversion costs a little processing time per file and is done once, at the end. The bigger cost is mistakes found late: a mixed-rate project means re-exporting and re-checking every file.

Going further

Leads to

Back to Local text-to-audio for games