prerequisite
Spectrograms
How a sound is pictured as frequencies over time, the view that audio models and audio editors both rely on.
Before this
This page assumes you are comfortable with:
Why you need this
A waveform shows when a sound is loud, not what it is made of. A spectrogram shows which frequencies are present at each moment, so you can see a muddy low end, a whine that should not be there, or a note that drifts. Many audio models also work with this kind of picture internally, which is part of why their mistakes look and sound the way they do.
The idea
Any sound is a sum of sines. A sine wave is the pure tone from Sound waves: one frequency, one amplitude. A mathematical result due to Fourier says that any sound, however complicated, can be written as many sine waves of different frequencies and strengths added together. This page states that fact without proving it. A violin note is a fundamental plus its overtones; a drum hit is a short burst of many frequencies at once; noise is every frequency at random strengths.
Frequency analysis of a short window. Take a short slice of the sound, called a window, of samples. A standard algorithm, the fast Fourier transform (FFT), answers one question about that slice: how strong is each frequency in it? It checks a set of evenly spaced frequencies, called bins. If the sample rate is samples per second, the window lasts seconds and the bins are spaced
hertz apart, from 0 Hz up to the Nyquist limit . For example, a 512-sample window at 16 kHz (16,000 samples per second) lasts ms and has bins Hz apart.
The spectrogram is many windows side by side. Slide the window along the sound, usually overlapping, and analyze each position. Draw each result as one thin vertical column: frequency up, brightness for strength. Put the columns in time order and you have a spectrogram:
- Time runs left to right.
- Frequency runs bottom to top.
- Brightness (or color) shows how strong each frequency is at that moment, usually on a decibel scale so quiet detail stays visible.
A steady note becomes a horizontal line. A chord becomes several stacked lines. A drum hit becomes a vertical stripe, because it is brief and contains many frequencies. A rising sweep becomes a line that climbs.
Window size is a trade. A long window holds many cycles of each frequency, so the bins are narrow and close notes separate cleanly, but anything that happens inside the window is smeared across its whole length. A short window pins events down in time but has wide bins, so nearby frequencies blur together. You cannot have both at once from one window. Editors let you pick the size: long for looking at pitch and harmony, short for looking at clicks and drum hits.
The mel scale. Pitch works on ratios (Sound waves explains octaves), but FFT bins are evenly spaced in hertz. So a plain spectrogram spends most of its rows on high frequencies, where a 100 Hz difference barely matters, and very few on low ones, where it is the difference between two notes. The mel scale fixes this by stretching low frequencies and squeezing high ones to roughly match how people hear pitch. One common formula is
which puts 1000 Hz at about 1000 mel, 2000 Hz at about 1521 mel, and 8000 Hz at about 2840 mel: the octave from 4000 to 8000 Hz gets fewer mel units than the octave from 1000 to 2000 Hz. A mel spectrogram groups the FFT bins into, say, 80 or 128 bands spaced evenly in mel. Many speech and music models use mel spectrograms, or learned representations that behave like them, because they are compact and spend detail where ears do. Neural audio codecs covers how generators compress sound further.
The demo synthesizes 1.5 s of sound at 16 kHz and draws the waveform on top and the spectrogram from 0 to 8000 Hz below, using 512-sample windows (32 ms) that start every 128 samples (8 ms). Pick a preset, hover to read the time and frequency under the pointer, and press Play to hear it.
Worked example
A C major chord. The demo's chord plays C4, E4, and G4, at about 261.63 Hz, 329.63 Hz, and 392.00 Hz, each with two overtones at twice and three times its frequency. That is nine frequencies, not three:
| Note | Fundamental | 2x | 3x |
|---|---|---|---|
| C4 | 261.63 Hz | 523.26 Hz | 784.89 Hz |
| E4 | 329.63 Hz | 659.26 Hz | 988.89 Hz |
| G4 | 392.00 Hz | 784.00 Hz | 1176.00 Hz |
Two of those lines nearly coincide: C4's third harmonic at 784.89 Hz and G4's second at 784.00 Hz are less than 1 Hz apart, far closer than any bin spacing, so they draw as one brighter line, and the spectrogram shows eight horizontal lines. The lines start together, glow brightest at the start, and fade slowly as the chord decays.
A drum hit. The demo's drum is a 60 Hz thump with a short burst of noise on top. You see a bright vertical stripe at 200 ms covering the whole height (the noise, every frequency at once) that vanishes within a few tens of milliseconds, and a short bright smudge at the very bottom (the thump) that lingers a little longer. A muddy kick in a generated track looks like that bottom smudge spread wide and long.
A 2048-sample window at 48 kHz. At 48,000 samples per second:
| Window at 48 kHz | Duration | Bin spacing | Good for |
|---|---|---|---|
| 256 samples | 5.3 ms | 187.5 Hz | drum attacks, clicks |
| 2048 samples | 42.7 ms | 23.4 Hz | notes and chords |
With 23.4 Hz bins, C4 and E4 (about 68 Hz apart) fall roughly three bins apart and show as separate lines. With 256-sample windows the bins are 187.5 Hz wide, so all three fundamentals of the chord land in one or two bins and blur together, but a drum attack lines up within about 5 ms.
In a game's audio pipeline
In generating and choosing, a spectrogram is a fast first check on a batch of takes: a hard horizontal line high up is a whine, a blur near the bottom is mud, and a missing band above some frequency hints that the model or a conversion cut the top off. In shaping it for the game, it shows you where a Lumen Clash card-play sound's snap lives so you can trim or layer it, and it shows a click at a loop seam as a thin vertical stripe. In setting up the studio, it explains why some models describe their output in frames rather than samples.
Common mistakes
- Reading one window size as the truth. A long window makes a crisp drum look smeared; a short one makes a clean chord look blurry. If a picture looks wrong, change the window before blaming the audio.
- Calling overtones extra notes. A single bright note shows a stack of lines. Count the evenly spaced ones as one note, or you will "hear" harmonies that are not there.
- Missing a click because it is thin. A loop-seam click is one narrow vertical stripe. Zoomed out, it can be thinner than a pixel. Zoom in on the seam.
- Ignoring the scale. A linear-frequency view crams the bass into a few rows; a mel or log view gives it room. Switch to a log or mel view before judging the low end.
Cost
One FFT on a window of samples takes time proportional to . A spectrogram of a clip with samples and a hop of samples between windows runs about windows, so the total is proportional to ; for a few seconds of audio that is a fraction of a second on any computer. Memory is one column of values per window. The real cost is the practice it takes to read the picture against what you hear.
Going further
- Neural audio codecs: how models compress audio into frames, often starting from a picture like this.
- Diffusion models: models that generate in a compressed space instead of raw samples.
- Seamless music loops: finding a click at a loop seam.
- Try it: open any song in an audio editor's spectrogram view, switch between window sizes, and watch drums sharpen while notes blur.
- The short-time Fourier transform, by name, for the formal version of a sliding window.