prerequisite

Neural audio codecs

How a neural network squeezes audio into a short sequence of tokens or latent frames that a generator can work with, and turns them back into sound.

Before this

This page assumes you are comfortable with:

Why you need this

Every music model in this cluster works on a compressed version of sound, not on the sound itself. The piece that does the compressing and uncompressing is a neural audio codec. When a generated track sounds smeared, metallic, or watery, the codec is often why, and knowing what it keeps and what it throws away tells you which problems a better prompt can fix and which it cannot.

The idea

Raw audio is too long for a model. A digital recording stores the height of the sound wave many times a second. The sample rate is how many of those measurements, called samples, are taken per second. At 48 kHz, which means 48,000 samples per second, a single second of stereo (two channels, left and right) is 96,000 numbers. A three-minute song is millions of numbers. A transformer, the kind of network that predicts one token at a time, would need millions of steps to write that out, and its attention over all previous positions would grow far too large. A diffusion model, which starts from noise and cleans it up over several steps, would have to clean up millions of numbers at every step.

A codec shrinks time. A neural audio codec is a pair of networks trained together:

  • The encoder reads a short slice of samples and summarizes it as one frame: a short list of numbers that describes what that slice sounds like.
  • The decoder reads frames and rebuilds samples that sound like the original. A decoder that turns a compressed description back into a waveform is often called a vocoder.

The two are trained so that encode-then-decode sounds as close as possible to the input. The frame rate is how many frames describe one second of sound, measured in Hz (per second). Codecs used for music generation run at a few to a few tens of frames per second, where the waveform had tens of thousands of samples.

Continuous latents or discrete tokens. A frame can be stored two ways.

Kind What a frame is Who uses it
Continuous latent A short list of ordinary decimal numbers, for example 64 of them Diffusion models, which add and remove noise from these numbers
Discrete token One whole number chosen from a fixed list, the codebook, like a word from a vocabulary Transformers, which predict the next token from a probability list

Turning a continuous frame into a token is called quantization: the encoder's numbers are snapped to the nearest entry in the codebook, and only that entry's index is kept.

Residual layers. One token per frame from a codebook of a few thousand entries cannot describe sound in fine detail. Many codecs fix this by stacking codebooks. The first codebook picks the closest entry. The second codebook encodes the error left over, the residual. A third encodes what the second missed, and so on. Early layers carry the broad shape of the sound (pitch, rhythm, which instrument), later layers carry fine detail (breath, string noise, the crack of a snare). A model can then predict the coarse layers first and fill in the fine ones after.

What is lost. A codec is lossy: the decoder invents detail that the frames did not fully specify. Typical artifacts:

Artifact What you hear Usual cause
Smearing Drum hits and consonants lose their sharp start A frame covers tens of milliseconds, so a sudden onset is averaged with what follows
Metallic or phasey tone A thin, robotic ring on vocals and cymbals The decoder guesses phase and high-frequency detail
Muddy low end Bass and kick blur into one boom Fine pitch detail in the low register is under-described
Warbling sustained notes A held note wobbles slightly Frame-to-frame values drift

No amount of prompt wording fixes a codec limit. A different seed or a different model can.

Worked example

How much does a codec shrink a three-minute song? The numbers below are real published figures, not an illustration: they use the codec that ACE-Step 1.5 describes in its technical report: a VAE (variational autoencoder, an encoder-decoder trained to give smooth, continuous latents) that "compresses 48kHz stereo audio into a compact 64-dimensional latent space at 25Hz", plus a tokenizer that compresses those 25 Hz latents into 5 Hz discrete codes from a codebook of about 64,000 entries. Sizes on this page use 1 MB = 1,000,000 bytes.

Step 1: raw samples. Three minutes is 180 s.

48,000×180=8,640,000 samples per channel48{,}000 \times 180 = 8{,}640{,}000 \text{ samples per channel}

Stereo doubles it to 17,280,000 numbers. Stored as 16-bit integers (2 bytes each), that is 34,560,000 bytes, or 34.56 MB.

Step 2: latent frames. At 25 frames per second:

25×180=4,500 frames25 \times 180 = 4{,}500 \text{ frames}

Each frame holds 64 numbers, so the whole song is 4,500×64=288,0004{,}500 \times 64 = 288{,}000 numbers.

Step 3: compare.

Measure Raw audio Latent Ratio
Time steps per second 48,000 25 1,920 to 1
Numbers for the whole song 17,280,000 288,000 60 to 1

The report's headline "1920x" is the first row: how many sample times collapse into one frame (48,000÷25=1,92048{,}000 \div 25 = 1{,}920). Counting every number, the shrink is 60 to 1, because each frame stores 64 numbers where each sample time stored 2. Both are true; they measure different things.

Step 4: tokens. At 5 tokens per second the song becomes 5×180=9005 \times 180 = 900 tokens, each one 200 ms of music. A codebook of 64,000 entries needs about 16 bits to name one entry (log⁡264,000≈15.97\log_2 64{,}000 \approx 15.97), so 900 tokens fit in about 1,800 bytes. A transformer writing 900 tokens is a job it can do in seconds; writing 8.64 million samples is not.

LeVo 2 (also published as SongGeneration 2) uses the same idea with different numbers. Its technical report describes a 48 kHz music codec with a 25 Hz frame rate and a 64-dimensional representation, a language model that predicts tokens, and a diffusion model plus VAE decoder that turn those tokens back into detailed sound.

In a game's audio pipeline

The codec lives inside stage 1, Set up the studio, even though you never choose it separately: picking a model picks its codec. It shows up again in stage 3, Generate and choose, when you listen for artifacts. Smeared transients matter most for game audio that must land on an input (a card play, a button click), and a metallic ring matters most on anything that loops, because a game repeats it hundreds of times. Choosing a local audio model compares models that sit on top of different codecs.

Common mistakes

  • Blaming the prompt for a codec artifact. You rewrite the prompt five times and every take still has the same thin ring on the cymbals. That is the decoder; try another seed or model instead.
  • Reading "1920x compression" as file size. You expect a generated file to be tiny. The codec's frames live only inside the model; the file you export is ordinary samples again, as large as any recording of that length.
  • Expecting sample-exact timing. You ask for a hit exactly on a beat and it lands a few tens of milliseconds late or soft. One frame at 25 Hz is 40 ms, and the model places events per frame, not per sample.
  • Assuming a higher export rate adds detail. You export at a higher sample rate and the top end still sounds dull. The decoder only rebuilds what its frames describe.

Cost

A codec costs almost nothing at generation time compared with the generator: encoding and decoding run once per clip, while a diffusion model runs its network once per step over all the frames. The real cost is quality: the frame rate and codebook set a ceiling on detail that no setting can raise. For a model on a token codec, generation time grows with the number of tokens, T=r×dT = r \times d, where rr is the token rate in Hz and dd is the duration in seconds; doubling a song's length doubles TT and roughly doubles the time for that stage. Memory for the frames themselves is small (288,000 numbers in the example above); the network weights, covered in Model size and memory, dominate.

Going further

  • The ACE-Step 1.5 technical report, its sections on the VAE and the tokenizer.
  • The LeVo 2 technical report, its sections on mixed tokens and dual-track tokens (vocals and accompaniment modeled as separate token streams).
  • Residual vector quantization, the general name for stacked codebooks.
  • Diffusion models, for how a generator works on continuous latents.

Leads to

Back to Local text-to-audio for games