prerequisite
Diffusion models
How a model creates something by starting from noise and removing a little of it at each step, steered by a text description.
Before this
Nothing beyond first-year college math. This is a starting page.
Why you need this
Many open music and sound-effect models are diffusion models. Their settings (steps, guidance, seed) make sense once you know what happens between pressing Generate and getting a clip, and then a harsh or mushy take points you at the setting to change.
The idea
Adding noise until nothing is left. Take a clean clip and add a little random noise, then a little more, over many steps, until only static is left. How much noise is present at each step is set by a fixed plan called the noise schedule.
Training a network to predict the noise. A neural network is a large function with millions or billions of adjustable numbers. In training it sees a huge number of noisy clips, each with the text that describes the clip, and must guess the noise: which part is static and which part is sound. Each wrong guess nudges its numbers. Afterward it can estimate the noise in a noisy clip, given a description.
Generation runs the process backwards. To make something new, start from pure noise, ask the network for its guess of the noise, subtract part of that guess, and repeat. Each pass is one step. Early guesses are rough, so the sampler removes only a share of the guess and asks again; as the noise thins, the guesses improve. The description steers every guess, which is how the words reach the sound.
Steps trade speed for quality. Each step is one full run of the network, so generation time grows in proportion to the step count. Too few steps leaves grit or smeared detail; more steps give cleaner results up to a point, past which they add only time. Some models are trained to work in very few steps; their docs give the counts.
Guidance pushes toward the prompt. The common method is classifier-free guidance. At each step the network makes two guesses: one with your prompt and one with no prompt at all. Their difference is what the prompt changes. The sampler starts from the no-prompt guess and moves toward the prompted one by a factor , the guidance scale:
At you get exactly the prompted guess. Above 1 the prompt's effect is exaggerated, which helps the result follow the prompt. Too much overshoots: values grow past what the audio can hold, and you hear harshness and clipping. Too little leaves a vague, washed-out result. Two network runs per step also roughly doubles the work.
Working in a latent space. One second of 48 kHz stereo audio is 96,000 numbers, which is slow to denoise at every step. So most audio diffusion models compress audio with a separate network into a much shorter list of numbers called a latent, remove noise there, and decode the result to sound at the end. As an illustration only, not any particular model's figure: if the latent had 25 frames per second of 64 numbers each, that is 1,600 numbers per second, 60 times fewer than the raw samples. Neural audio codecs covers that compression and what it loses.
The seed is the starting noise. The starting noise comes from a random number generator, and the seed is the number that starts it. Same seed, prompt, settings, and model version: same clip. Change only the seed and you get a different take.
The demo is a teaching model, not a real denoiser. The faint curve is a two-sine target. Step 0 is seeded noise; each step blends further toward a prediction made by pushing a weak no-prompt guess toward a prompted one by the guidance slider. At the last step, low guidance is quiet and mushy and high guidance clips, drawn in red. The readout gives the RMS distance from the target (a kind of average error). Play lets you hear the current step.
Worked example
A tiny signal of five numbers, the target the model is trying to reach:
Seed 7 gives the noise (mulberry32, roughly normal values scaled by 0.8 and rounded to two decimals). The starting point is target plus noise. A real model starts from noise so strong that the signal is gone entirely; here it is still faintly there so you can follow it.
The sampler runs 4 steps. Each step, the network guesses the noise present, and the sampler removes the guess divided by the steps left (4, 3, 2, 1); perfect guesses would remove a quarter of the original noise each time. The guesses are invented for this page: the true remaining noise plus a shrinking error, also from seed 7. Values are rounded to two decimals.
| Row | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| Start (target + noise) | -0.72 | 0.70 | 0.62 | -0.07 | -0.03 |
| Step 4: noise guess | -0.78 | 0.20 | -0.44 | -0.67 | 0.02 |
| Step 4: remove guess / 4 | -0.20 | 0.05 | -0.11 | -0.17 | 0.01 |
| After step 4 | -0.52 | 0.65 | 0.73 | 0.10 | -0.04 |
| Step 3: noise guess | -0.46 | 0.31 | -0.34 | -0.51 | -0.08 |
| Step 3: remove guess / 3 | -0.15 | 0.10 | -0.11 | -0.17 | -0.03 |
| After step 3 | -0.37 | 0.55 | 0.84 | 0.27 | -0.01 |
| Step 2: noise guess | -0.33 | 0.08 | -0.26 | -0.24 | -0.07 |
| Step 2: remove guess / 2 | -0.17 | 0.04 | -0.13 | -0.12 | -0.04 |
| After step 2 | -0.20 | 0.51 | 0.97 | 0.39 | 0.03 |
| Step 1: noise guess | -0.15 | 0.08 | -0.05 | -0.01 | 0.11 |
| Step 1: remove guess / 1 | -0.15 | 0.08 | -0.05 | -0.01 | 0.11 |
| After step 1 (result) | -0.05 | 0.43 | 1.02 | 0.40 | -0.08 |
| Target | 0.00 | 0.50 | 1.00 | 0.50 | 0.00 |
Follow position 1: , then , then , then . The result is off by at most 0.10, because the guesses were never perfect. The step 3 guess at position 2 (0.31 where 0.15 remained) is an early mistake that later steps partly repair. A different seed gives different noise and a slightly different result: another take.
Guidance with small numbers. Suppose at one step and one position the no-prompt guess is 0.2 and the prompted guess is 0.5.
| Guidance | Calculation | Guess used |
|---|---|---|
| 1 | 0.50 | |
| 3 | 1.10 | |
| 7 | 2.30 |
At the guess is 2.30, over four times the prompted guess; as a sample value, anything past 1.0 clips. Repeated over every step and number, that is how high guidance turns harsh.
In a game's audio pipeline
In generating and choosing, steps, guidance, and seed are the knobs on Seeds and sampling settings: fix the seed while you try steps and guidance, then vary the seed to get takes. For the Lumen Clash in-match loop, a take that sounds crunchy and over-bright suggests lowering guidance; one that wanders off the prompt suggests raising it. In shaping and shipping, the seed is the line in your take log and provenance record that lets you remake a clip.
Common mistakes
- Guidance too high. The symptom is harsh, distorted, or clipped audio and exaggerated features of the prompt. Lower guidance before blaming the prompt.
- Guidance too low. The result is vague, quiet, and ignores parts of the prompt. Raising it a little often fixes "it ignored the instruments I asked for".
- Too few steps for the model. Grainy texture, smeared transients, or metallic shimmer. Check the step count the model's docs recommend for its variant.
- Changing the seed while tuning settings. A new seed is a new take, so you cannot tell what the setting did. Hold the seed fixed while you compare.
- Assuming the same seed matches across versions. A new model version or a different tool can turn the same seed into different noise. Record the version with the seed.
Cost
Generation time is roughly the number of steps times the cost of one network run, doubled when classifier-free guidance runs the network twice per step. Memory holds the weights plus the latent and one step's working values; it does not grow with the step count. Working in a latent space is what makes minutes of audio practical. The human cost is listening: every setting you try multiplies the takes to compare.
Going further
- Seeds and sampling settings: steps, guidance, and seeds in practice.
- Neural audio codecs: the compressed space many diffusion models work in.
- Transformers and tokens: the other main way to generate, one token at a time.
- Choosing a local audio model: which models use this approach.
- The papers "Denoising Diffusion Probabilistic Models" (2020) and "Classifier-Free Diffusion Guidance" (Ho and Salimans), by name.
Leads to
- techniqueChoosing a local audio modelWhich open model to run for which job: full songs with vocals, instrumental game music, or sound effects, compared by quality, speed, memory, and license.
- techniqueSeeds and sampling settingsWhat the seed, number of steps, guidance strength, and duration settings do, and how to change one at a time.