technique

Seeds and sampling settings

What the seed, number of steps, guidance strength, and duration settings do, and how to change one at a time.

Before this

This page assumes you are comfortable with:

Why you need this

The prompt says what you want; the settings decide how the model gets there and whether you can get there again. A good lobby loop that you cannot regenerate is a lucky accident. A harsh, crackly take is often a setting, not the prompt. This page is stage 3 of the pipeline, generate and choose: the handful of numbers every music model exposes, what each one does to the sound, and the habit that makes them learnable.

The idea

A diffusion model (the kind that starts from random noise and removes a little at each step, steered by your prompt) has four settings that matter most. ACE-Step 1.5 names them as below; the names are from its inference docs, and the Gradio interface shows the same settings with labels such as Inference Steps, Guidance Scale, Seed, and Audio Duration.

Setting What it controls Turn it up and you get
seed The starting noise A different take; same seed, same start
inference_steps How many denoising steps Finer detail, slower generation
guidance_scale How hard each step pushes toward the prompt Closer to the prompt, then harsh and overdone
duration Length of the clip, in seconds A longer clip, more room for structure to wander

The seed

The seed is a whole number that starts a random number generator. The model uses it to draw the starting noise. Same seed, same noise. Same noise, same prompt, same settings, and the same model version: the same audio. Change only the seed and you get a new take of the same request. That is why every take you keep needs its seed written down.

In ACE-Step 1.5, a seed of -1 means "pick one at random". To reuse seeds you set use_random_seed to false, and a batch can take a list in seeds. The ACE-Step 1.5 tutorial names three sources of randomness: the starting noise (the seed), the planning language model's own sampling when thinking is on with lm_temperature above 0, and extra noise when infer_method is "sde". The default infer_method, "ode", is deterministic. So for a repeatable take, fix the seed and know whether the other two are in play. Results can also shift between different hardware or software versions, so record the model version too.

Steps

Each step removes some noise. More steps mean smaller, more careful moves and finer detail, at a cost in time that grows with the step count. ACE-Step 1.5 ships models built for different step counts:

Model Default steps Guidance Notes from the docs
turbo (acestep-v15-turbo) 8 (range 1 to 20) not used Distilled to work in 8 steps; the docs recommend shift 3.0
sft (acestep-v15-sft) 50 used More detail and better reading of the prompt; slightly less clarity
base (acestep-v15-base) 50 used All tasks, including three only it can do

Larger XL versions of all three have the same step counts. The inference docs recommend 32 to 64 steps for the base model, and the tutorial says 32 to 100; treat both as starting ranges. Turbo models were trained (distilled) to reach a finished result in few steps, so raising turbo's steps far above 8 is not the same as running sft at 50.

shift reshapes how the steps are spread between rough structure and fine detail. The tutorial describes a larger shift as "draw outline first then fill details". For turbo models the docs say to set it to 3.0 by hand, since the default 1.0 is not corrected for you.

Guidance

Guidance (classifier-free guidance, CFG) makes each step compare two predictions, one that ignores the prompt and one that follows it, and push past the difference. guidance_scale is how far. ACE-Step 1.5's default is 7.0, with a range of 1.0 to 15.0 and a typical range of 5.0 to 9.0.

  • Too low: the take drifts from the prompt. Instruments you named go missing, and the sound is soft and mushy.
  • Too high: everything is exaggerated. Transients crackle, bright sounds turn harsh, and the waveform can clip.

So the rule of thumb, with the seed held fixed: a crunchy or over-bright take wants lower guidance, and a take that wanders off the prompt wants higher guidance.

The demo above is a teaching model, not a real denoiser, but its guidance slider shows the same trade: low guidance is quiet and mushy, high guidance overshoots and clips (red).

Turbo models do not use guidance: the inference docs say the pipeline resets guidance_scale to 1.0 for them. The docs disagree about the sft model: the README's model table and the tutorial's model section mark sft as supporting guidance, while the tutorial's settings table and the Gradio guide say guidance works on the base model only. If changing guidance on sft makes no audible difference, that disagreement is why.

Duration

ACE-Step 1.5's duration runs from 10 to 600 s; -1 lets the model choose from the lyrics. The tutorial calls the setting guidance rather than a precise command ("actual generation may vary slightly"), says 30 to 60 s and 2 to 4 minutes are stable, and warns that very long generations can repeat or lose structure. For a loop, ask for a whole number of bars plus a little spare, and cut it in Seamless music loops.

LeVo 2 works differently: the SongGeneration README's input fields (lyrics, descriptions, prompt audio) include no seed, step, or guidance setting, and length follows from the section tags. Its tuning happens in the text, covered in Lyrics and song structure.

Change one thing at a time

The ACE-Step 1.5 tutorial puts it directly: tuning effects and random effects can be the same size, so fix the seed while you tune. If you change the seed and the guidance together and the take gets better, you have learned nothing. Change one setting, keep everything else, listen, write it down.

Worked example

Illustration only: these settings and listening notes show the method, not a record of real runs.

The Lumen Clash lobby loop at 120 BPM in 4/4. One bar is 2 s, so a request of 48 s holds 24 bars, enough to cut a clean 8-bar loop from the middle. Fixed for every row: model acestep-v15-sft, the base caption from Prompting for music, bpm 120, duration 48, thinking off and infer_method "ode" (so the seed is the only source of randomness).

Take seed inference_steps guidance_scale What changed Listen for
A 4127 50 7.0 Baseline Reference for the other three
B 4128 50 7.0 Seed only A different arrangement of the same request: melody, drum pattern, and section order move; mood and instruments should stay
C 4127 50 11.0 Guidance only Same arrangement as A but pushed: harsher highs, crackle on drum hits, peaks near clipping
D 4127 25 7.0 Steps only Same arrangement as A with half the denoising steps: grain or smear in sustained pads, softer detail

Reading the table: B against A tells you how much of what you hear comes from luck. C and D share A's seed, so any difference is the setting. If D sounds as good as A, keep 25 steps and save half the denoising time. If C sounds better than A, you might try 9.0 next, still on seed 4127.

The same comparison on turbo would vary seed and shift instead, at 8 steps: 8 is 16% of sft's 50.

In a game's audio pipeline

This is stage 3, generate and choose. The prompt comes from stage 2 (Prompting for music, Lyrics and song structure). Once one take is close, Variations and audio-to-audio fixes parts of it, and Curating takes runs batches and logs the seed and settings of every keeper. Those logged values become the provenance record in stage 5.

Common mistakes

  • Not recording the seed. The take you loved is gone; regenerating with "the same prompt" gives a different song.
  • Leaving seed at -1 while tuning. Every change comes with a new random take, and you credit the setting for what the seed did.
  • Raising guidance to force the prompt. Past the typical range you hear crackle and harsh treble rather than more of what you asked for.
  • Turbo with sft-style settings. Setting guidance on turbo does nothing; leaving shift at 1.0 ignores the docs' 3.0 for turbo.
  • Asking for exactly the loop length. The model's length and timing vary slightly; with no spare bars, the cut point has nowhere to go.
  • Changing model version mid-project. The same seed on a different model or version is a different take.

Cost

Generation time grows with steps and duration: sft at 50 steps does 6.25 times the denoising steps of turbo at 8. The ACE-Step 1.5 README reports a full song in under 2 seconds on an A100 data-center GPU and under 10 seconds on an RTX 3090, both NVIDIA cards; your own machine's time per take is worth measuring once and writing in the take log. Memory depends on the model, not these settings: the tutorial gives about 4.7 GB of weights for the 2B models and about 9 GB for the 4B XL ones. The real cost is listening: four 48 s takes is 192 s of audio, and comparing them properly means at least two passes. Each 48 s take at 48 kHz, 16-bit stereo is 9.2 MB as WAV (1 MB = 1,000,000 bytes).

Going further

  • Curating takes: batches, a listening checklist, and the take log.
  • Variations and audio-to-audio: repainting a bar without losing the rest.
  • Diffusion models: why steps and guidance behave this way.
  • The ACE-Step 1.5 tutorial's sections on inference settings and random factors.
  • Try it: pick one setting, keep a fixed seed, and generate three values of it in a row.

Leads to

Back to Local text-to-audio for games