technique
Curating takes
How to generate in batches, listen with a checklist, and keep a log so a good take can be found and remade later.
Before this
This page assumes you are comfortable with:
Why you need this
Generation is fast; listening is not. A model can hand you eight versions of the Lumen Clash lobby loop before you have finished hearing the first one, and most will be close but wrong in some small way. Curating is stage 3 of the pipeline, generate and choose: producing takes in batches, judging them against a fixed checklist, and writing down enough about each that the keeper can be found and made again.
The idea
A take is one generated clip. A batch is several takes made from one request, each with its own seed (the number that sets the starting noise, covered in Seeds and sampling settings). Curating has three parts.
1. Batch. ACE-Step 1.5's batch_size setting makes 1 to 8 takes at once; the inference docs give a default of 2. A larger batch costs more memory, and the Gradio guide says the interface clamps batch size to what your memory tier allows. Its AutoGen option starts the next batch, same settings with new seeds, while you listen to the current one, and Apply These Settings to UI copies a batch's settings back into the input fields so you can iterate on a good one.
ACE-Step 1.5 can also score takes: the Gradio guide describes a perplexity-based quality score, and the tutorial a lyrics alignment score that rates how well the singing lines up with the lyrics. Use scores to decide what to hear first, never to decide what to ship. A score cannot hear a loop seam or a drum fill in the wrong place.
2. Listen with a checklist. The same questions, in the same order, for every take. Without a list you judge each take against the last one you heard, and drift.
| Check | How to listen | What failure sounds like |
|---|---|---|
| Artifacts | Headphones, full volume range | Metallic swirl, warbling, crackle on hits |
| Tempo drift | Tap along, or loop 8 bars against a click at the set BPM | Beats slide early or late against the click |
| Low end | Small speakers and headphones | Bass and kick blur into one boomy smear |
| Prompt match | Read the prompt, then listen | Named instruments missing, wrong mood |
| Vocals (songs only) | Read the lyrics along | Wrong or mumbled words, skipped lines |
| Ending | The last 5 s | Cut mid-note, or a fade that never finishes |
| Loopability | Loop the best 8 bars many times | A jump in energy, key, or texture at the seam |
| Clipping | Peak meter and ears | Peaks at 0 dBFS, crackle on the loudest moments |
Two habits keep the list honest. Match loudness before comparing: a louder take sounds better for being louder, so set every finalist to the same integrated loudness first (see Decibels and loudness). Listen on the speakers players will use. A card game is often played on a phone or laptop; small speakers drop most of the bass, so a take whose groove lives in the bass line can fall apart there. Do one pass on headphones for artifacts and one on a phone speaker for the real experience.
3. Log it. A take log is a table with one row per take, written as you listen, not afterwards.
| Field | Why |
|---|---|
| Take ID | Matches the file name |
| Date | Model versions change; the date anchors the row |
| Model and version | acestep-v15-sft is a different take than acestep-v15-turbo at the same seed |
| Prompt | The prompt text, or the name of a prompt file kept beside the log |
| Seed | The one number that makes the take repeatable |
| Settings | Steps, guidance, duration, and anything else not at its default |
| Verdict | keep, maybe, or reject |
| Notes | Which checklist item failed, with a time ("drift at 0:31") |
Name files so they sort. Put the most general part first, use an ISO date (year-month-day sorts correctly), and pad numbers with zeros so take 10 sorts after take 09. A pattern: lumen-lobby_2026-10-02_t03_s4129.flac. Take ID and seed in the name mean a file that strays from its folder still points back to its row.
Worked example
Illustration only: these takes, seeds, and verdicts show the method, not a record of real runs.
The Lumen Clash lobby loop. One request, made as one batch of 8 on ACE-Step 1.5: model acestep-v15-sft, base lobby prompt from Prompting for music, bpm 120, duration 48 (24 bars of 4/4 at 2 s per bar), inference_steps 50, guidance_scale 7.0, seeds 4127 to 4134.
| Take | Seed | Pass 1 (headphones) | Pass 2 (checklist) | Pass 3 (loop test, phone) | Verdict |
|---|---|---|---|---|---|
| t01 | 4127 | ok | pads warble at 0:22 | reject | |
| t02 | 4128 | ok | clean, mood right | seam jumps in energy | maybe |
| t03 | 4129 | ok | clean, mood right | bars 3 to 10 loop clean, bass survives | keep |
| t04 | 4130 | piano missing | reject | ||
| t05 | 4131 | ok | tempo drifts late from 0:31 | reject | |
| t06 | 4132 | ok | clean | thin on phone, groove gone | maybe |
| t07 | 4133 | ends mid-note | reject | ||
| t08 | 4134 | ok | crackle on snare | reject |
How the keeper was chosen:
- Pass 1, all 8 takes, headphones. Anything wrong in the first listen goes: t04 (named instrument missing) and t07 (bad ending). Six remain.
- Pass 2, checklist, at matched loudness. t01, t05, and t08 each fail one item, noted with a time. Three remain.
- Pass 3, the best 8 bars looped on a phone speaker. t02's seam jumps, t06 loses its groove without bass. t03 holds.
t03 is the keeper. Its row gives everything needed to remake it: model, prompt, seed 4129, settings. t02 and t06 stay as "maybe": if t03 later fails in the game, they are the next listen, not a new batch.
Listening time: pass 1 is 8 takes of 48 s (384 s), pass 2 is 6 takes of 48 s (288 s), and pass 3 loops an 8-bar, 16 s section 8 times for each of 3 takes (384 s). Total 1,056 s, or 17.6 minutes, for one lobby loop.
In a game's audio pipeline
This is the end of stage 3, generate and choose. Settings come from Seeds and sampling settings; a near-miss take can be repaired with Variations and audio-to-audio instead of discarded. The keeper moves to stage 4 (Seamless music loops). Its log row becomes the start of the provenance record in stage 5 (Licensing and provenance).
Common mistakes
- Judging by the first 10 seconds. The problems live later: drift at 0:31, a bad ending. You ship a loop that stumbles every time it comes round.
- Comparing at different loudness. The louder take wins every time. In the game, after loudness matching, it is no better than the others.
- Writing the log afterwards. Seeds get swapped between rows, and the remake of "t03" is a different song.
- Trusting scores alone. A high-scoring take with a bad seam passes the screen and fails in the game.
- Only listening on studio headphones. The bass that made the take on headphones vanishes on the phone the player uses.
- Names that do not sort.
take1,take2,take10sorts as 1, 10, 2, and the keeper gets overwritten by a later "take3".
Cost
Generation is the small part. A batch of 8 runs in one request, limited by memory: each extra take in a batch needs its own working memory, so lower the batch on a smaller machine rather than the quality settings. Listening is the large part: 8 takes of 48 s is 384 s, 6.4 minutes, for a single pass, and the three-pass method above is 17.6 minutes. Disk is modest: 8 takes as 48 kHz, 16-bit stereo WAV are 73.7 MB (1 MB = 1,000,000 bytes); ACE-Step 1.5's default output format, FLAC, is lossless and smaller. Delete rejects once the log row is written; keep the log forever.
Going further
- Variations and audio-to-audio: fixing a "maybe" instead of rolling a new batch.
- Licensing and provenance: turning the log row into a provenance record.
- Decibels and loudness: why loudness matching comes before judging.
- The ACE-Step 1.5 tutorial's section on random factors, batches, and scoring.
- Try it: write your listening checklist on one card and use it unchanged for a whole batch.