technique

Variations and audio-to-audio

How to extend a clip, regenerate one part of it, or start from a reference recording instead of from silence.

Before this

This page assumes you are comfortable with:

Why you need this

The take you kept from Curating takes is rarely perfect. One bar has a glitch, or the game needs a tense version of the same tune for the last turns of a match. Rolling a whole new batch throws away everything that was right. Audio-to-audio features start from audio you already have, so you can fix a part, grow a clip, or make a variant that still belongs to the original. This is stage 3 of the pipeline, generate and choose.

The idea

Plain text-to-music starts every take from random noise. Audio-to-audio gives the model existing audio as well, and the model keeps some of it. What it keeps decides which feature you are using.

Job What is kept What is new ACE-Step 1.5 LeVo 2
Fix a section Everything outside a time range The range Repaint not documented
Make it longer All of it More time at the start or end Repaint, at the start or end not documented
Restyle it Melody, rhythm, chords, structure Style, sound, mood Cover (the interface calls it Remix) not documented
Borrow a sound Timbre, mix, performance feel Melody and structure reference_audio prompt_audio_path

The names are each model's own; the settings below are from the ACE-Step 1.5 inference docs and tutorial and from the SongGeneration README.

Repaint: regenerate a time range

Repaint regenerates the audio between two times and leaves the rest unchanged. In ACE-Step 1.5 you set task_type to "repaint", give the clip as src_audio, and set repainting_start and repainting_end in seconds (-1 means the end of the file). The caption describes what the new section should contain. The model listens to the audio on both sides of the range, so the new part is shaped to fit its neighbors.

The tutorial gives an operating range of 3 to 90 seconds for one repaint. It lists the uses: change the content or lyrics of a section, change a section's structure, continue a clip at its beginning or end, and smooth the join between two clips. Long pieces can be built from repeated repaint passes, each 3 to 90 s, each continuing from the last. That continuation is how you extend a clip in ACE-Step 1.5; there is no separate extend task. (The base model's complete task is described in the inference docs as "Complete/extend partial tracks", but the tutorial explains it as adding accompaniment to a single track, such as backing for a bare vocal, not as adding time.)

Choose ranges on bar lines. A bar at tempo TT BPM in 4/4 lasts 4×60/T4 \times 60 / T seconds. If the range starts and ends on bar lines, the new material enters on a downbeat (the first beat of a bar) and leaves on one, and the seams fall where music naturally changes.

Cover: keep the structure, change the style

Cover turns the source into semantic codes, the model's description of melody, rhythm, chords, and arrangement, and generates new audio from them under a new caption. Set task_type to "cover", the source as src_audio, and a caption for the new style. audio_cover_strength (0.0 to 1.0, default 1.0) sets how closely the structure is followed: the docs describe 1.0 as strong adherence, 0.5 as balanced, and 0.1 as loose, and suggest about 0.2 for style transfer. The tutorial calls one use a "retake lottery": the same structure, new interpretations, one per seed.

Reference audio: borrow a sound, not a tune

reference_audio guides the acoustic side: timbre, mixing, performance style. The tutorial explains the processing. The audio is converted to 48 kHz stereo; if shorter than 30 s it is repeated; three 10 s pieces from the front, middle, and back are joined into 30 s; and that is encoded in a way that drops melody and rhythm. So a reference sets how the result sounds, not what it plays.

LeVo 2's equivalent is prompt_audio_path. The SongGeneration README says only the first 10 seconds are used, recommends a song's chorus as the prompt for best musicality and structure, and says it can steer genre, instrumentation, rhythm, and voice. It warns against giving both a prompt audio and descriptions at once, since conflicts degrade the song. Without your own file, auto_prompt_audio_type picks a reference from a built-in library by style name ('Pop', 'Rock', 'Soundtrack', 'Auto', and others). The README describes no repaint, extend, or cover feature for LeVo 2.

Worked example

Illustration only: the takes, seeds, and settings below show the method, not a record of real runs.

Fixing a bad bar in the in-match loop. Suppose the in-match keeper is 24 bars of 4/4 at 150 BPM. One bar is 4×60/150=1.64 \times 60 / 150 = 1.6 s, so the take is 24×1.6=38.424 \times 1.6 = 38.4 s. Bar 9 has a drum fill that stumbles. With bars numbered from 1, bar nn starts at (n−1)×1.6(n - 1) \times 1.6 s.

Bars Start End Length
The bad bar 9 12.8 s 14.4 s 1.6 s, under the 3 s minimum
Repaint range 8 to 10 11.2 s 16.0 s 4.8 s

Repainting only bar 9 is below the tutorial's 3 s minimum, so widen the range by a bar on each side. That also moves the seams away from the problem.

Settings: task_type "repaint", src_audio the in-match keeper, repainting_start 11.2, repainting_end 16.0, caption the same as the original take (a new caption invites a new style in the middle of the loop), seeds 5201, 5202, and 5203. Judge each result by its seams: loop 2 s either side of 11.2 s and of 16.0 s, then play the whole 38.4 s once. Log the keeper's seed as a repaint of the original take, so the record shows both steps.

A tense variant of the lobby loop. The lobby keeper (from Curating takes: seed 4129, 120 BPM, 48 s) should get a darker cousin for late in a match, one that can crossfade with the original.

Setting Variant 1 Variant 2
task_type "cover" "cover"
src_audio lobby keeper lobby keeper
caption the lobby base prompt plus "tense, darker, driving low strings" same
bpm, duration 120, 48 120, 48
audio_cover_strength 0.7 0.4
seed 4301 4301

Same seed, one setting changed, as on Seeds and sampling settings. At 0.7 the melody and bar structure should survive, so bar 9 of the variant lines up under bar 9 of the original and a crossfade between them stays in time. At 0.4 the variant is freer and may sound more different, but its sections may no longer line up. For a crossfade, alignment matters more than novelty.

In a game's audio pipeline

This is stage 3, generate and choose, after Curating takes has picked a keeper. Repairs here save a new batch and a new listening session. The repaired loop and its variant then go to stage 4: Seamless music loops cuts them, and Stingers and transitions handles moving between the calm and tense versions.

Common mistakes

  • Ranges off the bar lines. The repainted part enters half a beat late; you hear a hiccup at the seam every time the loop comes round.
  • Ranges too short. A range under 3 s is outside the tutorial's stated operating range, and the model has little room to blend in; widen to whole bars.
  • A new caption for a repair. The repainted bars sound like a different song dropped into the middle.
  • Low cover strength for a variant meant to crossfade. The variant's bars drift out of line with the original, so the crossfade smears two rhythms.
  • Using a reference to copy a melody. Reference audio carries sound, not tune; the result plays something else in a similar tone.
  • Prompt audio and descriptions together in LeVo 2. If they disagree, the README warns, quality drops.

Cost

A repair costs one short generation and a short listen: the 4.8 s range is 12.5% of the 38.4 s take, and the seams are what you check. Listening to a whole new batch of 38.4 s takes would cost far more. Cover produces a full-length clip, so it costs about the same as a fresh take, but you start much closer to the goal. Memory can rise: the SongGeneration README lists SongGeneration-v2-large at 22 GB of GPU memory without prompt audio and 28 GB with it. Every variant is another file and another log row, so name variants after their source take.

Going further

  • Seamless music loops: cutting the repaired take into a loop.
  • Stingers and transitions: switching between the calm and tense loops.
  • The ACE-Step 1.5 tutorial's section on audio control, for repaint, cover, and reference audio in more depth.
  • The ACE-Step 1.5 base model's lego and complete tasks, for adding instrument tracks to existing audio.
  • Try it: pick a bar you want to change in any song, and work out its repaint range in seconds from the tempo.

Back to Local text-to-audio for games