technique
ComfyUI audio workflows
How a node-graph tool turns a generation into a repeatable, shareable workflow, using the built-in ACE-Step template.
Before this
This page assumes you are comfortable with:
Why you need this
A model's own web page is fine for one clip. A game needs dozens, made the same way, and remade months later when a loop needs one more variation. ComfyUI, a free node-graph tool that runs generative models locally, stores the whole recipe (model files, prompt, seed, every setting) as a workflow file you can reopen and rerun. This is the one page in the cluster that names ComfyUI and its node packs; elsewhere the pages describe what any tool must do.
The idea
A node graph. A workflow is a set of boxes, called nodes, joined by wires. Each node does one step: load a model, encode a prompt, make empty audio to fill, run the sampler, decode, save. A wire carries one piece of data from an output on one node to an input on another, and its color shows the data type (a model, conditioning, a latent, audio). Settings that are not wires are widgets inside the node: a text box, a number, a dropdown. When you press Run, ComfyUI works out which nodes the save node depends on and runs them in order.
The built-in ACE-Step 1.5 template. ACE-Step 1.5 runs on ComfyUI's built-in nodes; no extra install is needed beyond the model files. ComfyUI's official ACE-Step 1.5 tutorial says to update ComfyUI first and open the template from the Template Library. It names two versions: the recommended "ACE-Step 1.5 Music Generation AIO" template, which loads a single all-in-one checkpoint file, ace_step_1.5_turbo_aio.safetensors (9.34 GB, in the models/checkpoints folder), and a split version that loads the diffusion model, the two text-encoder models, and the VAE as four separate files. The template library also ships XL base, SFT, and turbo variants. ComfyUI's desktop app for Mac requires Apple Silicon and macOS 13 or later.
What each node in the AIO template does, in the order data flows:
| Node | What it does |
|---|---|
| CheckpointLoaderSimple | Loads the all-in-one file and outputs three things: the diffusion model (MODEL), the text side (CLIP, here ACE-Step's language-model planner and text encoder), and the VAE (the codec's decoder) |
| ModelSamplingAuraFlow | Sets the timestep shift of the diffusion schedule (the template uses 3, matching the ACE-Step docs' advice of 3.0 for turbo) |
| TextEncodeAceStepAudio1.5 | The brief: tags (the style text ACE-Step's own docs call the caption), lyrics, seed, bpm, duration, timesignature, language, keyscale, plus advanced settings for the planner: generate_audio_codes, cfg_scale, temperature, top_p, top_k, min_p |
| ConditioningZeroOut | Makes an empty copy of the conditioning to use as the "negative" input; this template has no negative prompt |
| EmptyAceStep1.5LatentAudio | Makes the blank latent to be filled: seconds and batch_size |
| KSampler | Runs the diffusion steps: seed, control_after_generate, steps, cfg, sampler_name, scheduler, denoise |
| VAEDecodeAudio | Turns the finished latent back into audio samples |
| SaveAudioAdvanced | Writes the file; the template saves FLAC into an audio subfolder of ComfyUI's output folder |
Two small Primitive nodes, titled "seed" and "Song Duration", feed the same value to two places each. The seed goes to the text encoder (where it drives the planner's sampling) and to the KSampler (where it sets the starting noise). The duration goes to the text encoder and to the empty latent. One number, one wire to each user: you cannot accidentally change one copy and not the other.
Where the settings live.
| You want to change | Node and widget |
|---|---|
| Prompt (style text) | TextEncodeAceStepAudio1.5, tags |
| Lyrics or an instrumental request | TextEncodeAceStepAudio1.5, lyrics |
| Seed | the "seed" Primitive node |
| Steps and guidance | KSampler, steps and cfg |
| Duration | the "Song Duration" Primitive node |
| Tempo, key, meter | TextEncodeAceStepAudio1.5, bpm, keyscale, timesignature |
| Takes per run | EmptyAceStep1.5LatentAudio, batch_size |
What the seed, steps, and guidance do to the sound is the subject of Seeds and sampling settings. One thing specific to this graph: there are two guidance numbers. The KSampler's cfg applies to the diffusion model, and the template sets it to 1 because, per the ACE-Step docs, turbo models do not use guidance. The text encoder's cfg_scale applies to the planner.
Saving a workflow so a result can be reproduced. Save or export the workflow from ComfyUI's menu; it is a JSON file holding every node, wire, and widget value. Set each seed's control_after_generate to fixed, or the seed changes after every run and the saved file no longer matches the take you kept. ComfyUI's documentation on workflow metadata says its built-in save nodes embed the workflow in images, video, and latent files; it does not list audio formats, so do not count on the FLAC carrying the recipe. Save the JSON next to the take, with the same name.
Community node packs. Models without built-in support arrive as custom nodes, folders of Python code installed into ComfyUI. For LeVo 2 / SongGeneration there are at least two packs. ComfyUI_SongGeneration (by smthemex) says it supports the v2 model, adds GGUF quantized files, and offloads layers so 12 GB VRAM users can run it; its stated test environment is Windows 11 with CUDA 12.4. ComfyUI_FL-SongGen lists only first-generation SongGeneration models, requires CUDA 11.8 or newer for GPU use, and notes that "Mac MPS may have limited support". Both describe their requirements in NVIDIA terms, so on a Mac expect to debug. Remember also that LeVo 2's license restricts it to research and education.
Worked example
The Lumen Clash lobby loop in the AIO template. Every value here is an illustration, not a record of a real run. The prompt and musical settings come from Prompting for music.
| Order | Node | Settings for the lobby loop | Template default |
|---|---|---|---|
| 1 | CheckpointLoaderSimple | ace_step_1.5_turbo_aio.safetensors |
same |
| 2 | ModelSamplingAuraFlow | shift 3 | 3 |
| 3 | "seed" Primitive | 4129, fixed | 31, fixed |
| 4 | "Song Duration" Primitive | 48 | 120 |
| 5 | TextEncodeAceStepAudio1.5 | tags: the lobby caption; lyrics: [Instrumental]; bpm 120; timesignature 4; language en; keyscale A minor; generate_audio_codes on; cfg_scale 2 |
a neo-soul song with lyrics; bpm 190; keyscale E minor |
| 6 | ConditioningZeroOut | none | none |
| 7 | EmptyAceStep1.5LatentAudio | seconds 48 (from the Primitive), batch_size 1 |
120, 1 |
| 8 | KSampler | steps 8, cfg 1, sampler_name euler, scheduler simple, denoise 1 |
same |
| 9 | VAEDecodeAudio | none | none |
| 10 | SaveAudioAdvanced | FLAC, filename prefix lumen-lobby |
FLAC, audio/ComfyUI |
How big is the latent? ComfyUI's ACE-Step 1.5 latent node makes frames of 64 numbers, which is 25 frames per second. For 48 s that is frames, or numbers: the whole loop, in the codec's compressed form, before decoding (see Neural audio codecs).
Turbo here, SFT in the take log. The take log on Curating takes records the lobby keeper on the acestep-v15-sft model at 50 steps and guidance 7.0. In ComfyUI, the matching move is a template whose loader names an SFT file; the XL SFT template, for example, sets the KSampler to 50 steps and cfg 7. Same seed on a different model is a different take, so log the model file name with every seed.
Disk. The split template's four files are 4.46 GB (diffusion model), 1.11 GB and 3.45 GB (two text encoders), and 321.8 MB (VAE): about 9.34 GB, the same as the all-in-one file. Sizes use 1 GB = 1,000,000,000 bytes.
In a game's audio pipeline
ComfyUI is the "which tool" answer in stage 1, Set up the studio, and the workbench for stage 3, Generate and choose: batches, fixed seeds, and one saved workflow per asset type. The saved JSON is also a provenance record for stage 5: it names the model file, the prompt, the seed, and every setting.
Common mistakes
- A seed set to randomize. You save the workflow, rerun it, and get a different track. Set
control_after_generatetofixedbefore you save. - Changing duration in one place. The planner plans 48 s but the latent holds 120 s, and the end is silence or mush. Use the shared Primitive node.
- Raising
cfgon a turbo model. The take turns harsh or distorted. Turbo expectscfg1; use an SFT or base model for guidance. - Trusting the audio file to hold the workflow. Months later the FLAC opens with no graph. Save the JSON yourself.
- Installing a CUDA-only node pack on a Mac. ComfyUI starts with red "missing node" boxes or import errors in the console. Check a pack's stated platforms first.
- A stale ComfyUI. The template opens with unknown nodes. Update first, as the tutorial says.
Cost
Learning the graph takes an afternoon; after that a new take is a few widget edits. Disk is about 9.34 GB for the ACE-Step 1.5 turbo files, more for each XL variant and each community model. Generation time is the model's own: ComfyUI's tutorial quotes a 4-minute song in about 1 s on an RTX 5090 and under 10 s on an RTX 3090; expect longer on a Mac. Each saved workflow is a small JSON file, a negligible cost for being able to remake any take exactly.
Going further
- ComfyUI's official ACE-Step 1.5 tutorial and the other audio templates in its Template Library.
- ComfyUI's documentation on workflows and on workflow metadata.
- Seeds and sampling settings, for what
steps,cfg, and the seed do. - Curating takes, for running batches through a saved workflow and logging the keepers.