technique

Choosing a local audio model

Which open model to run for which job: full songs with vocals, instrumental game music, or sound effects, compared by quality, speed, memory, and license.

Before this

This page assumes you are comfortable with:

Why you need this

The first decision in stage 1 is which model to run, and it decides almost everything after it: what you can ask for, how long you wait, whether it fits your machine, and whether you may ship what comes out. No single open model does songs, game loops, and sound effects equally well, so a game's audio set usually needs two or three.

The idea

Three jobs. Game audio splits into three kinds of request, and models specialize:

Job What you ask for What matters most
Song with vocals Lyrics plus a style, sung by a voice, minutes long Clear words, a real chorus, vocal quality
Instrumental music A mood, tempo, and instruments, no voice, 10 s to a few minutes Steady tempo, control over key and BPM, loopable sections
Sound effect A short event: a click, a whoosh, a card sliding Short length, a clean start, no music in it

Song models are trained on songs, so they are good at structure and voices and poor at a 300 ms click. Sound-effect models are trained on recordings of events and say in their own documentation that they cannot sing.

Two architectures. Every model here compresses audio with a codec first (see Neural audio codecs). What differs is how it generates the compressed version:

  • A diffusion model starts from random noise in the codec's latent space and removes noise over a number of steps, steered by the prompt (see Diffusion models). Few steps means fast.
  • A language-model generator writes discrete audio tokens one at a time, like a chatbot writing words. Long outputs take many steps.
  • Hybrids use a language model to plan the structure and a diffusion model to render the detail. Both music models below are hybrids.

The candidates, from their own documentation. These facts were taken from each model's published docs at the time of writing. Model documentation changes quickly; when it disagrees with this table, the documentation wins. VRAM is memory on an NVIDIA graphics card.

Model Job Architecture Size Memory (published) Speed (published) License
ACE-Step 1.5 Songs and instrumentals, 10 s to 10 minutes Language-model planner (0.6B, 1.7B, or 4B) plus a diffusion transformer (2B, or 4B for the XL variants) Core models about 10 GB on disk From under 4 GB VRAM (2B only, offloaded) to 24 GB or more for the XL model with the 4B planner Under 2 s per full song on an A100, under 10 s on an RTX 3090 MIT
LeVo 2 (SongGeneration 2), v2-large Songs with vocals, up to 4 min 30 s Hybrid: about 4B language model plus a diffusion renderer about 4B 22 GB VRAM without prompt audio, 28 GB with Real-time factor 0.82 on an H20 Academic, research, and education use only
LeVo 2, v2-medium Same, smaller Same smaller 12 GB / 18 GB Real-time factor 0.69 Same
Stable Audio Open 1.0 Sound effects, some music, up to 47 s Latent diffusion transformer 1B Not stated Not stated Stability AI Community License
Stable Audio 3 (small SFX, medium) Sound effects and music Latent diffusion transformer 0.6B (small SFX), 2B (medium) Not stated Under 2 s on an H200, a few seconds on a MacBook Pro M4 Stability AI Community License
MOSS-SoundEffect v2.0 Sound effects, up to 30 s, 48 kHz Diffusion transformer trained with flow matching 1.3B Not stated Not stated Apache 2.0

Sources: the ACE-Step 1.5 README, GPU compatibility guide, and install guide; the SongGeneration README and the LeVo 2 technical report; the Stable Audio Open 1.0, Stable Audio 3, and MOSS-SoundEffect v2.0 model cards. A real-time factor is generation time divided by audio length, so 0.82 means 82 seconds of work per 100 seconds of song, excluding model loading. Two more sound-effect models are out for a sold game: Sony AI's Woosh publishes its weights under a non-commercial Creative Commons license (CC BY-NC), and the TangoFlux model card says its checkpoints "are for non-commercial research use only".

What the licenses say, in one line each. MIT and Apache 2.0 let you use the model commercially. The Stability AI Community License is free for commercial use unless you or your organization earn more than USD 1 million a year; above that you need a separate license. The license file shipped with the SongGeneration code says you agree "to use the SongGeneration only for academic, research and education purposes, and refrain from using it for any commercial or production purposes under any circumstances." Put plainly: LeVo 2 is for research and education only under that license, so it is not a choice for shipping game audio, however good it sounds. Its README calls it "commercial-grade"; that describes quality, not permission. Licensing and provenance covers what these licenses mean for output files.

How the two music models differ in practice.

ACE-Step 1.5 LeVo 2
Strength Speed, small memory, fine control: separate settings for BPM, key, time signature, duration Vocal quality and lyric accuracy (its README reports a phoneme error rate of 8.55%)
Style input Free-form caption: tags or sentences Comma-separated tags only; sentences are explicitly discouraged
Instrumental An instrumental setting, or an [Instrumental] tag A pure-music mode
Runs on Apple Silicon Yes, officially, with an MLX backend Official code assumes CUDA; Mac use relies on community ports

Worked example

Pick a model for each item in the Lumen Clash audio set. The game already ships a lobby track; the plan below is for new material. Assume the game may be sold, so licenses matter.

Item Job Pick Why
Lobby music loop Instrumental, about 48 s ACE-Step 1.5 Set BPM and key directly; MIT license; seconds per take
In-match music loop Instrumental, same palette ACE-Step 1.5 Same model and base prompt keeps the soundtrack consistent
Win and lose stingers Instrumental, 2 to 4 s ACE-Step 1.5, then cut Its minimum request is 10 s, so generate a short phrase and trim it (see Stingers and transitions)
Card draw, card play, button click Sound effects under 1 s MOSS-SoundEffect v2.0, or Stable Audio Open 1.0 Built for effects; Apache 2.0, or the Community License if revenue stays under its threshold
A title song with vocals (optional) Song ACE-Step 1.5 for shipping; LeVo 2 only for private sketches LeVo 2's license excludes production use

Does it fit the M5 Max, 128 GB unified memory? Unified memory is one pool shared by the processor and the graphics cores, so VRAM figures translate roughly to a share of that pool. ACE-Step 1.5's top tier asks for 24 GB or more. MOSS-SoundEffect v2.0 publishes no memory figure, so use the rule of thumb from Model size and memory: 1.3B parameters at 16-bit (2 bytes each) is 2.6 GB of weights, plus a 20 percent working margin, about 3.1 GB. That is an estimate, not a published number. Together that is about 27 GB, leaving roughly 100 GB for the system and your other apps. Even LeVo 2 large at its published 28 GB would fit alongside.

How long will a song take? With LeVo 2 large's published real-time factor of 0.82 on an H20, a 3-minute (180 s) song is about 0.82×180=147.60.82 \times 180 = 147.6 s of pure inference on that card. The factor was measured on a data-center NVIDIA card; it does not transfer to a Mac.

In a game's audio pipeline

This is the first question of stage 1, Set up the studio. The answer feeds Running audio models on Apple Silicon and ComfyUI audio workflows, which get the chosen model running, and stage 5, where Licensing and provenance records which model and license produced each file.

Common mistakes

  • One model for everything. You ask a song model for a button click and get two seconds of music with a drum fill. Use a sound-effect model for effects.
  • Reading the license last. You build a soundtrack, then find the model's terms exclude commercial or production use. Check the license before the first take.
  • Comparing speed across machines. You expect "under 10 s on an RTX 3090" on a laptop and wait much longer. Published speeds name their hardware; yours will differ.
  • Trusting a third-party summary over the docs. A blog says a model needs 10 GB; the model's README says 22 GB for the variant you downloaded. Older and newer variants of the same family differ.
  • Ignoring the minimum length. You request 2 s from a music model that accepts 10 s or more and get an error or a padded clip. Generate long, then cut.

Cost

Choosing costs reading time, not compute: an hour with the READMEs and licenses saves a soundtrack you cannot ship. Running costs differ by orders of magnitude. ACE-Step 1.5 turns a full song around in seconds on a recent NVIDIA card; LeVo 2 large takes most of the song's own length on a data-center card and needs about 22 to 28 GB. Disk: about 10 GB for ACE-Step 1.5's core models, more for each extra variant. Sound-effect models are smaller (0.6B to 1.3B parameters), so they cost the least memory and disk; your listening time across many short takes is the larger cost.

Going further

Leads to

Back to Local text-to-audio for games