Topic hub

Local text-to-audio for games

How to turn a written description into game music, songs, and sound effects with open models running on your own machine, from choosing a model to shipping the files, with the sound and machine-learning basics underneath.

The problem

A game needs a lot of audio: music that loops under the menu, a different loop for play, short cues for winning and losing, and dozens of small sounds for every card, button, and hit. Buying or commissioning all of it is expensive, and a hosted generator sends every prompt to someone else's server and puts its own terms on what comes back.

Open audio models now run on a well-equipped laptop. You type a description, or a set of lyrics, and a few seconds to a few minutes later you have a clip. What you do not have yet is a game asset. A raw clip rarely loops, its loudness does not match the next clip, its ending trails off, and you may not know what you are allowed to do with it. This cluster covers the whole trip: picking and running a model, writing the description, getting a good take and keeping a record of it, shaping it into loops, stingers, and effects, and shipping it with a clear record of where it came from.

Underneath every technique is a small set of first-year ideas about sound and about how these models work. Each one has its own page. Follow the "Before this" links downward until you reach something you already know, then read back up.

The pipeline

Every page points back to this picture of how a piece of game audio gets made.

Stage Question it answers Techniques that live here
1. Set up the studio Which model, on which machine, through which tool? Choosing a local audio model, Running audio models on Apple Silicon, ComfyUI audio workflows
2. Write the brief How do I describe the sound I want? Prompting for music, Lyrics and song structure, Prompting sound effects
3. Generate and choose How do I get good takes reliably, and pick the right one? Seeds and sampling settings, Variations and audio-to-audio, Curating takes
4. Shape it for the game How does a raw clip become a usable game asset? Seamless music loops, Stingers and transitions, Editing and layering sound effects, Loudness for games
5. Ship it What format does it go into the game in, and am I allowed to use it? Exporting game audio, Licensing and provenance

Stage 3 is where the time goes. Generation is fast; listening is not. Stage 4 is where most clips that sounded fine on their own fail inside a game, because a game plays them on repeat, back to back, and on top of each other.

Which page for which job

You have never run a model. Read Choosing a local audio model, then Running audio models on Apple Silicon, then ComfyUI audio workflows. Those three get one clip out of a model on your own machine.

You want a song with vocals, the kind of thing a hosted song generator produces. Prompting for music, then Lyrics and song structure. Music basics underneath both. For a song that ships in a game, use ACE-Step 1.5: LeVo 2 sings well, but its license rules out commercial and production use.

You want the music for a game. Prompting for music, Seeds and sampling settings, Curating takes, then Seamless music loops and Stingers and transitions.

You want sound effects. Prompting sound effects, then Editing and layering sound effects, then Loudness for games.

Something sounds wrong and you do not know why. The usual suspects, in order: a click every time the music loops (Seamless music loops), one sound much louder than the rest (Loudness for games), a harsh or crackly take (Seeds and sampling settings, on guidance), a gap at the loop point after export (Exporting game audio), and muddy, smeared detail (Neural audio codecs explains what the model's compression throws away).

You are about to ship. Exporting game audio and Licensing and provenance, in that order, before any file goes into a release. Better still, read the licensing page before you pick a model: a model whose license forbids commercial use makes every take from it unusable in a released game.

The running example across the cluster is the audio set for Lumen Clash, a digital card game: a lobby music loop, an in-match music loop, win and lose stingers, and short card and UI sounds for draw, play, and a button click. The prompts, seeds, and settings shown in worked examples are illustrations, not records of real runs.

The basics underneath

None of these needs anything beyond first-year college material, and most need less.

You need For
Sound waves Frequency, pitch, amplitude, and timbre: what the model is ultimately producing.
Digital audio and sampling Sample rate, bit depth, aliasing, and how big an audio file is.
Decibels and loudness Why levels are logarithmic, what dBFS and LUFS mean, and why peak is not loudness.
Spectrograms Seeing a sound as frequencies over time, the picture both models and editors work with.
Music basics for prompting Tempo, bars, key, chords, and song sections: the words a music model listens for.
Transformers and tokens How a model generates one token at a time, for words or for audio.
Diffusion models How a model starts from noise and removes it step by step, steered by your prompt.
Neural audio codecs How audio is squeezed into tokens or latent frames a generator can handle, and what gets lost.
Model size and memory Parameters, bytes per parameter, quantization, and whether a model fits your machine.

Conventions used across this cluster

So the pages agree with each other:

  • Sample rates are 44.1 kHz or 48 kHz, meaning 44,100 or 48,000 samples per second.
  • Levels: dBFS for peaks, where 0 dBFS is the loudest a file can hold; LUFS for how loud a whole clip sounds; dB for a change in level.
  • Tempo is in BPM, time signatures are written 4/4, and musical sections are measured in bars. Everything else is in seconds, or milliseconds under a second.
  • Sizes use 1 MB = 1,000,000 bytes.
  • Models are named with their version: ACE-Step 1.5, LeVo 2 (also published as SongGeneration 2). "2B" means two billion parameters.
  • The hands-on machine is a MacBook Pro with an M5 Max and 128 GB of unified memory. Where a step differs on an NVIDIA graphics card, the page says so in one sentence.
  • Model facts come from the models' own published documentation at the time of writing. These models change quickly; when a number on a page and the model's current documentation disagree, the documentation wins.

One tool per job. Technique pages describe what a tool must do rather than which one to use, except the ComfyUI page. Names only, no links: names stay stable, download pages do not.

Job Pick Why Also fine
Running music models with a repeatable workflow ComfyUI ACE-Step 1.5 runs on its built-in nodes, workflows save as files you can rerun, and it runs on Apple Silicon. The model's own web interface
Trying SongGeneration songs on a Mac, for study only SongGen-Mac A Mac-specific build of the SongGeneration interface, patched to run on Apple Silicon. Its own documentation lists the first-generation SongGeneration models, not LeVo 2; the Mac port of LeVo 2 itself is the MLX conversion described on Running audio models on Apple Silicon. The SongGeneration license allows academic, research, and education use only, so nothing made with either goes into a shipped game; see Licensing and provenance. The official SongGeneration code on an NVIDIA machine with enough VRAM
Editing, trimming, fades, and loop checking Audacity Free, loop playback, zoom to the sample, loudness normalization built in. Ocenaudio (simpler), Reaper (paid, full production)
Measuring loudness Youlean Loudness Meter Reads integrated LUFS and true peak, the two numbers the loudness page asks for. Has a free version. The loudness tools inside your editor
Batch format conversion FFmpeg Converts a folder of masters to the game's format with the same settings every time. Your editor's batch export

Settings that matter in any tool. Keep a lossless master of every keeper (WAV or FLAC). Pick one sample rate for the whole game and convert everything to it once. Note the seed and settings of every take you keep, before you close the tool.

What not to use. A hosted generator for anything you plan to ship without reading its terms first, and any "AI master" or "enhance" button run over a clip you have not listened to. Both can change what you are allowed to do with the file, or what it sounds like, without telling you.

Prices and versions are deliberately absent. "Free" or "paid" is all this page will say, because anything more specific goes stale.

How to read this cluster

The learning path below is sorted so that each row depends only on the rows above it. If you already work with audio, skip the sound basics and start at Choosing a local audio model. If you are here for one song, start at Lyrics and song structure and follow its links back only as far as you need.

Learning path

Each row depends only on rows above it. Read top to bottom, or jump to a technique and follow its "Before this" links downward.