Topic hub
Local text-to-audio for games
How to turn a written description into game music, songs, and sound effects with open models running on your own machine, from choosing a model to shipping the files, with the sound and machine-learning basics underneath.
The problem
A game needs a lot of audio: music that loops under the menu, a different loop for play, short cues for winning and losing, and dozens of small sounds for every card, button, and hit. Buying or commissioning all of it is expensive, and a hosted generator sends every prompt to someone else's server and puts its own terms on what comes back.
Open audio models now run on a well-equipped laptop. You type a description, or a set of lyrics, and a few seconds to a few minutes later you have a clip. What you do not have yet is a game asset. A raw clip rarely loops, its loudness does not match the next clip, its ending trails off, and you may not know what you are allowed to do with it. This cluster covers the whole trip: picking and running a model, writing the description, getting a good take and keeping a record of it, shaping it into loops, stingers, and effects, and shipping it with a clear record of where it came from.
Underneath every technique is a small set of first-year ideas about sound and about how these models work. Each one has its own page. Follow the "Before this" links downward until you reach something you already know, then read back up.
The pipeline
Every page points back to this picture of how a piece of game audio gets made.
| Stage | Question it answers | Techniques that live here |
|---|---|---|
| 1. Set up the studio | Which model, on which machine, through which tool? | Choosing a local audio model, Running audio models on Apple Silicon, ComfyUI audio workflows |
| 2. Write the brief | How do I describe the sound I want? | Prompting for music, Lyrics and song structure, Prompting sound effects |
| 3. Generate and choose | How do I get good takes reliably, and pick the right one? | Seeds and sampling settings, Variations and audio-to-audio, Curating takes |
| 4. Shape it for the game | How does a raw clip become a usable game asset? | Seamless music loops, Stingers and transitions, Editing and layering sound effects, Loudness for games |
| 5. Ship it | What format does it go into the game in, and am I allowed to use it? | Exporting game audio, Licensing and provenance |
Stage 3 is where the time goes. Generation is fast; listening is not. Stage 4 is where most clips that sounded fine on their own fail inside a game, because a game plays them on repeat, back to back, and on top of each other.
Which page for which job
You have never run a model. Read Choosing a local audio model, then Running audio models on Apple Silicon, then ComfyUI audio workflows. Those three get one clip out of a model on your own machine.
You want a song with vocals, the kind of thing a hosted song generator produces. Prompting for music, then Lyrics and song structure. Music basics underneath both. For a song that ships in a game, use ACE-Step 1.5: LeVo 2 sings well, but its license rules out commercial and production use.
You want the music for a game. Prompting for music, Seeds and sampling settings, Curating takes, then Seamless music loops and Stingers and transitions.
You want sound effects. Prompting sound effects, then Editing and layering sound effects, then Loudness for games.
Something sounds wrong and you do not know why. The usual suspects, in order: a click every time the music loops (Seamless music loops), one sound much louder than the rest (Loudness for games), a harsh or crackly take (Seeds and sampling settings, on guidance), a gap at the loop point after export (Exporting game audio), and muddy, smeared detail (Neural audio codecs explains what the model's compression throws away).
You are about to ship. Exporting game audio and Licensing and provenance, in that order, before any file goes into a release. Better still, read the licensing page before you pick a model: a model whose license forbids commercial use makes every take from it unusable in a released game.
The running example across the cluster is the audio set for Lumen Clash, a digital card game: a lobby music loop, an in-match music loop, win and lose stingers, and short card and UI sounds for draw, play, and a button click. The prompts, seeds, and settings shown in worked examples are illustrations, not records of real runs.
The basics underneath
None of these needs anything beyond first-year college material, and most need less.
| You need | For |
|---|---|
| Sound waves | Frequency, pitch, amplitude, and timbre: what the model is ultimately producing. |
| Digital audio and sampling | Sample rate, bit depth, aliasing, and how big an audio file is. |
| Decibels and loudness | Why levels are logarithmic, what dBFS and LUFS mean, and why peak is not loudness. |
| Spectrograms | Seeing a sound as frequencies over time, the picture both models and editors work with. |
| Music basics for prompting | Tempo, bars, key, chords, and song sections: the words a music model listens for. |
| Transformers and tokens | How a model generates one token at a time, for words or for audio. |
| Diffusion models | How a model starts from noise and removes it step by step, steered by your prompt. |
| Neural audio codecs | How audio is squeezed into tokens or latent frames a generator can handle, and what gets lost. |
| Model size and memory | Parameters, bytes per parameter, quantization, and whether a model fits your machine. |
Conventions used across this cluster
So the pages agree with each other:
- Sample rates are 44.1 kHz or 48 kHz, meaning 44,100 or 48,000 samples per second.
- Levels: dBFS for peaks, where 0 dBFS is the loudest a file can hold; LUFS for how loud a whole clip sounds; dB for a change in level.
- Tempo is in BPM, time signatures are written 4/4, and musical sections are measured in bars. Everything else is in seconds, or milliseconds under a second.
- Sizes use 1 MB = 1,000,000 bytes.
- Models are named with their version: ACE-Step 1.5, LeVo 2 (also published as SongGeneration 2). "2B" means two billion parameters.
- The hands-on machine is a MacBook Pro with an M5 Max and 128 GB of unified memory. Where a step differs on an NVIDIA graphics card, the page says so in one sentence.
- Model facts come from the models' own published documentation at the time of writing. These models change quickly; when a number on a page and the model's current documentation disagree, the documentation wins.
Recommended software
One tool per job. Technique pages describe what a tool must do rather than which one to use, except the ComfyUI page. Names only, no links: names stay stable, download pages do not.
| Job | Pick | Why | Also fine |
|---|---|---|---|
| Running music models with a repeatable workflow | ComfyUI | ACE-Step 1.5 runs on its built-in nodes, workflows save as files you can rerun, and it runs on Apple Silicon. | The model's own web interface |
| Trying SongGeneration songs on a Mac, for study only | SongGen-Mac | A Mac-specific build of the SongGeneration interface, patched to run on Apple Silicon. Its own documentation lists the first-generation SongGeneration models, not LeVo 2; the Mac port of LeVo 2 itself is the MLX conversion described on Running audio models on Apple Silicon. The SongGeneration license allows academic, research, and education use only, so nothing made with either goes into a shipped game; see Licensing and provenance. | The official SongGeneration code on an NVIDIA machine with enough VRAM |
| Editing, trimming, fades, and loop checking | Audacity | Free, loop playback, zoom to the sample, loudness normalization built in. | Ocenaudio (simpler), Reaper (paid, full production) |
| Measuring loudness | Youlean Loudness Meter | Reads integrated LUFS and true peak, the two numbers the loudness page asks for. Has a free version. | The loudness tools inside your editor |
| Batch format conversion | FFmpeg | Converts a folder of masters to the game's format with the same settings every time. | Your editor's batch export |
Settings that matter in any tool. Keep a lossless master of every keeper (WAV or FLAC). Pick one sample rate for the whole game and convert everything to it once. Note the seed and settings of every take you keep, before you close the tool.
What not to use. A hosted generator for anything you plan to ship without reading its terms first, and any "AI master" or "enhance" button run over a clip you have not listened to. Both can change what you are allowed to do with the file, or what it sounds like, without telling you.
Prices and versions are deliberately absent. "Free" or "paid" is all this page will say, because anything more specific goes stale.
How to read this cluster
The learning path below is sorted so that each row depends only on the rows above it. If you already work with audio, skip the sound basics and start at Choosing a local audio model. If you are here for one song, start at Lyrics and song structure and follow its links back only as far as you need.
Learning path
Each row depends only on rows above it. Read top to bottom, or jump to a technique and follow its "Before this" links downward.
- prerequisiteSound wavesWhat sound physically is: pressure waves with a frequency you hear as pitch and an amplitude you hear as loudness.
- prerequisiteDiffusion modelsHow a model creates something by starting from noise and removing a little of it at each step, steered by a text description.
- prerequisiteModel size and memoryHow many parameters a model has, how many bytes each one takes, and whether it fits in your machine's memory.
- prerequisiteTransformers and tokensHow a language model generates one token at a time, and how the same idea generates audio once sound is turned into tokens.
- prerequisiteDecibels and loudnessWhy audio levels are measured on a log scale, what dBFS and LUFS mean, and how loudness differs from peak level.
- prerequisiteDigital audio and samplingHow a sound wave becomes a list of numbers: sample rate, bit depth, channels, aliasing, and how big an audio file is.
- prerequisiteMusic basics for promptingTempo, bars, time signature, key, chords, and song sections: the musical vocabulary a music model understands.
- prerequisiteSpectrogramsHow a sound is pictured as frequencies over time, the view that audio models and audio editors both rely on.
- techniqueLoudness for gamesHow to make every music track and sound effect sit at a consistent level, with headroom so nothing clips when sounds overlap.
- techniqueSeamless music loopsHow to cut a generated track into a loop that plays forever without a click, a hiccup, or a tempo jump.
- prerequisiteNeural audio codecsHow a neural network squeezes audio into a short sequence of tokens or latent frames that a generator can work with, and turns them back into sound.
- techniqueExporting game audioWhich file formats, sample rates, and compression settings to use so game audio is small, loads fast, and loops cleanly.
- techniqueStingers and transitionsShort musical cues for winning, losing, and changing scenes, and how to move between loops without a jarring cut.
- techniqueComfyUI audio workflowsHow a node-graph tool turns a generation into a repeatable, shareable workflow, using the built-in ACE-Step template.
- techniquePrompting for musicHow to describe game music so the model delivers it: genre, mood, instruments, tempo, key, and what models tend to ignore.
- techniquePrompting sound effectsHow to describe one-shot sounds, UI clicks, impacts, and ambience so a model produces something short, clean, and usable.
- techniqueRunning audio models on Apple SiliconWhat unified memory, the Metal backend, and MLX mean for running music models on a Mac, and how to get a first clip out of one.
- techniqueSeeds and sampling settingsWhat the seed, number of steps, guidance strength, and duration settings do, and how to change one at a time.
- techniqueEditing and layering sound effectsHow to trim, fade, and combine generated sounds into crisp game effects.
- techniqueLyrics and song structureHow to write lyrics and section tags so a model sings a full song with verses and a chorus where you want them.