technique

Prompting for music

How to describe game music so the model delivers it: genre, mood, instruments, tempo, key, and what models tend to ignore.

Before this

This page assumes you are comfortable with:

Why you need this

The prompt is the brief you hand the model in stage 2, and a vague brief gets a generic track. Game music has needs a song does not: no vocals, a steady tempo you can cut into bars, a mood that suits one screen, and a family resemblance across every track in the game. This page shows how to write prompts that ask for exactly that, in the form each model's documentation says it expects.

The idea

A prompt is a description, not a command. A music model learned which words go with which sounds from labeled examples. Words that appeared often in its training labels (genre names, instrument names, mood words) steer it strongly. Words that appeared rarely (chord progressions, bar counts, "loopable") steer it weakly or not at all.

Tags or sentences: each model says which it wants.

Model What the docs call the style text Form Where tempo and key go
ACE-Step 1.5 caption, up to 512 characters Simple style words, comma-separated tags, or full sentences; its tutorial says the model was trained to handle all three Separate settings: bpm (30 to 300), keyscale (for example "A minor"), timesignature, duration (10 to 600 s). The tutorial says not to write tempo, BPM, or key into the caption.
LeVo 2 descriptions Comma-separated tags only. The SongGeneration README warns: "Do not write full descriptive sentences or natural language paragraphs." The README lists four tag dimensions: gender, genre, emotion, instrument. It names no tempo or key field; one community integration's guide adds a "the bpm is 120" phrase, which the official README does not mention.

LeVo 2's license limits it to research and education, so for a shipping game the prompts on this page are written for ACE-Step 1.5; the LeVo 2 column is here so you can read its documentation.

Order of importance. Put the strongest information first and stop before the prompt contradicts itself:

  1. Genre or style: "cinematic fantasy", "lo-fi hip hop", "chiptune". One or two, not five.
  2. Mood: "calm", "mysterious", "triumphant", "tense".
  3. Instrumentation: name three or four instruments you want to hear. A missing instrument is the most common way a take fails.
  4. Texture and production: ACE-Step 1.5's tutorial lists timbre words (warm, bright, crisp, airy, punchy) and production words (lo-fi, high-fidelity, live recording) as dimensions that shape the mix.
  5. Tempo, key, meter: in ACE-Step 1.5, as settings. The tutorial calls these guidance rather than precise commands: the model samples around the value you give.

Instrumental only. ACE-Step 1.5 has an instrumental setting, and its tutorial says to use an [Instrumental] tag in the lyrics, alone or with section tags. LeVo 2 has a pure-music mode (--bgm). Ask for it explicitly; a song model with no lyrics may still hum or chant.

Negative words. Neither the ACE-Step 1.5 tutorial nor the SongGeneration README describes a negative-prompt field (a list of things to avoid). Writing "no vocals" or "no drums" into the caption mentions the very thing you do not want, and models often react to the word rather than the "no". Use the instrumental setting for vocals, and simply leave out instruments you do not want.

Consistency across a soundtrack. Write one base prompt that fixes genre, palette, and production for the whole game. Each track adds a few mood and energy words to it and changes its settings. Same base, same model, and same key for tracks that can follow each other make the soundtrack sound like one composer wrote it.

Weak versus strong.

Weak prompt Problem Stronger version Why it is better
"good background music" No genre, mood, or instruments; the model picks for you "cinematic fantasy, calm, warm piano, soft synth pads, round bass" Genre, mood, and named instruments
"epic fast slow music" Contradiction; the tutorial warns against it "orchestral, driving, urgent strings" with bpm 150 One direction; tempo in its setting
"lofi beat 85 bpm in D minor" Tempo and key buried in the caption "lo-fi hip hop, relaxed, dusty drums, mellow electric piano" with bpm 85, keyscale "D minor" Settings carry the numbers
"a seamless 8-bar loop" Loop structure is not something the model was labeled with Ask for the music; cut the loop yourself See Seamless music loops
A 600-character paragraph Over the 512-character caption limit The same ideas in about 150 characters Fits; the key words carry the weight

Worked example

These prompts, seeds, and settings are illustrations, not records of real runs. Model: ACE-Step 1.5.

The base prompt for every Lumen Clash music track:

cinematic fantasy, warm felt piano, soft synth pads, round electric bass,
light percussion, polished modern production

Written on one line, that is 118 characters, well under the 512 limit. The lobby caption below comes to 151 characters and the in-match caption to 163.

The two loops share the base, the key, and the meter, and differ in mood words and tempo:

Setting Lobby loop In-match loop
caption base + ", calm, hopeful, spacious, gentle" base + ", tense, driving, urgent, pulsing low strings"
instrumental on on
keyscale A minor A minor
timesignature 4 4
bpm 120 150
duration 48 s 38.4 s
seed 4129 5110

Why these tempos and durations. From Music basics for prompting, one bar of 4/4 lasts 4×60/BPM4 \times 60 / \text{BPM} seconds.

Lobby, 120 BPM In-match, 150 BPM
One beat 60/120=0.560 / 120 = 0.5 s 60/150=0.460 / 150 = 0.4 s
One bar 2 s 1.6 s
8 bars 16 s 12.8 s
24 bars (the request) 48 s 38.4 s

At 120 BPM, 8 bars land on a round 16 s, so the lobby request of 48 s holds exactly three 8-bar phrases. At 150 BPM, 8 bars come to 12.8 s; that is not round, but 24 bars still come to an exact 38.4 s, and a loop is cut in samples anyway. Requesting a whole number of bars means the model's own phrasing is more likely to line up with the bar lines you will cut on.

Both loops use A minor, so a crossfade between them, or a win stinger in A minor, does not clash. The lobby's slower tempo and mood words make it the calm member of the family; the in-match loop keeps every base instrument and adds pulsing low strings for pressure.

The LeVo 2 equivalent, for comparison only. The same lobby brief in the SongGeneration README's format would be a tag list, descriptions: "cinematic, calm, piano, synthesizer, bass", run with the pure-music mode. It has no field for 120 BPM or A minor.

In a game's audio pipeline

This is stage 2, Write the brief. The base prompt is written once per game and reused for every music track. Lyrics, for songs with vocals, are Lyrics and song structure. In stage 3 you keep the prompt fixed and vary seeds and settings (Seeds and sampling settings), and the caption you used goes into every row of the take log.

Common mistakes

  • Tempo in the caption. You write "120 bpm" in the text and the take comes back at some other tempo. Use the bpm setting.
  • "No vocals" in the prompt. The model sings or hums anyway. Turn on the instrumental setting instead.
  • A new prompt per track. Each track sounds good alone, but the soundtrack sounds like a playlist from five games. Reuse a base prompt.
  • Too many genres. "Jazz metal chiptune orchestral" produces mush. One or two genre words.
  • Naming no instruments. The model chooses its own, and your piano theme arrives as a guitar. Name three or four.
  • Asking for structure the model ignores. "Loopable", "8 bars", and chord names rarely change the result. Cut loops yourself.

Cost

Writing a base prompt takes minutes; testing it takes most of an afternoon, because you judge a prompt by listening to several takes, not one. Each prompt change resets what you know, so change the prompt rarely and settings often. With ACE-Step 1.5 the generation time per take is small on a capable machine; your listening time dominates: a 48 s take costs 48 s per listen, and you will listen to each candidate several times. A shared base prompt cuts later cost, because a new track starts from a known-good palette rather than from zero.

Going further

  • The ACE-Step 1.5 tutorial, its caption and metadata sections.
  • The SongGeneration README's description input format and its list of predefined tags.
  • Lyrics and song structure, for songs with vocals.
  • Seeds and sampling settings, for getting good takes from a good prompt.

Leads to

Back to Local text-to-audio for games