technique
Lyrics and song structure
How to write lyrics and section tags so a model sings a full song with verses and a chorus where you want them.
Before this
This page assumes you are comfortable with:
Why you need this
A full song with vocals, the kind hosted services such as Mureka produce, has a shape: an intro, verses that tell, a chorus that repeats, an ending. A song model does not invent that shape reliably on its own. You give it the shape as section tags inside the lyrics, and the words as lines the model can fit to the beat. This page is stage 2 of the pipeline, writing the brief, for the vocal case. The style prompt that sits beside the lyrics is covered in Prompting for music.
The idea
A song model takes two texts. One describes the sound (genre, mood, instruments, voice). The other is the lyrics, which double as a timeline: each section tag says "a new part starts here", and the lines under it are what gets sung in that part. Musicians call this a song structure: the order of sections such as verse, chorus, and bridge, each a whole number of bars (a bar is one measure, four beats in 4/4).
The model fits syllables to beats. A line is sung over a stretch of bars, so the number of syllables (spoken beats in a word: "shuf-fle" is two) decides whether the line sounds relaxed or rushed. Lines in the same position of each verse should have about the same count, so the melody can repeat.
The two models in this cluster spell all of this differently. Use each one's own syntax exactly.
ACE-Step 1.5
The ACE-Step 1.5 tutorial calls the style text the caption and the timeline the lyrics. In the inference docs these are the caption setting (up to 512 characters) and the lyrics setting (up to 4,096 characters).
| Rule | What the tutorial says |
|---|---|
| Section tags | Square brackets, capitalized: [Intro], [Verse] or [Verse 1], [Pre-Chorus], [Chorus], [Bridge], [Outro] |
| Instrumental parts | [Instrumental], or named ones such as [Guitar Solo], [Piano Interlude]; [Instrumental] alone gives a song with no vocals |
| Performance hints | One modifier after a dash: [Chorus - anthemic], [Bridge - whispered]. Do not stack several; the model may sing the tag text |
| Section breaks | A blank line between sections |
| Line length | 6 to 10 syllables per line usually works best |
| Intensity | UPPERCASE words are sung harder |
| Backing vocals | Words in parentheses become background vocals or harmonies |
| Consistency | Instruments and moods in tags must agree with the caption; conflicts lower quality |
Tempo, key, and language go in their own settings, not in the caption: bpm, keyscale, timesignature, and vocal_language (an ISO 639-1 code such as en). The README lists support for more than 50 languages.
LeVo 2 (SongGeneration 2)
One warning before the syntax: the SongGeneration license allows academic, research, and education use only, and rules out commercial or production use. Learn from its format and compare its singing, but a song for a released game comes from ACE-Step 1.5. Licensing and provenance quotes the clause.
The SongGeneration README calls the lyrics field gt_lyric and the style field descriptions.
| Rule | What the README says |
|---|---|
| Lyric sections | [verse], [chorus], [bridge], lower case; these must contain lyrics |
| Instrumental sections | [intro-short], [intro-medium], [inst-short], [inst-medium], [outro-short], [outro-medium]; these must not contain lyrics. Short is about 0 to 10 s, medium about 10 to 20 s |
| Section breaks | A semicolon ; between every section, all on one line |
| Sentence breaks | A period . between sentences; in English the last sentence of a section ends with a period before the ; |
| Punctuation | English half-width punctuation only |
| Descriptions | Comma-separated tags, not sentences, from up to four dimensions: gender, genre, emotion, instrument |
The README describes the field with capitalized examples ([Verse]) in one place but lists the labels in lower case and uses lower case in every full example; follow the lower-case examples. It also warns that lyric formatting strongly affects quality. The largest published version, SongGeneration-v2-large, lists a maximum length of 4 minutes 30 seconds and lyrics in Chinese, English, Spanish, Japanese, and other languages. By default it sings over accompaniment; the --bgm, --vocal, and --separate options produce music only, vocals only, or the two as separate tracks.
Why sections get repeated or skipped
| What you hear | Likely cause | Fix |
|---|---|---|
| The last chorus is missing or cut off | The requested duration is too short for the tags | Compute the length from bars (below), or let ACE-Step 1.5 choose: a duration of -1 sets it from the lyrics |
| A verse repeats, or a long instrumental fills the end | Duration far longer than the lyrics need | Shorten the duration or add a tagged instrumental section where you want space |
| A line is crammed or stretched | Its syllable count is far from its neighbors' | Even out the counts within a section |
| Verse words leak into the chorus | Blurred boundaries, or a missing tag or separator | One tag per section, one break per section, in the model's syntax |
| Tag words are sung | Too many modifiers in one tag | One modifier at most |
Language and pronunciation. Set the language explicitly when you know it. Invented words, like a game's name, can be sung oddly. If a take mispronounces one, spell it the way it sounds ("Loo-men") in the lyrics and keep the real spelling everywhere else. Vocal style belongs in the style text ("female vocal, breathy") and, for ACE-Step 1.5, optionally in a short tag such as [whispered].
Worked example
Illustration only: these lyrics, prompts, seeds, and settings show what you would write, not a record of real runs.
A short title-screen song for Lumen Clash, in 4/4 at 100 BPM. One bar is 4 beats at 0.6 s per beat, so 2.4 s.
| Section | Bars | Length | Starts at |
|---|---|---|---|
| Intro | 4 | 9.6 s | 0.0 s |
| Verse 1 | 8 | 19.2 s | 9.6 s |
| Chorus | 8 | 19.2 s | 28.8 s |
| Verse 2 | 8 | 19.2 s | 48.0 s |
| Chorus | 8 | 19.2 s | 67.2 s |
| Outro | 4 | 9.6 s | 86.4 s |
| Total | 40 | 96.0 s |
Each 8-bar section gets 4 lines, so each line has 2 bars (8 beats, 4.8 s). Every line below has 7 syllables (counting "every" as two, "ev-ry"), not counting the parenthesized echo, about one per beat, which leaves room to breathe. The 4-bar intro and outro are 9.6 s, inside LeVo 2's "short" range of about 0 to 10 s.
ACE-Step 1.5 version. Settings: caption "upbeat synth-pop, female vocal, bright, warm synth pads, punchy drums, catchy chorus"; bpm 100; keyscale "C Major"; timesignature 4 (the inference docs give 4/4 as 4); vocal_language "en"; duration 96; seed 2207.
[Intro - synth]
[Verse 1]
Shuffle up the morning light
Every card a little spark
Draw the one you need tonight
Hold it up against the dark
[Chorus]
Lumen, Lumen, let it clash (let it clash)
Throw your light against the light
Every turn a brighter flash
Play it to the end of night
[Verse 2]
Silver edges, golden face
Every rule a spoken vow
Find your footing, take your place
Who is shining brighter now
[Chorus - anthemic]
Lumen, Lumen, let it clash (let it clash)
Throw your light against the light
Every turn a BRIGHTER FLASH
Play it to the end of night
[Outro - fade out]
The caption says "synth" and "catchy chorus", and the tags say synth and anthemic, so caption and lyrics agree. The parenthesized echo becomes a backing vocal; the uppercase words lift the last chorus.
LeVo 2 version, for comparison only (its license keeps it out of the shipped game). descriptions: "female, synth-pop, bright, synthesizer, drum machine." The gt_lyric is one line (wrapped here for reading). Commas inside lines are dropped so that the period is the only separator inside a section, the one the README's rules name.
[intro-short] ; [verse] Shuffle up the morning light. Every card a little spark. Draw the one you need tonight. Hold it up against the dark. ; [chorus] Lumen Lumen let it clash. Throw your light against the light. Every turn a brighter flash. Play it to the end of night. ; [verse] Silver edges golden face. Every rule a spoken vow. Find your footing take your place. Who is shining brighter now. ; [chorus] Lumen Lumen let it clash. Throw your light against the light. Every turn a brighter flash. Play it to the end of night. ; [outro-short]
The README's input format has no duration or seed field, so the length follows from the sections, and the take is identified by its idx, the name it gives the output file. Record that name in the take log.
In a game's audio pipeline
This is stage 2, write the brief. The style text comes from Prompting for music. The settings around the lyrics (seed, steps, duration) are stage 3, in Seeds and sampling settings, and fixing one bad line without regenerating the song is Variations and audio-to-audio. A title-screen song usually also needs a clean loop or ending, which is stage 4.
Common mistakes
- Using one model's tags in the other.
[Verse 1]and blank lines in LeVo 2, or semicolons in ACE-Step 1.5. The model sings the separators or ignores the structure. - Lyrics inside an instrumental tag in LeVo 2. The README says those sections must stay empty; you hear words over the intro or a garbled section.
- Lines of 4 and 14 syllables side by side. One line drags, the next is rapped at double speed.
- A sentence as the description for LeVo 2. The README asks for comma-separated tags; sentences degrade the result.
- Duration shorter than the structure. The song stops mid-chorus.
- Conflicting caption and tags. A "piano ballad" caption with a
[Guitar Solo - distorted]tag gives a muddled section.
Cost
Writing and checking a two-verse lyric takes about as long as reading it aloud a few times; count syllables by clapping. Generation cost depends on length: a 96 s song is a fraction of either model's maximum (ACE-Step 1.5 goes to 600 s, LeVo 2 large to 4 minutes 30 seconds), and memory needs are on Choosing a local audio model. The larger cost is listening: each take must be heard all the way through, because a skipped chorus at 67 s is invisible in the first 10 seconds. A 96 s song at 48 kHz, 16-bit stereo is 18.4 MB as WAV (1 MB = 1,000,000 bytes) before compression.
Going further
- Seeds and sampling settings: making a good take repeatable.
- Variations and audio-to-audio: repainting a single section.
- The ACE-Step 1.5 tutorial's lyrics section, for its full lists of vocal and energy tags.
- The SongGeneration README's input guide and its sample input files.
- Try it: write a verse for a game you know, clap the syllables of each line, and even them out.