technique
Prompting sound effects
How to describe one-shot sounds, UI clicks, impacts, and ambience so a model produces something short, clean, and usable.
Before this
This page assumes you are comfortable with:
- prerequisiteSound wavesWhat sound physically is: pressure waves with a frequency you hear as pitch and an amplitude you hear as loudness.
- techniqueChoosing a local audio modelWhich open model to run for which job: full songs with vocals, instrumental game music, or sound effects, compared by quality, speed, memory, and license.
Why you need this
A card game needs dozens of tiny sounds, and each one plays hundreds of times per session. A sound-effect model can draft them in seconds, but only if the description says what the sound is, how long it lasts, and what it should leave out. This page is stage 2 of the pipeline, writing the brief, for sounds rather than music. Which model to run is covered in Choosing a local audio model.
The idea
A sound effect has a physical cause. Something (the source) does something (the action) to something made of a material, in a space, for a length of time. A model trained on labeled recordings has heard thousands of captions written that way, so the closer your prompt reads like a caption of a real recording, the better it lands.
Write the prompt from five slots:
| Slot | Question | Example words |
|---|---|---|
| Source | What makes the sound? | a playing card, a wooden button, a coin |
| Action | What happens to it? | slides, snaps down, taps, flips |
| Material | What surface or texture? | on felt, on a wooden table, stiff paper |
| Space | Where is it heard? | close-up, dry, no room sound |
| Length | How long, and how does it end? | short, single hit, quick decay |
The published model cards write their examples the same way. The TangoFlux README's example prompt is "Hammer slowly hitting the wooden table": source, action, material in six words. The Stable Audio Open 1.0 model card adds two cautions worth knowing: the model was trained on English descriptions only, and it is better at sound effects than at music.
Ask for a dry sound. Reverb is the room's echo after the sound itself. A recording made in a hall carries that hall forever. In a game you add the space yourself, so the same "card play" works in a quiet menu and in a crowded arena. Words like "dry", "close-up", and "no reverb" push the model toward a clean source. Avoid naming a big space ("in a cathedral") unless you want it baked in.
Ask for one event. Models like to fill time. A prompt without a count can return three taps in a row, or a tap followed by a rattle. "A single tap" and "one card" set the count.
Why short sounds come back long
These models generate a clip of a length you set, and most have a floor on that length. The TangoFlux README says you must pass a duration, between 1 and 30 seconds. Stable Audio Open 1.0 generates up to 47 seconds, and the smaller Stable Audio Open Small up to 11 seconds; both set the length with a seconds_total setting. ACE-Step 1.5 is a music model whose duration setting starts at 10 seconds, so it is a poor fit for a 60 ms click.
A UI click lasts a few tens of milliseconds. If you ask for 1 s, the model must put something in the remaining 940 ms. Usually that is silence or a decay, but sometimes it is a second hit, a hum, or noise. Plan to ask for the shortest duration the model allows and trim the rest in Editing and layering sound effects. Describing the ending ("quick decay, then silence") makes the trim easier.
One-shots and ambience beds
A one-shot is a single event with a start and an end: a click, a hit, a card snap. An ambience bed is a continuous background texture that loops under a scene: a quiet room tone, wind, crowd murmur. They need opposite prompts.
| One-shot | Ambience bed | |
|---|---|---|
| Length asked | Shortest allowed | Longest useful (the loop length) |
| Key words | single, short, sharp attack, quick decay | steady, continuous, no sudden events, even level |
| Space | dry | the space is the point; name it |
| Failure | extra hits, long tail | a loud event in the middle that repeats every loop |
Families of related sounds
A game's UI sounds should feel like one set. Write a base prompt shared by the family, then change one slot per sound. Keep the space slot and the material slot fixed across the family, and the sounds will share a texture. Generate each member several times; keep the take log described in Curating takes.
Weak versus strong
| Weak prompt | Strong prompt | Why the strong one works |
|---|---|---|
| card sound | a single playing card sliding off a deck onto felt, close-up, dry, short | names source, action, material, space, and length |
| click | a single small plastic button click, close-up, dry, sharp attack, quick decay | one event, a material, a shape for the ending |
| epic magical card play with sparkles and whoosh | a stiff card snapped down onto a wooden table, close-up, dry, single hit | one physical event; add the sparkle as a separate layer later |
| tavern ambience | quiet room tone with distant indistinct murmur, steady, continuous, no sudden sounds | asks for evenness, which a loop needs |
| UI sound, not too long, no echo please | short dry UI tap | models read keywords, not requests; "no echo please" can be read as "echo" |
The last row matters. To a text encoder, a word mentioned is a word heard. Describe what you want ("dry") rather than what you do not want ("no echo"), unless the model's docs list a separate negative-prompt setting for the words to avoid.
Worked example
Illustration only: these prompts, seeds, and settings are examples of what you would type, not records of real runs. The model here is TangoFlux, which accepts durations from 1 to 30 s and outputs 44.1 kHz stereo (44,100 samples per second per channel). The README reports that 50 steps give its best results and 25 steps give similar quality faster. One warning: the TangoFlux model card says its checkpoints "are for non-commercial research use only", so treat this as practice. For sounds that ship in the game, send the same prompts to a model whose license allows it, such as MOSS-SoundEffect v2.0 (Apache 2.0); Choosing a local audio model compares them. Its length limits and sample rate differ, so recheck the sample counts below against its model card.
The Lumen Clash card and UI family shares one base: "close-up, dry, single, quick decay".
| Sound | Target length | Prompt (base appended) | Duration asked | Seed |
|---|---|---|---|---|
| Card draw | 250 ms | a single playing card sliding off a deck, stiff paper on felt | 1 s | 3101 |
| Card play | 400 ms | a single stiff card snapped down onto a wooden table | 1 s | 3102 |
| Button click | 60 ms | a single small plastic button click, sharp attack | 1 s | 3103 |
How much of each clip you keep, at 44,100 samples per second:
| Sound | Kept samples | Share of the 1 s clip (44,100 samples) |
|---|---|---|
| Card draw, 250 ms | 11,025 | 25% |
| Card play, 400 ms | 17,640 | 40% |
| Button click, 60 ms | 2,646 | 6% |
The click is the risky one. 94% of the generated clip is something you will throw away, and that is where stray second clicks appear. Two habits help. Generate several seeds per sound (say 3101 to 3108, eight takes), because you are choosing the best first 60 ms, not the best second. And listen with a loop playback mode, so you hear the start of the clip over and over the way a player will.
If the card play keeps coming back with a ring after the snap, change one slot at a time: the material first ("onto felt" instead of "onto a wooden table"), then the length words. Changing several slots at once leaves you unable to tell which one fixed it.
In a game's audio pipeline
This is stage 2, write the brief, for the sound-effect half of the set. The prompts and seeds you choose here go into the take log in stage 3 (Curating takes). Stage 4 trims the padding, fades the edges, and layers the click under a whoosh (Editing and layering sound effects), then brings every effect to a consistent level (Loudness for games).
Common mistakes
- Writing a mood, not an event. "Satisfying card sound" gives the model nothing physical. You hear a vague swish with no clear attack.
- Asking for wet sounds. A prompt that names a room or says "epic" brings reverb. In the game the sound smears into the music and cannot be made dry again.
- No count. The model returns two or three hits. In the game the button "double-clicks".
- Stacking effects into one prompt. "Card snap with sparkle and whoosh" returns a blurred mix. Generate each layer separately and combine them later.
- Using a music model for a click. A 10-second minimum and a musical training set give you a short melody or a drum fill instead of a click.
- An event inside an ambience loop. A single cough in a room-tone bed becomes a cough every 20 seconds.
Cost
Generation is quick: the TangoFlux README reports about 3 seconds for a 30-second clip on a single A40 data-center GPU, and a 1-second request is less work. Memory is modest next to the music models, since TangoFlux's base model has 515M parameters (515 million). The real cost is listening. Eight takes of three sounds is 24 clips; each one needs a few loop-mode listens of its first second, plus a trim. Generated files are small before trimming: one second of 44.1 kHz 16-bit stereo is 176,400 bytes, about 0.18 MB (1 MB = 1,000,000 bytes). The shipped click, trimmed to 60 ms and converted to mono, is 5,292 bytes. Check each model's license on Licensing and provenance before shipping: the Woosh weights from Sony AI and the TangoFlux checkpoints, for example, are both published for non-commercial use only.
Going further
- Editing and layering sound effects: trimming, fades, and stacking two generated layers.
- Curating takes: the take log that keeps the seed of the click you chose.
- The Stable Audio Open 1.0 and Stable Audio Open Small model cards, for their stated limits.
- The Woosh technical report from Sony AI, which compares open sound-effect models.
- Try it: write a five-slot prompt for a sound in a game you know, and list which slot you would vary to make a family.