technique

Prompting sound effects

How to describe one-shot sounds, UI clicks, impacts, and ambience so a model produces something short, clean, and usable.

Before this

This page assumes you are comfortable with:

Why you need this

A card game needs dozens of tiny sounds, and each one plays hundreds of times per session. A sound-effect model can draft them in seconds, but only if the description says what the sound is, how long it lasts, and what it should leave out. This page is stage 2 of the pipeline, writing the brief, for sounds rather than music. Which model to run is covered in Choosing a local audio model.

The idea

A sound effect has a physical cause. Something (the source) does something (the action) to something made of a material, in a space, for a length of time. A model trained on labeled recordings has heard thousands of captions written that way, so the closer your prompt reads like a caption of a real recording, the better it lands.

Write the prompt from five slots:

Slot Question Example words
Source What makes the sound? a playing card, a wooden button, a coin
Action What happens to it? slides, snaps down, taps, flips
Material What surface or texture? on felt, on a wooden table, stiff paper
Space Where is it heard? close-up, dry, no room sound
Length How long, and how does it end? short, single hit, quick decay

The published model cards write their examples the same way. The TangoFlux README's example prompt is "Hammer slowly hitting the wooden table": source, action, material in six words. The Stable Audio Open 1.0 model card adds two cautions worth knowing: the model was trained on English descriptions only, and it is better at sound effects than at music.

Ask for a dry sound. Reverb is the room's echo after the sound itself. A recording made in a hall carries that hall forever. In a game you add the space yourself, so the same "card play" works in a quiet menu and in a crowded arena. Words like "dry", "close-up", and "no reverb" push the model toward a clean source. Avoid naming a big space ("in a cathedral") unless you want it baked in.

Ask for one event. Models like to fill time. A prompt without a count can return three taps in a row, or a tap followed by a rattle. "A single tap" and "one card" set the count.

Why short sounds come back long

These models generate a clip of a length you set, and most have a floor on that length. The TangoFlux README says you must pass a duration, between 1 and 30 seconds. Stable Audio Open 1.0 generates up to 47 seconds, and the smaller Stable Audio Open Small up to 11 seconds; both set the length with a seconds_total setting. ACE-Step 1.5 is a music model whose duration setting starts at 10 seconds, so it is a poor fit for a 60 ms click.

A UI click lasts a few tens of milliseconds. If you ask for 1 s, the model must put something in the remaining 940 ms. Usually that is silence or a decay, but sometimes it is a second hit, a hum, or noise. Plan to ask for the shortest duration the model allows and trim the rest in Editing and layering sound effects. Describing the ending ("quick decay, then silence") makes the trim easier.

One-shots and ambience beds

A one-shot is a single event with a start and an end: a click, a hit, a card snap. An ambience bed is a continuous background texture that loops under a scene: a quiet room tone, wind, crowd murmur. They need opposite prompts.

One-shot Ambience bed
Length asked Shortest allowed Longest useful (the loop length)
Key words single, short, sharp attack, quick decay steady, continuous, no sudden events, even level
Space dry the space is the point; name it
Failure extra hits, long tail a loud event in the middle that repeats every loop

A game's UI sounds should feel like one set. Write a base prompt shared by the family, then change one slot per sound. Keep the space slot and the material slot fixed across the family, and the sounds will share a texture. Generate each member several times; keep the take log described in Curating takes.

Weak versus strong

Weak prompt Strong prompt Why the strong one works
card sound a single playing card sliding off a deck onto felt, close-up, dry, short names source, action, material, space, and length
click a single small plastic button click, close-up, dry, sharp attack, quick decay one event, a material, a shape for the ending
epic magical card play with sparkles and whoosh a stiff card snapped down onto a wooden table, close-up, dry, single hit one physical event; add the sparkle as a separate layer later
tavern ambience quiet room tone with distant indistinct murmur, steady, continuous, no sudden sounds asks for evenness, which a loop needs
UI sound, not too long, no echo please short dry UI tap models read keywords, not requests; "no echo please" can be read as "echo"

The last row matters. To a text encoder, a word mentioned is a word heard. Describe what you want ("dry") rather than what you do not want ("no echo"), unless the model's docs list a separate negative-prompt setting for the words to avoid.

Worked example

Illustration only: these prompts, seeds, and settings are examples of what you would type, not records of real runs. The model here is TangoFlux, which accepts durations from 1 to 30 s and outputs 44.1 kHz stereo (44,100 samples per second per channel). The README reports that 50 steps give its best results and 25 steps give similar quality faster. One warning: the TangoFlux model card says its checkpoints "are for non-commercial research use only", so treat this as practice. For sounds that ship in the game, send the same prompts to a model whose license allows it, such as MOSS-SoundEffect v2.0 (Apache 2.0); Choosing a local audio model compares them. Its length limits and sample rate differ, so recheck the sample counts below against its model card.

The Lumen Clash card and UI family shares one base: "close-up, dry, single, quick decay".

Sound Target length Prompt (base appended) Duration asked Seed
Card draw 250 ms a single playing card sliding off a deck, stiff paper on felt 1 s 3101
Card play 400 ms a single stiff card snapped down onto a wooden table 1 s 3102
Button click 60 ms a single small plastic button click, sharp attack 1 s 3103

How much of each clip you keep, at 44,100 samples per second:

Sound Kept samples Share of the 1 s clip (44,100 samples)
Card draw, 250 ms 11,025 25%
Card play, 400 ms 17,640 40%
Button click, 60 ms 2,646 6%

The click is the risky one. 94% of the generated clip is something you will throw away, and that is where stray second clicks appear. Two habits help. Generate several seeds per sound (say 3101 to 3108, eight takes), because you are choosing the best first 60 ms, not the best second. And listen with a loop playback mode, so you hear the start of the clip over and over the way a player will.

If the card play keeps coming back with a ring after the snap, change one slot at a time: the material first ("onto felt" instead of "onto a wooden table"), then the length words. Changing several slots at once leaves you unable to tell which one fixed it.

In a game's audio pipeline

This is stage 2, write the brief, for the sound-effect half of the set. The prompts and seeds you choose here go into the take log in stage 3 (Curating takes). Stage 4 trims the padding, fades the edges, and layers the click under a whoosh (Editing and layering sound effects), then brings every effect to a consistent level (Loudness for games).

Common mistakes

  • Writing a mood, not an event. "Satisfying card sound" gives the model nothing physical. You hear a vague swish with no clear attack.
  • Asking for wet sounds. A prompt that names a room or says "epic" brings reverb. In the game the sound smears into the music and cannot be made dry again.
  • No count. The model returns two or three hits. In the game the button "double-clicks".
  • Stacking effects into one prompt. "Card snap with sparkle and whoosh" returns a blurred mix. Generate each layer separately and combine them later.
  • Using a music model for a click. A 10-second minimum and a musical training set give you a short melody or a drum fill instead of a click.
  • An event inside an ambience loop. A single cough in a room-tone bed becomes a cough every 20 seconds.

Cost

Generation is quick: the TangoFlux README reports about 3 seconds for a 30-second clip on a single A40 data-center GPU, and a 1-second request is less work. Memory is modest next to the music models, since TangoFlux's base model has 515M parameters (515 million). The real cost is listening. Eight takes of three sounds is 24 clips; each one needs a few loop-mode listens of its first second, plus a trim. Generated files are small before trimming: one second of 44.1 kHz 16-bit stereo is 176,400 bytes, about 0.18 MB (1 MB = 1,000,000 bytes). The shipped click, trimmed to 60 ms and converted to mono, is 5,292 bytes. Check each model's license on Licensing and provenance before shipping: the Woosh weights from Sony AI and the TangoFlux checkpoints, for example, are both published for non-commercial use only.

Going further

  • Editing and layering sound effects: trimming, fades, and stacking two generated layers.
  • Curating takes: the take log that keeps the seed of the click you chose.
  • The Stable Audio Open 1.0 and Stable Audio Open Small model cards, for their stated limits.
  • The Woosh technical report from Sony AI, which compares open sound-effect models.
  • Try it: write a five-slot prompt for a sound in a game you know, and list which slot you would vary to make a family.

Leads to

Back to Local text-to-audio for games