technique

Editing and layering sound effects

How to trim, fade, and combine generated sounds into crisp game effects.

Before this

This page assumes you are comfortable with:

Why you need this

A sound-effect model gives you a clip, not an effect. The clip usually starts late, ends with a long hiss, and has either too much body or too little snap. A card-play sound has to land the instant the card does, sound the same size every time without sounding identical, and stay out of the music's way. This page is stage 4 of the pipeline: the edits that turn generated clips into effects a player feels as responsive.

The idea

Five edits, in this order.

1. Trim the start. Generated clips often open with silence or room noise before the sound begins. The point where the sound begins is its onset (or attack). Cut so the file starts a millisecond or two before the onset. A game plays a sound the moment the player acts; every millisecond of silence at the start is delay the player feels as lag.

2. Trim the end. Cut where the sound has decayed into the noise floor (the faint hiss that is always there). Long tails waste file size and smear into the next sound.

3. Fade both edges. Audio is stored as samples, measurements of the wave at a fixed rate, here 48 kHz, meaning 48,000 samples per second. Cutting a wave anywhere but at zero leaves a sudden jump at the file's edge, which plays as a click. A fade-in of about 1 ms and a fade-out of tens of milliseconds ramp the level from and to zero. The fade-in must be short, or it softens the attack you just trimmed toward.

4. Layer. Many good effects are two or three sounds played together, each doing one job. A card play might be a soft whoosh (the motion) followed by a slap (the contact). Layers are lined up by their onsets, not by their file starts, and each gets its own level in dB, a change in level, where -6 dB is about half the amplitude.

5. Shape the tone in words, then by ear. EQ (equalization) turns frequency ranges up or down. For small game sounds, two moves cover most cases:

  • Cut the mud. Low frequencies (the rumble and boom) take up space the music needs and vanish on phone speakers anyway. Turn them down, or cut them below the point where the sound stops needing them.
  • Keep the snap. The bright, high detail of a click or a slap is what makes it feel crisp. Do not dull it; if anything, a small lift there helps the sound read over music.

Finally, vary at playback. The same sample played twenty times in a row sounds like a machine gun. The game can change each playback slightly: a small random pitch change, a small random level change, or a pick from two or three variants. Changing pitch by ss semitones (a semitone is one step on a piano keyboard) multiplies the playback rate by 2s/122^{s/12}, which also changes the length.

Worked example

The clips, seeds, and levels here are illustrations, not records of real runs. The Lumen Clash card-play sound is built from two generated layers at 48 kHz.

The raw layers.

Layer Seed Raw length Silence before onset Peak
Whoosh (paper moving through air) 2201 600 ms 60 ms -4 dBFS
Slap (card landing on felt) 2207 900 ms 140 ms -3 dBFS

dBFS is the level of a sample relative to the loudest a file can hold, 0 dBFS.

Trim and fade. The slap's 140 ms of lead silence is 0.140×48,000=6,7200.140 \times 48{,}000 = 6{,}720 samples. Left in, it is 8.4 frames of delay at 60 frames per second: a player sees the card land and hears it a beat late. Trim to 2 ms (96 samples) before each onset. Fade in over 1 ms (48 samples). After trimming, the whoosh is 540 ms and the slap 760 ms.

Line up the onsets. The whoosh should lead, as the card moves before it lands. Place the whoosh onset at 0 ms and the slap onset at 80 ms (3,840 samples).

Layer Starts at Ends at (trimmed) Starts at, in samples
Whoosh 0 ms 540 ms 0
Slap 80 ms 840 ms 3,840

Set levels. The slap is the point of the sound, so it stays at 0 dB. The whoosh is support, so it drops by 8 dB: a factor of 10−8/20=0.39810^{-8/20} = 0.398, taking its peak from -4 dBFS to -12 dBFS.

Check the combined peak. In the worst case, both peaks land at the same moment with the same sign, and their amplitudes add. The slap's -3 dBFS is 0.708 and the whoosh's -12 dBFS is 0.251:

0.708+0.251=0.959⇒20log⁡100.959=−0.36 dBFS0.708 + 0.251 = 0.959 \quad\Rightarrow\quad 20 \log_{10} 0.959 = -0.36 \text{ dBFS}

That is under 0 dBFS but leaves almost no room. Turning the whole effect down 2 dB brings the worst case to -2.36 dBFS. Loudness for games sets the final level; this step only makes sure the layers do not clip each other.

End it. The slap's tail is decayed into the noise floor by 500 ms, so the mix ends at 500 ms with a 30 ms fade-out (1,440 samples) from 470 ms. The finished card-play sound is 500 ms, mono.

Vary it. At playback, pick a pitch shift between -1 and +1 semitone from a seeded generator. That is a rate between 2−1/12=0.9442^{-1/12} = 0.944 and 21/12=1.0592^{1/12} = 1.059, so the 500 ms sound plays between about 472 ms and 530 ms.

function mulberry32(a) {
  return () => {
    a |= 0; a = (a + 0x6d2b79f5) | 0;
    let t = Math.imul(a ^ (a >>> 15), 1 | a);
    t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
    return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
  };
}
const rand = mulberry32(7);
for (let i = 0; i < 5; i++) {
  const semis = rand() * 2 - 1;                 // -1 to +1 semitone
  const rate = 2 ** (semis / 12);               // use as source.playbackRate.value
  console.log(semis.toFixed(2), rate.toFixed(3), (500 / rate).toFixed(0) + " ms");
}

In a game's audio pipeline

Editing follows Prompting sound effects and curating takes: only keepers are edited. Do every edit on the lossless master and record it (trim points, fades, layer offsets, gains) in the file's provenance record from Licensing and provenance, so the effect can be rebuilt if a layer is regenerated. Loudness for games comes next, then Exporting game audio, where most effects ship as mono.

Common mistakes

  • Lead silence left in. Symptom: buttons and cards feel laggy even though the frame rate is fine.
  • No fade on a cut edge. Symptom: a tick at the start or end of the sound, most audible on quiet, low sounds.
  • Fade-in too long. Symptom: the click loses its click; everything sounds soft and distant.
  • Layers aligned by file start, not onset. Symptom: a flam, two hits close together instead of one.
  • Too much low end. Symptom: the effect sounds big alone but muddies the music in game, and disappears on a phone.
  • Identical repeats. Symptom: dealing five cards sounds like one sample stuttering.

Cost

No model time beyond the layers themselves, which are short and quick to generate; a batch of several takes per layer is cheap. The cost is the maker's editing time: trimming and fading a one-shot is quick once you know where the onset is, while building a layered effect and checking it against music and on a phone speaker takes longer, per effect. File size drops with every edit. With 1 MB = 1,000,000 bytes, the finished 500 ms mono effect at 16-bit is 0.5×48,000×2=48,0000.5 \times 48{,}000 \times 2 = 48{,}000 bytes (0.048 MB) uncompressed, smaller than the raw 900 ms slap alone even in mono (86,400 bytes), and half what it would be in stereo. Pitch variation at playback costs nothing in file size, which makes it cheaper than shipping extra variants.

Going further

  • Transient shaping: adjusting the attack of a sound separately from its tail.
  • Sound families: generating several takes of one prompt to ship as round-robin variants.
  • Prompting sound effects, for asking for dry, short sounds that need less editing.
  • Ducking: briefly lowering music when an important effect plays, as an alternative to making the effect louder.

Back to Local text-to-audio for games