prerequisite

Transformers and tokens

How a language model generates one token at a time, and how the same idea generates audio once sound is turned into tokens.

Before this

Nothing beyond first-year college math. This is a starting page.

Why you need this

Several open music models write a song the way a chatbot writes a reply: one small piece at a time, each chosen from a list of likely next pieces. Knowing how that loop works explains what a "temperature" setting does, why long songs take longer than short ones in a fairly predictable way, and why a model can lose track of a chorus it sang two minutes ago.

The idea

Tokens. A model cannot work with raw text or raw sound directly. It works with tokens: small units drawn from a fixed list called the vocabulary. For text, a token is often a word or a piece of a word ("card", "play", "ing"). Each token in the vocabulary has a number, its ID, and the model only ever sees and produces those IDs. A typical text vocabulary has tens of thousands of entries.

Next-token prediction. A language model is a function that reads a sequence of tokens and outputs, for every token in the vocabulary, a score for how well it would fit next. The scores are called logits. They can be any real number; higher means more likely. To turn them into probabilities, the model applies the softmax function: raise ee (about 2.718) to the power of each score, then divide each result by the total so they add up to 1. With scores z1,…,zVz_1, \dots, z_V for a vocabulary of VV tokens,

pi=ezi∑j=1Vezjp_i = \frac{e^{z_i}}{\sum_{j=1}^{V} e^{z_j}}

With two tokens scored 1 and 0, e1≈2.718e^1 \approx 2.718 and e0=1e^0 = 1, so the probabilities are 2.718/3.718≈0.732.718 / 3.718 \approx 0.73 and 1/3.718≈0.271 / 3.718 \approx 0.27.

Sampling and temperature. The model then picks one token at random, weighted by those probabilities. Always picking the single most likely token gives safe, repetitive output; sampling gives variety. Temperature, written TT, controls how much variety. Each score is divided by TT before the softmax:

pi=ezi/T∑j=1Vezj/Tp_i = \frac{e^{z_i / T}}{\sum_{j=1}^{V} e^{z_j / T}}

A temperature below 1 stretches the gaps between scores, so the favorite gets even more of the probability. A temperature above 1 shrinks the gaps, so unlikely tokens get picked more often. The random pick comes from a seeded generator, so the same seed and settings repeat the same choices.

Context window. The model can only read a limited number of recent tokens at once, called its context window. Anything earlier is simply not visible when it chooses the next token.

Attention. The network inside most of these models is a transformer. Its key step, attention, lets each position in the sequence look back at every earlier position and decide how much each one matters right now. When a model is about to sing the second chorus, attention is what lets it find the first chorus and reuse its melody. Attention compares every position with every other, so its work grows with the square of the sequence length.

Autoregressive generation. The model generates by looping: predict, pick one token, append it to the sequence, and predict again with the longer sequence. This is called autoregressive generation, meaning each output depends on the outputs before it. The loop cannot skip ahead, because step 500 needs the token from step 499. So the time grows with the number of tokens: twice the length takes at least twice as many steps.

The same loop for audio. Sound can be turned into tokens too. A separate network, an audio codec, cuts sound into short frames and gives each frame one or more token IDs from its own vocabulary; Neural audio codecs explains how. A music model then predicts audio tokens instead of word pieces, often after reading your prompt and lyrics as text tokens first. The codec's decoder turns the finished token sequence back into samples. Some models are not autoregressive at all; Diffusion models covers the main alternative.

Worked example

Suppose a tiny model has a 4-token vocabulary and has just read "The card". Its scores for the next token are:

Token Score zz
glows 2.0
burns 1.0
falls 0.5
sings -0.5

Temperature 1. Exponentiate each score: e2.0≈7.389e^{2.0} \approx 7.389, e1.0≈2.718e^{1.0} \approx 2.718, e0.5≈1.649e^{0.5} \approx 1.649, e−0.5≈0.607e^{-0.5} \approx 0.607. The total is about 12.363. Divide each by the total.

Temperature 0.5. Divide each score by 0.5 first, which doubles them: 4.0, 2.0, 1.0, -1.0. Exponentiate: 54.598, 7.389, 2.718, 0.368. Total about 65.073.

Temperature 1.5. Divide each score by 1.5: about 1.333, 0.667, 0.333, -0.333. Exponentiate: 3.794, 1.948, 1.396, 0.717. Total about 7.854.

Token T=0.5T = 0.5 T=1T = 1 T=1.5T = 1.5
glows 0.839 0.598 0.483
burns 0.114 0.220 0.248
falls 0.042 0.133 0.178
sings 0.006 0.049 0.091

Each column adds to 1, give or take rounding in the last digit. At temperature 0.5, "glows" wins about 84 times in 100 and "sings" almost never appears. At 1.5, "glows" wins less than half the time and "sings" turns up about 9 times in 100. The scores did not change; only how boldly the model samples from them.

The same calculation in a browser console:

const scores = [2.0, 1.0, 0.5, -0.5]; // glows, burns, falls, sings
function softmax(z, T) {
  const e = z.map((v) => Math.exp(v / T));
  const total = e.reduce((a, b) => a + b, 0);
  return e.map((v) => +(v / total).toFixed(3));
}
console.log(softmax(scores, 0.5)); // [0.839, 0.114, 0.042, 0.006]
console.log(softmax(scores, 1.5)); // [0.483, 0.248, 0.178, 0.091]

The same arithmetic runs in a music model, over a vocabulary of audio tokens instead of four words. A low temperature tends toward safe, repetitive phrases; a high one toward surprising turns and, past some point, wrong notes and garbled sound.

Length. As an illustration only (not any particular model's figure), suppose a codec makes 50 tokens per second of audio. A 3-minute song is then 180×50=9,000180 \times 50 = 9{,}000 tokens, so 9,000 trips around the loop, each one reading a longer sequence than the last.

In a game's audio pipeline

In setting up the studio, knowing whether a model is autoregressive tells you how its generation time grows with song length. In generating and choosing, temperature and related sampling settings are among the controls on Seeds and sampling settings; a seed fixes the random picks, so a good Lumen Clash take can be remade. In writing the brief, the context window and attention explain why a long song with lyrics can drift: the model may no longer see the opening when it writes the ending.

Common mistakes

  • Raising temperature to "make it more creative" without limit. Past a point the unlikely tokens are unlikely for a reason. The symptom in music is off-key notes, garbled vocals, or a sudden change of style mid-take.
  • Setting temperature very low for safety. The model falls into loops. You hear the same short phrase or drum fill repeating long after it should have moved on.
  • Changing temperature and seed together. Two changes at once means you cannot tell which one made the take better or worse.
  • Expecting a model to remember beyond its context window. A long generation can forget the key or the chorus melody from its start. The symptom is a song whose ending sounds like a different song.
  • Assuming every audio model works this way. Some generate the whole clip at once by removing noise. Their settings have different names, and temperature may not exist at all.

Cost

For autoregressive generation of LL tokens, the model runs LL steps one after another. With attention, step tt looks back over tt earlier tokens, so the total attention work grows roughly with L2L^2; implementations store the earlier steps' results in memory (often called a cache) to avoid recomputing them, and that memory grows with LL. In practice this means generation time for a song grows a bit faster than its length, and very long generations need noticeably more memory. Listening time grows with length too, and higher temperatures mean more takes to throw away.

Going further

  • Neural audio codecs: how sound becomes tokens in the first place.
  • Diffusion models: the other main way to generate audio.
  • Seeds and sampling settings: the settings that control the loop in practice.
  • The 2017 paper "Attention Is All You Need", by name, which introduced the transformer.
  • Try it: run the console sample above with temperatures of your own, such as 0.1 and 5, and watch the probabilities collapse onto one token or flatten toward an even split.

Leads to

Back to Local text-to-audio for games