prerequisite

Model size and memory

How many parameters a model has, how many bytes each one takes, and whether it fits in your machine's memory.

Before this

Nothing beyond first-year college math. This is a starting page.

Why you need this

Before a model can make a single sound, its numbers have to be loaded into memory. If they do not fit, the model either refuses to start, crashes partway through a song, or crawls. A few lines of arithmetic tell you in advance whether a model will run on your machine, which version of it to download, and how much disk space it will take.

The idea

Parameters are the stored numbers. A neural network is a long chain of multiplications and additions. The numbers it multiplies by are fixed after training and are called parameters, or weights. They are what training produced, and they are what you download. Model sizes are given as a count of parameters: "2B" means two billion.

Each parameter takes some bytes. A byte is 8 bits. How many bytes one parameter takes depends on the number format it is stored in:

Format Bits per parameter Bytes per parameter
32-bit float 32 4
16-bit float 16 2
8-bit integer 8 1
4-bit integer 4 0.5

So the memory for the weights alone is

weight bytes=parameters×bytes per parameter\text{weight bytes} = \text{parameters} \times \text{bytes per parameter}

On this page 1 GB means 1,000,000,000 bytes (and 1 MB means 1,000,000). A 1B model at 16-bit is 1,000,000,000×2=2,000,000,0001{,}000{,}000{,}000 \times 2 = 2{,}000{,}000{,}000 bytes, or 2 GB. Most models are trained in 32-bit or 16-bit and published in 16-bit.

Quantization stores them with fewer bits. Quantization rounds each weight to one of a small set of levels so it fits in fewer bits. With 4 bits there are only 24=162^4 = 16 levels. To make those 16 levels useful, weights are handled in small groups: each group stores one scale, and each weight stores a small whole number that is multiplied by the scale to get the weight back.

Take a group of four weights, [0.12, −0.40, 0.33, 0.05][0.12,\ -0.40,\ 0.33,\ 0.05]. A signed 4-bit integer runs from −8-8 to 77. Choose the scale so the largest weight maps to 7: 0.40/7≈0.05710.40 / 7 \approx 0.0571. Divide each weight by the scale and round:

Weight Weight / scale Stored integer Restored (integer x scale) Error
0.12 2.10 2 0.114 -0.006
-0.40 -7.00 -7 -0.400 0.000
0.33 5.77 6 0.343 0.013
0.05 0.88 1 0.057 0.007

Every weight is now slightly off. The scales cost memory too: storing one 16-bit scale for every 32 weights adds half a bit per weight, so a "4-bit" model in that layout really uses about 4.5 bits per parameter. Formats differ; the model's files list their real sizes.

What quantization costs in quality. Billions of small rounding errors add up. At 8-bit the effect is usually hard to hear. At 4-bit it depends on the model and how it was quantized; in audio it can show up as extra noise, duller detail, or more takes that go wrong. A quantized version is a different model for the purposes of your records: note which one you used.

Memory beyond the weights. Running a model needs more than its weights. It also holds the working values of the current step (often called activations), the audio or latent being generated, any cache an autoregressive model keeps, and the other networks in the pipeline, such as a text encoder that reads the prompt and a decoder that turns the result into sound. The extra grows with clip length and batch size. Published memory requirements already include it; when all you have is a parameter count, add a margin. This page uses 20 percent as a planning margin, a rule of thumb rather than any model's figure.

VRAM versus unified memory. A separate graphics card has its own memory, called VRAM, and the model must fit there to run at full speed. An 8 GB card, the size used in the example below, is common. Apple Silicon Macs instead have unified memory: one pool shared by the CPU and the GPU, so a model can use most of it. The hands-on machine for this cluster, an M5 Max with 128 GB unified memory, can hold models that would never fit on a consumer card, though the operating system and your other apps still need their share.

Offloading. When a model does not fit in VRAM, some tools keep part of it in ordinary system memory and move pieces onto the card as each is needed. That works, but every move crosses the slower link between the CPU's memory and the card, often many times per step, so generation can slow down several times over.

Worked example

Two hypothetical models, 2B and 4B parameters, each at 16-bit and 4-bit. The extra column adds the 20 percent planning margin.

Model Format Weights With 20% margin Fits 8 GB VRAM? Fits 128 GB unified?
2B 16-bit 2B x 2 = 4 GB 4.8 GB Yes Yes
2B 4-bit 2B x 0.5 = 1 GB 1.2 GB Yes Yes
4B 16-bit 4B x 2 = 8 GB 9.6 GB No Yes
4B 4-bit 4B x 0.5 = 2 GB 2.4 GB Yes Yes

The 4B model at 16-bit needs 8 GB for weights alone, which already fills an 8 GB card before a single step runs. On that card you would use the 4-bit version (2.4 GB with margin) or offload part of the 16-bit one and accept slow generation. With the group scales described above, the 4-bit files are closer to 4×109×4.5/8=2.254 \times 10^9 \times 4.5 / 8 = 2.25 GB than 2 GB, which still fits easily.

On the M5 Max with 128 GB, every row fits with room to spare. The 4B model at 16-bit takes 9.6 GB, under a tenth of the pool, so you can choose the unquantized version and keep an audio editor and a browser open beside it. For a large music model, that difference is the whole reason to run it at full precision on a Mac rather than a quantized version on a card. Real models' sizes and requirements are on Choosing a local audio model, from their own documentation.

In a game's audio pipeline

This is the first check in setting up the studio: before downloading anything for the Lumen Clash soundtrack, compare the model's published memory needs with your machine. It sets which version you download (full or quantized), how long a clip you can ask for in one go, and how many takes you can batch. In shipping it, the version you used, including its quantization, belongs in the provenance record, because a quantized model can produce different audio from the same seed.

Common mistakes

  • Counting only the weights. The model loads, then fails with an out-of-memory error partway through a long clip. Leave room for working memory, or ask for shorter clips.
  • Confusing bits and bytes. A 4B model at 16-bit is 8 GB, not 64 GB. Divide bits by 8 before multiplying.
  • Mixing GB conventions. A card sold as "8 GB" holds 8×2308 \times 2^{30} bytes, about 8.59 GB in this page's units, and the display takes part of it. Plan with what the system reports as free, not the box.
  • Quantizing and then comparing to old takes. The same seed and prompt on the 4-bit version will not give the 16-bit version's clip. Record the version, or you cannot remake a keeper.
  • Leaving offloading on without noticing. Generation that should take seconds takes minutes. Check whether the tool reports moving weights to system memory.

Cost

Memory needed grows linearly with parameter count and with bits per parameter: double either and the weights double. Disk space for the download is about the size of the weights, so a few 16-bit models of a few billion parameters take tens of GB, and keeping both full and quantized versions adds up. Loading time grows with size too. Quantization saves memory and disk but costs some quality and your time checking that the takes still sound right. Offloading saves memory at a large cost in generation time.

Going further

Leads to

Back to Local text-to-audio for games