technique
Running audio models on Apple Silicon
What unified memory, the Metal backend, and MLX mean for running music models on a Mac, and how to get a first clip out of one.
Before this
This page assumes you are comfortable with:
Why you need this
Most open audio models are written and tested on NVIDIA graphics cards, and their instructions assume one. A Mac can run many of them, and a Mac with a lot of memory can run models that do not fit on most single NVIDIA cards, but only if you know which parts of a project assume NVIDIA and how to work around them. This page gets you from a chosen model to a first clip on the hands-on machine: an M5 Max, 128 GB unified memory.
The idea
Unified memory. On a PC with an NVIDIA card, the card has its own memory, VRAM, separate from the computer's main memory. A model must fit in VRAM to run fast; a 24 GB card cannot hold a model that needs 28 GB without moving parts of it back and forth, which is slow. Apple Silicon has one pool of memory shared by the CPU (the general-purpose processor) and the GPU (the graphics processor that does the heavy math). On a 128 GB machine, a model that needs 28 GB of VRAM on NVIDIA can sit in that pool with room to spare. The catch: macOS, your browser, and your audio editor draw from the same pool, so "fits" means fits alongside everything else you have open.
Two ways to use the Mac's GPU.
| Route | What it is | Who uses it |
|---|---|---|
| PyTorch with the Metal backend (MPS) | PyTorch, the most common machine-learning library, can run on the Mac GPU through Apple's Metal graphics interface. The device is named mps where NVIDIA code says cuda. |
Most research code, after small changes |
| MLX | Apple's own array library, designed for Apple Silicon's unified memory | Ports and conversions made specifically for Macs |
Why some projects need a Mac-specific fork. Code written for NVIDIA often contains CUDA-only paths: CUDA is NVIDIA's programming platform, and some libraries ship only CUDA builds. Typical blockers are a hard-coded cuda device, an attention library such as flash attention that only builds for NVIDIA, and audio-loading libraries with platform-specific builds. A port or fork patches those paths to mps or rewrites the heavy parts in MLX. PyTorch also has a fallback switch, the PYTORCH_ENABLE_MPS_FALLBACK=1 environment variable, which runs any operation the Metal backend lacks on the CPU instead of failing, at a cost in speed.
Pin the Python version. Each model's docs name the Python versions it supports, and its dependencies are pinned to match. ACE-Step 1.5's install guide requires Python 3.11 to 3.12 (a stable release, not a pre-release). The SongGeneration README asks for Python 3.8.12 or newer. Give each model its own virtual environment (a private folder of Python packages for one project) so their pinned dependencies never collide.
Where weights go, and how big they are. Weights download on first run. ACE-Step 1.5 puts them in a checkpoints folder inside the project (about 10 GB for the core models, per its install guide); the ACESTEP_CHECKPOINTS_DIR setting points several installs at one shared folder. Code that downloads through the Hugging Face library uses its cache, by default in a .cache/huggingface/hub folder in your home directory, moved with HF_HOME or HF_HUB_CACHE.
The offline setting. The Hugging Face library's documentation says that if HF_HUB_OFFLINE is set, "no HTTP calls will be made to the Hugging Face Hub", only cached files are used, and if no cached file exists, an error is raised. People set it to stop slow update checks and then forget it is in their shell's startup file. A first run that fails with an offline or missing-file error, before any generation starts, is very often this. Unset it for the download, or download the weights explicitly, then set it again.
Getting one clip from ACE-Step 1.5. ACE-Step 1.5 supports Apple Silicon officially. Its install guide describes, in outline:
- Use Python 3.11 or 3.12 and the package manager the project names to install its dependencies.
- Start the project's macOS launch script for the web interface. The guide says the macOS scripts set the MLX backend for its language-model planner automatically.
- Wait for the first-run download of the core models.
- Open the model's own web interface on port 7860 of your machine, write a caption, set a duration, and generate.
The project also publishes a macOS package with dependencies preinstalled. The guide names MLX for the planner; the README separately lists MPS among the supported devices.
Getting one clip from LeVo 2. Its license limits it to academic, research, and education use, so treat what follows as a way to study the model, not to make shipping game audio (see Choosing a local audio model). The official SongGeneration code asks for CUDA 11.8 or newer. Its README's steps are: download the shared runtime folders and the checkpoint for your variant (for example v2-large), write a JSON Lines input file where each line has an idx (output name), gt_lyric (tagged lyrics), and optional descriptions (style tags), then run its generation script. It offers a --low_mem flag for out-of-memory errors and --not_use_flash_attn for machines without flash attention. On a Mac you need a community port. Two kinds exist: one patches the PyTorch code to Metal; its README asks for at least an M-series Pro chip with 24 GB, recommends 32 GB or more, and lists SongGeneration "Base" and "Large" without saying whether they are the v2 checkpoints. Another converts the v2 language model to MLX: the converted v2-large token generator is 9.5 GB at 16-bit, 5.0 GB at 8-bit, and 2.7 GB at 4-bit, while turning tokens into audio still runs the official PyTorch decoder on Metal. Its card reports a 12 s test clip taking about 1 minute to generate tokens and 73.27 s to decode, on the medium model.
On an NVIDIA machine, by contrast, the official instructions work as written: CUDA, the stated Python version, and enough VRAM.
Worked example
The memory budget for LeVo 2 large on the M5 Max. Sizes on this page use 1 GB = 1,000,000,000 bytes.
The SongGeneration README gives 22 GB without prompt audio and 28 GB with it, measured on NVIDIA cards. Plan for the larger figure:
| Item | Memory |
|---|---|
| Total unified memory | 128 GB |
| LeVo 2 large, with prompt audio | 28 GB |
| Left for everything else | GB |
| ACE-Step 1.5 at its top tier (24 GB or more) open at the same time | 24 GB |
| Left for macOS, a browser, an audio editor | GB |
On paper both music models fit at once with plenty of room. Practice may differ: the Metal-patched community port reports total memory plus swap (disk used as overflow memory) of around 80 GB during generation with its large model, even on 64 GB Macs. If that holds on your machine, GB is left, enough for an editor but not for ACE-Step's top tier at the same time. The two figures disagree because one is NVIDIA VRAM and the other is a port's observation on Macs; trust a measurement on your own machine over either.
To measure: watch the system's memory pressure graph while one generation runs, with nothing else open, and write down the peak. That number, not a README's, is your budget.
In a game's audio pipeline
This is stage 1, Set up the studio: making sure the model you picked in Choosing a local audio model actually runs on your machine. Once one clip comes out, ComfyUI audio workflows turns the run into a saved, repeatable workflow, and stage 3 starts.
Common mistakes
- Following NVIDIA instructions on a Mac. The install fails while building an attention library. Look for the project's Mac instructions, a disable flag, or a port.
- A forgotten offline setting. The first run stops with an offline-mode or file-not-found error before generating anything. Check whether
HF_HUB_OFFLINEis set. - One environment for every model. Installing a second model breaks the first, because each pins different library versions. One virtual environment per model.
- Counting only the model. The model fits, but the machine swaps and stutters because a browser and editor use the same pool. Budget everything that is open.
- Running out of disk, not memory. A download stops halfway with a disk-full error. Weights, caches, and swap all need free space; the Metal port above asks for 70 GB free for its large model because of swap.
Cost
Setup is a one-time cost of an afternoon per model, longer for a model that needs a port. Disk: about 10 GB for ACE-Step 1.5's core models, and tens of GB for LeVo 2 and its runtime files, plus swap headroom. Generation on a Mac is slower than on a recent NVIDIA card: the MLX port's 12 s test clip took over two minutes end to end, and the Metal port's README reports about 12 minutes for a song of about 2 min 30 s with its large model. Memory is where 128 GB pays off: models that need 22 to 28 GB of VRAM, more than most consumer NVIDIA cards have, fit outright.
Going further
- The ACE-Step 1.5 install guide and GPU compatibility guide, macOS sections.
- The Hugging Face library's environment variable reference, for the cache and offline settings.
- PyTorch's notes on the MPS backend, for which operations fall back to the CPU.
- Apple's MLX documentation, for what a conversion to MLX involves.