KoboldCpp generates music with ACE-Step 1.5 and ACE-Step XL. You describe a song, and it generates the audio.
Quick start
Section titled “Quick start”Music needs three model files, plus an optional fourth. A template downloads them for you.
- In the launcher, click Get Help.
- Choose Newbie Templates, pick
LowSpec-MusicGenand click Load Template. - Click Launch. KoboldCpp downloads the files.
- Open
http://localhost:5001/musicui.
The model files
Section titled “The model files”Download the files from huggingface.co/koboldcpp/music and load them on the Audio tab:
| File | Launcher field | Flag | Required |
|---|---|---|---|
Music LLM, e.g. acestep-5Hz-lm | MusicLLM: | --musicllm | No. Without it, choose Main LLM as Lyrics Planner in MusicUI to plan the song with the loaded text model. /api/extra/music/prepare needs it. |
Embedding model, e.g. Qwen3-Embedding-0.6B | MusicEmbeds: | --musicembeddings | Yes |
Diffusion model (DiT), e.g. acestep-v15-turbo | MusicDiffuser: | --musicdiffusion | Yes |
| VAE | MusicVAE: | --musicvae | Yes |
The diffusion model does not load without the embedding model and the VAE. KoboldCpp stops with "Invalid config: Music Diffusion requires Music embedding and Music VAE models!".
On the command line:
koboldcpp --musicllm music-lm.gguf --musicembeddings music-embed.gguf --musicdiffusion music-dit.gguf --musicvae music-vae.ggufUsing it
Section titled “Using it”- MusicUI (
/musicui) is the bundled interface for music and speech. It supports reference audio uploads and MP3 output. It also has a TTS tab. - KoboldAI Lite: Add File > Launch MusicUI.
- API: the endpoints below.
| Endpoint | Returns |
|---|---|
POST /api/extra/music/prepare | JSON with the song plan (metadata and lyrics). Add "gen_codes": true to also get audio codes. |
POST /api/extra/music/generate | WAV or MP3 audio |
The API works in two steps:
- Send
caption(a description of the song),lyricsandduration(in seconds) to/api/extra/music/prepare. The plan you get back contains the rewritten caption, the lyrics,bpm,duration,keyscale,timesignature,vocal_language,seedand further generation settings. - Send that plan to
/api/extra/music/generate. It returns the audio.
With --password set, both endpoints need the password.
Advanced options
Section titled “Advanced options”| Launcher field | Flag | What it does |
|---|---|---|
| Music Low VRAM | --musiclowvram | Keeps the music models out of VRAM when idle and swaps them in when needed. |
LowSpec-MusicGen turns Music Low VRAM on. Measured with that template on an RTX 3090 and a 10-second instrumental:
| Music Low VRAM | Peak VRAM | Prepare + generate |
|---|---|---|
| On (template default) | 5644 MiB | 6.1 s + 11.1 s |
| Off | 11330 MiB | 0.8 s + 2.1 s |
With only MusicLLM: set and no other music file, KoboldCpp loads the music LLM alone. It can then only plan songs: /api/extra/music/prepare works, but /api/extra/music/generate returns an empty response instead of audio.