KoboldCpp does not include a model. Most GGUF models are on Hugging Face. Download one .gguf file in the quant you want, not the whole repository.
The quickest start is a newbie template: in the launcher, click Get Help and pick one under Newbie Templates. It downloads a suitable model for you. See Choose a model.
Recommended models
Section titled “Recommended models”| Model | Good for | Download | File size |
|---|---|---|---|
| Qwen3-VL-8B | The recommended all-rounder | Q4_K_S | 4.8 GB |
| Gemma3-4B | Lightweight and fast | Q4_K_M | 2.5 GB |
| L3-8B-Stheno-v3.2 | Creative writing and roleplay | Q4_K_S | 4.7 GB |
For image input, also download the model's vision projector (mmproj): for Qwen3-VL-8B or for Gemma3-4B. See Vision.
Finding more models
Section titled “Finding more models”- Search for "GGUF" on Hugging Face.
- Bartowski's Hugging Face page has GGUF files of many popular models.
- The official koboldcpp organization has vision projectors (
mmproj), image, Whisper, TTS and music models, and launcher templates.
A model repository usually lists one file per quant. Pick one with GGUF and quantization, and check that it fits with How much memory a model needs.
Download from the launcher
Section titled “Download from the launcher”HF Search next to GGUF Text Model: on the Quick Launch tab (and next to Text Model: on Loaded Files) searches Hugging Face from inside the launcher.
- Click HF Search. The Model File Browser opens.
- Type a model name and click Search Huggingface.
- Pick a repository, then a
.gguffile. Sizes are shown in GiB. A Q4 quant is preselected when the repository has one. - Click Confirm Selection. The file's address appears in the model field.
- Click Launch. KoboldCpp downloads the file first, then loads it.
You can also paste a Hugging Face download link into any model field, or pass it on the command line (koboldcpp stands for your KoboldCpp file; see Command line):
koboldcpp --model https://huggingface.co/ggml-org/gemma-3-4b-it-GGUF/resolve/main/gemma-3-4b-it-Q4_K_M.gguf- Links containing
/blob/main/are changed to/resolve/main/for you. - Files are saved to the folder set in Download Dir: on the Loaded Files tab (
--downloaddir). Without it, they go to the current working folder; on Windows, to the exe's folder if the working folder isSystem32orSysWOW64. - A file that is already there is reused, not downloaded again.
- KoboldCpp downloads with
aria2c(bundled on Windows),curlorwget. If none is available, it prints "Please install aria2, curl, or wget."
Split GGUF files
Section titled “Split GGUF files”Large models are often split into several files named like model-00001-of-00003.gguf.
- Local files: keep all parts in the same folder and select the first part (
-00001-of-…). KoboldCpp loads the others. Selecting another part fails with "model must be loaded with the first split". - Download links: give the link to the first part. For the text model, KoboldCpp downloads all parts. HF Search lists only the first part of each split model.
- The automatic GPU layer estimate is less accurate for split files. KoboldCpp prints "Multi-Part GGUF detected. Layer estimates may not be very accurate". See GPU layers.