KoboldCpp runs models in the GGUF format (files ending in .gguf). A GGUF file contains the model's weights and its metadata in one file. Most GGUF models work in KoboldCpp. A model with a new architecture needs a KoboldCpp version that supports it.
Most models come in several versions of different sizes, called quants. The quant is part of the file name, for example Q4_K_M in gemma-3-4b-it-Q4_K_M.gguf.
Choosing a quant
Section titled “Choosing a quant”- Start with Q4_K_M or Q4_K_S. The beginner models recommended by KoboldCpp use these.
- Within the K-quants (
Q2_Kup toQ6_K) of the same model, a bigger file means better quality and more memory. Pick the biggest one that fits your memory together with your context. See How much memory a model needs. - Compare file sizes only within one quant family.
Q4_1is bigger thanQ4_K_Mbut has lower quality, andQ5_1is bigger thanQ5_K_Mbut has lower quality. - Below Q4, quality drops quickly (see the table below).
Quant names
Section titled “Quant names”| Family | Names | Notes |
|---|---|---|
| K-quants | Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K | The usual choice. The number is the approximate bits per weight. |
| Legacy quants | Q4_0, Q4_1, Q5_0, Q5_1, Q8_0 | Older types. Q8_0 is the largest quant and the closest to the original. |
| I-quants | IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ2_M, IQ3_XXS, IQ3_XS, IQ3_S, IQ3_M, IQ4_XS, IQ4_NL | The number is the approximate bits per weight. For the same number, the suffix gives the size: XXS is the smallest, then XS, S and M. |
| Unquantized | F16, BF16 | 16 bits per weight, about twice the size of Q8_0. |
Q3_K, Q4_K and Q5_K without a suffix are other names for Q3_K_M, Q4_K_M and Q5_K_M.
_S, _M and _L (small, medium, large) mix quant types inside one file: some tensors get a larger type. That is why Q4_K_S and Q4_K_M differ in size although both are "Q4_K".
Size and quality loss per quant
Section titled “Size and quality loss per quant”llama.cpp's quantize tool lists these reference values for Llama-3-8B. "Quality loss" is the increase in perplexity compared with the original model; lower is better.
| Quant | File size | Quality loss |
|---|---|---|
Q2_K | 2.96 GB | +3.5199 |
Q3_K_S | 3.41 GB | +1.6321 |
Q3_K_M | 3.74 GB | +0.6569 |
Q3_K_L | 4.03 GB | +0.5562 |
Q4_0 | 4.34 GB | +0.4685 |
Q4_K_S | 4.37 GB | +0.2689 |
Q4_K_M | 4.58 GB | +0.1754 |
Q4_1 | 4.78 GB | +0.4511 |
Q5_0 | 5.21 GB | +0.1316 |
Q5_K_S | 5.21 GB | +0.1049 |
Q5_K_M | 5.33 GB | +0.0569 |
Q5_1 | 5.65 GB | +0.1062 |
Q6_K | 6.14 GB | +0.0217 |
Q8_0 | 7.96 GB | +0.0026 |
Other models have other file sizes; the table shows how the quants compare with each other.
Inspect a model file
Section titled “Inspect a model file”On the Extra tab, Analyze Model reads the metadata, weight types and tensor names of a GGUF or safetensors file. On the command line (koboldcpp stands for your KoboldCpp file; see Command line):
koboldcpp --analyze mymodel.ggufOther formats
Section titled “Other formats”KoboldCpp does not load text models in safetensors or PyTorch .bin format directly. For most models, someone has already published GGUF files; see Where to get models.
To convert a model yourself:
- Download the conversion and quantization tools.
- Run
convert_hf_to_gguf.pyon the model to get a GGUF file. It needs Python 3 with the packages from KoboldCpp'srequirements.txt, including PyTorch andtransformers. - Run
quantize_gguf.exeon that file to make a smaller quant. The tools contain only the Windows exe; on Linux and macOS, buildquantize_gguffrom the KoboldCpp source withmake quantize_gguf.
For example, to make a Q4_K_M quant of a model downloaded into the folder mymodel:
python convert_hf_to_gguf.py mymodel --outfile mymodel-16bit.ggufquantize_gguf.exe mymodel-16bit.gguf mymodel-Q4_K_M.gguf Q4_K_Mconvert_hf_to_gguf.py <model folder>writes a 16-bit GGUF by default.--outtypepicks another type:f32,f16,bf16orq8_0. For vision models, a second run with--mmprojwrites the projector file; this works only for some models.quantize_gguf <input> <output> <type>takes the quant type last. Run it without arguments to list all types.
Advanced: bits per weight
Section titled “Advanced: bits per weight”These are tensor types, not file quants. Each type stores weights in fixed-size blocks, which give its exact bits per weight. A file quant such as Q4_K_M mixes its base type (Q4_K) with larger types for some tensors, so the file's average is slightly higher.
| Type | Bits per weight |
|---|---|
IQ1_S | 1.5625 |
IQ1_M | 1.75 |
IQ2_XXS | 2.0625 |
IQ2_XS | 2.3125 |
IQ2_S | 2.5625 |
Q2_K | 2.625 |
IQ3_XXS | 3.0625 |
Q3_K | 3.4375 |
IQ3_S | 3.4375 |
IQ4_XS | 4.25 |
Q4_0 | 4.5 |
Q4_K | 4.5 |
IQ4_NL | 4.5 |
Q4_1 | 5.0 |
Q5_0 | 5.5 |
Q5_K | 5.5 |
Q5_1 | 6.0 |
Q6_K | 6.5625 |
Q8_0 | 8.5 |
F16, BF16 | 16 |
Advanced: legacy GGML models
Section titled “Advanced: legacy GGML models”KoboldCpp still loads the older GGML formats that came before GGUF, usually .bin files. Supported are old llama-family files (GGML, GGMF, GGJT) and old GPT-J, GPT-2, RWKV, GPT-NeoX and MPT files.
- KoboldCpp shows a warning when you load one.
- Some newer features might be unavailable. Reconvert or re-download the model as GGUF if you can.
- GPU offload for these formats is built only into the CUDA and ROCm libraries. With Vulkan they run on the CPU.
- For non-llama legacy formats, the batch size is capped at 256.