A model needs memory for its weights (roughly the file size), for the context (the KV cache, which grows with Context Size:) and for working buffers. When that is more than your graphics card or PC has, loading fails.
See How much memory a model needs.
"cudaMalloc failed: out of memory" (NVIDIA)
Section titled “"cudaMalloc failed: out of memory" (NVIDIA)”ggml_backend_cuda_buffer_type_alloc_buffer: allocating … MiB on device 0: cudaMalloc failed: out of memoryFollowed by one of these:
unable to allocate CUDA0 bufferfailed to allocate buffer for kv cachefailed to allocate compute pp buffersThe layers on the graphics card, the context and the buffers together need more VRAM than is free. Try these in order, and relaunch after each:
Leave GPU Layers: at
-1(--gpulayers -1). KoboldCpp then fits as many layers as possible into VRAM. If you set a number yourself, lower it.Lower Context Size: (
--contextsize). The default is 16384. See Context size.Close other programs that use the graphics card. KoboldCpp warns you about them:
Note: KoboldCpp has detected that a significant amount of GPU VRAM (… MB) is currently used by another application.Quantize the context: set Quantize KV Cache: on the Context tab to
q8_0, orq4_0for more savings (--quantkv q8_0). Keep Use FlashAttention on; without it, only part of the cache is quantized, and it can even use more VRAM.Use a smaller quant of the model, or a smaller model. See Choosing a quant.
Lower Batch Size: on the Hardware tab (
--batchsize, default 512).As a last resort,
--lowvramkeeps the context in system RAM. It is slow and not recommended.
More options: GPU layers and Saving VRAM.
"ErrorOutOfDeviceMemory" (Vulkan)
Section titled “"ErrorOutOfDeviceMemory" (Vulkan)”ggml_vulkan: Device memory allocation of size … failed.vk::Device::allocateMemory: ErrorOutOfDeviceMemoryalloc_tensor_range: failed to allocate Vulkan0 buffer of size …The Vulkan backend ran out of VRAM. The fixes are the same as for NVIDIA: put fewer layers on the graphics card, lower the context size, close other programs.
Some Vulkan devices cannot allocate more than about 2 GB in one piece, even with more VRAM free. For image generation, these flags help: --sdvaeauto, --sdquant, and --sdconvdirect vaeonly.
See Image generation.
KoboldCpp is killed while loading
Section titled “KoboldCpp is killed while loading”The console shows Terminated, or KoboldCpp ends without an error, while system memory (RAM) is almost full. The model does not fit into RAM.
- Use a smaller quant or a smaller model.
- Close other programs.
- Turn on Use MMAP (
--usemmap). The system then reads the model file from disk as needed, instead of loading it all into RAM first. - Lower Context Size:.
Advanced: "not enough space in the scratch memory"
Section titled “Advanced: "not enough space in the scratch memory"”Messages like not enough space in the context's memory pool or scratch memory come from old GGML models and very old versions. Current versions show the CUDA or Vulkan errors above. Update KoboldCpp and use a GGUF model.