Skip to content
KoboldCpp
GitHub

Release files and system requirements

Each release on GitHub contains one ready-to-run file per system and GPU type. You need one of them. For a quick pick, see Download KoboldCpp.

All files come from the latest release. The Source code archives that GitHub adds to every release are not programs; don't download them to run KoboldCpp.

FileSystemGPU backendsCPU
koboldcpp.exeWindows x64CUDA 12, VulkanAVX2 (old-CPU modes included)
koboldcpp-nocuda.exeWindows x64VulkanAVX2 (old-CPU modes included)
koboldcpp-oldpc.exeWindows x64CUDA 11, VulkanAVX1 (no-AVX modes included)
koboldcpp-linux-x64Linux x64CUDA 12, VulkanAVX2 (old-CPU modes included)
koboldcpp-linux-x64-nocudaLinux x64VulkanAVX2 (old-CPU modes included)
koboldcpp-linux-x64-oldpcLinux x64CUDA 11, VulkanAVX1 (no-AVX modes included)
koboldcpp-mac-arm64macOS, Apple SiliconMetalnot applicable

Every file can also run on the CPU alone.

  • Main builds (koboldcpp.exe, koboldcpp-linux-x64): for NVIDIA graphics cards on newer PCs.
  • nocuda builds: the same without CUDA, and much smaller. For AMD or Intel graphics, no graphics card, or when you don't need CUDA. AMD users should try the Vulkan backend in this build first.
  • oldpc builds: CUDA 11 and AVX1, for older processors and older NVIDIA cards, or when the main build crashes. They are also the only builds with GPU support (Vulkan) on processors without AVX.
  • Mac build: for Apple Silicon (M-series) Macs. Intel Macs have no ready-made file; see macOS.

There are no ready-made files for Windows on ARM or Linux on ARM. On ARM Linux, build from source.

KoboldCpp has no fixed RAM or VRAM minimum; what you need depends on the model. The launcher's newbie templates give a rough guide: LowSpec recommends 6 GB VRAM, MidSpec 12 GB, HighSpec 24 GB. See Choose a model.

More: How much memory a model needs.

The main and nocuda builds run at full speed on processors with AVX2. For older processors, every x64 build includes slower compatibility modes, and KoboldCpp usually switches to them automatically when it detects that your CPU lacks AVX2 or AVX.

See Advanced: processor modes.

  • Windows: x64. Windows 7 is not recommended; see Windows 7.
  • Linux: x64 with glibc 2.29 or newer, for example Ubuntu 20.04 or newer. ldd --version shows your glibc version.
  • macOS: Apple Silicon only, macOS 14 (Sonoma) or newer.

You can pick a mode yourself, in the launcher's backend list or with a flag:

Launcher backendFlagCPU needsIn builds
Use CUDA--usecudaAVX2 (main), AVX1 (oldpc)main, oldpc
Use Vulkan--usevulkanAVX2main, nocuda
Use CPU--usecpuAVX2main, nocuda
Use CPU (Old CPU)--noavx2AVX1all x64 builds
Use Vulkan (Old CPU)--usevulkan --noavx2AVX1all x64 builds
Use Vulkan (Older CPU)--usevulkan --failsafeSSE3 and SSSE3oldpc only
Failsafe Mode (Older CPU)--failsafeany x86-64 CPUall x64 builds

The launcher only lists backends that your build contains.

  • --noavx2 is a slower compatibility mode for processors without AVX2.
  • --failsafe is an extremely slow mode that runs on any x86-64 processor. It implies --noavx2. Without --usevulkan it uses no GPU.
  • On a processor with no AVX at all that should still use a graphics card, use an oldpc build and select Use Vulkan (Older CPU) yourself. The launcher may pick Use Vulkan (Old CPU), which needs AVX.

More: Old PCs and CPU-only.

BuildCUDA toolkitMinimum NVIDIA driverCompute capability compiled
Main (koboldcpp.exe, koboldcpp-linux-x64)12.1531.14 (Windows), 530.30.02 (Linux)5.0, 6.1, 7.5, 8.0
oldpc (koboldcpp-oldpc.exe, koboldcpp-linux-x64-oldpc)11.4472.50 (Windows), 470.42.01 (Linux)3.5, 5.0, 6.1, 7.5
  • The code is compiled as PTX, so newer cards compile it on first launch. That first start prints "Initializing CUDA/HIP, please wait, the following step may take a few minutes (only for first launch)".
  • Kepler cards with compute capability 3.5 or higher (for example K40 and K80) are only covered by the oldpc builds. Compute capability 3.0 cards are not covered by any build; try Vulkan instead.
  • If nvidia-smi reports "CUDA Version: 11.x" or "12.0", use the oldpc build, or update the driver.

koboldcpp-mac-arm64 uses Metal. It also contains a CPU-only fallback that uses Accelerate without Metal; turn it on with --failsafe. See macOS.

Rolling builds are experimental pre-release builds that are updated automatically. They use the same file names as a normal release.

  • They may be unstable, and sometimes don't start at all.
  • They can support a new model architecture before the next release does.
  • For stability and proper support, use the official releases.

An official ROCm build for AMD graphics cards on Linux, koboldcpp-linux-x64-rocm, is published under the rocm-rolling tag. The official short link is https://koboldai.org/cpplinuxrocm.

  • It is very experimental.
  • It needs x64 Linux with glibc from Ubuntu 22.04 or newer.
  • It brings its own ROCm 7 libraries. You don't need to install ROCm; the normal amdgpu kernel driver is enough.
  • It is updated separately from the releases and can lag behind them: the file tested for these docs reported version 1.121 while the latest release was 1.122.1.
  • Support is limited, because the maintainers don't have AMD devices themselves.
  • There is no official ROCm build for Windows.

It is built for the Radeon RX 9000 series (gfx1200, gfx1201), RX 7000 series (gfx1100 to gfx1102), RX 6600 to RX 6950 XT (gfx1030 to gfx1032) and RX 5600/5700 (gfx1010), for older cards (RX 400/500 series: gfx803, Vega: gfx900, Radeon VII: gfx906) and for Instinct datacenter cards (gfx906, gfx908, gfx90a, gfx942).

To use ROCm, set Backend: to Use hipBLAS (ROCm), or pass --usecuda.

For most AMD users, the Vulkan backend in the nocuda build is the recommended start. If generation is slow with Vulkan, try the ROCm build; see GPU backends: ROCm.

Scripts that download older file names no longer work. The files were renamed and the old names were removed in v1.94:

Old nameCurrent name
koboldcpp_cu12.exekoboldcpp.exe
koboldcpp_nocuda.exekoboldcpp-nocuda.exe
koboldcpp_oldcpu.exekoboldcpp-oldpc.exe
koboldcpp-linux-x64-cuda1210koboldcpp-linux-x64
koboldcpp-linux-x64-cuda1150koboldcpp-linux-x64-oldpc

To always get the newest file, download from this address, with the file name at the end:

https://github.com/LostRuins/koboldcpp/releases/latest/download/<file name>

The release notes link conversion and quantization tools for making GGUF files.