The model loads, but its replies are wrong. Most settings on this page are in KoboldAI Lite's settings, not in the launcher.
The AI writes gibberish
Section titled “The AI writes gibberish”Replies are random characters, repeated symbols or nonsense words. Known causes are graphics driver bugs (often AMD cards with Vulkan), a badly converted model file, and bugs in a KoboldCpp version with specific models.
- Update your graphics driver.
- Try another backend. With AMD on Vulkan, try Use Vulkan (Old CPU). To rule out the graphics card, try Use CPU.
- Download the model again, from a well-known uploader.
- Update KoboldCpp.
- Try a different Batch Size: on the Hardware tab (
--batchsize). - Turn off Use FlashAttention (
--noflashattention). - If you set a custom RoPE configuration, remove it. The default scales automatically.
If a model still produces gibberish, report it with the model name and your settings.
The AI never stops, or keeps rambling
Section titled “The AI never stops, or keeps rambling”The model's end token (EOS) is banned, or the instruct format does not match the model. The reply then runs until the output length is used up.
- In KoboldAI Lite, open Settings > Misc. Under Advanced Settings, set EOS token ban to Auto (the default) or Unban.
- In instruct mode, check that the instruct tags match the model. Lite's default placeholder tags adapt to the model automatically.
The old --unbantokens flag no longer exists; KoboldCpp exits with unrecognized arguments if you pass it.
The AI replies as me in chat mode
Section titled “The AI replies as me in chat mode”The model writes your lines as well as its own.
- Retry the reply.
- In KoboldAI Lite, open Settings > General and turn off Allow multiline replies.
- Use a model that was trained for chat.
- Check the character card and the start of the chat: consistent, correctly spelled names and a few good example messages help.
Replies are cut short
Section titled “Replies are cut short”In other apps that use the API, replies stop after a fixed length when the app does not send a length. KoboldCpp then uses its default, Default Gen Amt: on the Context tab (--defaultgenamt, 2048 tokens). It is capped at half the context size.
- Set the reply length (
max_tokensor "max length") in your app. - Or raise Default Gen Amt:, and raise Context Size: if half of it is below what you need.
- Check that Prompt Limit: (
--genlimit) is not set; it caps every reply.
In KoboldAI Lite chat mode, turning off Allow multiline replies (Settings > General) also shortens replies to one line.
Badly formatted replies over the Chat Completions API
Section titled “Badly formatted replies over the Chat Completions API”Chat template heuristics failed to identify chat completions format. Alpaca will be used.KoboldCpp could not tell which chat format the model uses and fell back to Alpaca, which may not match.
- Turn on Use Jinja on the launcher's Context tab (
--jinja). KoboldCpp then uses the chat template stored in the model file. - Or pick the format yourself: Chat Adapter: on the Loaded Files tab, with Pick Premade (
--chatcompletionsadapter).
See Chat formatting.
The whole prompt is processed again on every message
Section titled “The whole prompt is processed again on every message”Normally KoboldCpp reuses the processed text and, when the context is full, shifts it (ContextShift). On models with sliding window attention (SWA), such as Gemma 3, SWA is on by default and turns ContextShift off:
SWA Mode is ENABLED!Note that using SWA Mode cannot be used with Context Shifting!- Keep Use FastForwarding and Use ContextShift on (the defaults).
- To get ContextShift back on an SWA model, turn off Allow SWA on the Context tab (
--noswa). This uses much more memory for the context. - Parallel Requests: above 1 also turns ContextShift off. Set it to 1 if you don't need parallel requests.
The AI ignores my image or audio
Section titled “The AI ignores my image or audio”Warning: Media excluded - Context size too low or not enough mtmd tokens! (needed …)Media will be IGNORED! You probably want to relaunch with a larger context size!The image or audio did not fit into the context. Relaunch with a larger Context Size:. Long audio clips need a lot of context.
Without a vision file (mmproj) loaded, the model cannot see images at all. See Vision and audio input.
Warnings about the context size
Section titled “Warnings about the context size”Warning! Request max_context_length=32768 exceeds allocated context size of 16384. It will be reduced to fit.Warning: You are trying to generate text with max_length (10000) near or exceeding max_context_length limit (16384).Your app asked for more context than KoboldCpp reserved at launch, or for a reply almost as long as the whole context. Most of the story is then cut off, and replies lose coherence.
- Relaunch with a larger Context Size: (
--contextsize). An app's context setting cannot go above the launch value. - Or lower the context size and reply length in your app.
See Context size.
"Quantized KV was used without flash attention"
Section titled “"Quantized KV was used without flash attention"”Warning: Quantized KV was used without flash attention! This is NOT RECOMMENDED!With Use FlashAttention off, a quantized context cache (Quantize KV Cache:, --quantkv) only partly works and can even use more VRAM. Turn Use FlashAttention back on, or set Quantize KV Cache: to f16.
Advanced: sampler settings
Section titled “Advanced: sampler settings”Extreme sampler values also produce bad output. KoboldAI Lite's defaults are a good reset point:
| Setting | Lite default |
|---|---|
| Temperature | 0.75 |
| Repetition penalty | 1.05 |
| Top-P | 0.92 |
| Sampler order | [6,0,1,3,4,2,5] |
An API request with an invalid order logs ERROR: sampler_order must be a list of integers.
See KoboldAI Lite.