KoboldCpp turns spoken audio into text with Whisper. You can talk to the AI in KoboldAI Lite, or send audio files to the transcription API.
Quick start
Section titled “Quick start”Started with a Chatbot newbie template? It already loads a Whisper model. Skip to step 4.
- Download a Whisper model (
.binfile) from huggingface.co/koboldcpp/whisper. A good first pick is whisper-base.en-q5_1.bin, the one theLowSpec-Chatbottemplate uses. - In the launcher, open the Audio tab and select the file in Whisper Model (Speech-To-Text):.
- Click Launch.
- In KoboldAI Lite, open Settings > Media and set Voice Input under Audio Input.
On the command line:
koboldcpp --model mymodel.gguf --whispermodel whisper-model.binVoice input in Lite
Section titled “Voice input in Lite”Settings > Media > Audio Input > Voice Input has these modes:
| Mode | How it works |
|---|---|
| Detect Voice | Hands-free: Lite detects when you speak. |
| Push-To-Talk | Lite records while you hold the voice button. |
| Toggle-To-Talk | The voice button switches recording on and off. |
The same section has Suppress Non-Speech, Language and Delay. The Add File menu also has a Microphone option.
Browsers allow the microphone only on HTTPS or on the PC itself (localhost). From another device, use the Remote Tunnel link or HTTPS.
Audio formats
Section titled “Audio formats”Whisper in KoboldCpp reads WAV, MP3 and FLAC. Convert other formats, such as WebM or OGG, before you send them. Some apps send WebM by default.
| Endpoint | Format |
|---|---|
POST /api/extra/transcribe | KoboldCpp |
POST /v1/audio/transcriptions | OpenAI |
Send the audio as base64 in audio_data, or as a multipart file upload. The answer is {"text": "..."}.
Optional fields:
| Field | Meaning |
|---|---|
language (or langcode) | Language code. Default auto. |
prompt | Text that guides the transcription. |
suppress_non_speech | Ignore sounds that are not speech. |
A multipart upload reads only language and prompt. Use JSON for langcode and suppress_non_speech.
With --password set, these endpoints need the password.
Without a Whisper model
Section titled “Without a Whisper model”If no Whisper model is loaded but the text model has an audio-capable mmproj, KoboldCpp transcribes with the text model instead. See Vision and audio input.