How to Set Up a Fully Local Speech-to-Speech Stack with llama.cpp Instead of OpenAI API
You can replace the OpenAI API backend in the Hugging Face speech-to-speech pipeline with a local llama.cpp server by setting --openai_base_url to your local endpoint and using a dummy API key, enabling completely offline voice conversations.
The huggingface/speech-to-speech repository orchestrates a three-stage pipeline that converts speech to text, processes it through a language model, and synthesizes the response back to speech. While the LLM stage defaults to OpenAI's remote API, the architecture supports any OpenAI-compatible endpoint, including a local llama.cpp server. This guide explains how to set up a fully local stack with llama.cpp instead of OpenAI API for speech-to-speech processing, eliminating external dependencies and keeping all data on your machine.
Architectural Overview
The speech-to-speech pipeline consists of three distinct stages, each with configurable backends:
| Stage | Default | Local Alternative | Key Source Files |
|---|---|---|---|
| STT | Whisper / Parakeet | Same (already local) | src/speech_to_speech/STT/* |
| LLM | OpenAI API (responses-api or chat-completions) |
llama.cpp server | src/speech_to_speech/LLM/base_openai_compatible_language_model.py, src/speech_to_speech/arguments_classes/module_arguments.py |
| TTS | Qwen3-TTS, Kokoro | Same (already local) | src/speech_to_speech/TTS/* |
The repository selects the LLM backend based on the llm_backend argument defined in src/speech_to_speech/arguments_classes/module_arguments.py. When set to "responses-api" or "chat-completions", the pipeline instantiates BaseOpenAICompatibleLanguageModel from src/speech_to_speech/LLM/base_openai_compatible_language_model.py, which sends HTTP requests to the URL stored in RuntimeConfig.openai_base_url. By pointing this URL to a local llama.cpp server, all LLM inference runs locally without API keys.
Install and Run llama.cpp
To serve models locally, you must build llama.cpp with the server component enabled and start it in OpenAI-compatible mode.
First, clone and build the repository:
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make LLAMA_BUILD_SERVER=1
Next, download a GGML format model. For example, download a quantized Llama-2-7B model:
mkdir -p models/7B
curl -L -o models/7B/ggml-model-q4_0.bin \
https://huggingface.co/TheBloke/Llama-2-7B-GGML/resolve/main/llama-2-7b.Q4_0.ggml.bin
Start the server with the OpenAI-compatible API enabled:
./main -m models/7B/ggml-model-q4_0.bin \
-c 2048 \
--host 127.0.0.1 \
--port 5000 \
--api \
--chat-template default
The server now exposes /chat/completions and /completions endpoints at http://127.0.0.1:5000/v1, mimicking the OpenAI API specification.
Configure Speech-to-Speech for Local LLM
The pipeline accepts specific CLI arguments to redirect LLM traffic from OpenAI's servers to your local instance.
CLI Arguments
The two critical flags are:
--llm_backend: Set toresponses-apiorchat-completionsto use the OpenAI-compatible client--openai_base_url: Override the defaulthttps://api.openai.com/v1with your local server address
Run the pipeline with:
python -m speech_to_speech.scripts.listen_and_play \
--llm_backend responses-api \
--openai_base_url http://127.0.0.1:5000/v1 \
--openai_api_key dummy \
--stt_backend whisper \
--tts_backend qwen3_tts
The --openai_api_key parameter requires a non-empty string, but llama.cpp ignores this value when running locally.
Environment Variables
For persistent configuration, export the environment variables that the argument parser reads in src/speech_to_speech/arguments_classes/module_arguments.py:
export S2S_LLM_BACKEND=responses-api
export S2S_OPENAI_BASE_URL=http://127.0.0.1:5000/v1
export S2S_OPENAI_API_KEY=dummy
With these variables set, you can launch the pipeline without repeating the flags:
python -m speech_to_speech.scripts.listen_and_play \
--stt_backend whisper \
--tts_backend qwen3_tts
Code Implementation Details
Understanding the source code path helps troubleshoot connection issues.
src/speech_to_speech/arguments_classes/module_arguments.py
This file defines the llm_backend argument and parsing logic for openai_base_url and openai_api_key. It maps CLI flags and environment variables to the runtime configuration.
src/speech_to_speech/LLM/base_openai_compatible_language_model.py
This class implements the HTTP client that communicates with the OpenAI-compatible endpoint. It constructs POST requests to {openai_base_url}/chat/completions (or /responses depending on the backend mode) using the openai_api_key for the Authorization header. When you point openai_base_url to your local llama.cpp server, these requests never leave your machine.
src/speech_to_speech/LLM/language_model.py
This file provides alternative local backends like transformers and mlx-lm. While these run models directly in Python, using llama.cpp via the OpenAI-compatible interface often provides better performance and broader model support through GGML quantization.
The entry point script initializes the pipeline. It parses arguments via module_arguments.py and instantiates the appropriate LLM handler based on your llm_backend selection.
Complete Working Example
This full example demonstrates building llama.cpp, starting the server, and running the speech-to-speech pipeline locally:
# Build llama.cpp with server support
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make LLAMA_BUILD_SERVER=1
# Download a quantized model
mkdir -p models/7B
curl -L -o models/7B/ggml-model-q4_0.bin \
https://huggingface.co/TheBloke/Llama-2-7B-GGML/resolve/main/llama-2-7b.Q4_0.ggml.bin
# Terminal 1: Start the local LLM server
./main -m models/7B/ggml-model-q4_0.bin \
-c 2048 \
--host 127.0.0.1 \
--port 5000 \
--api \
--chat-template default
# Terminal 2: Run speech-to-speech with local backend
export S2S_LLM_BACKEND=responses-api
export S2S_OPENAI_BASE_URL=http://127.0.0.1:5000/v1
export S2S_OPENAI_API_KEY=dummy
python -m speech_to_speech.scripts.listen_and_play \
--stt_backend whisper \
--tts_backend qwen3_tts \
--mic_device default
The pipeline now processes voice input through Whisper, sends the text to your local llama.cpp instance, and synthesizes the response using the local TTS backend, creating a fully offline voice conversation system.
Performance Tips and Troubleshooting
Model Compatibility
Ensure your GGML model supports the chat template format you intend to use. Start llama.cpp with --chat-template default for most instruction-tuned models, or specify a custom JSON template file for specialized models.
GPU Acceleration
By default, llama.cpp runs on CPU. For NVIDIA GPU support, compile with CUDA enabled:
make LLAMA_CUBLAS=1 LLAMA_BUILD_SERVER=1
Then launch with --gpu-layer-count to offload layers to the GPU:
./main -m models/7B/ggml-model-q4_0.bin --gpu-layer-count 35 --api --port 5000
Timeout Configuration
If the local model generates responses slowly, increase the client timeout in BaseOpenAICompatibleLanguageModel by setting the --llm_timeout flag (in seconds):
python -m speech_to_speech.scripts.listen_and_play \
--llm_backend responses-api \
--openai_base_url http://127.0.0.1:5000/v1 \
--llm_timeout 60
Summary
- The huggingface/speech-to-speech pipeline uses
BaseOpenAICompatibleLanguageModelto communicate with LLM backends via an OpenAI-compatible REST API. - llama.cpp provides this same interface when started with the
--apiflag, allowing you to replace remote OpenAI calls with local inference. - Set
--openai_base_urlto your local server (e.g.,http://127.0.0.1:5000/v1) and use any non-empty string for--openai_api_keyto route traffic locally. - Configure persistent settings using the
S2S_LLM_BACKEND,S2S_OPENAI_BASE_URL, andS2S_OPENAI_API_KEYenvironment variables. - All three pipeline stages (STT, LLM, TTS) can run locally, creating a privacy-preserving, offline voice assistant.
Frequently Asked Questions
Does llama.cpp require a valid OpenAI API key?
No. The llama.cpp server does not validate API keys, but the speech-to-speech client code requires a non-empty string for the openai_api_key parameter. You can pass any dummy value like "dummy" or "local" when using a local server, as implemented in src/speech_to_speech/LLM/base_openai_compatible_language_model.py.
Which LLM backend option should I use with llama.cpp?
Use either responses-api or chat-completions for the --llm_backend flag. Both options instantiate the BaseOpenAICompatibleLanguageModel class that communicates via HTTP. The responses-api option uses the newer OpenAI responses format, while chat-completions uses the traditional chat completion endpoint. Both work with llama.cpp's OpenAI-compatible server.
Can I use GPU acceleration with llama.cpp in this setup?
Yes. Compile llama.cpp with LLAMA_CUBLAS=1 for NVIDIA GPUs or LLAMA_METAL=1 for Apple Silicon, then start the server with the --gpu-layer-count parameter to offload model layers from CPU to GPU. The speech-to-speech pipeline connects to the server via HTTP, so it remains agnostic to whether the model runs on CPU or GPU.
Where is the OpenAI-compatible client implemented in the source code?
The client logic resides in src/speech_to_speech/LLM/base_openai_compatible_language_model.py. This file defines the BaseOpenAICompatibleLanguageModel class that constructs HTTP POST requests to the endpoint specified by RuntimeConfig.openai_base_url, handling authentication headers and response parsing for both streaming and non-streaming modes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →