How to Set Up a Fully Local Speech-to-Speech Stack with llama.cpp Instead of OpenAI API

You can replace the OpenAI API backend in the Hugging Face speech-to-speech pipeline with a local llama.cpp server by setting --openai_base_url to your local endpoint and using a dummy API key, enabling completely offline voice conversations.

The huggingface/speech-to-speech repository orchestrates a three-stage pipeline that converts speech to text, processes it through a language model, and synthesizes the response back to speech. While the LLM stage defaults to OpenAI's remote API, the architecture supports any OpenAI-compatible endpoint, including a local llama.cpp server. This guide explains how to set up a fully local stack with llama.cpp instead of OpenAI API for speech-to-speech processing, eliminating external dependencies and keeping all data on your machine.

Architectural Overview

The speech-to-speech pipeline consists of three distinct stages, each with configurable backends:

Stage Default Local Alternative Key Source Files
STT Whisper / Parakeet Same (already local) src/speech_to_speech/STT/*
LLM OpenAI API (responses-api or chat-completions) llama.cpp server src/speech_to_speech/LLM/base_openai_compatible_language_model.py, src/speech_to_speech/arguments_classes/module_arguments.py
TTS Qwen3-TTS, Kokoro Same (already local) src/speech_to_speech/TTS/*

The repository selects the LLM backend based on the llm_backend argument defined in src/speech_to_speech/arguments_classes/module_arguments.py. When set to "responses-api" or "chat-completions", the pipeline instantiates BaseOpenAICompatibleLanguageModel from src/speech_to_speech/LLM/base_openai_compatible_language_model.py, which sends HTTP requests to the URL stored in RuntimeConfig.openai_base_url. By pointing this URL to a local llama.cpp server, all LLM inference runs locally without API keys.

Install and Run llama.cpp

To serve models locally, you must build llama.cpp with the server component enabled and start it in OpenAI-compatible mode.

First, clone and build the repository:

git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make LLAMA_BUILD_SERVER=1

Next, download a GGML format model. For example, download a quantized Llama-2-7B model:

mkdir -p models/7B
curl -L -o models/7B/ggml-model-q4_0.bin \
     https://huggingface.co/TheBloke/Llama-2-7B-GGML/resolve/main/llama-2-7b.Q4_0.ggml.bin

Start the server with the OpenAI-compatible API enabled:

./main -m models/7B/ggml-model-q4_0.bin \
       -c 2048 \
       --host 127.0.0.1 \
       --port 5000 \
       --api \
       --chat-template default

The server now exposes /chat/completions and /completions endpoints at http://127.0.0.1:5000/v1, mimicking the OpenAI API specification.

Configure Speech-to-Speech for Local LLM

The pipeline accepts specific CLI arguments to redirect LLM traffic from OpenAI's servers to your local instance.

CLI Arguments

The two critical flags are:

  • --llm_backend: Set to responses-api or chat-completions to use the OpenAI-compatible client
  • --openai_base_url: Override the default https://api.openai.com/v1 with your local server address

Run the pipeline with:

python -m speech_to_speech.scripts.listen_and_play \
    --llm_backend responses-api \
    --openai_base_url http://127.0.0.1:5000/v1 \
    --openai_api_key dummy \
    --stt_backend whisper \
    --tts_backend qwen3_tts

The --openai_api_key parameter requires a non-empty string, but llama.cpp ignores this value when running locally.

Environment Variables

For persistent configuration, export the environment variables that the argument parser reads in src/speech_to_speech/arguments_classes/module_arguments.py:

export S2S_LLM_BACKEND=responses-api
export S2S_OPENAI_BASE_URL=http://127.0.0.1:5000/v1
export S2S_OPENAI_API_KEY=dummy

With these variables set, you can launch the pipeline without repeating the flags:

python -m speech_to_speech.scripts.listen_and_play \
    --stt_backend whisper \
    --tts_backend qwen3_tts

Code Implementation Details

Understanding the source code path helps troubleshoot connection issues.

src/speech_to_speech/arguments_classes/module_arguments.py

This file defines the llm_backend argument and parsing logic for openai_base_url and openai_api_key. It maps CLI flags and environment variables to the runtime configuration.

src/speech_to_speech/LLM/base_openai_compatible_language_model.py

This class implements the HTTP client that communicates with the OpenAI-compatible endpoint. It constructs POST requests to {openai_base_url}/chat/completions (or /responses depending on the backend mode) using the openai_api_key for the Authorization header. When you point openai_base_url to your local llama.cpp server, these requests never leave your machine.

src/speech_to_speech/LLM/language_model.py

This file provides alternative local backends like transformers and mlx-lm. While these run models directly in Python, using llama.cpp via the OpenAI-compatible interface often provides better performance and broader model support through GGML quantization.

scripts/listen_and_play.py

The entry point script initializes the pipeline. It parses arguments via module_arguments.py and instantiates the appropriate LLM handler based on your llm_backend selection.

Complete Working Example

This full example demonstrates building llama.cpp, starting the server, and running the speech-to-speech pipeline locally:


# Build llama.cpp with server support

git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make LLAMA_BUILD_SERVER=1

# Download a quantized model

mkdir -p models/7B
curl -L -o models/7B/ggml-model-q4_0.bin \
     https://huggingface.co/TheBloke/Llama-2-7B-GGML/resolve/main/llama-2-7b.Q4_0.ggml.bin

# Terminal 1: Start the local LLM server

./main -m models/7B/ggml-model-q4_0.bin \
       -c 2048 \
       --host 127.0.0.1 \
       --port 5000 \
       --api \
       --chat-template default

# Terminal 2: Run speech-to-speech with local backend

export S2S_LLM_BACKEND=responses-api
export S2S_OPENAI_BASE_URL=http://127.0.0.1:5000/v1
export S2S_OPENAI_API_KEY=dummy

python -m speech_to_speech.scripts.listen_and_play \
    --stt_backend whisper \
    --tts_backend qwen3_tts \
    --mic_device default

The pipeline now processes voice input through Whisper, sends the text to your local llama.cpp instance, and synthesizes the response using the local TTS backend, creating a fully offline voice conversation system.

Performance Tips and Troubleshooting

Model Compatibility

Ensure your GGML model supports the chat template format you intend to use. Start llama.cpp with --chat-template default for most instruction-tuned models, or specify a custom JSON template file for specialized models.

GPU Acceleration

By default, llama.cpp runs on CPU. For NVIDIA GPU support, compile with CUDA enabled:

make LLAMA_CUBLAS=1 LLAMA_BUILD_SERVER=1

Then launch with --gpu-layer-count to offload layers to the GPU:

./main -m models/7B/ggml-model-q4_0.bin --gpu-layer-count 35 --api --port 5000

Timeout Configuration

If the local model generates responses slowly, increase the client timeout in BaseOpenAICompatibleLanguageModel by setting the --llm_timeout flag (in seconds):

python -m speech_to_speech.scripts.listen_and_play \
    --llm_backend responses-api \
    --openai_base_url http://127.0.0.1:5000/v1 \
    --llm_timeout 60

Summary

  • The huggingface/speech-to-speech pipeline uses BaseOpenAICompatibleLanguageModel to communicate with LLM backends via an OpenAI-compatible REST API.
  • llama.cpp provides this same interface when started with the --api flag, allowing you to replace remote OpenAI calls with local inference.
  • Set --openai_base_url to your local server (e.g., http://127.0.0.1:5000/v1) and use any non-empty string for --openai_api_key to route traffic locally.
  • Configure persistent settings using the S2S_LLM_BACKEND, S2S_OPENAI_BASE_URL, and S2S_OPENAI_API_KEY environment variables.
  • All three pipeline stages (STT, LLM, TTS) can run locally, creating a privacy-preserving, offline voice assistant.

Frequently Asked Questions

Does llama.cpp require a valid OpenAI API key?

No. The llama.cpp server does not validate API keys, but the speech-to-speech client code requires a non-empty string for the openai_api_key parameter. You can pass any dummy value like "dummy" or "local" when using a local server, as implemented in src/speech_to_speech/LLM/base_openai_compatible_language_model.py.

Which LLM backend option should I use with llama.cpp?

Use either responses-api or chat-completions for the --llm_backend flag. Both options instantiate the BaseOpenAICompatibleLanguageModel class that communicates via HTTP. The responses-api option uses the newer OpenAI responses format, while chat-completions uses the traditional chat completion endpoint. Both work with llama.cpp's OpenAI-compatible server.

Can I use GPU acceleration with llama.cpp in this setup?

Yes. Compile llama.cpp with LLAMA_CUBLAS=1 for NVIDIA GPUs or LLAMA_METAL=1 for Apple Silicon, then start the server with the --gpu-layer-count parameter to offload model layers from CPU to GPU. The speech-to-speech pipeline connects to the server via HTTP, so it remains agnostic to whether the model runs on CPU or GPU.

Where is the OpenAI-compatible client implemented in the source code?

The client logic resides in src/speech_to_speech/LLM/base_openai_compatible_language_model.py. This file defines the BaseOpenAICompatibleLanguageModel class that constructs HTTP POST requests to the endpoint specified by RuntimeConfig.openai_base_url, handling authentication headers and response parsing for both streaming and non-streaming modes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →