# How to Set Up a Fully Local Speech-to-Speech Stack with llama.cpp Instead of OpenAI API

> Create a local speech-to-speech stack with llama.cpp by replacing the OpenAI API. Learn how to set up your offline voice conversations easily and efficiently.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: how-to-guide
- Published: 2026-07-11

---

**You can replace the OpenAI API backend in the Hugging Face speech-to-speech pipeline with a local llama.cpp server by setting `--openai_base_url` to your local endpoint and using a dummy API key, enabling completely offline voice conversations.**

The `huggingface/speech-to-speech` repository orchestrates a three-stage pipeline that converts speech to text, processes it through a language model, and synthesizes the response back to speech. While the LLM stage defaults to OpenAI's remote API, the architecture supports any OpenAI-compatible endpoint, including a local llama.cpp server. This guide explains how to set up a fully local stack with llama.cpp instead of OpenAI API for speech-to-speech processing, eliminating external dependencies and keeping all data on your machine.

## Architectural Overview

The speech-to-speech pipeline consists of three distinct stages, each with configurable backends:

| Stage | Default | Local Alternative | Key Source Files |
|-------|---------|-----------------|------------------|
| **STT** | Whisper / Parakeet | Same (already local) | `src/speech_to_speech/STT/*` |
| **LLM** | OpenAI API (`responses-api` or `chat-completions`) | **llama.cpp** server | [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py), [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py) |
| **TTS** | Qwen3-TTS, Kokoro | Same (already local) | `src/speech_to_speech/TTS/*` |

The repository selects the LLM backend based on the `llm_backend` argument defined in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py). When set to `"responses-api"` or `"chat-completions"`, the pipeline instantiates `BaseOpenAICompatibleLanguageModel` from [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py), which sends HTTP requests to the URL stored in `RuntimeConfig.openai_base_url`. By pointing this URL to a local llama.cpp server, all LLM inference runs locally without API keys.

## Install and Run llama.cpp

To serve models locally, you must build llama.cpp with the server component enabled and start it in OpenAI-compatible mode.

First, clone and build the repository:

```bash
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make LLAMA_BUILD_SERVER=1

```

Next, download a GGML format model. For example, download a quantized Llama-2-7B model:

```bash
mkdir -p models/7B
curl -L -o models/7B/ggml-model-q4_0.bin \
     https://huggingface.co/TheBloke/Llama-2-7B-GGML/resolve/main/llama-2-7b.Q4_0.ggml.bin

```

Start the server with the OpenAI-compatible API enabled:

```bash
./main -m models/7B/ggml-model-q4_0.bin \
       -c 2048 \
       --host 127.0.0.1 \
       --port 5000 \
       --api \
       --chat-template default

```

The server now exposes `/chat/completions` and `/completions` endpoints at `http://127.0.0.1:5000/v1`, mimicking the OpenAI API specification.

## Configure Speech-to-Speech for Local LLM

The pipeline accepts specific CLI arguments to redirect LLM traffic from OpenAI's servers to your local instance.

### CLI Arguments

The two critical flags are:

- **`--llm_backend`**: Set to `responses-api` or `chat-completions` to use the OpenAI-compatible client
- **`--openai_base_url`**: Override the default `https://api.openai.com/v1` with your local server address

Run the pipeline with:

```bash
python -m speech_to_speech.scripts.listen_and_play \
    --llm_backend responses-api \
    --openai_base_url http://127.0.0.1:5000/v1 \
    --openai_api_key dummy \
    --stt_backend whisper \
    --tts_backend qwen3_tts

```

The `--openai_api_key` parameter requires a non-empty string, but llama.cpp ignores this value when running locally.

### Environment Variables

For persistent configuration, export the environment variables that the argument parser reads in [`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py):

```bash
export S2S_LLM_BACKEND=responses-api
export S2S_OPENAI_BASE_URL=http://127.0.0.1:5000/v1
export S2S_OPENAI_API_KEY=dummy

```

With these variables set, you can launch the pipeline without repeating the flags:

```bash
python -m speech_to_speech.scripts.listen_and_play \
    --stt_backend whisper \
    --tts_backend qwen3_tts

```

## Code Implementation Details

Understanding the source code path helps troubleshoot connection issues.

**[`src/speech_to_speech/arguments_classes/module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/module_arguments.py)**

This file defines the `llm_backend` argument and parsing logic for `openai_base_url` and `openai_api_key`. It maps CLI flags and environment variables to the runtime configuration.

**[`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py)**

This class implements the HTTP client that communicates with the OpenAI-compatible endpoint. It constructs POST requests to `{openai_base_url}/chat/completions` (or `/responses` depending on the backend mode) using the `openai_api_key` for the Authorization header. When you point `openai_base_url` to your local llama.cpp server, these requests never leave your machine.

**[`src/speech_to_speech/LLM/language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/language_model.py)**

This file provides alternative local backends like `transformers` and `mlx-lm`. While these run models directly in Python, using llama.cpp via the OpenAI-compatible interface often provides better performance and broader model support through GGML quantization.

**[`scripts/listen_and_play.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play.py)**

The entry point script initializes the pipeline. It parses arguments via [`module_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/module_arguments.py) and instantiates the appropriate LLM handler based on your `llm_backend` selection.

## Complete Working Example

This full example demonstrates building llama.cpp, starting the server, and running the speech-to-speech pipeline locally:

```bash

# Build llama.cpp with server support

git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make LLAMA_BUILD_SERVER=1

# Download a quantized model

mkdir -p models/7B
curl -L -o models/7B/ggml-model-q4_0.bin \
     https://huggingface.co/TheBloke/Llama-2-7B-GGML/resolve/main/llama-2-7b.Q4_0.ggml.bin

# Terminal 1: Start the local LLM server

./main -m models/7B/ggml-model-q4_0.bin \
       -c 2048 \
       --host 127.0.0.1 \
       --port 5000 \
       --api \
       --chat-template default

# Terminal 2: Run speech-to-speech with local backend

export S2S_LLM_BACKEND=responses-api
export S2S_OPENAI_BASE_URL=http://127.0.0.1:5000/v1
export S2S_OPENAI_API_KEY=dummy

python -m speech_to_speech.scripts.listen_and_play \
    --stt_backend whisper \
    --tts_backend qwen3_tts \
    --mic_device default

```

The pipeline now processes voice input through Whisper, sends the text to your local llama.cpp instance, and synthesizes the response using the local TTS backend, creating a fully offline voice conversation system.

## Performance Tips and Troubleshooting

**Model Compatibility**

Ensure your GGML model supports the chat template format you intend to use. Start llama.cpp with `--chat-template default` for most instruction-tuned models, or specify a custom JSON template file for specialized models.

**GPU Acceleration**

By default, llama.cpp runs on CPU. For NVIDIA GPU support, compile with CUDA enabled:

```bash
make LLAMA_CUBLAS=1 LLAMA_BUILD_SERVER=1

```

Then launch with `--gpu-layer-count` to offload layers to the GPU:

```bash
./main -m models/7B/ggml-model-q4_0.bin --gpu-layer-count 35 --api --port 5000

```

**Timeout Configuration**

If the local model generates responses slowly, increase the client timeout in `BaseOpenAICompatibleLanguageModel` by setting the `--llm_timeout` flag (in seconds):

```bash
python -m speech_to_speech.scripts.listen_and_play \
    --llm_backend responses-api \
    --openai_base_url http://127.0.0.1:5000/v1 \
    --llm_timeout 60

```

## Summary

- The **huggingface/speech-to-speech** pipeline uses `BaseOpenAICompatibleLanguageModel` to communicate with LLM backends via an OpenAI-compatible REST API.
- **llama.cpp** provides this same interface when started with the `--api` flag, allowing you to replace remote OpenAI calls with local inference.
- Set `--openai_base_url` to your local server (e.g., `http://127.0.0.1:5000/v1`) and use any non-empty string for `--openai_api_key` to route traffic locally.
- Configure persistent settings using the `S2S_LLM_BACKEND`, `S2S_OPENAI_BASE_URL`, and `S2S_OPENAI_API_KEY` environment variables.
- All three pipeline stages (STT, LLM, TTS) can run locally, creating a privacy-preserving, offline voice assistant.

## Frequently Asked Questions

### Does llama.cpp require a valid OpenAI API key?

No. The llama.cpp server does not validate API keys, but the `speech-to-speech` client code requires a non-empty string for the `openai_api_key` parameter. You can pass any dummy value like `"dummy"` or `"local"` when using a local server, as implemented in [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py).

### Which LLM backend option should I use with llama.cpp?

Use either `responses-api` or `chat-completions` for the `--llm_backend` flag. Both options instantiate the `BaseOpenAICompatibleLanguageModel` class that communicates via HTTP. The `responses-api` option uses the newer OpenAI responses format, while `chat-completions` uses the traditional chat completion endpoint. Both work with llama.cpp's OpenAI-compatible server.

### Can I use GPU acceleration with llama.cpp in this setup?

Yes. Compile llama.cpp with `LLAMA_CUBLAS=1` for NVIDIA GPUs or `LLAMA_METAL=1` for Apple Silicon, then start the server with the `--gpu-layer-count` parameter to offload model layers from CPU to GPU. The speech-to-speech pipeline connects to the server via HTTP, so it remains agnostic to whether the model runs on CPU or GPU.

### Where is the OpenAI-compatible client implemented in the source code?

The client logic resides in [`src/speech_to_speech/LLM/base_openai_compatible_language_model.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/LLM/base_openai_compatible_language_model.py). This file defines the `BaseOpenAICompatibleLanguageModel` class that constructs HTTP POST requests to the endpoint specified by `RuntimeConfig.openai_base_url`, handling authentication headers and response parsing for both streaming and non-streaming modes.