# Speech-to-Speech Input and Output Formats: Complete Audio Specification for the Hugging Face Realtime API

> Discover the complete audio specification for Hugging Face speech-to-speech models. Learn about input and output formats: 16 kHz PCM, mono, 16-bit signed, base64-encoded chunks.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: api-reference
- Published: 2026-08-02

---

**The speech-to-speech pipeline accepts and produces raw linear PCM audio at 16 kHz, mono, 16-bit signed little-endian, transmitted as base64-encoded chunks over the OpenAI-compatible Realtime protocol.**

The **huggingface/speech-to-speech** repository implements a fully local speech-to-speech system that mirrors the OpenAI Realtime API. Understanding the exact audio input and output formats is essential for building compatible clients. This guide covers the PCM specifications, transport encoding, and where each conversion happens in the codebase.

---

## Input Audio Format: What the Server Expects

The `RealtimeService` accepts audio through the `input_audio_buffer.append` event. Two sampling rates are supported:

| Parameter | Default | Alternative |
|-----------|---------|-------------|
| **Sample rate** | 16 kHz | 24 kHz (auto-resampled) |
| **Channels** | Mono (1) | — |
| **Bit depth** | 16-bit signed | — |
| **Endianness** | Little-endian | — |
| **Transport** | Base64-encoded PCM chunks | — |

### How Input Audio Is Processed

When your client sends audio to the server, the following happens in [`src/speech_to_speech/api/openai_realtime/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/README.md):

1. **Base64 decoding** — The `audio` field in `input_audio_buffer.append` is decoded from base64 to raw bytes
2. **Resampling** — 24 kHz input is resampled to 16 kHz using internal resampling logic; 16 kHz passes through unchanged
3. **Frame splitting** — Audio is chunked into **512-sample frames** (32 ms at 16 kHz) for the VAD (Voice Activity Detector)

The 512-sample frame size is hardcoded for low-latency processing. As noted in the Realtime Engine documentation:

> "Inbound audio: Client sends `input_audio_buffer.append` with base64 PCM. … Resampled to 16 kHz … split into 512-sample frames for the VAD."

---

## Output Audio Format: What the Server Sends Back

Synthesized speech follows the same PCM specification as the internal processing:

- **16 kHz sample rate**
- **Mono channel**
- **16-bit signed little-endian**

The TTS handler writes PCM chunks to `send_audio_chunks_queue`, and the WebSocket router encodes each chunk as a `response.output_audio.delta` event. This design ensures zero additional resampling overhead on the output path.

### Output Audio Flow

1. **[`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py)** generates 16 kHz PCM bytes
2. **[`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py)** pulls from `send_audio_chunks_queue`
3. **Base64 encoding** wraps the PCM for JSON transport
4. **Client receives** `response.output_audio.delta` events

---

## Why 16 kHz? The Design Rationale

The pipeline standardizes on **16 kHz** for all internal processing (VAD → STT → LLM → TTS). This sampling rate provides:

- **Low latency** — Smaller buffer sizes and faster inference
- **Adequate quality** — Sufficient for voice assistant intelligibility
- **Model compatibility** — Matches the training data of most open-source speech models

The optional **24 kHz input support** exists solely for compatibility with the official OpenAI Realtime schema. The server handles the conversion transparently, so developers can use either rate without changing client logic.

---

## Code Examples: Working with PCM Audio

### Recording and Sending 16 kHz PCM

```python
import sounddevice as sd
import base64
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="unused",
)

def send_pcm_chunk(chunk: bytes):
    """Send 512-sample 16-kHz PCM chunk (16-bit signed little-endian)."""
    b64 = base64.b64encode(chunk).decode()
    client.realtime.connect(model="local").send({
        "type": "input_audio_buffer.append",
        "audio": b64
    })

# Record 32ms chunks: 512 samples @ 16 kHz, 16-bit = 1024 bytes

with sd.RawInputStream(
    samplerate=16000,
    channels=1,
    dtype="int16",
    blocksize=512
) as stream:
    while True:
        data = stream.read(512)[0]
        send_pcm_chunk(data)

```

Key parameters in this snippet:
- `samplerate=16000` — matches the pipeline's internal rate
- `dtype="int16"` — 16-bit signed PCM
- `blocksize=512` — aligns with the VAD frame size

### Receiving and Playing Synthesized Audio

```python
import base64
import numpy as np
import sounddevice as sd

def play_pcm_chunk(b64_audio: str):
    """Decode base64 PCM and play at 16 kHz."""
    raw = base64.b64decode(b64_audio)
    samples = np.frombuffer(raw, dtype=np.int16)
    sd.play(samples, samplerate=16000, blocking=False)

with client.realtime.connect(model="local") as conn:
    for event in conn:
        if event["type"] == "response.output_audio.delta":
            play_pcm_chunk(event["audio"])

```

### Starting the Realtime Server

```bash
speech-to-speech --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3

```

---

## Implementation Details: Where Formats Are Handled

| Component | File Path | Role |
|-----------|-----------|------|
| **Protocol specification** | [`src/speech_to_speech/api/openai_realtime/README.md`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/README.md) | Documents `input_audio_buffer.append` and `response.output_audio.delta` formats |
| **WebSocket audio routing** | [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) | Decodes base64, resamples 24→16 kHz, chunks to 512 samples |
| **TTS output generation** | [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) | Produces 16 kHz PCM for `send_audio_chunks_queue` |
| **Client reference** | [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) | Full client implementation with correct PCM handling |
| **Inter-component queues** | [`src/speech_to_speech/pipeline/messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/pipeline/messages.py) | Defines `recv_audio_chunks_queue` and `send_audio_chunks_queue` for 16 kHz PCM transport |

The queue definitions in [`messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/messages.py) enforce the 16 kHz convention throughout the pipeline. All handlers read from and write to these queues, ensuring consistent audio format across VAD, speech recognition, language model, and text-to-speech components.

---

## Summary

- **Input format**: 16 kHz or 24 kHz mono 16-bit PCM, base64-encoded in `input_audio_buffer.append` events
- **Output format**: 16 kHz mono 16-bit PCM, base64-encoded in `response.output_audio.delta` events
- **Internal standard**: 16 kHz for all processing, with automatic resampling of 24 kHz input
- **Frame size**: 512 samples (32 ms) for VAD processing
- **Transport**: WebSocket with JSON events containing base64 audio data

---

## Frequently Asked Questions

### Does the server support MP3 or WAV file uploads?

No. The speech-to-speech pipeline only accepts **raw PCM audio** over the Realtime WebSocket protocol. File-based ingestion would require client-side conversion to the streaming PCM format described above.

### What happens if I send 44.1 kHz or 48 kHz audio?

The server does not automatically resample arbitrary rates. Send **16 kHz** or **24 kHz** PCM for guaranteed compatibility. Other sample rates may cause pitch distortion or processing errors in the VAD stage.

### Can I change the internal 16 kHz processing rate?

No. The 16 kHz rate is baked into multiple components: the VAD frame size (512 samples), STT model configurations, and TTS output generation. Changing it would require modifications across [`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py), [`messages.py`](https://github.com/huggingface/speech-to-speech/blob/main/messages.py), and all handler implementations.

### Why base64 encoding instead of binary WebSocket frames?

The OpenAI Realtime protocol specification uses JSON events with base64-encoded audio fields. The huggingface/speech-to-speech implementation maintains strict compatibility with this schema, allowing drop-in replacement of OpenAI's hosted service with your local server.