Speech-to-Speech Input and Output Formats: Complete Audio Specification for the Hugging Face Realtime API
The speech-to-speech pipeline accepts and produces raw linear PCM audio at 16 kHz, mono, 16-bit signed little-endian, transmitted as base64-encoded chunks over the OpenAI-compatible Realtime protocol.
The huggingface/speech-to-speech repository implements a fully local speech-to-speech system that mirrors the OpenAI Realtime API. Understanding the exact audio input and output formats is essential for building compatible clients. This guide covers the PCM specifications, transport encoding, and where each conversion happens in the codebase.
Input Audio Format: What the Server Expects
The RealtimeService accepts audio through the input_audio_buffer.append event. Two sampling rates are supported:
| Parameter | Default | Alternative |
|---|---|---|
| Sample rate | 16 kHz | 24 kHz (auto-resampled) |
| Channels | Mono (1) | — |
| Bit depth | 16-bit signed | — |
| Endianness | Little-endian | — |
| Transport | Base64-encoded PCM chunks | — |
How Input Audio Is Processed
When your client sends audio to the server, the following happens in src/speech_to_speech/api/openai_realtime/README.md:
- Base64 decoding — The
audiofield ininput_audio_buffer.appendis decoded from base64 to raw bytes - Resampling — 24 kHz input is resampled to 16 kHz using internal resampling logic; 16 kHz passes through unchanged
- Frame splitting — Audio is chunked into 512-sample frames (32 ms at 16 kHz) for the VAD (Voice Activity Detector)
The 512-sample frame size is hardcoded for low-latency processing. As noted in the Realtime Engine documentation:
"Inbound audio: Client sends
input_audio_buffer.appendwith base64 PCM. … Resampled to 16 kHz … split into 512-sample frames for the VAD."
Output Audio Format: What the Server Sends Back
Synthesized speech follows the same PCM specification as the internal processing:
- 16 kHz sample rate
- Mono channel
- 16-bit signed little-endian
The TTS handler writes PCM chunks to send_audio_chunks_queue, and the WebSocket router encodes each chunk as a response.output_audio.delta event. This design ensures zero additional resampling overhead on the output path.
Output Audio Flow
qwen3_tts_handler.pygenerates 16 kHz PCM byteswebsocket_streamer.pypulls fromsend_audio_chunks_queue- Base64 encoding wraps the PCM for JSON transport
- Client receives
response.output_audio.deltaevents
Why 16 kHz? The Design Rationale
The pipeline standardizes on 16 kHz for all internal processing (VAD → STT → LLM → TTS). This sampling rate provides:
- Low latency — Smaller buffer sizes and faster inference
- Adequate quality — Sufficient for voice assistant intelligibility
- Model compatibility — Matches the training data of most open-source speech models
The optional 24 kHz input support exists solely for compatibility with the official OpenAI Realtime schema. The server handles the conversion transparently, so developers can use either rate without changing client logic.
Code Examples: Working with PCM Audio
Recording and Sending 16 kHz PCM
import sounddevice as sd
import base64
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8765/v1",
websocket_base_url="ws://localhost:8765/v1",
api_key="unused",
)
def send_pcm_chunk(chunk: bytes):
"""Send 512-sample 16-kHz PCM chunk (16-bit signed little-endian)."""
b64 = base64.b64encode(chunk).decode()
client.realtime.connect(model="local").send({
"type": "input_audio_buffer.append",
"audio": b64
})
# Record 32ms chunks: 512 samples @ 16 kHz, 16-bit = 1024 bytes
with sd.RawInputStream(
samplerate=16000,
channels=1,
dtype="int16",
blocksize=512
) as stream:
while True:
data = stream.read(512)[0]
send_pcm_chunk(data)
Key parameters in this snippet:
samplerate=16000— matches the pipeline's internal ratedtype="int16"— 16-bit signed PCMblocksize=512— aligns with the VAD frame size
Receiving and Playing Synthesized Audio
import base64
import numpy as np
import sounddevice as sd
def play_pcm_chunk(b64_audio: str):
"""Decode base64 PCM and play at 16 kHz."""
raw = base64.b64decode(b64_audio)
samples = np.frombuffer(raw, dtype=np.int16)
sd.play(samples, samplerate=16000, blocking=False)
with client.realtime.connect(model="local") as conn:
for event in conn:
if event["type"] == "response.output_audio.delta":
play_pcm_chunk(event["audio"])
Starting the Realtime Server
speech-to-speech --mode realtime \
--stt parakeet-tdt \
--llm_backend responses-api \
--tts qwen3
Implementation Details: Where Formats Are Handled
| Component | File Path | Role |
|---|---|---|
| Protocol specification | src/speech_to_speech/api/openai_realtime/README.md |
Documents input_audio_buffer.append and response.output_audio.delta formats |
| WebSocket audio routing | src/speech_to_speech/connections/websocket_streamer.py |
Decodes base64, resamples 24→16 kHz, chunks to 512 samples |
| TTS output generation | src/speech_to_speech/TTS/qwen3_tts_handler.py |
Produces 16 kHz PCM for send_audio_chunks_queue |
| Client reference | scripts/listen_and_play_realtime.py |
Full client implementation with correct PCM handling |
| Inter-component queues | src/speech_to_speech/pipeline/messages.py |
Defines recv_audio_chunks_queue and send_audio_chunks_queue for 16 kHz PCM transport |
The queue definitions in messages.py enforce the 16 kHz convention throughout the pipeline. All handlers read from and write to these queues, ensuring consistent audio format across VAD, speech recognition, language model, and text-to-speech components.
Summary
- Input format: 16 kHz or 24 kHz mono 16-bit PCM, base64-encoded in
input_audio_buffer.appendevents - Output format: 16 kHz mono 16-bit PCM, base64-encoded in
response.output_audio.deltaevents - Internal standard: 16 kHz for all processing, with automatic resampling of 24 kHz input
- Frame size: 512 samples (32 ms) for VAD processing
- Transport: WebSocket with JSON events containing base64 audio data
Frequently Asked Questions
Does the server support MP3 or WAV file uploads?
No. The speech-to-speech pipeline only accepts raw PCM audio over the Realtime WebSocket protocol. File-based ingestion would require client-side conversion to the streaming PCM format described above.
What happens if I send 44.1 kHz or 48 kHz audio?
The server does not automatically resample arbitrary rates. Send 16 kHz or 24 kHz PCM for guaranteed compatibility. Other sample rates may cause pitch distortion or processing errors in the VAD stage.
Can I change the internal 16 kHz processing rate?
No. The 16 kHz rate is baked into multiple components: the VAD frame size (512 samples), STT model configurations, and TTS output generation. Changing it would require modifications across websocket_streamer.py, messages.py, and all handler implementations.
Why base64 encoding instead of binary WebSocket frames?
The OpenAI Realtime protocol specification uses JSON events with base64-encoded audio fields. The huggingface/speech-to-speech implementation maintains strict compatibility with this schema, allowing drop-in replacement of OpenAI's hosted service with your local server.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →