Speech-to-Speech Input and Output Formats: Complete Audio Specification for the Hugging Face Realtime API

The speech-to-speech pipeline accepts and produces raw linear PCM audio at 16 kHz, mono, 16-bit signed little-endian, transmitted as base64-encoded chunks over the OpenAI-compatible Realtime protocol.

The huggingface/speech-to-speech repository implements a fully local speech-to-speech system that mirrors the OpenAI Realtime API. Understanding the exact audio input and output formats is essential for building compatible clients. This guide covers the PCM specifications, transport encoding, and where each conversion happens in the codebase.


Input Audio Format: What the Server Expects

The RealtimeService accepts audio through the input_audio_buffer.append event. Two sampling rates are supported:

Parameter Default Alternative
Sample rate 16 kHz 24 kHz (auto-resampled)
Channels Mono (1) —
Bit depth 16-bit signed —
Endianness Little-endian —
Transport Base64-encoded PCM chunks —

How Input Audio Is Processed

When your client sends audio to the server, the following happens in src/speech_to_speech/api/openai_realtime/README.md:

  1. Base64 decoding — The audio field in input_audio_buffer.append is decoded from base64 to raw bytes
  2. Resampling — 24 kHz input is resampled to 16 kHz using internal resampling logic; 16 kHz passes through unchanged
  3. Frame splitting — Audio is chunked into 512-sample frames (32 ms at 16 kHz) for the VAD (Voice Activity Detector)

The 512-sample frame size is hardcoded for low-latency processing. As noted in the Realtime Engine documentation:

"Inbound audio: Client sends input_audio_buffer.append with base64 PCM. … Resampled to 16 kHz … split into 512-sample frames for the VAD."


Output Audio Format: What the Server Sends Back

Synthesized speech follows the same PCM specification as the internal processing:

  • 16 kHz sample rate
  • Mono channel
  • 16-bit signed little-endian

The TTS handler writes PCM chunks to send_audio_chunks_queue, and the WebSocket router encodes each chunk as a response.output_audio.delta event. This design ensures zero additional resampling overhead on the output path.

Output Audio Flow

  1. qwen3_tts_handler.py generates 16 kHz PCM bytes
  2. websocket_streamer.py pulls from send_audio_chunks_queue
  3. Base64 encoding wraps the PCM for JSON transport
  4. Client receives response.output_audio.delta events

Why 16 kHz? The Design Rationale

The pipeline standardizes on 16 kHz for all internal processing (VAD → STT → LLM → TTS). This sampling rate provides:

  • Low latency — Smaller buffer sizes and faster inference
  • Adequate quality — Sufficient for voice assistant intelligibility
  • Model compatibility — Matches the training data of most open-source speech models

The optional 24 kHz input support exists solely for compatibility with the official OpenAI Realtime schema. The server handles the conversion transparently, so developers can use either rate without changing client logic.


Code Examples: Working with PCM Audio

Recording and Sending 16 kHz PCM

import sounddevice as sd
import base64
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="unused",
)

def send_pcm_chunk(chunk: bytes):
    """Send 512-sample 16-kHz PCM chunk (16-bit signed little-endian)."""
    b64 = base64.b64encode(chunk).decode()
    client.realtime.connect(model="local").send({
        "type": "input_audio_buffer.append",
        "audio": b64
    })

# Record 32ms chunks: 512 samples @ 16 kHz, 16-bit = 1024 bytes

with sd.RawInputStream(
    samplerate=16000,
    channels=1,
    dtype="int16",
    blocksize=512
) as stream:
    while True:
        data = stream.read(512)[0]
        send_pcm_chunk(data)

Key parameters in this snippet:

  • samplerate=16000 — matches the pipeline's internal rate
  • dtype="int16" — 16-bit signed PCM
  • blocksize=512 — aligns with the VAD frame size

Receiving and Playing Synthesized Audio

import base64
import numpy as np
import sounddevice as sd

def play_pcm_chunk(b64_audio: str):
    """Decode base64 PCM and play at 16 kHz."""
    raw = base64.b64decode(b64_audio)
    samples = np.frombuffer(raw, dtype=np.int16)
    sd.play(samples, samplerate=16000, blocking=False)

with client.realtime.connect(model="local") as conn:
    for event in conn:
        if event["type"] == "response.output_audio.delta":
            play_pcm_chunk(event["audio"])

Starting the Realtime Server

speech-to-speech --mode realtime \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3

Implementation Details: Where Formats Are Handled

Component File Path Role
Protocol specification src/speech_to_speech/api/openai_realtime/README.md Documents input_audio_buffer.append and response.output_audio.delta formats
WebSocket audio routing src/speech_to_speech/connections/websocket_streamer.py Decodes base64, resamples 24→16 kHz, chunks to 512 samples
TTS output generation src/speech_to_speech/TTS/qwen3_tts_handler.py Produces 16 kHz PCM for send_audio_chunks_queue
Client reference scripts/listen_and_play_realtime.py Full client implementation with correct PCM handling
Inter-component queues src/speech_to_speech/pipeline/messages.py Defines recv_audio_chunks_queue and send_audio_chunks_queue for 16 kHz PCM transport

The queue definitions in messages.py enforce the 16 kHz convention throughout the pipeline. All handlers read from and write to these queues, ensuring consistent audio format across VAD, speech recognition, language model, and text-to-speech components.


Summary

  • Input format: 16 kHz or 24 kHz mono 16-bit PCM, base64-encoded in input_audio_buffer.append events
  • Output format: 16 kHz mono 16-bit PCM, base64-encoded in response.output_audio.delta events
  • Internal standard: 16 kHz for all processing, with automatic resampling of 24 kHz input
  • Frame size: 512 samples (32 ms) for VAD processing
  • Transport: WebSocket with JSON events containing base64 audio data

Frequently Asked Questions

Does the server support MP3 or WAV file uploads?

No. The speech-to-speech pipeline only accepts raw PCM audio over the Realtime WebSocket protocol. File-based ingestion would require client-side conversion to the streaming PCM format described above.

What happens if I send 44.1 kHz or 48 kHz audio?

The server does not automatically resample arbitrary rates. Send 16 kHz or 24 kHz PCM for guaranteed compatibility. Other sample rates may cause pitch distortion or processing errors in the VAD stage.

Can I change the internal 16 kHz processing rate?

No. The 16 kHz rate is baked into multiple components: the VAD frame size (512 samples), STT model configurations, and TTS output generation. Changing it would require modifications across websocket_streamer.py, messages.py, and all handler implementations.

Why base64 encoding instead of binary WebSocket frames?

The OpenAI Realtime protocol specification uses JSON events with base64-encoded audio fields. The huggingface/speech-to-speech implementation maintains strict compatibility with this schema, allowing drop-in replacement of OpenAI's hosted service with your local server.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →