Speech-to-Speech Pipeline Audio Formats and Sample Rates: Complete Input/Output Guide

The Hugging Face speech-to-speech pipeline strictly expects 16-bit signed integer PCM (int16) mono audio at a 16 kHz sample rate for all input and output operations.

All stages of the huggingface/speech-to-speech pipeline are hardcoded to operate at 16 kHz with 16-bit PCM encoding. Whether you are streaming audio from a microphone, receiving data over a WebSocket, or generating speech through the TTS module, the system enforces these specifications through the PIPELINE_SR = 16000 constant defined in s2s_pipeline.py.

Input Audio Requirements

The pipeline accepts audio from multiple sources—local microphones, socket connections, and WebSocket streams—but normalizes everything to the same internal format before processing.

Microphone and Local Audio Streamer

When capturing audio locally, the system initializes the audio interface with rigid parameters. In src/speech_to_speech/connections/local_audio_streamer.py, the sd.Stream is configured with:

sd.Stream(samplerate=16000, dtype="int16", channels=1)

This ensures that raw audio enters the pipeline as contiguous blocks of 16-bit integers at 16 kHz in mono configuration. The VAD (Voice Activity Detection) handler arguments in src/speech_to_speech/arguments_classes/vad_arguments.py reinforce this by defaulting sample_rate to 16000, preventing downstream handlers from processing mismatched sample rates.

WebSocket Input Format

For network-based input, the WebSocket streamer expects raw PCM bytes formatted as 16-bit signed integers in little-endian byte order. The implementation in src/speech_to_speech/connections/websocket_streamer.py specifically processes incoming data in chunks of 512 samples (exactly 1024 bytes):


# From websocket_streamer.py - splits received bytes into 512-sample chunks

audio_chunk = np.frombuffer(data, dtype=np.int16)

Any client sending audio to the pipeline must stream int16 PCM at 16 kHz and should structure transmissions to align with these 1024-byte boundaries for optimal latency.

Internal Pipeline Processing

Once inside the pipeline, audio maintains a consistent format across all intermediate queues. The global constant PIPELINE_SR = 16000 defined at line 49-50 of src/speech_to_speech/s2s_pipeline.py serves as the single source of truth for the processing sample rate.

Audio items flow through the system as either:

  • Raw bytes objects containing PCM data
  • np.ndarray objects with dtype=np.int16

This standardization eliminates format conversion overhead between pipeline stages, allowing the VAD, LLM, and TTS handlers to operate on homogeneous data structures.

Output Audio Specifications

Text-to-Speech handlers generate audio at various native sample rates depending on the underlying model, but all outputs are resampled to the pipeline standard before reaching the output queue.

TTS Handler Output

Different TTS implementations handle the conversion to 16 kHz internally:

  • Qwen3-TTS: The handler in src/speech_to_speech/TTS/qwen3_tts_handler.py rescales generated audio to PIPELINE_SR (lines 69-71), ensuring the output matches the pipeline's 16 kHz requirement regardless of the model's internal generation rate.

  • Pocket-TTS: This handler generates audio at a native 24 kHz but explicitly resamples to 16 kHz before yielding. The implementation in src/speech_to_speech/TTS/pocket_tts_handler.py (lines 33-40) converts the higher-rate output to int16 PCM at the pipeline rate:


# Pocket-TTS resampling logic

if self.sample_rate != PIPELINE_SR:
    audio = resample(audio, orig_sr=self.sample_rate, target_sr=PIPELINE_SR)

WebSocket Output Format

When transmitting audio back to clients, the WebSocket streamer converts any np.ndarray or custom audio objects to raw PCM bytes. As implemented in src/speech_to_speech/connections/websocket_streamer.py (lines 90-99), the streamer buffers audio until at least 100 ms of data (approximately 3200 bytes or 1600 samples) is ready, then transmits the raw 16-bit little-endian PCM without headers or compression.

Practical Implementation Example

To feed microphone audio into the pipeline correctly, configure your audio callback to output the exact format the VAD handler expects:

import sounddevice as sd
import numpy as np
from speech_to_speech.pipeline.queue_types import AudioInItem

# Pipeline constants

SAMPLE_RATE = 16000
BLOCK_SIZE = 512  # Matches VAD chunk size exactly

def audio_callback(indata, frames, time, status):
    """Capture microphone audio in pipeline-compatible format."""
    # Ensure contiguous int16 array

    pcm_array = np.ascontiguousarray(indata, dtype=np.int16)
    
    # Convert to bytes for queue insertion

    pcm_bytes = pcm_array.tobytes()
    
    # Place in input queue (assumes input_queue is defined)

    input_queue.put(pcm_bytes)

# Start audio stream with pipeline specifications

with sd.Stream(
    samplerate=SAMPLE_RATE,
    dtype='int16',
    channels=1,
    blocksize=BLOCK_SIZE,
    callback=audio_callback
):
    sd.sleep(5000)  # Record for 5 seconds

This example aligns the microphone's block size (512 samples) with the VAD handler's processing window, minimizing latency while maintaining the required 16 kHz int16 PCM format throughout the capture chain.

Summary

  • Input Format: 16-bit signed integer PCM (int16), mono, 16 kHz sample rate
  • WebSocket Input: Raw PCM bytes in 1024-byte chunks (512 samples of int16)
  • Internal Processing: All handlers use PIPELINE_SR = 16000 defined in s2s_pipeline.py
  • TTS Output: Handlers resample from native rates (e.g., 24 kHz) to 16 kHz int16 PCM before yielding
  • WebSocket Output: Buffered raw PCM bytes (100ms/3200byte chunks) at 16 kHz
  • Critical Files: vad_arguments.py, local_audio_streamer.py, websocket_streamer.py, pocket_tts_handler.py

Frequently Asked Questions

What happens if I send audio at 44.1 kHz or 48 kHz instead of 16 kHz?

The pipeline does not perform automatic input resampling. Sending audio at sample rates other than 16 kHz will cause the VAD handler to process the data incorrectly, resulting in distorted voice detection or pipeline failures. You must resample external audio to 16 kHz before inserting it into the input queue.

Does the pipeline support stereo audio or floating-point formats?

No. The local_audio_streamer.py explicitly configures channels=1 and dtype="int16". The WebSocket handler expects np.int16 arrays. Stereo audio or float32 formats will raise shape or dtype errors when handlers attempt to process the buffers.

Why does the WebSocket streamer use 512-sample chunks?

The 512-sample window (defined in websocket_streamer.py) aligns with the VAD handler's processing granularity, ensuring that voice activity detection can analyze audio in real-time without additional buffering delays. This represents 32 milliseconds of audio at 16 kHz, balancing latency with processing efficiency.

Can I change the pipeline sample rate by modifying the configuration?

While PIPELINE_SR is defined as a constant in s2s_pipeline.py, changing it would require updating multiple hardcoded dependencies across vad_arguments.py, the TTS handlers, and the WebSocket streamers. The system is designed around 16 kHz for ASR and TTS model compatibility, and altering this rate is not supported without extensive code modifications.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →