# Speech-to-Speech Pipeline Audio Formats and Sample Rates: Complete Input/Output Guide

> Learn the Hugging Face speech-to-speech pipeline's audio format requirements. Understand the expected 16-bit int16 mono audio at 16 kHz for seamless input and output.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: api-reference
- Published: 2026-07-30

---

**The Hugging Face speech-to-speech pipeline strictly expects 16-bit signed integer PCM (`int16`) mono audio at a 16 kHz sample rate for all input and output operations.**

All stages of the `huggingface/speech-to-speech` pipeline are hardcoded to operate at **16 kHz** with **16-bit PCM encoding**. Whether you are streaming audio from a microphone, receiving data over a WebSocket, or generating speech through the TTS module, the system enforces these specifications through the `PIPELINE_SR = 16000` constant defined in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py).

## Input Audio Requirements

The pipeline accepts audio from multiple sources—local microphones, socket connections, and WebSocket streams—but normalizes everything to the same internal format before processing.

### Microphone and Local Audio Streamer

When capturing audio locally, the system initializes the audio interface with rigid parameters. In [`src/speech_to_speech/connections/local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/local_audio_streamer.py), the `sd.Stream` is configured with:

```python
sd.Stream(samplerate=16000, dtype="int16", channels=1)

```

This ensures that raw audio enters the pipeline as contiguous blocks of **16-bit integers at 16 kHz** in mono configuration. The VAD (Voice Activity Detection) handler arguments in [`src/speech_to_speech/arguments_classes/vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/arguments_classes/vad_arguments.py) reinforce this by defaulting `sample_rate` to `16000`, preventing downstream handlers from processing mismatched sample rates.

### WebSocket Input Format

For network-based input, the WebSocket streamer expects raw PCM bytes formatted as **16-bit signed integers in little-endian byte order**. The implementation in [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) specifically processes incoming data in chunks of **512 samples** (exactly **1024 bytes**):

```python

# From websocket_streamer.py - splits received bytes into 512-sample chunks

audio_chunk = np.frombuffer(data, dtype=np.int16)

```

Any client sending audio to the pipeline must stream `int16` PCM at 16 kHz and should structure transmissions to align with these 1024-byte boundaries for optimal latency.

## Internal Pipeline Processing

Once inside the pipeline, audio maintains a consistent format across all intermediate queues. The global constant `PIPELINE_SR = 16000` defined at line 49-50 of [`src/speech_to_speech/s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/s2s_pipeline.py) serves as the single source of truth for the processing sample rate.

Audio items flow through the system as either:
- Raw `bytes` objects containing PCM data
- `np.ndarray` objects with `dtype=np.int16`

This standardization eliminates format conversion overhead between pipeline stages, allowing the VAD, LLM, and TTS handlers to operate on homogeneous data structures.

## Output Audio Specifications

Text-to-Speech handlers generate audio at various native sample rates depending on the underlying model, but all outputs are resampled to the pipeline standard before reaching the output queue.

### TTS Handler Output

Different TTS implementations handle the conversion to 16 kHz internally:

- **Qwen3-TTS**: The handler in [`src/speech_to_speech/TTS/qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/qwen3_tts_handler.py) rescales generated audio to `PIPELINE_SR` (lines 69-71), ensuring the output matches the pipeline's 16 kHz requirement regardless of the model's internal generation rate.

- **Pocket-TTS**: This handler generates audio at a native **24 kHz** but explicitly resamples to 16 kHz before yielding. The implementation in [`src/speech_to_speech/TTS/pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/TTS/pocket_tts_handler.py) (lines 33-40) converts the higher-rate output to `int16` PCM at the pipeline rate:

```python

# Pocket-TTS resampling logic

if self.sample_rate != PIPELINE_SR:
    audio = resample(audio, orig_sr=self.sample_rate, target_sr=PIPELINE_SR)

```

### WebSocket Output Format

When transmitting audio back to clients, the WebSocket streamer converts any `np.ndarray` or custom audio objects to raw PCM bytes. As implemented in [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) (lines 90-99), the streamer buffers audio until at least **100 ms** of data (approximately **3200 bytes** or **1600 samples**) is ready, then transmits the raw **16-bit little-endian PCM** without headers or compression.

## Practical Implementation Example

To feed microphone audio into the pipeline correctly, configure your audio callback to output the exact format the VAD handler expects:

```python
import sounddevice as sd
import numpy as np
from speech_to_speech.pipeline.queue_types import AudioInItem

# Pipeline constants

SAMPLE_RATE = 16000
BLOCK_SIZE = 512  # Matches VAD chunk size exactly

def audio_callback(indata, frames, time, status):
    """Capture microphone audio in pipeline-compatible format."""
    # Ensure contiguous int16 array

    pcm_array = np.ascontiguousarray(indata, dtype=np.int16)
    
    # Convert to bytes for queue insertion

    pcm_bytes = pcm_array.tobytes()
    
    # Place in input queue (assumes input_queue is defined)

    input_queue.put(pcm_bytes)

# Start audio stream with pipeline specifications

with sd.Stream(
    samplerate=SAMPLE_RATE,
    dtype='int16',
    channels=1,
    blocksize=BLOCK_SIZE,
    callback=audio_callback
):
    sd.sleep(5000)  # Record for 5 seconds

```

This example aligns the microphone's block size (512 samples) with the VAD handler's processing window, minimizing latency while maintaining the required **16 kHz int16 PCM** format throughout the capture chain.

## Summary

- **Input Format**: 16-bit signed integer PCM (`int16`), mono, 16 kHz sample rate
- **WebSocket Input**: Raw PCM bytes in 1024-byte chunks (512 samples of int16)
- **Internal Processing**: All handlers use `PIPELINE_SR = 16000` defined in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py)
- **TTS Output**: Handlers resample from native rates (e.g., 24 kHz) to 16 kHz int16 PCM before yielding
- **WebSocket Output**: Buffered raw PCM bytes (100ms/3200byte chunks) at 16 kHz
- **Critical Files**: [`vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_arguments.py), [`local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/local_audio_streamer.py), [`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py), [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py)

## Frequently Asked Questions

### What happens if I send audio at 44.1 kHz or 48 kHz instead of 16 kHz?

The pipeline does not perform automatic input resampling. Sending audio at sample rates other than 16 kHz will cause the VAD handler to process the data incorrectly, resulting in distorted voice detection or pipeline failures. You must resample external audio to 16 kHz before inserting it into the input queue.

### Does the pipeline support stereo audio or floating-point formats?

No. The [`local_audio_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/local_audio_streamer.py) explicitly configures `channels=1` and `dtype="int16"`. The WebSocket handler expects `np.int16` arrays. Stereo audio or float32 formats will raise shape or dtype errors when handlers attempt to process the buffers.

### Why does the WebSocket streamer use 512-sample chunks?

The 512-sample window (defined in [`websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/websocket_streamer.py)) aligns with the VAD handler's processing granularity, ensuring that voice activity detection can analyze audio in real-time without additional buffering delays. This represents 32 milliseconds of audio at 16 kHz, balancing latency with processing efficiency.

### Can I change the pipeline sample rate by modifying the configuration?

While `PIPELINE_SR` is defined as a constant in [`s2s_pipeline.py`](https://github.com/huggingface/speech-to-speech/blob/main/s2s_pipeline.py), changing it would require updating multiple hardcoded dependencies across [`vad_arguments.py`](https://github.com/huggingface/speech-to-speech/blob/main/vad_arguments.py), the TTS handlers, and the WebSocket streamers. The system is designed around 16 kHz for ASR and TTS model compatibility, and altering this rate is not supported without extensive code modifications.