# How the Speech-to-Speech Pipeline Handles Audio Format Conversion from Client PCM to 16kHz

> Discover how the huggingface speech-to-speech pipeline seamlessly converts client PCM audio to 16kHz using TorchAudio Resample transform before processing and after transmission.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: internals
- Published: 2026-07-10

---

**The Speech-to-Speech pipeline automatically resamples any client PCM audio to 16 kHz using TorchAudio's `Resample` transform in the `VADHandler` before processing, and down-samples TTS output back to 16 kHz before transmission.**

The `huggingface/speech-to-speech` repository implements a real-time voice conversation pipeline that standardizes all audio to **16 kHz, 16-bit PCM** regardless of the client's original sampling rate. Whether audio arrives via raw TCP sockets or WebSocket connections, the pipeline automatically resamples and normalizes the stream before speech-to-text processing begins.

## Receiving Client PCM Streams

The pipeline accepts raw PCM bytes through two primary connection handlers. Each wraps incoming audio in an `AudioInItem` dataclass that preserves the client's original sample rate.

### SocketReceiver for TCP Connections

In [`src/speech_to_speech/connections/socket_receiver.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/socket_receiver.py), the receiver reads raw bytes and constructs an `AudioInItem` containing the audio buffer and its declared sample rate (often 8 kHz or 44.1 kHz). This occurs around line 30, where the receiver loops through incoming socket data and packages it for the processing queue.

### WebSocketStreamer for WebSocket Connections

Similarly, [`src/speech_to_speech/connections/websocket_streamer.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/connections/websocket_streamer.py) handles WebSocket clients. At approximately line 154, incoming messages are converted to `AudioInItem` instances, preserving the source sample rate metadata before passing the object downstream.

## Resampling to 16 kHz in the VADHandler

The `VADHandler` serves as the gatekeeper for audio format validation. Located in [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py), this component checks the `sample_rate` field of each `AudioInItem`.

If the rate differs from 16000 Hz, the handler invokes the `resample_audio` utility from [`src/speech_to_speech/utils/utils.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/utils.py) (around line 120). This function utilizes **TorchAudio's `Resample`** transform for standard PyTorch backends, or the MLX-based resampler when running on Apple Silicon, converting the waveform to exactly 16 kHz.

## Preparing Audio for STT Processing

After resampling, the audio requires normalization for the Whisper-based speech-to-text models. In [`src/speech_to_speech/STT/whisper_stt_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/STT/whisper_stt_handler.py) (approximately line 85), the pipeline converts the 16-bit PCM data to **float32** format and normalizes values to the [-1, 1] range. This standardization ensures compatibility across different STT implementations including `WhisperSTTHandler` and `MLXAudioWhisperSTTHandler`.

## Down-Sampling TTS Output to 16 kHz

Text-to-speech handlers in the pipeline generate audio at **24 kHz** by default. Before returning audio to the client, each handler must down-sample to the required 16 kHz format.

The [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py) (around line 140), [`kokoro_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/kokoro_handler.py) (around line 330), and [`qwen3_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/qwen3_tts_handler.py) (around line 210) all call the same `resample_audio` utility, specifying a source rate of 24000 Hz and target rate of 16000 Hz. This ensures consistent output sampling regardless of the specific TTS model used.

## Practical Code Examples

The following example demonstrates how an 8 kHz client stream is processed:

```python

# Client sends 8 kHz PCM over TCP

# SocketReceiver creates AudioInItem(sample_rate=8000, audio=buffer)

# VADHandler detects mismatch and calls:

from speech_to_speech.utils.utils import resample_audio
resampled = resample_audio(audio_tensor, orig_freq=8000, new_freq=16000)

```

This example shows TTS output processing before transmission:

```python

# PocketTTSHandler generates 24 kHz audio

# Before sending to client, it down-samples:

final_audio = resample_audio(generated_audio, src_rate=24000, tgt_rate=16000)

```

## Summary

- The pipeline receives raw PCM through `SocketReceiver` or `WebSocketStreamer`, wrapping data in `AudioInItem` objects that preserve original sample rates.
- The `VADHandler` validates incoming rates and invokes `resample_audio` from [`utils.py`](https://github.com/huggingface/speech-to-speech/blob/main/utils.py) to convert any rate to 16 kHz using TorchAudio.
- STT handlers normalize resampled audio to float32 [-1, 1] range before transcription.
- TTS handlers generate at 24 kHz but down-sample to 16 kHz using the same utility before transmission.
- This architecture guarantees consistent **16 kHz, 16-bit PCM** throughout the entire pipeline.

## Frequently Asked Questions

### What happens if a client sends audio at 44.1 kHz?

The `VADHandler` detects that the `sample_rate` field in the `AudioInItem` does not equal 16000. It automatically calls `resample_audio` with `orig_freq=44100` and `new_freq=16000`, converting the stream to the required format before any VAD or STT processing occurs.

### Does the pipeline support different bit depths for PCM input?

The pipeline expects 16-bit PCM from clients. While the connection handlers receive raw bytes, the normalization step in STT handlers converts these to float32. Any bit depth conversion from 8-bit or 24-bit must be handled client-side before transmission, as the `resample_audio` function assumes 16-bit input tensors.

### Why does the TTS generate at 24 kHz only to down-sample to 16 kHz?

Higher generation rates (24 kHz) improve speech quality and model fidelity during synthesis. The subsequent down-sampling to 16 kHz in [`pocket_tts_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/pocket_tts_handler.py) and other TTS modules ensures network bandwidth consistency and matches the pipeline's standardized input format, simplifying the client-side audio playback requirements.

### Can I use a different target sample rate than 16 kHz?

The 16 kHz requirement is hardcoded throughout the pipeline, specifically in the `VADHandler` validation logic and the `resample_audio` calls in TTS handlers. Modifying this would require changes to [`src/speech_to_speech/VAD/vad_handler.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/VAD/vad_handler.py), [`src/speech_to_speech/utils/utils.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/utils/utils.py), and all TTS handler files to update the target frequency constants.