How the Speech-to-Speech Pipeline Handles Audio Format Conversion from Client PCM to 16kHz
The Speech-to-Speech pipeline automatically resamples any client PCM audio to 16 kHz using TorchAudio's Resample transform in the VADHandler before processing, and down-samples TTS output back to 16 kHz before transmission.
The huggingface/speech-to-speech repository implements a real-time voice conversation pipeline that standardizes all audio to 16 kHz, 16-bit PCM regardless of the client's original sampling rate. Whether audio arrives via raw TCP sockets or WebSocket connections, the pipeline automatically resamples and normalizes the stream before speech-to-text processing begins.
Receiving Client PCM Streams
The pipeline accepts raw PCM bytes through two primary connection handlers. Each wraps incoming audio in an AudioInItem dataclass that preserves the client's original sample rate.
SocketReceiver for TCP Connections
In src/speech_to_speech/connections/socket_receiver.py, the receiver reads raw bytes and constructs an AudioInItem containing the audio buffer and its declared sample rate (often 8 kHz or 44.1 kHz). This occurs around line 30, where the receiver loops through incoming socket data and packages it for the processing queue.
WebSocketStreamer for WebSocket Connections
Similarly, src/speech_to_speech/connections/websocket_streamer.py handles WebSocket clients. At approximately line 154, incoming messages are converted to AudioInItem instances, preserving the source sample rate metadata before passing the object downstream.
Resampling to 16 kHz in the VADHandler
The VADHandler serves as the gatekeeper for audio format validation. Located in src/speech_to_speech/VAD/vad_handler.py, this component checks the sample_rate field of each AudioInItem.
If the rate differs from 16000 Hz, the handler invokes the resample_audio utility from src/speech_to_speech/utils/utils.py (around line 120). This function utilizes TorchAudio's Resample transform for standard PyTorch backends, or the MLX-based resampler when running on Apple Silicon, converting the waveform to exactly 16 kHz.
Preparing Audio for STT Processing
After resampling, the audio requires normalization for the Whisper-based speech-to-text models. In src/speech_to_speech/STT/whisper_stt_handler.py (approximately line 85), the pipeline converts the 16-bit PCM data to float32 format and normalizes values to the [-1, 1] range. This standardization ensures compatibility across different STT implementations including WhisperSTTHandler and MLXAudioWhisperSTTHandler.
Down-Sampling TTS Output to 16 kHz
Text-to-speech handlers in the pipeline generate audio at 24 kHz by default. Before returning audio to the client, each handler must down-sample to the required 16 kHz format.
The pocket_tts_handler.py (around line 140), kokoro_handler.py (around line 330), and qwen3_tts_handler.py (around line 210) all call the same resample_audio utility, specifying a source rate of 24000 Hz and target rate of 16000 Hz. This ensures consistent output sampling regardless of the specific TTS model used.
Practical Code Examples
The following example demonstrates how an 8 kHz client stream is processed:
# Client sends 8 kHz PCM over TCP
# SocketReceiver creates AudioInItem(sample_rate=8000, audio=buffer)
# VADHandler detects mismatch and calls:
from speech_to_speech.utils.utils import resample_audio
resampled = resample_audio(audio_tensor, orig_freq=8000, new_freq=16000)
This example shows TTS output processing before transmission:
# PocketTTSHandler generates 24 kHz audio
# Before sending to client, it down-samples:
final_audio = resample_audio(generated_audio, src_rate=24000, tgt_rate=16000)
Summary
- The pipeline receives raw PCM through
SocketReceiverorWebSocketStreamer, wrapping data inAudioInItemobjects that preserve original sample rates. - The
VADHandlervalidates incoming rates and invokesresample_audiofromutils.pyto convert any rate to 16 kHz using TorchAudio. - STT handlers normalize resampled audio to float32 [-1, 1] range before transcription.
- TTS handlers generate at 24 kHz but down-sample to 16 kHz using the same utility before transmission.
- This architecture guarantees consistent 16 kHz, 16-bit PCM throughout the entire pipeline.
Frequently Asked Questions
What happens if a client sends audio at 44.1 kHz?
The VADHandler detects that the sample_rate field in the AudioInItem does not equal 16000. It automatically calls resample_audio with orig_freq=44100 and new_freq=16000, converting the stream to the required format before any VAD or STT processing occurs.
Does the pipeline support different bit depths for PCM input?
The pipeline expects 16-bit PCM from clients. While the connection handlers receive raw bytes, the normalization step in STT handlers converts these to float32. Any bit depth conversion from 8-bit or 24-bit must be handled client-side before transmission, as the resample_audio function assumes 16-bit input tensors.
Why does the TTS generate at 24 kHz only to down-sample to 16 kHz?
Higher generation rates (24 kHz) improve speech quality and model fidelity during synthesis. The subsequent down-sampling to 16 kHz in pocket_tts_handler.py and other TTS modules ensures network bandwidth consistency and matches the pipeline's standardized input format, simplifying the client-side audio playback requirements.
Can I use a different target sample rate than 16 kHz?
The 16 kHz requirement is hardcoded throughout the pipeline, specifically in the VADHandler validation logic and the resample_audio calls in TTS handlers. Modifying this would require changes to src/speech_to_speech/VAD/vad_handler.py, src/speech_to_speech/utils/utils.py, and all TTS handler files to update the target frequency constants.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →