Supported Output Audio Formats in Hugging Face Speech-to-Speech
The Hugging Face speech-to-speech repository supports two output audio formats: WAV (audio/wav) and raw PCM (audio/pcm), both operating at 16 kHz with 16-bit little-endian encoding.
The Hugging Face speech-to-speech repository implements an OpenAI Realtime-compatible server for streaming synthesized audio. When configuring audio output, developers must select one of the supported output audio formats to ensure proper decoding by client applications.
Where Output Audio Formats Are Defined
The supported output audio formats are hardcoded in the runtime configuration builder, CLI scripts, and test utilities that interface with the OpenAI Realtime API.
Runtime Configuration Builder
In src/speech_to_speech/api/openai_realtime/runtime_config.py, the RealtimeAudioConfig is constructed with the output format dictionary containing the type and rate keys.
CLI Argument Mapping
In scripts/listen_and_play_realtime.py, the helper function maybe_pcm_format constructs the format dictionary for PCM output:
def maybe_pcm_format(rate: int) -> Optional[dict]:
# Our local pipeline defaults to 16 kHz internally when format is omitted,
# but we also support the OpenAI realtime audio format schema.
if rate == 16000:
return {"type": "audio/pcm", "rate": 16000}
When the user specifies --file-format=WAVE, the script sets the WAV format:
if args.file_format.upper() == "WAVE":
output_cfg["format"] = {"type": "audio/wav", "rate": 16000}
Test Fixtures and Validation
In tests/openai_realtime/conftest.py, the AudioPCM model validates the PCM format:
from openai.types.realtime.realtime_audio_formats import AudioPCM
fmt = AudioPCM.model_construct(rate=16000, type="audio/pcm")
In tests/openai_realtime/test_realtime_service.py, the test_session_update_nested_audio_format test verifies that the server correctly propagates the chosen output format in its session responses.
Configuring WAV Output
The WAV format (audio/wav) wraps PCM data in a standard container header. This format is suitable for direct file writing or playback in standard media players.
To request WAV output via CLI:
python -m speech_to_speech.scripts.listen_and_play_realtime \
--recv-rate 16000 \
--file-format=WAVE
Configuring PCM Output
The PCM format (audio/pcm) streams raw 16-bit little-endian integer samples without container overhead. This minimizes latency for real-time streaming applications but requires the client to interpret the raw byte stream using the known 16 kHz sample rate.
To request PCM output via CLI:
python -m speech_to_speech.scripts.listen_and_play_realtime \
--recv-rate 16000 \
--data-format=LEI16@16000
Programmatic Format Construction
For Python applications, construct format objects using the OpenAI Realtime types and pass them to the service configuration:
from openai.types.realtime.realtime_audio_config_output import RealtimeAudioConfigOutput
# WAV output configuration
wav_cfg = RealtimeAudioConfigOutput(
format={"type": "audio/wav", "rate": 16000}
)
# PCM output configuration
pcm_cfg = RealtimeAudioConfigOutput(
format={"type": "audio/pcm", "rate": 16000}
)
Technical Specifications
Both supported output audio formats share identical audio characteristics:
- Sample rate: 16 kHz
- Bit depth: 16-bit
- Encoding: Little-endian signed integers (
LEI16) - Channels: Mono
The internal pipeline processes audio at 16 kHz regardless of format selection. The choice between WAV and PCM only affects whether the server includes a container wrapper (WAV) or transmits raw samples (PCM).
Summary
- The Hugging Face speech-to-speech repository supports WAV (
audio/wav) and PCM (audio/pcm) output formats. - Both formats use 16 kHz, 16-bit little-endian mono encoding.
- Configure formats via CLI arguments in
scripts/listen_and_play_realtime.pyor programmatically viaRealtimeAudioConfigOutput. - WAV includes a container header suitable for file writing; PCM provides raw byte streams for low-latency applications.
Frequently Asked Questions
What is the difference between WAV and PCM output in speech-to-speech?
WAV output (audio/wav) includes a standard WAV file header followed by PCM data, making it suitable for writing directly to .wav files. PCM output (audio/pcm) transmits raw 16-bit little-endian samples without headers, requiring the client to know the sample rate (16 kHz) and bit depth (16-bit) to interpret the stream correctly.
Can I use a sample rate other than 16 kHz for output?
No. The internal pipeline in scripts/listen_and_play_realtime.py and the runtime configuration enforce 16 kHz. The maybe_pcm_format function explicitly checks for rate == 16000, and the CLI arguments default to this rate. Attempting to use other rates will result in configuration errors or fallback to the default 16 kHz processing.
How do I switch between WAV and PCM in my application?
Pass --file-format=WAVE for WAV output or --data-format=LEI16@16000 for PCM output when running listen_and_play_realtime.py. Programmatically, instantiate RealtimeAudioConfigOutput with format={"type": "audio/wav", "rate": 16000} or format={"type": "audio/pcm", "rate": 16000} and pass this configuration to the Realtime service session.
Where is the output format validated in the test suite?
The tests/openai_realtime/test_realtime_service.py file contains test_session_update_nested_audio_format, which verifies that the server correctly mirrors the selected output format in its session responses. Additionally, tests/openai_realtime/conftest.py provides the AudioPCM fixture used to validate PCM format construction across the test suite.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →