Supported Output Audio Formats in Hugging Face Speech-to-Speech

The Hugging Face speech-to-speech repository supports two output audio formats: WAV (audio/wav) and raw PCM (audio/pcm), both operating at 16 kHz with 16-bit little-endian encoding.

The Hugging Face speech-to-speech repository implements an OpenAI Realtime-compatible server for streaming synthesized audio. When configuring audio output, developers must select one of the supported output audio formats to ensure proper decoding by client applications.

Where Output Audio Formats Are Defined

The supported output audio formats are hardcoded in the runtime configuration builder, CLI scripts, and test utilities that interface with the OpenAI Realtime API.

Runtime Configuration Builder

In src/speech_to_speech/api/openai_realtime/runtime_config.py, the RealtimeAudioConfig is constructed with the output format dictionary containing the type and rate keys.

CLI Argument Mapping

In scripts/listen_and_play_realtime.py, the helper function maybe_pcm_format constructs the format dictionary for PCM output:

def maybe_pcm_format(rate: int) -> Optional[dict]:
    # Our local pipeline defaults to 16 kHz internally when format is omitted,

    # but we also support the OpenAI realtime audio format schema.

    if rate == 16000:
        return {"type": "audio/pcm", "rate": 16000}

When the user specifies --file-format=WAVE, the script sets the WAV format:

if args.file_format.upper() == "WAVE":
    output_cfg["format"] = {"type": "audio/wav", "rate": 16000}

Test Fixtures and Validation

In tests/openai_realtime/conftest.py, the AudioPCM model validates the PCM format:

from openai.types.realtime.realtime_audio_formats import AudioPCM

fmt = AudioPCM.model_construct(rate=16000, type="audio/pcm")

In tests/openai_realtime/test_realtime_service.py, the test_session_update_nested_audio_format test verifies that the server correctly propagates the chosen output format in its session responses.

Configuring WAV Output

The WAV format (audio/wav) wraps PCM data in a standard container header. This format is suitable for direct file writing or playback in standard media players.

To request WAV output via CLI:

python -m speech_to_speech.scripts.listen_and_play_realtime \
    --recv-rate 16000 \
    --file-format=WAVE

Configuring PCM Output

The PCM format (audio/pcm) streams raw 16-bit little-endian integer samples without container overhead. This minimizes latency for real-time streaming applications but requires the client to interpret the raw byte stream using the known 16 kHz sample rate.

To request PCM output via CLI:

python -m speech_to_speech.scripts.listen_and_play_realtime \
    --recv-rate 16000 \
    --data-format=LEI16@16000

Programmatic Format Construction

For Python applications, construct format objects using the OpenAI Realtime types and pass them to the service configuration:

from openai.types.realtime.realtime_audio_config_output import RealtimeAudioConfigOutput

# WAV output configuration

wav_cfg = RealtimeAudioConfigOutput(
    format={"type": "audio/wav", "rate": 16000}
)

# PCM output configuration

pcm_cfg = RealtimeAudioConfigOutput(
    format={"type": "audio/pcm", "rate": 16000}
)

Technical Specifications

Both supported output audio formats share identical audio characteristics:

  • Sample rate: 16 kHz
  • Bit depth: 16-bit
  • Encoding: Little-endian signed integers (LEI16)
  • Channels: Mono

The internal pipeline processes audio at 16 kHz regardless of format selection. The choice between WAV and PCM only affects whether the server includes a container wrapper (WAV) or transmits raw samples (PCM).

Summary

  • The Hugging Face speech-to-speech repository supports WAV (audio/wav) and PCM (audio/pcm) output formats.
  • Both formats use 16 kHz, 16-bit little-endian mono encoding.
  • Configure formats via CLI arguments in scripts/listen_and_play_realtime.py or programmatically via RealtimeAudioConfigOutput.
  • WAV includes a container header suitable for file writing; PCM provides raw byte streams for low-latency applications.

Frequently Asked Questions

What is the difference between WAV and PCM output in speech-to-speech?

WAV output (audio/wav) includes a standard WAV file header followed by PCM data, making it suitable for writing directly to .wav files. PCM output (audio/pcm) transmits raw 16-bit little-endian samples without headers, requiring the client to know the sample rate (16 kHz) and bit depth (16-bit) to interpret the stream correctly.

Can I use a sample rate other than 16 kHz for output?

No. The internal pipeline in scripts/listen_and_play_realtime.py and the runtime configuration enforce 16 kHz. The maybe_pcm_format function explicitly checks for rate == 16000, and the CLI arguments default to this rate. Attempting to use other rates will result in configuration errors or fallback to the default 16 kHz processing.

How do I switch between WAV and PCM in my application?

Pass --file-format=WAVE for WAV output or --data-format=LEI16@16000 for PCM output when running listen_and_play_realtime.py. Programmatically, instantiate RealtimeAudioConfigOutput with format={"type": "audio/wav", "rate": 16000} or format={"type": "audio/pcm", "rate": 16000} and pass this configuration to the Realtime service session.

Where is the output format validated in the test suite?

The tests/openai_realtime/test_realtime_service.py file contains test_session_update_nested_audio_format, which verifies that the server correctly mirrors the selected output format in its session responses. Additionally, tests/openai_realtime/conftest.py provides the AudioPCM fixture used to validate PCM format construction across the test suite.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →