# Supported Output Audio Formats in Hugging Face Speech-to-Speech

> Discover supported output audio formats in Hugging Face Speech-to-Speech. Learn about WAV and raw PCM outputs with 16 kHz 16-bit little-endian encoding for your audio projects.

- Repository: [Hugging Face/speech-to-speech](https://github.com/huggingface/speech-to-speech)
- Tags: api-reference
- Published: 2026-07-07

---

**The Hugging Face speech-to-speech repository supports two output audio formats: WAV (`audio/wav`) and raw PCM (`audio/pcm`), both operating at 16 kHz with 16-bit little-endian encoding.**

The Hugging Face **speech-to-speech** repository implements an OpenAI Realtime-compatible server for streaming synthesized audio. When configuring audio output, developers must select one of the supported output audio formats to ensure proper decoding by client applications.

## Where Output Audio Formats Are Defined

The supported output audio formats are hardcoded in the runtime configuration builder, CLI scripts, and test utilities that interface with the OpenAI Realtime API.

### Runtime Configuration Builder

In [`src/speech_to_speech/api/openai_realtime/runtime_config.py`](https://github.com/huggingface/speech-to-speech/blob/main/src/speech_to_speech/api/openai_realtime/runtime_config.py), the `RealtimeAudioConfig` is constructed with the output format dictionary containing the `type` and `rate` keys.

### CLI Argument Mapping

In [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py), the helper function `maybe_pcm_format` constructs the format dictionary for PCM output:

```python
def maybe_pcm_format(rate: int) -> Optional[dict]:
    # Our local pipeline defaults to 16 kHz internally when format is omitted,

    # but we also support the OpenAI realtime audio format schema.

    if rate == 16000:
        return {"type": "audio/pcm", "rate": 16000}

```

When the user specifies `--file-format=WAVE`, the script sets the WAV format:

```python
if args.file_format.upper() == "WAVE":
    output_cfg["format"] = {"type": "audio/wav", "rate": 16000}

```

### Test Fixtures and Validation

In [`tests/openai_realtime/conftest.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/conftest.py), the `AudioPCM` model validates the PCM format:

```python
from openai.types.realtime.realtime_audio_formats import AudioPCM

fmt = AudioPCM.model_construct(rate=16000, type="audio/pcm")

```

In [`tests/openai_realtime/test_realtime_service.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_realtime_service.py), the `test_session_update_nested_audio_format` test verifies that the server correctly propagates the chosen output format in its session responses.

## Configuring WAV Output

The **WAV** format (`audio/wav`) wraps PCM data in a standard container header. This format is suitable for direct file writing or playback in standard media players.

To request WAV output via CLI:

```bash
python -m speech_to_speech.scripts.listen_and_play_realtime \
    --recv-rate 16000 \
    --file-format=WAVE

```

## Configuring PCM Output

The **PCM** format (`audio/pcm`) streams raw 16-bit little-endian integer samples without container overhead. This minimizes latency for real-time streaming applications but requires the client to interpret the raw byte stream using the known 16 kHz sample rate.

To request PCM output via CLI:

```bash
python -m speech_to_speech.scripts.listen_and_play_realtime \
    --recv-rate 16000 \
    --data-format=LEI16@16000

```

## Programmatic Format Construction

For Python applications, construct format objects using the OpenAI Realtime types and pass them to the service configuration:

```python
from openai.types.realtime.realtime_audio_config_output import RealtimeAudioConfigOutput

# WAV output configuration

wav_cfg = RealtimeAudioConfigOutput(
    format={"type": "audio/wav", "rate": 16000}
)

# PCM output configuration

pcm_cfg = RealtimeAudioConfigOutput(
    format={"type": "audio/pcm", "rate": 16000}
)

```

## Technical Specifications

Both supported output audio formats share identical audio characteristics:

- **Sample rate**: 16 kHz
- **Bit depth**: 16-bit
- **Encoding**: Little-endian signed integers (`LEI16`)
- **Channels**: Mono

The internal pipeline processes audio at 16 kHz regardless of format selection. The choice between WAV and PCM only affects whether the server includes a container wrapper (WAV) or transmits raw samples (PCM).

## Summary

- The Hugging Face **speech-to-speech** repository supports **WAV** (`audio/wav`) and **PCM** (`audio/pcm`) output formats.
- Both formats use 16 kHz, 16-bit little-endian mono encoding.
- Configure formats via CLI arguments in [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) or programmatically via `RealtimeAudioConfigOutput`.
- WAV includes a container header suitable for file writing; PCM provides raw byte streams for low-latency applications.

## Frequently Asked Questions

### What is the difference between WAV and PCM output in speech-to-speech?

WAV output (`audio/wav`) includes a standard WAV file header followed by PCM data, making it suitable for writing directly to `.wav` files. PCM output (`audio/pcm`) transmits raw 16-bit little-endian samples without headers, requiring the client to know the sample rate (16 kHz) and bit depth (16-bit) to interpret the stream correctly.

### Can I use a sample rate other than 16 kHz for output?

No. The internal pipeline in [`scripts/listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/scripts/listen_and_play_realtime.py) and the runtime configuration enforce 16 kHz. The `maybe_pcm_format` function explicitly checks for `rate == 16000`, and the CLI arguments default to this rate. Attempting to use other rates will result in configuration errors or fallback to the default 16 kHz processing.

### How do I switch between WAV and PCM in my application?

Pass `--file-format=WAVE` for WAV output or `--data-format=LEI16@16000` for PCM output when running [`listen_and_play_realtime.py`](https://github.com/huggingface/speech-to-speech/blob/main/listen_and_play_realtime.py). Programmatically, instantiate `RealtimeAudioConfigOutput` with `format={"type": "audio/wav", "rate": 16000}` or `format={"type": "audio/pcm", "rate": 16000}` and pass this configuration to the Realtime service session.

### Where is the output format validated in the test suite?

The [`tests/openai_realtime/test_realtime_service.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/test_realtime_service.py) file contains `test_session_update_nested_audio_format`, which verifies that the server correctly mirrors the selected output format in its session responses. Additionally, [`tests/openai_realtime/conftest.py`](https://github.com/huggingface/speech-to-speech/blob/main/tests/openai_realtime/conftest.py) provides the `AudioPCM` fixture used to validate PCM format construction across the test suite.