# Supertonic WAV Output Format and Audio Parameters: A Deep Dive into the TTS Audio Pipeline

> Explore the Supertonic WAV output format and audio parameters. Understand its PCM encoding, 16-bit depth, and configurable sample rates for high-quality TTS audio generation.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: deep-dive
- Published: 2026-05-14

---

**Supertonic generates standard PCM-encoded mono WAV files with 16-bit integer depth and a configurable sample rate (typically 22050 Hz or 24000 Hz) derived from the ONNX model configuration, using a canonical Rust implementation shared across all language bindings.**

The supertone-inc/supertonic repository provides an open-source text-to-speech (TTS) engine that synthesizes speech through a diffusion-based neural pipeline. Understanding the Supertonic WAV output format and audio parameters is essential for integrating generated audio into media workflows, as the library enforces strict RIFF WAVE standards across its nine language implementations. Every binding ultimately relies on the same low-level audio encoding logic, ensuring consistent file structure regardless of whether you use Rust, Python, or Node.js.

## Core WAV Implementation in Rust

The definitive audio serialization logic resides in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) within the `write_wav_file` function (lines 99–113). This routine creates a **RIFF WAVE** container using the `hound` crate to produce portable, standard-compliant files.

### The write_wav_file Function

The implementation follows a four-stage process:

1. **Create WavSpec**: Constructs a format descriptor with `channels: 1` (mono), `bits_per_sample: 16`, `sample_format: Int` (signed integer PCM), and `sample_rate` pulled from the model configuration (lines 99–104).
2. **Initialize Writer**: Opens a `WavWriter` with `WavWriter::create(filename, spec)?` (line 106).
3. **Sample Processing**: Iterates through the floating-point waveform, clamps each value to the range **‑1.0 … +1.0**, scales by **32767.0**, and casts to `i16` before writing (lines 108–111):
   ```rust
   let clamped = sample.max(-1.0).min(1.0);
   let val = (clamped * 32767.0) as i16;
   writer.write_sample(val)?;
   ```

4. **Finalization**: Commits the header and data with `writer.finalize()?` (line 113).

All language bindings either call this Rust function directly via foreign-function interfaces or replicate its behavior to maintain byte-identical output.

### Parameter Sources

The audio characteristics are not hard-coded arbitrarily but derived from the model configuration:

- **Sample Rate**: Loaded from [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json) under the key `ae.sample_rate` via the `Config` struct (lines 34–38 in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)). This value drives both the inference pipeline timing and the WAV header.
- **Bit Depth**: Fixed at **16 bits** (signed integer), offering standard CD-quality dynamic range without the storage overhead of 24-bit formats.
- **Channels**: Fixed at **1** (mono), as the TTS vocoder generates single-channel waveforms.

## Audio Parameter Specifications

Supertonic exposes a consistent audio profile across every language wrapper, though only the sample rate varies between model checkpoints.

### Sample Rate Configuration

The `sample_rate` parameter is determined at model load time and propagated through the inference graph. Typical ONNX releases ship with **22050 Hz** or **24000 Hz** configurations. The value is stored on the `TextToSpeech` object (accessible as `tts.sample_rate` in Rust/Python or `textToSpeech.sampleRate` in Swift/JavaScript) and written into the WAV header's `dwSamplesPerSec` field.

### Bit Depth and Channel Layout

Every output file uses **16-bit PCM integer encoding** with **mono channel layout**. This decision balances fidelity with file size and ensures universal playback compatibility across browsers, mobile devices, and professional DAWs. The format is set via `SampleFormat::Int` in the `WavSpec` struct, with no support for floating-point WAV variants or multi-channel surround formats.

## Cross-Language Wrapper Implementations

All nine language bindings converge on the same audio parameters, though implementation strategies vary between native code and WASM bridges.

### Python and Node.js

- **Python**: The [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) module provides `write_wav_file` (line 215), a thin wrapper around the `soundfile` library that passes `samplerate=self.sample_rate` to ensure header consistency with the model configuration.
- **Node.js**: The [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js) implementation (line 158) calls the Rust-compiled WASM export `writeWavFile`, marshaling the Float32Array through memory buffers before returning a Node `Buffer` containing the complete RIFF structure.

### Java, Go, and C++

- **Java**: [`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java) constructs a `WavFile` using standard `javax.sound.sampled` APIs (line 817), explicitly setting the audio format to 16-bit mono PCM with the config-driven sample rate.
- **Go**: [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) defines a `SampleRate` field in the `Config` struct (line 40) that feeds into the WASM encoder, ensuring the Go layer respects the JSON configuration without reimplementing the bit-manipulation logic.
- **C++**: [`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp) manually constructs the 44-byte RIFF header (line 955), calculating `byte_rate` as `sample_rate * 2` (16-bit mono) and validating against the `sample_rate` value extracted from the configuration JSON.

### Swift, C#, and Flutter

- **Swift**: [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) invokes a C-bridge to the native `writeWavFile` function (line 94), passing the `sampleRate` property from the loaded model.
- **C#**: [`csharp/Helper.cs`](https://github.com/supertone-inc/supertonic/blob/main/csharp/Helper.cs) utilizes NAudio's `WaveFileWriter` (line 554), setting `SampleRate = cfgs.ae.sample_rate` and `WaveFormat = new WaveFormat(sampleRate, 16, 1)`.
- **Flutter**: `flutter/lib/helper.dart` forwards the `sampleRate` value (line 280) through platform channels to the underlying native implementation, maintaining the same parameter chain as the other bindings.

## The TTS Pipeline: From Text to WAV

Understanding the Supertonic WAV output format requires context on how the raw audio buffer is generated. The pipeline executes six distinct stages before `write_wav_file` serializes the result:

1. **Text Processing**: `UnicodeProcessor.call` converts input strings to integer ID sequences.
2. **Duration Prediction**: The `dp_ort` ONNX model outputs phoneme durations scaled by the configuration's sample rate.
3. **Latent Sampling**: `sample_noisy_latent` initializes a noise tensor sized according to the target audio length in samples.
4. **Iterative Denoising**: `vector_est_ort` refines the latent representation over `total_step` diffusion iterations.
5. **Vocoding**: `vocoder_ort` synthesizes the final floating-point waveform `Vec<f32>` (or equivalent).
6. **File Encoding**: `write_wav_file` clamps samples to [-1.0, 1.0], quantizes to 16-bit integers, and writes the RIFF container.

The sample rate governs the temporal scaling in steps 2–5 and establishes the timebase for the final WAV header.

## Practical Code Examples

### Rust Example

```rust
use supertonic::rust::helper::{load_text_to_speech, load_voice_style, write_wav_file};

fn main() -> anyhow::Result<()> {
    let tts = load_text_to_speech("./model/onnx", false)?;
    let style = load_voice_style(&["./model/style1.json".into()], false)?;

    let (wav, duration) = tts.call(
        "Hello, world!",
        "en",
        &style,
        10,    // total diffusion steps
        1.0,   // speed factor
        0.3,   // silence duration (seconds)
    )?;

    write_wav_file("hello.wav", &wav, tts.sample_rate)?;
    println!("Saved {:.2}s at {} Hz", duration, tts.sample_rate);
    Ok(())
}

```

Source: [`rust/src/example_onnx.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/example_onnx.rs) (lines 100–130)

### Python Example

```python
from supertonic.py.helper import load_text_to_speech, load_voice_style, write_wav_file

tts = load_text_to_speech("./model/onnx", use_gpu=False)
style = load_voice_style(["./model/style1.json"])

wav, duration = tts.call(
    "Hello, world!",
    "en",
    style,
    total_step=10,
    speed=1.0,
    silence_duration=0.3,
)

write_wav_file("hello.wav", wav, tts.sample_rate)

```

Source: [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) (lines 100–115)

### Node.js and Swift Examples

**Node.js** (via WASM):

```javascript
import { TextToSpeech, loadVoiceStyle, writeWavFile } from './web/helper.js';

const tts = new TextToSpeech(cfg, await loadVoiceStyle(['./model/style1.json']));
const { wav, duration } = await tts.call('Hello, world!', 'en', style, 10, 1.0, 0.3);
const buffer = writeWavFile(wav, tts.sampleRate); // Returns Node Buffer

```

Source: [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js) (lines 215–218)

**Swift**:

```swift
let tts = try loadTextToSpeech(onnxDir: "./model/onnx", useGPU: false)
let style = try loadVoiceStyle(paths: ["./model/style1.json"], verbose: false)
let (wav, duration) = try tts.call(
    text: "Hello, world!",
    lang: "en",
    style: style,
    totalStep: 10,
    speed: 1.0,
    silenceDuration: 0.3
)
try writeWavFile(outputPath: "hello.wav", wav: wav, sampleRate: tts.sampleRate)

```

Source: [`swift/Sources/ExampleONNX.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/ExampleONNX.swift) (lines 115–151)

## Summary

- Supertonic outputs **standard RIFF WAV files** with 16-bit signed integer PCM encoding and mono channel configuration.
- The **sample rate** is model-dependent (typically 22050 Hz or 24000 Hz) and sourced from [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json) → `ae.sample_rate`.
- All language bindings share the **canonical Rust implementation** in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs), ensuring consistent file headers across Python, Node.js, Java, Go, C++, Swift, C#, and Flutter.
- Audio samples are **clamped to [-1.0, 1.0]** and scaled by **32767.0** before quantization to prevent clipping and maximize dynamic range.
- The `sample_rate` property is exposed on all high-level TTS objects, allowing runtime inspection of the configured audio parameters.

## Frequently Asked Questions

### What audio codec does Supertonic use for output files?

Supertonic uses uncompressed **PCM audio** packaged in the RIFF WAVE format. Specifically, it writes 16-bit signed integer samples with a single audio channel (mono). The implementation does not use MP3, AAC, or Ogg compression, ensuring lossless quality suitable for further editing or waveform analysis.

### Can I change the bit depth or channel count in the output WAV?

No. The `write_wav_file` function in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) hard-codes `bits_per_sample: 16` and `channels: 1` in the `WavSpec` struct (lines 99–104). To generate 24-bit or stereo files, you must post-process the generated WAV using external tools like FFmpeg or SoX, or modify the source code to construct a different `WavSpec` before calling the writer.

### Where does Supertonic get the sample rate value?

The sample rate is read from the ONNX model's configuration file, [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json), specifically from the field `ae.sample_rate`. This value is parsed by the `Config` struct during initialization (lines 34–38 of [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)) and propagated to the `TextToSpeech` instance. Changing the model checkpoint automatically updates the output WAV header's sample rate without code changes.

### Why does the pipeline clamp samples to -1.0 and 1.0 before writing?

The clamping operation prevents integer overflow during the conversion from floating-point neural network outputs (which can occasionally exceed nominal amplitude ranges) to 16-bit integers. By enforcing `sample.max(-1.0).min(1.0)` before multiplying by 32767.0, Supertonic ensures that the final `i16` values stay within the valid range of -32768 to 32767, avoiding wrap-around distortion in the audio file.