Supertonic WAV Output Format and Audio Parameters: A Deep Dive into the TTS Audio Pipeline

Supertonic generates standard PCM-encoded mono WAV files with 16-bit integer depth and a configurable sample rate (typically 22050 Hz or 24000 Hz) derived from the ONNX model configuration, using a canonical Rust implementation shared across all language bindings.

The supertone-inc/supertonic repository provides an open-source text-to-speech (TTS) engine that synthesizes speech through a diffusion-based neural pipeline. Understanding the Supertonic WAV output format and audio parameters is essential for integrating generated audio into media workflows, as the library enforces strict RIFF WAVE standards across its nine language implementations. Every binding ultimately relies on the same low-level audio encoding logic, ensuring consistent file structure regardless of whether you use Rust, Python, or Node.js.

Core WAV Implementation in Rust

The definitive audio serialization logic resides in rust/src/helper.rs within the write_wav_file function (lines 99–113). This routine creates a RIFF WAVE container using the hound crate to produce portable, standard-compliant files.

The write_wav_file Function

The implementation follows a four-stage process:

  1. Create WavSpec: Constructs a format descriptor with channels: 1 (mono), bits_per_sample: 16, sample_format: Int (signed integer PCM), and sample_rate pulled from the model configuration (lines 99–104).

  2. Initialize Writer: Opens a WavWriter with WavWriter::create(filename, spec)? (line 106).

  3. Sample Processing: Iterates through the floating-point waveform, clamps each value to the range ‑1.0 … +1.0, scales by 32767.0, and casts to i16 before writing (lines 108–111):

    let clamped = sample.max(-1.0).min(1.0);
    let val = (clamped * 32767.0) as i16;
    writer.write_sample(val)?;
  4. Finalization: Commits the header and data with writer.finalize()? (line 113).

All language bindings either call this Rust function directly via foreign-function interfaces or replicate its behavior to maintain byte-identical output.

Parameter Sources

The audio characteristics are not hard-coded arbitrarily but derived from the model configuration:

  • Sample Rate: Loaded from tts.json under the key ae.sample_rate via the Config struct (lines 34–38 in rust/src/helper.rs). This value drives both the inference pipeline timing and the WAV header.
  • Bit Depth: Fixed at 16 bits (signed integer), offering standard CD-quality dynamic range without the storage overhead of 24-bit formats.
  • Channels: Fixed at 1 (mono), as the TTS vocoder generates single-channel waveforms.

Audio Parameter Specifications

Supertonic exposes a consistent audio profile across every language wrapper, though only the sample rate varies between model checkpoints.

Sample Rate Configuration

The sample_rate parameter is determined at model load time and propagated through the inference graph. Typical ONNX releases ship with 22050 Hz or 24000 Hz configurations. The value is stored on the TextToSpeech object (accessible as tts.sample_rate in Rust/Python or textToSpeech.sampleRate in Swift/JavaScript) and written into the WAV header's dwSamplesPerSec field.

Bit Depth and Channel Layout

Every output file uses 16-bit PCM integer encoding with mono channel layout. This decision balances fidelity with file size and ensures universal playback compatibility across browsers, mobile devices, and professional DAWs. The format is set via SampleFormat::Int in the WavSpec struct, with no support for floating-point WAV variants or multi-channel surround formats.

Cross-Language Wrapper Implementations

All nine language bindings converge on the same audio parameters, though implementation strategies vary between native code and WASM bridges.

Python and Node.js

  • Python: The py/helper.py module provides write_wav_file (line 215), a thin wrapper around the soundfile library that passes samplerate=self.sample_rate to ensure header consistency with the model configuration.
  • Node.js: The nodejs/helper.js implementation (line 158) calls the Rust-compiled WASM export writeWavFile, marshaling the Float32Array through memory buffers before returning a Node Buffer containing the complete RIFF structure.

Java, Go, and C++

  • Java: java/Helper.java constructs a WavFile using standard javax.sound.sampled APIs (line 817), explicitly setting the audio format to 16-bit mono PCM with the config-driven sample rate.
  • Go: go/helper.go defines a SampleRate field in the Config struct (line 40) that feeds into the WASM encoder, ensuring the Go layer respects the JSON configuration without reimplementing the bit-manipulation logic.
  • C++: cpp/helper.cpp manually constructs the 44-byte RIFF header (line 955), calculating byte_rate as sample_rate * 2 (16-bit mono) and validating against the sample_rate value extracted from the configuration JSON.

Swift, C#, and Flutter

  • Swift: swift/Sources/Helper.swift invokes a C-bridge to the native writeWavFile function (line 94), passing the sampleRate property from the loaded model.
  • C#: csharp/Helper.cs utilizes NAudio's WaveFileWriter (line 554), setting SampleRate = cfgs.ae.sample_rate and WaveFormat = new WaveFormat(sampleRate, 16, 1).
  • Flutter: flutter/lib/helper.dart forwards the sampleRate value (line 280) through platform channels to the underlying native implementation, maintaining the same parameter chain as the other bindings.

The TTS Pipeline: From Text to WAV

Understanding the Supertonic WAV output format requires context on how the raw audio buffer is generated. The pipeline executes six distinct stages before write_wav_file serializes the result:

  1. Text Processing: UnicodeProcessor.call converts input strings to integer ID sequences.
  2. Duration Prediction: The dp_ort ONNX model outputs phoneme durations scaled by the configuration's sample rate.
  3. Latent Sampling: sample_noisy_latent initializes a noise tensor sized according to the target audio length in samples.
  4. Iterative Denoising: vector_est_ort refines the latent representation over total_step diffusion iterations.
  5. Vocoding: vocoder_ort synthesizes the final floating-point waveform Vec<f32> (or equivalent).
  6. File Encoding: write_wav_file clamps samples to [-1.0, 1.0], quantizes to 16-bit integers, and writes the RIFF container.

The sample rate governs the temporal scaling in steps 2–5 and establishes the timebase for the final WAV header.

Practical Code Examples

Rust Example

use supertonic::rust::helper::{load_text_to_speech, load_voice_style, write_wav_file};

fn main() -> anyhow::Result<()> {
    let tts = load_text_to_speech("./model/onnx", false)?;
    let style = load_voice_style(&["./model/style1.json".into()], false)?;

    let (wav, duration) = tts.call(
        "Hello, world!",
        "en",
        &style,
        10,    // total diffusion steps
        1.0,   // speed factor
        0.3,   // silence duration (seconds)
    )?;

    write_wav_file("hello.wav", &wav, tts.sample_rate)?;
    println!("Saved {:.2}s at {} Hz", duration, tts.sample_rate);
    Ok(())
}

Source: rust/src/example_onnx.rs (lines 100–130)

Python Example

from supertonic.py.helper import load_text_to_speech, load_voice_style, write_wav_file

tts = load_text_to_speech("./model/onnx", use_gpu=False)
style = load_voice_style(["./model/style1.json"])

wav, duration = tts.call(
    "Hello, world!",
    "en",
    style,
    total_step=10,
    speed=1.0,
    silence_duration=0.3,
)

write_wav_file("hello.wav", wav, tts.sample_rate)

Source: py/example_onnx.py (lines 100–115)

Node.js and Swift Examples

Node.js (via WASM):

import { TextToSpeech, loadVoiceStyle, writeWavFile } from './web/helper.js';

const tts = new TextToSpeech(cfg, await loadVoiceStyle(['./model/style1.json']));
const { wav, duration } = await tts.call('Hello, world!', 'en', style, 10, 1.0, 0.3);
const buffer = writeWavFile(wav, tts.sampleRate); // Returns Node Buffer

Source: web/main.js (lines 215–218)

Swift:

let tts = try loadTextToSpeech(onnxDir: "./model/onnx", useGPU: false)
let style = try loadVoiceStyle(paths: ["./model/style1.json"], verbose: false)
let (wav, duration) = try tts.call(
    text: "Hello, world!",
    lang: "en",
    style: style,
    totalStep: 10,
    speed: 1.0,
    silenceDuration: 0.3
)
try writeWavFile(outputPath: "hello.wav", wav: wav, sampleRate: tts.sampleRate)

Source: swift/Sources/ExampleONNX.swift (lines 115–151)

Summary

  • Supertonic outputs standard RIFF WAV files with 16-bit signed integer PCM encoding and mono channel configuration.
  • The sample rate is model-dependent (typically 22050 Hz or 24000 Hz) and sourced from tts.json → ae.sample_rate.
  • All language bindings share the canonical Rust implementation in rust/src/helper.rs, ensuring consistent file headers across Python, Node.js, Java, Go, C++, Swift, C#, and Flutter.
  • Audio samples are clamped to [-1.0, 1.0] and scaled by 32767.0 before quantization to prevent clipping and maximize dynamic range.
  • The sample_rate property is exposed on all high-level TTS objects, allowing runtime inspection of the configured audio parameters.

Frequently Asked Questions

What audio codec does Supertonic use for output files?

Supertonic uses uncompressed PCM audio packaged in the RIFF WAVE format. Specifically, it writes 16-bit signed integer samples with a single audio channel (mono). The implementation does not use MP3, AAC, or Ogg compression, ensuring lossless quality suitable for further editing or waveform analysis.

Can I change the bit depth or channel count in the output WAV?

No. The write_wav_file function in rust/src/helper.rs hard-codes bits_per_sample: 16 and channels: 1 in the WavSpec struct (lines 99–104). To generate 24-bit or stereo files, you must post-process the generated WAV using external tools like FFmpeg or SoX, or modify the source code to construct a different WavSpec before calling the writer.

Where does Supertonic get the sample rate value?

The sample rate is read from the ONNX model's configuration file, tts.json, specifically from the field ae.sample_rate. This value is parsed by the Config struct during initialization (lines 34–38 of rust/src/helper.rs) and propagated to the TextToSpeech instance. Changing the model checkpoint automatically updates the output WAV header's sample rate without code changes.

Why does the pipeline clamp samples to -1.0 and 1.0 before writing?

The clamping operation prevents integer overflow during the conversion from floating-point neural network outputs (which can occasionally exceed nominal amplitude ranges) to 16-bit integers. By enforcing sample.max(-1.0).min(1.0) before multiplying by 32767.0, Supertonic ensures that the final i16 values stay within the valid range of -32768 to 32767, avoiding wrap-around distortion in the audio file.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →