Supertonic WAV Output Format and Audio Parameters: A Deep Dive into the TTS Audio Pipeline
Supertonic generates standard PCM-encoded mono WAV files with 16-bit integer depth and a configurable sample rate (typically 22050 Hz or 24000 Hz) derived from the ONNX model configuration, using a canonical Rust implementation shared across all language bindings.
The supertone-inc/supertonic repository provides an open-source text-to-speech (TTS) engine that synthesizes speech through a diffusion-based neural pipeline. Understanding the Supertonic WAV output format and audio parameters is essential for integrating generated audio into media workflows, as the library enforces strict RIFF WAVE standards across its nine language implementations. Every binding ultimately relies on the same low-level audio encoding logic, ensuring consistent file structure regardless of whether you use Rust, Python, or Node.js.
Core WAV Implementation in Rust
The definitive audio serialization logic resides in rust/src/helper.rs within the write_wav_file function (lines 99–113). This routine creates a RIFF WAVE container using the hound crate to produce portable, standard-compliant files.
The write_wav_file Function
The implementation follows a four-stage process:
-
Create WavSpec: Constructs a format descriptor with
channels: 1(mono),bits_per_sample: 16,sample_format: Int(signed integer PCM), andsample_ratepulled from the model configuration (lines 99–104). -
Initialize Writer: Opens a
WavWriterwithWavWriter::create(filename, spec)?(line 106). -
Sample Processing: Iterates through the floating-point waveform, clamps each value to the range ‑1.0 … +1.0, scales by 32767.0, and casts to
i16before writing (lines 108–111):let clamped = sample.max(-1.0).min(1.0); let val = (clamped * 32767.0) as i16; writer.write_sample(val)?; -
Finalization: Commits the header and data with
writer.finalize()?(line 113).
All language bindings either call this Rust function directly via foreign-function interfaces or replicate its behavior to maintain byte-identical output.
Parameter Sources
The audio characteristics are not hard-coded arbitrarily but derived from the model configuration:
- Sample Rate: Loaded from
tts.jsonunder the keyae.sample_ratevia theConfigstruct (lines 34–38 inrust/src/helper.rs). This value drives both the inference pipeline timing and the WAV header. - Bit Depth: Fixed at 16 bits (signed integer), offering standard CD-quality dynamic range without the storage overhead of 24-bit formats.
- Channels: Fixed at 1 (mono), as the TTS vocoder generates single-channel waveforms.
Audio Parameter Specifications
Supertonic exposes a consistent audio profile across every language wrapper, though only the sample rate varies between model checkpoints.
Sample Rate Configuration
The sample_rate parameter is determined at model load time and propagated through the inference graph. Typical ONNX releases ship with 22050 Hz or 24000 Hz configurations. The value is stored on the TextToSpeech object (accessible as tts.sample_rate in Rust/Python or textToSpeech.sampleRate in Swift/JavaScript) and written into the WAV header's dwSamplesPerSec field.
Bit Depth and Channel Layout
Every output file uses 16-bit PCM integer encoding with mono channel layout. This decision balances fidelity with file size and ensures universal playback compatibility across browsers, mobile devices, and professional DAWs. The format is set via SampleFormat::Int in the WavSpec struct, with no support for floating-point WAV variants or multi-channel surround formats.
Cross-Language Wrapper Implementations
All nine language bindings converge on the same audio parameters, though implementation strategies vary between native code and WASM bridges.
Python and Node.js
- Python: The
py/helper.pymodule provideswrite_wav_file(line 215), a thin wrapper around thesoundfilelibrary that passessamplerate=self.sample_rateto ensure header consistency with the model configuration. - Node.js: The
nodejs/helper.jsimplementation (line 158) calls the Rust-compiled WASM exportwriteWavFile, marshaling the Float32Array through memory buffers before returning a NodeBuffercontaining the complete RIFF structure.
Java, Go, and C++
- Java:
java/Helper.javaconstructs aWavFileusing standardjavax.sound.sampledAPIs (line 817), explicitly setting the audio format to 16-bit mono PCM with the config-driven sample rate. - Go:
go/helper.godefines aSampleRatefield in theConfigstruct (line 40) that feeds into the WASM encoder, ensuring the Go layer respects the JSON configuration without reimplementing the bit-manipulation logic. - C++:
cpp/helper.cppmanually constructs the 44-byte RIFF header (line 955), calculatingbyte_rateassample_rate * 2(16-bit mono) and validating against thesample_ratevalue extracted from the configuration JSON.
Swift, C#, and Flutter
- Swift:
swift/Sources/Helper.swiftinvokes a C-bridge to the nativewriteWavFilefunction (line 94), passing thesampleRateproperty from the loaded model. - C#:
csharp/Helper.csutilizes NAudio'sWaveFileWriter(line 554), settingSampleRate = cfgs.ae.sample_rateandWaveFormat = new WaveFormat(sampleRate, 16, 1). - Flutter:
flutter/lib/helper.dartforwards thesampleRatevalue (line 280) through platform channels to the underlying native implementation, maintaining the same parameter chain as the other bindings.
The TTS Pipeline: From Text to WAV
Understanding the Supertonic WAV output format requires context on how the raw audio buffer is generated. The pipeline executes six distinct stages before write_wav_file serializes the result:
- Text Processing:
UnicodeProcessor.callconverts input strings to integer ID sequences. - Duration Prediction: The
dp_ortONNX model outputs phoneme durations scaled by the configuration's sample rate. - Latent Sampling:
sample_noisy_latentinitializes a noise tensor sized according to the target audio length in samples. - Iterative Denoising:
vector_est_ortrefines the latent representation overtotal_stepdiffusion iterations. - Vocoding:
vocoder_ortsynthesizes the final floating-point waveformVec<f32>(or equivalent). - File Encoding:
write_wav_fileclamps samples to [-1.0, 1.0], quantizes to 16-bit integers, and writes the RIFF container.
The sample rate governs the temporal scaling in steps 2–5 and establishes the timebase for the final WAV header.
Practical Code Examples
Rust Example
use supertonic::rust::helper::{load_text_to_speech, load_voice_style, write_wav_file};
fn main() -> anyhow::Result<()> {
let tts = load_text_to_speech("./model/onnx", false)?;
let style = load_voice_style(&["./model/style1.json".into()], false)?;
let (wav, duration) = tts.call(
"Hello, world!",
"en",
&style,
10, // total diffusion steps
1.0, // speed factor
0.3, // silence duration (seconds)
)?;
write_wav_file("hello.wav", &wav, tts.sample_rate)?;
println!("Saved {:.2}s at {} Hz", duration, tts.sample_rate);
Ok(())
}
Source: rust/src/example_onnx.rs (lines 100–130)
Python Example
from supertonic.py.helper import load_text_to_speech, load_voice_style, write_wav_file
tts = load_text_to_speech("./model/onnx", use_gpu=False)
style = load_voice_style(["./model/style1.json"])
wav, duration = tts.call(
"Hello, world!",
"en",
style,
total_step=10,
speed=1.0,
silence_duration=0.3,
)
write_wav_file("hello.wav", wav, tts.sample_rate)
Source: py/example_onnx.py (lines 100–115)
Node.js and Swift Examples
Node.js (via WASM):
import { TextToSpeech, loadVoiceStyle, writeWavFile } from './web/helper.js';
const tts = new TextToSpeech(cfg, await loadVoiceStyle(['./model/style1.json']));
const { wav, duration } = await tts.call('Hello, world!', 'en', style, 10, 1.0, 0.3);
const buffer = writeWavFile(wav, tts.sampleRate); // Returns Node Buffer
Source: web/main.js (lines 215–218)
Swift:
let tts = try loadTextToSpeech(onnxDir: "./model/onnx", useGPU: false)
let style = try loadVoiceStyle(paths: ["./model/style1.json"], verbose: false)
let (wav, duration) = try tts.call(
text: "Hello, world!",
lang: "en",
style: style,
totalStep: 10,
speed: 1.0,
silenceDuration: 0.3
)
try writeWavFile(outputPath: "hello.wav", wav: wav, sampleRate: tts.sampleRate)
Source: swift/Sources/ExampleONNX.swift (lines 115–151)
Summary
- Supertonic outputs standard RIFF WAV files with 16-bit signed integer PCM encoding and mono channel configuration.
- The sample rate is model-dependent (typically 22050 Hz or 24000 Hz) and sourced from
tts.json→ae.sample_rate. - All language bindings share the canonical Rust implementation in
rust/src/helper.rs, ensuring consistent file headers across Python, Node.js, Java, Go, C++, Swift, C#, and Flutter. - Audio samples are clamped to [-1.0, 1.0] and scaled by 32767.0 before quantization to prevent clipping and maximize dynamic range.
- The
sample_rateproperty is exposed on all high-level TTS objects, allowing runtime inspection of the configured audio parameters.
Frequently Asked Questions
What audio codec does Supertonic use for output files?
Supertonic uses uncompressed PCM audio packaged in the RIFF WAVE format. Specifically, it writes 16-bit signed integer samples with a single audio channel (mono). The implementation does not use MP3, AAC, or Ogg compression, ensuring lossless quality suitable for further editing or waveform analysis.
Can I change the bit depth or channel count in the output WAV?
No. The write_wav_file function in rust/src/helper.rs hard-codes bits_per_sample: 16 and channels: 1 in the WavSpec struct (lines 99–104). To generate 24-bit or stereo files, you must post-process the generated WAV using external tools like FFmpeg or SoX, or modify the source code to construct a different WavSpec before calling the writer.
Where does Supertonic get the sample rate value?
The sample rate is read from the ONNX model's configuration file, tts.json, specifically from the field ae.sample_rate. This value is parsed by the Config struct during initialization (lines 34–38 of rust/src/helper.rs) and propagated to the TextToSpeech instance. Changing the model checkpoint automatically updates the output WAV header's sample rate without code changes.
Why does the pipeline clamp samples to -1.0 and 1.0 before writing?
The clamping operation prevents integer overflow during the conversion from floating-point neural network outputs (which can occasionally exceed nominal amplitude ranges) to 16-bit integers. By enforcing sample.max(-1.0).min(1.0) before multiplying by 32767.0, Supertonic ensures that the final i16 values stay within the valid range of -32768 to 32767, avoiding wrap-around distortion in the audio file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →