Supertonic Model Architecture Explained: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder
Supertonic uses a four-stage ONNX pipeline—consisting of a duration predictor, text encoder, vector-estimator (flow-matching diffusion model), and vocoder—to convert text into 16-bit PCM audio entirely on-device.
The supertone-inc/supertonic repository implements a fully on-device text-to-speech (TTS) system designed for ONNX Runtime inference. Unlike server-dependent TTS services, Supertonic bundles its neural network into four discrete ONNX sub-graphs that execute locally in a strict sequential order. The architecture is formally described in the research paper SupertonicTTS: Towards Highly Efficient and Streamained Text-to-Speech System referenced in the repository's README.
The Four-Component ONNX Pipeline
Supertonic's architecture decomposes the generation process into specialized models, each exported as a standalone ONNX file. The load_onnx_all function in py/helper.py (lines 297-300) initializes four onnxruntime.InferenceSession objects that together form the complete inference graph.
| Component | ONNX File | Architectural Role |
|---|---|---|
| Duration Predictor | duration_predictor.onnx |
Regresses the temporal length (in seconds) for each input phoneme or token. |
| Text Encoder | text_encoder.onnx |
Maps token IDs and style vectors into a latent conditioning representation for the diffusion process. |
| Vector-Estimator | vector_estimator.onnx |
Implements a flow-matching diffusion model that iteratively denoises a Gaussian latent toward a clean audio representation. |
| Vocoder | vocoder.onnx |
Synthesizes the final 16-bit PCM waveform from the refined latent representation. |
These models are injected into the TextToSpeech class constructor at py/helper.py (lines 40-48), which orchestrates the end-to-end inference loop.
Data Flow Through the TextToSpeech Pipeline
The _infer method inside TextToSpeech manages the hand-off between components. The process follows seven distinct stages from raw string to playable audio.
Text Preprocessing and Tokenization
Input strings first pass through the UnicodeProcessor, which converts raw text into language-tagged token IDs and attention masks. This occurs via self.text_processor inside the _infer method (lines 88-90 of py/helper.py). The processor also handles language-specific normalization required for the encoder's vocabulary.
Duration Prediction and Adjustment
The token IDs and a style-specific duration vector (style.dp) feed into the duration predictor. The raw output (dur_onnx) is divided by a configurable speed factor (default 1.05) to yield the final duration in seconds (lines 90-94). This adjustment allows real-time control over speaking rate without retraining the model.
Latent Initialization
sample_noisy_latent generates a Gaussian noise tensor whose temporal dimensions match the predicted audio length, converted to samples using the model's sample rate (lines 61-70). This noisy_latent serves as the starting point for the iterative diffusion process.
Text Embedding Generation
The text encoder consumes token IDs, the text-style vector (style.ttl), and the attention mask to produce text_emb_onnx (lines 94-98). This embedding conditions the subsequent diffusion steps, ensuring the generated audio aligns with the phonetic and prosodic content of the input.
Flow-Matching Diffusion via Vector-Estimator
The vector-estimator implements the core generative model using flow-matching principles. For each denoising step (total_step), the current latent (xt), text embedding, style vectors, masks, and step counters are passed through the estimator (lines 100-113). The model returns an updated latent that progressively refines the noisy prior toward a clean audio representation. The default configuration typically uses 8 to 20 steps depending on latency requirements.
Waveform Synthesis
After the final diffusion iteration, the refined latent tensor feeds into the vocoder, which outputs a NumPy array (wav) containing 16-bit PCM samples (lines 114-115). The vocoder acts as a neural upsampler, converting the compressed latent representation into a standard audio waveform ready for playback or file encoding.
Post-Processing and Chunking
The pipeline supports long-form synthesis by concatenating multiple chunks, inserting silence intervals, and returning both the waveform and total duration (lines 126-144). This allows the generation of arbitrarily long utterances without exceeding model memory constraints.
Practical Implementation Examples
The Supertonic repository provides two abstraction layers for invoking the architecture: a high-level Python SDK and direct ONNX Runtime sessions.
High-Level SDK Usage
For rapid prototyping, the TTS class handles model downloading, voice style loading, and inference automatically:
from supertonic import TTS
# auto_download fetches ONNX assets on first run
tts = TTS(auto_download=True)
style = tts.get_voice_style(voice_name="M1")
wav, duration = tts.synthesize(
"A gentle breeze moved through the open window while everyone listened to the story.",
voice_style=style,
lang="en",
)
tts.save_audio(wav, "output.wav")
print(f"Generated {duration:.2f}s of audio")
Source: Quick-Start section of the repository README (lines 33-48).
Direct ONNX Runtime Inference
For production deployments requiring fine-grained control, instantiate the TextToSpeech class directly with pre-loaded ONNX sessions:
import os
import json
import numpy as np
import onnxruntime as ort
from py.helper import (
load_cfgs, load_onnx_all, load_text_processor,
load_voice_style, TextToSpeech,
)
onnx_dir = "../assets/onnx"
cfgs = load_cfgs(onnx_dir)
# Explicitly load all four model components
dp_ort, txt_enc_ort, vec_est_ort, voc_ort = load_onnx_all(
onnx_dir,
ort.SessionOptions(),
["CPUExecutionProvider"]
)
text_processor = load_text_processor(onnx_dir)
style = load_voice_style([os.path.join("../assets/voice_styles", "M1.json")])
# Inject the four ONNX sessions into the pipeline
tts = TextToSpeech(
cfgs,
text_processor,
dp_ort,
txt_enc_ort,
vec_est_ort,
voc_ort
)
wav, dur = tts(
text="Hello world, this is a Supertonic demo!",
lang="en",
style=style,
total_step=8,
speed=1.05,
)
Source: TextToSpeech constructor in py/helper.py (lines 40-48); voice-style loader logic (lines 39-68).
Summary
- Supertonic implements a four-stage ONNX pipeline comprising a duration predictor, text encoder, vector-estimator, and vocoder, all loaded via
load_onnx_allinpy/helper.py. - The vector-estimator uses flow-matching diffusion to iteratively denoise latent representations conditioned by the text encoder's embeddings.
- Duration prediction adjusts for speaking rate via a configurable speed factor (default 1.05) applied to the raw duration predictor outputs.
- All inference occurs locally through ONNX Runtime sessions, enabling offline, privacy-preserving text-to-speech synthesis on consumer hardware.
Frequently Asked Questions
What is the vector-estimator in Supertonic?
The vector-estimator is an ONNX sub-graph implementing a flow-matching diffusion model. It takes a noisy Gaussian latent and, over multiple inference steps (controlled by total_step), refines it into a clean latent representation suitable for vocoding. According to the source code in py/helper.py (lines 100-113), it consumes the text embedding, style vectors, and step counter to guide the denoising trajectory.
How does the duration predictor work?
The duration predictor (duration_predictor.onnx) is a regression network that estimates the temporal length (in seconds) for each input token. It receives token IDs concatenated with a style-specific duration vector (style.dp). The raw predictions are divided by the speed parameter (default 1.05) in _infer to produce the final timing values used to initialize the diffusion latent, as implemented in py/helper.py (lines 90-94).
Can Supertonic run without installing the high-level Python SDK?
Yes. The repository exposes raw ONNX Runtime interfaces through py/helper.py and py/example_onnx.py. You can manually construct onnxruntime.InferenceSession objects for all four components and pass them to the TextToSpeech class, bypassing the TTS SDK entirely for custom inference pipelines or edge deployments.
What audio format does the Supertonic vocoder output?
The vocoder emits a NumPy array containing 16-bit PCM samples. This raw waveform can be written to standard audio formats (WAV, FLAC) using libraries like soundfile or scipy.io.wavfile. The output shape is (1, samples) representing mono audio at the model's native sample rate.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →