# Supertonic Model Architecture Explained: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder

> Explore the Supertonic model architecture and its four key components: duration predictor, text encoder, vector estimator, and vocoder. Convert text to audio on-device with this ONNX pipeline.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: architecture
- Published: 2026-05-14

---

**Supertonic uses a four-stage ONNX pipeline—consisting of a duration predictor, text encoder, vector-estimator (flow-matching diffusion model), and vocoder—to convert text into 16-bit PCM audio entirely on-device.**

The `supertone-inc/supertonic` repository implements a fully on-device text-to-speech (TTS) system designed for ONNX Runtime inference. Unlike server-dependent TTS services, Supertonic bundles its neural network into four discrete ONNX sub-graphs that execute locally in a strict sequential order. The architecture is formally described in the research paper *SupertonicTTS: Towards Highly Efficient and Streamained Text-to-Speech System* referenced in the repository's README.

## The Four-Component ONNX Pipeline

Supertonic's architecture decomposes the generation process into specialized models, each exported as a standalone ONNX file. The `load_onnx_all` function in **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 297-300) initializes four `onnxruntime.InferenceSession` objects that together form the complete inference graph.

| Component | ONNX File | Architectural Role |
|-----------|-----------|-------------------|
| **Duration Predictor** | `duration_predictor.onnx` | Regresses the temporal length (in seconds) for each input phoneme or token. |
| **Text Encoder** | `text_encoder.onnx` | Maps token IDs and style vectors into a latent conditioning representation for the diffusion process. |
| **Vector-Estimator** | `vector_estimator.onnx` | Implements a flow-matching diffusion model that iteratively denoises a Gaussian latent toward a clean audio representation. |
| **Vocoder** | `vocoder.onnx` | Synthesizes the final 16-bit PCM waveform from the refined latent representation. |

These models are injected into the `TextToSpeech` class constructor at **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 40-48), which orchestrates the end-to-end inference loop.

## Data Flow Through the TextToSpeech Pipeline

The `_infer` method inside `TextToSpeech` manages the hand-off between components. The process follows seven distinct stages from raw string to playable audio.

### Text Preprocessing and Tokenization

Input strings first pass through the `UnicodeProcessor`, which converts raw text into language-tagged token IDs and attention masks. This occurs via `self.text_processor` inside the `_infer` method (lines 88-90 of **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)**). The processor also handles language-specific normalization required for the encoder's vocabulary.

### Duration Prediction and Adjustment

The token IDs and a style-specific duration vector (`style.dp`) feed into the **duration predictor**. The raw output (`dur_onnx`) is divided by a configurable speed factor (default **1.05**) to yield the final duration in seconds (lines 90-94). This adjustment allows real-time control over speaking rate without retraining the model.

### Latent Initialization

`sample_noisy_latent` generates a Gaussian noise tensor whose temporal dimensions match the predicted audio length, converted to samples using the model's sample rate (lines 61-70). This `noisy_latent` serves as the starting point for the iterative diffusion process.

### Text Embedding Generation

The **text encoder** consumes token IDs, the text-style vector (`style.ttl`), and the attention mask to produce `text_emb_onnx` (lines 94-98). This embedding conditions the subsequent diffusion steps, ensuring the generated audio aligns with the phonetic and prosodic content of the input.

### Flow-Matching Diffusion via Vector-Estimator

The **vector-estimator** implements the core generative model using flow-matching principles. For each denoising step (`total_step`), the current latent (`xt`), text embedding, style vectors, masks, and step counters are passed through the estimator (lines 100-113). The model returns an updated latent that progressively refines the noisy prior toward a clean audio representation. The default configuration typically uses 8 to 20 steps depending on latency requirements.

### Waveform Synthesis

After the final diffusion iteration, the refined latent tensor feeds into the **vocoder**, which outputs a NumPy array (`wav`) containing 16-bit PCM samples (lines 114-115). The vocoder acts as a neural upsampler, converting the compressed latent representation into a standard audio waveform ready for playback or file encoding.

### Post-Processing and Chunking

The pipeline supports long-form synthesis by concatenating multiple chunks, inserting silence intervals, and returning both the waveform and total duration (lines 126-144). This allows the generation of arbitrarily long utterances without exceeding model memory constraints.

## Practical Implementation Examples

The Supertonic repository provides two abstraction layers for invoking the architecture: a high-level Python SDK and direct ONNX Runtime sessions.

### High-Level SDK Usage

For rapid prototyping, the `TTS` class handles model downloading, voice style loading, and inference automatically:

```python
from supertonic import TTS

# auto_download fetches ONNX assets on first run

tts = TTS(auto_download=True)
style = tts.get_voice_style(voice_name="M1")

wav, duration = tts.synthesize(
    "A gentle breeze moved through the open window while everyone listened to the story.",
    voice_style=style,
    lang="en",
)

tts.save_audio(wav, "output.wav")
print(f"Generated {duration:.2f}s of audio")

```

*Source:* Quick-Start section of the repository README (lines 33-48).

### Direct ONNX Runtime Inference

For production deployments requiring fine-grained control, instantiate the `TextToSpeech` class directly with pre-loaded ONNX sessions:

```python
import os
import json
import numpy as np
import onnxruntime as ort
from py.helper import (
    load_cfgs, load_onnx_all, load_text_processor,
    load_voice_style, TextToSpeech,
)

onnx_dir = "../assets/onnx"
cfgs = load_cfgs(onnx_dir)

# Explicitly load all four model components

dp_ort, txt_enc_ort, vec_est_ort, voc_ort = load_onnx_all(
    onnx_dir, 
    ort.SessionOptions(), 
    ["CPUExecutionProvider"]
)

text_processor = load_text_processor(onnx_dir)
style = load_voice_style([os.path.join("../assets/voice_styles", "M1.json")])

# Inject the four ONNX sessions into the pipeline

tts = TextToSpeech(
    cfgs, 
    text_processor, 
    dp_ort, 
    txt_enc_ort, 
    vec_est_ort, 
    voc_ort
)

wav, dur = tts(
    text="Hello world, this is a Supertonic demo!",
    lang="en",
    style=style,
    total_step=8,
    speed=1.05,
)

```

*Source:* `TextToSpeech` constructor in **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 40-48); voice-style loader logic (lines 39-68).

## Summary

- **Supertonic implements a four-stage ONNX pipeline** comprising a duration predictor, text encoder, vector-estimator, and vocoder, all loaded via `load_onnx_all` in **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)**.
- **The vector-estimator uses flow-matching diffusion** to iteratively denoise latent representations conditioned by the text encoder's embeddings.
- **Duration prediction adjusts for speaking rate** via a configurable speed factor (default 1.05) applied to the raw duration predictor outputs.
- **All inference occurs locally** through ONNX Runtime sessions, enabling offline, privacy-preserving text-to-speech synthesis on consumer hardware.

## Frequently Asked Questions

### What is the vector-estimator in Supertonic?

The **vector-estimator** is an ONNX sub-graph implementing a flow-matching diffusion model. It takes a noisy Gaussian latent and, over multiple inference steps (controlled by `total_step`), refines it into a clean latent representation suitable for vocoding. According to the source code in **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 100-113), it consumes the text embedding, style vectors, and step counter to guide the denoising trajectory.

### How does the duration predictor work?

The **duration predictor** (`duration_predictor.onnx`) is a regression network that estimates the temporal length (in seconds) for each input token. It receives token IDs concatenated with a style-specific duration vector (`style.dp`). The raw predictions are divided by the speed parameter (default 1.05) in `_infer` to produce the final timing values used to initialize the diffusion latent, as implemented in **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** (lines 90-94).

### Can Supertonic run without installing the high-level Python SDK?

Yes. The repository exposes raw ONNX Runtime interfaces through **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** and **[`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py)**. You can manually construct `onnxruntime.InferenceSession` objects for all four components and pass them to the `TextToSpeech` class, bypassing the `TTS` SDK entirely for custom inference pipelines or edge deployments.

### What audio format does the Supertonic vocoder output?

The **vocoder** emits a NumPy array containing **16-bit PCM** samples. This raw waveform can be written to standard audio formats (WAV, FLAC) using libraries like `soundfile` or `scipy.io.wavfile`. The output shape is `(1, samples)` representing mono audio at the model's native sample rate.