# Differences Between the ONNX Models in Supertonic TTS: duration_predictor, text_encoder, vector_estimator, and vocoder

> Understand the distinct roles of Supertonic ONNX models: duration_predictor, text_encoder, vector_estimator, and vocoder. Learn how each contributes to the TTS pipeline from tokens to audio.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: deep-dive
- Published: 2026-06-12

---

**The four ONNX models in Supertonic represent distinct stages of a neural text-to-speech pipeline: duration_predictor calculates token timings, text_encoder generates semantic embeddings, vector_estimator refines latent representations through a diffusion process, and vocoder synthesizes the final audio waveform.**

Supertonic by supertone-inc implements a multi-stage neural TTS architecture where each component is encapsulated as an independent ONNX model. These models are orchestrated together in the `TextToSpeech` class defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), allowing developers to swap or retrain individual stages without modifying the surrounding Python code. Understanding the specific role, inputs, and outputs of each model is essential for debugging, optimizing, or extending the synthesis pipeline.

## Overview of the Four ONNX Models

### duration_predictor

The **duration_predictor** determines how long each text token (character) should be spoken in seconds. It takes `text_ids` (tokenized character IDs), `style_dp` (a speaker-specific style vector), and `text_mask` as inputs. The model outputs `dur_onnx`, a 1-D float array containing the predicted duration for each token. In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), this model is invoked via `dp_ort.run()` at approximately line 190 within the `TextToSpeech._infer` method.

### text_encoder

The **text_encoder** transforms discrete token IDs into dense semantic embeddings that condition the acoustic model. Its inputs include `text_ids`, `style_ttl` (tone/style vector), and `text_mask`. The output is `text_emb_onnx`, a 2-D tensor with shape `[batch, embed_dim]` that captures the linguistic content of the input text. This embedding serves as the primary conditioning signal for the latent refinement stage. The encoder runs via `text_enc_ort.run()` at approximately line 194.

### vector_estimator

The **vector_estimator** implements a diffusion-like refinement process that denoises a latent representation into a structured acoustic encoding. It accepts `noisy_latent` (the initial random tensor), `text_emb` (from the text encoder), `style_ttl`, `text_mask`, `latent_mask`, `current_step`, and `total_step` as inputs. The model outputs an updated `xt` tensor, with the process iterating for a configurable number of steps (controlled by `total_step`). In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), this appears as `vector_est_ort.run()` inside a loop spanning lines approximately 200-213.

### vocoder

The **vocoder** converts the final refined latent tensor into audible raw audio. It takes a single input, `latent` (the output from the vector estimator), and produces `wav_tts`, a float32 waveform tensor with shape `[1, samples]`. This represents the final audio output ready for playback or saving to disk. The vocoder executes via `vocoder_ort.run()` at approximately line 214.

## How the Models Connect in the Pipeline

The four models execute sequentially within the `TextToSpeech._infer` method, with each stage feeding into the next:

1. **Tokenization** – The `UnicodeProcessor` class creates `text_ids` and `text_mask` from raw text input.
2. **Duration Prediction** – The `duration_predictor` generates token-level timings, which are scaled by a speed factor to control speech rate.
3. **Text Encoding** – The `text_encoder` produces semantic embeddings that guide the acoustic generation.
4. **Latent Synthesis** – A random noisy latent is created via `sample_noisy_latent`, then iteratively refined by the `vector_estimator` across multiple diffusion steps.
5. **Waveform Generation** – The final latent representation passes through the `vocoder` to produce the output waveform.

All four models share the same configuration dictionary (`cfgs["ae"]`) and are loaded simultaneously by the `load_onnx_all` function (lines approximately 297-305 in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)). Because they are independent ONNX files, you can replace individual models or adjust their parameters without affecting the pipeline structure.

## Working with the Models in Code

Below is a complete example demonstrating how to load the four ONNX models and run inference through each stage manually. This assumes the ONNX files reside in `./models` and that dependencies (`onnxruntime`, `numpy`) are installed.

```python
import os, json, numpy as np, onnxruntime as ort
from py.helper import (
    load_cfgs, load_text_processor, load_onnx, load_onnx_all,
    UnicodeProcessor, Style, TextToSpeech,
)

onnx_dir = "./models"

# ----------------------------------------------------------------------

# 1️⃣ Load all four ONNX sessions (the same code the library uses)

# ----------------------------------------------------------------------

cfgs = load_cfgs(onnx_dir)
dp_ort, text_enc_ort, vector_est_ort, vocoder_ort = load_onnx_all(
    onnx_dir, ort.SessionOptions(), ["CPUExecutionProvider"]
)

# ----------------------------------------------------------------------

# 2️⃣ Prepare a tiny utterance

# ----------------------------------------------------------------------

text = "Hello world."
lang = "en"
unicode_path = os.path.join(onnx_dir, "unicode_indexer.json")
processor = UnicodeProcessor(unicode_path)

text_ids, text_mask = processor([text], [lang])
style = Style(
    # Minimal style tensors – normally loaded from a voice‑style file

    ttl=np.zeros((1, cfgs["ttl"]["latent_dim"], cfgs["ttl"]["seq_len"]), dtype=np.float32),
    dp=np.zeros((1, cfgs["dp"]["latent_dim"], cfgs["dp"]["seq_len"]), dtype=np.float32),
)

# ----------------------------------------------------------------------

# 3️⃣ Duration prediction (optional – you can skip it)

# ----------------------------------------------------------------------

dur_raw, *_ = dp_ort.run(
    None,
    {"text_ids": text_ids, "style_dp": style.dp, "text_mask": text_mask},
)
print("Predicted token durations (s):", dur_raw.squeeze())

# ----------------------------------------------------------------------

# 4️⃣ Text encoding

# ----------------------------------------------------------------------

text_emb, *_ = text_enc_ort.run(
    None,
    {"text_ids": text_ids, "style_ttl": style.ttl, "text_mask": text_mask},
)
print("Text embedding shape:", text_emb.shape)

# ----------------------------------------------------------------------

# 5️⃣ Latent refinement with the vector estimator

# ----------------------------------------------------------------------

# (this mirrors the loop inside TextToSpeech._infer)

latent, latent_mask = TextToSpeech(
    cfgs, processor, dp_ort, text_enc_ort, vector_est_ort, vocoder_ort
).sample_noisy_latent(np.array([dur_raw.squeeze().sum()]))
total_steps = 4
total_step_np = np.full((1,), total_steps, dtype=np.float32)

for step in range(total_steps):
    current_step = np.full((1,), step, dtype=np.float32)
    latent, *_ = vector_est_ort.run(
        None,
        {
            "noisy_latent": latent,
            "text_emb": text_emb,
            "style_ttl": style.ttl,
            "text_mask": text_mask,
            "latent_mask": latent_mask,
            "current_step": current_step,
            "total_step": total_step_np,
        },
    )

# ----------------------------------------------------------------------

# 6️⃣ Vocoder – final waveform

# ----------------------------------------------------------------------

wav, *_ = vocoder_ort.run(None, {"latent": latent})
wav_np = wav.squeeze()
print("Waveform samples:", wav_np.shape[0])

# ----------------------------------------------------------------------

# 7️⃣ End‑to‑end convenience (the library’s public API)

# ----------------------------------------------------------------------

tts = TextToSpeech(cfgs, processor, dp_ort, text_enc_ort, vector_est_ort, vocoder_ort)
wave, dur = tts(text, lang, style, total_step=4)
print("Finished! Waveform length (samples):", wave.shape[1])

```

The snippet above demonstrates three usage patterns: loading individual models with `load_onnx_all`, manually invoking each stage with specific inputs, and using the high-level `TextToSpeech` wrapper that handles the entire pipeline. For reference implementations in other languages, see [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js) (JavaScript), [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) (iOS), [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) (Go), [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) (Rust), [`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp) (C++), [`csharp/Helper.cs`](https://github.com/supertone-inc/supertonic/blob/main/csharp/Helper.cs) (C#), and `flutter/lib/helper.dart` (Dart).

## Summary

- **duration_predictor** calculates how long each character should be spoken, outputting a 1-D float array of durations used to scale the latent representation.
- **text_encoder** converts token IDs into dense semantic embeddings (`text_emb_onnx`) that provide linguistic conditioning for the acoustic model.
- **vector_estimator** iteratively denoises a random latent tensor through a diffusion process, using text embeddings and style vectors to guide the refinement.
- **vocoder** transforms the final latent tensor into a raw float32 waveform (`wav_tts`) ready for audio playback.
- All four models are loaded together via `load_onnx_all` in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) and executed sequentially in `TextToSpeech._infer`, but their independent ONNX format allows modular replacement and debugging.

## Frequently Asked Questions

### What is the difference between the vector_estimator and the vocoder?

The **vector_estimator** operates in the latent space, iteratively refining a noisy tensor into a structured acoustic representation through a diffusion-like process. It requires multiple inference steps (controlled by `total_step`) and conditioning inputs including text embeddings and style vectors. The **vocoder**, in contrast, performs a single-shot conversion from the final latent tensor to a raw audio waveform, serving as the final decoder that produces audible output.

### Can I replace the duration_predictor with my own timing model?

Yes. Since the `duration_predictor` is an independent ONNX file, you can substitute it with a custom model provided it accepts the same inputs (`text_ids`, `style_dp`, `text_mask`) and outputs a 1-D float array of durations compatible with the `dur_onnx` format expected by the pipeline. The `TextToSpeech` class in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) loads the model via `load_onnx_all`, making it straightforward to point to a different ONNX file path.

### Why does the vector_estimator require multiple inference steps while the other models run once?

The **vector_estimator** implements a diffusion or iterative refinement process where the latent tensor is gradually denoised across `total_step` iterations. Each step requires a separate call to `vector_est_ort.run()` with an updated `current_step` parameter,不同于 the feed-forward architectures of the duration_predictor, text_encoder, and vocoder, which produce deterministic outputs in a single pass. This multi-step approach improves the quality and stability of the generated acoustic latents.

### How do I debug intermediate outputs between the ONNX models?

You can manually invoke each model using the `run()` method on the ONNX Runtime sessions (`dp_ort`, `text_enc_ort`, `vector_est_ort`, `vocoder_ort`) as shown in the code example above. This allows you to inspect the `dur_onnx`, `text_emb_onnx`, and intermediate `xt` tensors before they pass to the next stage. The [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) file provides a minimal demonstration of this approach, isolate specific stages for verification.