Differences Between the ONNX Models in Supertonic TTS: duration_predictor, text_encoder, vector_estimator, and vocoder

The four ONNX models in Supertonic represent distinct stages of a neural text-to-speech pipeline: duration_predictor calculates token timings, text_encoder generates semantic embeddings, vector_estimator refines latent representations through a diffusion process, and vocoder synthesizes the final audio waveform.

Supertonic by supertone-inc implements a multi-stage neural TTS architecture where each component is encapsulated as an independent ONNX model. These models are orchestrated together in the TextToSpeech class defined in py/helper.py, allowing developers to swap or retrain individual stages without modifying the surrounding Python code. Understanding the specific role, inputs, and outputs of each model is essential for debugging, optimizing, or extending the synthesis pipeline.

Overview of the Four ONNX Models

duration_predictor

The duration_predictor determines how long each text token (character) should be spoken in seconds. It takes text_ids (tokenized character IDs), style_dp (a speaker-specific style vector), and text_mask as inputs. The model outputs dur_onnx, a 1-D float array containing the predicted duration for each token. In py/helper.py, this model is invoked via dp_ort.run() at approximately line 190 within the TextToSpeech._infer method.

text_encoder

The text_encoder transforms discrete token IDs into dense semantic embeddings that condition the acoustic model. Its inputs include text_ids, style_ttl (tone/style vector), and text_mask. The output is text_emb_onnx, a 2-D tensor with shape [batch, embed_dim] that captures the linguistic content of the input text. This embedding serves as the primary conditioning signal for the latent refinement stage. The encoder runs via text_enc_ort.run() at approximately line 194.

vector_estimator

The vector_estimator implements a diffusion-like refinement process that denoises a latent representation into a structured acoustic encoding. It accepts noisy_latent (the initial random tensor), text_emb (from the text encoder), style_ttl, text_mask, latent_mask, current_step, and total_step as inputs. The model outputs an updated xt tensor, with the process iterating for a configurable number of steps (controlled by total_step). In py/helper.py, this appears as vector_est_ort.run() inside a loop spanning lines approximately 200-213.

vocoder

The vocoder converts the final refined latent tensor into audible raw audio. It takes a single input, latent (the output from the vector estimator), and produces wav_tts, a float32 waveform tensor with shape [1, samples]. This represents the final audio output ready for playback or saving to disk. The vocoder executes via vocoder_ort.run() at approximately line 214.

How the Models Connect in the Pipeline

The four models execute sequentially within the TextToSpeech._infer method, with each stage feeding into the next:

  1. Tokenization – The UnicodeProcessor class creates text_ids and text_mask from raw text input.
  2. Duration Prediction – The duration_predictor generates token-level timings, which are scaled by a speed factor to control speech rate.
  3. Text Encoding – The text_encoder produces semantic embeddings that guide the acoustic generation.
  4. Latent Synthesis – A random noisy latent is created via sample_noisy_latent, then iteratively refined by the vector_estimator across multiple diffusion steps.
  5. Waveform Generation – The final latent representation passes through the vocoder to produce the output waveform.

All four models share the same configuration dictionary (cfgs["ae"]) and are loaded simultaneously by the load_onnx_all function (lines approximately 297-305 in py/helper.py). Because they are independent ONNX files, you can replace individual models or adjust their parameters without affecting the pipeline structure.

Working with the Models in Code

Below is a complete example demonstrating how to load the four ONNX models and run inference through each stage manually. This assumes the ONNX files reside in ./models and that dependencies (onnxruntime, numpy) are installed.

import os, json, numpy as np, onnxruntime as ort
from py.helper import (
    load_cfgs, load_text_processor, load_onnx, load_onnx_all,
    UnicodeProcessor, Style, TextToSpeech,
)

onnx_dir = "./models"

# ----------------------------------------------------------------------

# 1️⃣ Load all four ONNX sessions (the same code the library uses)

# ----------------------------------------------------------------------

cfgs = load_cfgs(onnx_dir)
dp_ort, text_enc_ort, vector_est_ort, vocoder_ort = load_onnx_all(
    onnx_dir, ort.SessionOptions(), ["CPUExecutionProvider"]
)

# ----------------------------------------------------------------------

# 2️⃣ Prepare a tiny utterance

# ----------------------------------------------------------------------

text = "Hello world."
lang = "en"
unicode_path = os.path.join(onnx_dir, "unicode_indexer.json")
processor = UnicodeProcessor(unicode_path)

text_ids, text_mask = processor([text], [lang])
style = Style(
    # Minimal style tensors – normally loaded from a voice‑style file

    ttl=np.zeros((1, cfgs["ttl"]["latent_dim"], cfgs["ttl"]["seq_len"]), dtype=np.float32),
    dp=np.zeros((1, cfgs["dp"]["latent_dim"], cfgs["dp"]["seq_len"]), dtype=np.float32),
)

# ----------------------------------------------------------------------

# 3️⃣ Duration prediction (optional – you can skip it)

# ----------------------------------------------------------------------

dur_raw, *_ = dp_ort.run(
    None,
    {"text_ids": text_ids, "style_dp": style.dp, "text_mask": text_mask},
)
print("Predicted token durations (s):", dur_raw.squeeze())

# ----------------------------------------------------------------------

# 4️⃣ Text encoding

# ----------------------------------------------------------------------

text_emb, *_ = text_enc_ort.run(
    None,
    {"text_ids": text_ids, "style_ttl": style.ttl, "text_mask": text_mask},
)
print("Text embedding shape:", text_emb.shape)

# ----------------------------------------------------------------------

# 5️⃣ Latent refinement with the vector estimator

# ----------------------------------------------------------------------

# (this mirrors the loop inside TextToSpeech._infer)

latent, latent_mask = TextToSpeech(
    cfgs, processor, dp_ort, text_enc_ort, vector_est_ort, vocoder_ort
).sample_noisy_latent(np.array([dur_raw.squeeze().sum()]))
total_steps = 4
total_step_np = np.full((1,), total_steps, dtype=np.float32)

for step in range(total_steps):
    current_step = np.full((1,), step, dtype=np.float32)
    latent, *_ = vector_est_ort.run(
        None,
        {
            "noisy_latent": latent,
            "text_emb": text_emb,
            "style_ttl": style.ttl,
            "text_mask": text_mask,
            "latent_mask": latent_mask,
            "current_step": current_step,
            "total_step": total_step_np,
        },
    )

# ----------------------------------------------------------------------

# 6️⃣ Vocoder – final waveform

# ----------------------------------------------------------------------

wav, *_ = vocoder_ort.run(None, {"latent": latent})
wav_np = wav.squeeze()
print("Waveform samples:", wav_np.shape[0])

# ----------------------------------------------------------------------

# 7️⃣ End‑to‑end convenience (the library’s public API)

# ----------------------------------------------------------------------

tts = TextToSpeech(cfgs, processor, dp_ort, text_enc_ort, vector_est_ort, vocoder_ort)
wave, dur = tts(text, lang, style, total_step=4)
print("Finished! Waveform length (samples):", wave.shape[1])

The snippet above demonstrates three usage patterns: loading individual models with load_onnx_all, manually invoking each stage with specific inputs, and using the high-level TextToSpeech wrapper that handles the entire pipeline. For reference implementations in other languages, see nodejs/helper.js (JavaScript), swift/Sources/Helper.swift (iOS), go/helper.go (Go), rust/src/helper.rs (Rust), cpp/helper.cpp (C++), csharp/Helper.cs (C#), and flutter/lib/helper.dart (Dart).

Summary

  • duration_predictor calculates how long each character should be spoken, outputting a 1-D float array of durations used to scale the latent representation.
  • text_encoder converts token IDs into dense semantic embeddings (text_emb_onnx) that provide linguistic conditioning for the acoustic model.
  • vector_estimator iteratively denoises a random latent tensor through a diffusion process, using text embeddings and style vectors to guide the refinement.
  • vocoder transforms the final latent tensor into a raw float32 waveform (wav_tts) ready for audio playback.
  • All four models are loaded together via load_onnx_all in py/helper.py and executed sequentially in TextToSpeech._infer, but their independent ONNX format allows modular replacement and debugging.

Frequently Asked Questions

What is the difference between the vector_estimator and the vocoder?

The vector_estimator operates in the latent space, iteratively refining a noisy tensor into a structured acoustic representation through a diffusion-like process. It requires multiple inference steps (controlled by total_step) and conditioning inputs including text embeddings and style vectors. The vocoder, in contrast, performs a single-shot conversion from the final latent tensor to a raw audio waveform, serving as the final decoder that produces audible output.

Can I replace the duration_predictor with my own timing model?

Yes. Since the duration_predictor is an independent ONNX file, you can substitute it with a custom model provided it accepts the same inputs (text_ids, style_dp, text_mask) and outputs a 1-D float array of durations compatible with the dur_onnx format expected by the pipeline. The TextToSpeech class in py/helper.py loads the model via load_onnx_all, making it straightforward to point to a different ONNX file path.

Why does the vector_estimator require multiple inference steps while the other models run once?

The vector_estimator implements a diffusion or iterative refinement process where the latent tensor is gradually denoised across total_step iterations. Each step requires a separate call to vector_est_ort.run() with an updated current_step parameter,不同于 the feed-forward architectures of the duration_predictor, text_encoder, and vocoder, which produce deterministic outputs in a single pass. This multi-step approach improves the quality and stability of the generated acoustic latents.

How do I debug intermediate outputs between the ONNX models?

You can manually invoke each model using the run() method on the ONNX Runtime sessions (dp_ort, text_enc_ort, vector_est_ort, vocoder_ort) as shown in the code example above. This allows you to inspect the dur_onnx, text_emb_onnx, and intermediate xt tensors before they pass to the next stage. The py/example_onnx.py file provides a minimal demonstration of this approach, isolate specific stages for verification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →