How Supertonic’s ONNX Inference Pipeline Works: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder Explained

Supertonic’s ONNX inference pipeline converts text to speech through four sequential stages—duration prediction, text encoding, iterative latent refinement via diffusion, and waveform generation—each executed by separate ONNX models orchestrated in py/helper.py.

The supertone-inc/supertonic repository implements a high-performance text-to-speech (TTS) engine using a modular ONNX architecture. This pipeline decouples the core synthesis components into independent runtime sessions, enabling efficient CPU-based inference while maintaining architectural flexibility. Understanding how the duration predictor, text encoder, vector estimator, and vocoder interact is essential for optimizing synthesis quality and runtime performance.

The Four-Stage ONNX Architecture

Supertonic’s inference stack is defined in py/helper.py and encapsulated by the TextToSpeech class. The pipeline processes phonetic information through four distinct ONNX graphs, with each stage feeding specific tensors into the next.

Stage 1: Duration Predictor

The duration predictor (duration_predictor.onnx) estimates the temporal length of the target utterance. It consumes tokenized text representations alongside a speaker-style vector and outputs a raw duration prediction in seconds.

During inference, this prediction is divided by the user-supplied speed factor to accommodate faster or slower speech rates. In the TextToSpeech._infer method, this execution occurs via dp_ort.run at lines 90–93, producing the temporal foundation that governs subsequent latent generation.

Stage 2: Text Encoder

The text encoder (text_encoder.onnx) projects discrete token IDs into a continuous latent-space representation (text_emb_onnx). This embedding carries phonetic and linguistic features required for acoustic generation.

Like the duration predictor, the encoder receives the speaker’s tone-level (TTL) style vector to ensure prosodic consistency. The inference session text_enc_ort.run executes at lines 94–98 in helper.py, transforming the masked token sequence into a high-dimensional embedding that guides the diffusion process.

Stage 3: Vector Estimator (Diffusion Denoiser)

The vector estimator (vector_estimator.onnx) implements an iterative diffusion-based refinement process. Starting from a random Gaussian latent generated by sample_noisy_latent, the model progressively denoises the representation across total_step iterations.

Each iteration calls vector_est_ort.run (lines 100–113), accepting the current latent, text embedding, style tensors, masks, and step counters as inputs. This stage is computationally intensive, as the denoising loop executes sequentially until reaching the final, clean latent representation ready for audio generation.

Stage 4: Vocoder

The vocoder (vocoder.onnx) serves as the final rendering stage, converting the denoised latent into a raw audio waveform. This session executes once per inference batch via vocoder_ort.run at lines 114–115, outputting a NumPy array of audio samples at the configured sample rate.

End-to-End Inference Flow

The complete synthesis process involves several preparatory steps before the four ONNX stages execute.

Preprocessing and Tokenization

The UnicodeProcessor class normalizes input strings, strips emojis, and wraps text with language tags (e.g., <en>…</en>). It then maps Unicode code points to integer token IDs using the unicode_indexer.json mapping, producing the text_ids and binary text_mask tensors required by the encoder and duration predictor.

Style Vector Loading

Voice characteristics are injected via load_voice_style, which reads JSON voice-style files from disk and packs the tone-level (TTL) and duration-level (DP) vectors into a Style object. These tensors condition both the duration predictor and the diffusion process to match the target speaker’s prosody.

ONNX Session Initialization

The load_text_to_speech function constructs four onnxruntime.InferenceSession objects (dp_ort, text_enc_ort, vector_est_ort, vocoder_ort) targeting the ONNX assets directory. By default, all sessions run on CPU, though the architecture supports GPU acceleration through session configuration.

The _infer Method Execution

The core inference logic resides in TextToSpeech._infer:

  1. Duration Prediction: Generates temporal boundaries for the utterance.
  2. Text Encoding: Creates conditioned embeddings from tokenized input.
  3. Latent Initialization: sample_noisy_latent creates a Gaussian tensor shaped (batch, latent_dim, latent_len), masked to the predicted audio length.
  4. Diffusion Loop: Iterates total_step times, calling the vector estimator to refine the latent.
  5. Waveform Synthesis: Passes the final latent to the vocoder, producing the output wav.

In single-utterance mode (__call__), long texts are automatically chunked, synthesized separately, and concatenated with brief silences. Batch mode (batch) returns raw waveforms and durations directly without chunking.

Implementation Examples

Single Utterance Synthesis

from helper import load_text_to_speech, load_voice_style

# Initialize the four-model ONNX pipeline

tts = load_text_to_speech("../assets/onnx", use_gpu=False)

# Load speaker style vectors (TTL + DP)

style = load_voice_style(["../assets/voice_styles/M1.json"], verbose=True)

# Generate speech with 8 diffusion steps

wav, duration = tts(
    text="Hello, world! This is Supertonic speaking.",
    lang="en",
    style=style,
    total_step=8,
    speed=1.05,  # 5% faster than predicted

)

# Export to WAV

import soundfile as sf
sf.write("hello.wav", wav[0], tts.sample_rate)

Batch Processing

from helper import load_text_to_speech, load_voice_style

tts = load_text_to_speech("../assets/onnx")
style = load_voice_style(
    ["../assets/voice_styles/M1.json", "../assets/voice_styles/F2.json"]
)

# Process multiple texts with different languages

texts = [
    "Good morning, how are you?",
    "¡Buenos días! ¿Cómo estás?"
]
langs = ["en", "es"]

# Batch inference returns arrays per input

wavs, durs = tts.batch(texts, langs, style, total_step=8, speed=1.0)

# Save outputs

import soundfile as sf
for i, w in enumerate(wavs):
    sf.write(f"out_{i}.wav", w, tts.sample_rate)

Summary

  • Supertonic’s ONNX inference pipeline separates TTS functionality into four specialized models: duration predictor, text encoder, vector estimator, and vocoder.
  • The duration predictor establishes temporal boundaries, while the text encoder generates linguistically informed embeddings.
  • The vector estimator performs iterative diffusion denoising over total_step iterations to refine the acoustic latent.
  • The vocoder completes the chain by rendering the latent into a final waveform.
  • All stages are orchestrated in py/helper.py by the TextToSpeech class, supporting both single-utterance chunking and batch processing modes.

Frequently Asked Questions

What is the role of the Vector Estimator in Supertonic’s pipeline?

The Vector Estimator implements a diffusion-based denoising process that refines a random Gaussian latent into a structured acoustic representation. It runs iteratively for total_step cycles, using the text embedding and speaker style vectors to condition the final output, effectively determining the timbre and quality of the synthesized speech.

How does the Duration Predictor handle speech speed adjustments?

The Duration Predictor outputs raw duration estimates in seconds, which are immediately divided by the user-supplied speed parameter in TextToSpeech._infer. Values greater than 1.0 shorten the predicted duration for faster speech, while values below 1.0 extend it for slower, more deliberate pronunciation.

Can I run the Supertonic ONNX pipeline on GPU?

While the load_text_to_speech function initializes onnxruntime.InferenceSession objects for CPU by default, the underlying ONNX Runtime supports GPU execution. You can modify the session creation in py/helper.py to pass providers=['CUDAExecutionProvider'] or other GPU-specific configurations to the InferenceSession constructor.

What text preprocessing steps does Supertonic apply before ONNX inference?

Supertonic’s UnicodeProcessor normalizes Unicode input, strips unsupported emoji characters, wraps text with language identifiers (e.g., <en>), and maps characters to integer token IDs using unicode_indexer.json. This produces the text_ids and text_mask tensors fed into both the duration predictor and text encoder.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →