# How Supertonic’s ONNX Inference Pipeline Works: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder Explained

> Explore Supertonic's ONNX inference pipeline. Understand how duration predictor, text encoder, vector estimator, and vocoder work together for advanced text to speech generation.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: internals
- Published: 2026-06-15

---

**Supertonic’s ONNX inference pipeline converts text to speech through four sequential stages—duration prediction, text encoding, iterative latent refinement via diffusion, and waveform generation—each executed by separate ONNX models orchestrated in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py).**

The `supertone-inc/supertonic` repository implements a high-performance text-to-speech (TTS) engine using a modular ONNX architecture. This pipeline decouples the core synthesis components into independent runtime sessions, enabling efficient CPU-based inference while maintaining architectural flexibility. Understanding how the **duration predictor**, **text encoder**, **vector estimator**, and **vocoder** interact is essential for optimizing synthesis quality and runtime performance.

## The Four-Stage ONNX Architecture

Supertonic’s inference stack is defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) and encapsulated by the `TextToSpeech` class. The pipeline processes phonetic information through four distinct ONNX graphs, with each stage feeding specific tensors into the next.

### Stage 1: Duration Predictor

The **duration predictor** (`duration_predictor.onnx`) estimates the temporal length of the target utterance. It consumes tokenized text representations alongside a speaker-style vector and outputs a raw duration prediction in seconds. 

During inference, this prediction is divided by the user-supplied `speed` factor to accommodate faster or slower speech rates. In the `TextToSpeech._infer` method, this execution occurs via `dp_ort.run` at lines 90–93, producing the temporal foundation that governs subsequent latent generation.

### Stage 2: Text Encoder

The **text encoder** (`text_encoder.onnx`) projects discrete token IDs into a continuous latent-space representation (`text_emb_onnx`). This embedding carries phonetic and linguistic features required for acoustic generation.

Like the duration predictor, the encoder receives the speaker’s tone-level (TTL) style vector to ensure prosodic consistency. The inference session `text_enc_ort.run` executes at lines 94–98 in [`helper.py`](https://github.com/supertone-inc/supertonic/blob/main/helper.py), transforming the masked token sequence into a high-dimensional embedding that guides the diffusion process.

### Stage 3: Vector Estimator (Diffusion Denoiser)

The **vector estimator** (`vector_estimator.onnx`) implements an iterative diffusion-based refinement process. Starting from a random Gaussian latent generated by `sample_noisy_latent`, the model progressively denoises the representation across `total_step` iterations.

Each iteration calls `vector_est_ort.run` (lines 100–113), accepting the current latent, text embedding, style tensors, masks, and step counters as inputs. This stage is computationally intensive, as the denoising loop executes sequentially until reaching the final, clean latent representation ready for audio generation.

### Stage 4: Vocoder

The **vocoder** (`vocoder.onnx`) serves as the final rendering stage, converting the denoised latent into a raw audio waveform. This session executes once per inference batch via `vocoder_ort.run` at lines 114–115, outputting a NumPy array of audio samples at the configured sample rate.

## End-to-End Inference Flow

The complete synthesis process involves several preparatory steps before the four ONNX stages execute.

### Preprocessing and Tokenization

The `UnicodeProcessor` class normalizes input strings, strips emojis, and wraps text with language tags (e.g., `<en>…</en>`). It then maps Unicode code points to integer token IDs using the [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) mapping, producing the `text_ids` and binary `text_mask` tensors required by the encoder and duration predictor.

### Style Vector Loading

Voice characteristics are injected via `load_voice_style`, which reads JSON voice-style files from disk and packs the tone-level (TTL) and duration-level (DP) vectors into a `Style` object. These tensors condition both the duration predictor and the diffusion process to match the target speaker’s prosody.

### ONNX Session Initialization

The `load_text_to_speech` function constructs four `onnxruntime.InferenceSession` objects (`dp_ort`, `text_enc_ort`, `vector_est_ort`, `vocoder_ort`) targeting the ONNX assets directory. By default, all sessions run on CPU, though the architecture supports GPU acceleration through session configuration.

### The `_infer` Method Execution

The core inference logic resides in `TextToSpeech._infer`:

1. **Duration Prediction**: Generates temporal boundaries for the utterance.
2. **Text Encoding**: Creates conditioned embeddings from tokenized input.
3. **Latent Initialization**: `sample_noisy_latent` creates a Gaussian tensor shaped `(batch, latent_dim, latent_len)`, masked to the predicted audio length.
4. **Diffusion Loop**: Iterates `total_step` times, calling the vector estimator to refine the latent.
5. **Waveform Synthesis**: Passes the final latent to the vocoder, producing the output `wav`.

In single-utterance mode (`__call__`), long texts are automatically chunked, synthesized separately, and concatenated with brief silences. Batch mode (`batch`) returns raw waveforms and durations directly without chunking.

## Implementation Examples

### Single Utterance Synthesis

```python
from helper import load_text_to_speech, load_voice_style

# Initialize the four-model ONNX pipeline

tts = load_text_to_speech("../assets/onnx", use_gpu=False)

# Load speaker style vectors (TTL + DP)

style = load_voice_style(["../assets/voice_styles/M1.json"], verbose=True)

# Generate speech with 8 diffusion steps

wav, duration = tts(
    text="Hello, world! This is Supertonic speaking.",
    lang="en",
    style=style,
    total_step=8,
    speed=1.05,  # 5% faster than predicted

)

# Export to WAV

import soundfile as sf
sf.write("hello.wav", wav[0], tts.sample_rate)

```

### Batch Processing

```python
from helper import load_text_to_speech, load_voice_style

tts = load_text_to_speech("../assets/onnx")
style = load_voice_style(
    ["../assets/voice_styles/M1.json", "../assets/voice_styles/F2.json"]
)

# Process multiple texts with different languages

texts = [
    "Good morning, how are you?",
    "¡Buenos días! ¿Cómo estás?"
]
langs = ["en", "es"]

# Batch inference returns arrays per input

wavs, durs = tts.batch(texts, langs, style, total_step=8, speed=1.0)

# Save outputs

import soundfile as sf
for i, w in enumerate(wavs):
    sf.write(f"out_{i}.wav", w, tts.sample_rate)

```

## Summary

- **Supertonic’s ONNX inference pipeline** separates TTS functionality into four specialized models: duration predictor, text encoder, vector estimator, and vocoder.
- The **duration predictor** establishes temporal boundaries, while the **text encoder** generates linguistically informed embeddings.
- The **vector estimator** performs iterative diffusion denoising over `total_step` iterations to refine the acoustic latent.
- The **vocoder** completes the chain by rendering the latent into a final waveform.
- All stages are orchestrated in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) by the `TextToSpeech` class, supporting both single-utterance chunking and batch processing modes.

## Frequently Asked Questions

### What is the role of the Vector Estimator in Supertonic’s pipeline?

The Vector Estimator implements a diffusion-based denoising process that refines a random Gaussian latent into a structured acoustic representation. It runs iteratively for `total_step` cycles, using the text embedding and speaker style vectors to condition the final output, effectively determining the timbre and quality of the synthesized speech.

### How does the Duration Predictor handle speech speed adjustments?

The Duration Predictor outputs raw duration estimates in seconds, which are immediately divided by the user-supplied `speed` parameter in `TextToSpeech._infer`. Values greater than 1.0 shorten the predicted duration for faster speech, while values below 1.0 extend it for slower, more deliberate pronunciation.

### Can I run the Supertonic ONNX pipeline on GPU?

While the `load_text_to_speech` function initializes `onnxruntime.InferenceSession` objects for CPU by default, the underlying ONNX Runtime supports GPU execution. You can modify the session creation in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) to pass `providers=['CUDAExecutionProvider']` or other GPU-specific configurations to the InferenceSession constructor.

### What text preprocessing steps does Supertonic apply before ONNX inference?

Supertonic’s `UnicodeProcessor` normalizes Unicode input, strips unsupported emoji characters, wraps text with language identifiers (e.g., `<en>`), and maps characters to integer token IDs using [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json). This produces the `text_ids` and `text_mask` tensors fed into both the duration predictor and text encoder.