# How Supertonic's ONNX Inference Pipeline Works Internally

> Explore Supertonic's ONNX inference pipeline. Discover how it uses four ONNX models, a Rust core, and Python bindings to convert text to audio through an iterative denoising loop.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: internals
- Published: 2026-06-12

---

**Supertonic's ONNX inference pipeline processes text-to-speech synthesis by chaining four ONNX models—duration predictor, text encoder, vector estimator, and vocoder—through a Rust core with Python bindings, converting normalized Unicode text into raw audio via an iterative denoising loop.**

Supertonic is an open-source text-to-speech engine developed by Supertone Inc. that leverages ONNX Runtime for high-performance inference. The pipeline is implemented primarily in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) with a mirrored Python implementation in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), enabling language-agnostic deployment while maintaining consistent behavior across both environments.

## Configuration and Model Loading

The pipeline initialization begins by reading [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json) to extract sample-rate and chunk-size parameters through the `load_cfgs` function (lines 46-52). Subsequently, the `load_text_to_speech` function (lines 805-819) loads four ONNX model files into `ort::Session` objects:

- `duration_predictor.onnx`
- `text_encoder.onnx`
- `vector_estimator.onnx` (denoiser)
- `vocoder.onnx`

These sessions remain cached throughout the synthesis process to avoid reloading overhead.

## Unicode Text Processing

Input strings undergo normalization through the `UnicodeProcessor::call` method, which chains `preprocess_text` (lines 89-112) to apply NFKD normalization, strip emojis and punctuation, and inject language tags. Each character is then mapped to integer IDs using `text_to_unicode_values` (lines 115-117), which references the pre-computed [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) dictionary.

A binary mask indicating valid token positions is generated via `get_text_mask` (lines 32-35), ensuring the model ignores padding positions during attention computations.

## Duration Prediction

The `duration_predictor` receives three inputs: the `text_ids` tensor, a style-specific DP (duration predictor) tensor, and the text mask. The `dp_ort.run` execution (lines 602-607) returns raw duration values for each token. These durations are then rescaled by a user-specified speed factor in the `apply speed` logic (lines 611-614), allowing temporal stretching or compression of the output speech.

## Text Encoding

With durations established, the `text_encoder` processes the same `text_ids`, a style-specific TTL tensor, and the mask through `text_enc_ort.run` (lines 618-622). This produces a dense `text_emb` tensor that captures phonetic and linguistic context for the subsequent denoising phase.

## Latent Sampling and Denoising Loop

The pipeline generates a zero-mean Gaussian latent tensor sized to match the predicted total duration via `sample_noisy_latent` (lines 37-88). This latent representation undergoes refinement through an iterative denoising process (defaulting to 8 steps).

For each step, the `vector_estimator` executes `vector_est_ort.run` (lines 442-466) within the loop, consuming the current latent tensor, step index, text embedding, style tensors, and both text and latent masks. This diffusion-inspired process gradually denoises the latent representation into a structured audio embedding.

## Vocoder and Waveform Generation

The final denoised latent tensor feeds directly into the `vocoder` via `vocoder_ort.run` (lines 670-674). Unlike intermediate stages that manipulate latent representations, the vocoder outputs a raw floating-point waveform representing the actual audio signal.

## Chunking and Silence Insertion

For long inputs, `chunk_text` (lines 30-49) segments text at paragraph or sentence boundaries while respecting an abbreviation list to prevent spurious splits. In single-speaker mode, the pipeline optionally inserts short silence intervals (default 0.3 seconds) between chunks through the `call` method's silence handling logic (lines 93-115), creating natural pauses in extended utterances.

## Batch Processing vs Single-Speaker Mode

Supertonic supports two execution modes. When the `--batch` flag is active, `TextToSpeech::batch` (lines 220-226) bypasses chunking and processes each `(text, style, lang)` triple independently, generating parallel outputs. In standard single-speaker mode, the pipeline iterates over chunks sequentially, concatenating results with optional silence padding.

## Output Writing

The final floating-point waveform is persisted to disk using `write_wav_file` (lines 94-108), which utilizes the `hound` crate to encode 16-bit PCM WAV files. This occurs after all processing stages complete, ensuring the entire pipeline executes before I/O operations begin.

## Running the Inference Pipeline

The following commands demonstrate complete inference workflows for both Rust and Python implementations.

### Rust CLI Example

```bash
cargo run --release --bin example_onnx -- \
  --onnx-dir ../assets/onnx \
  --text "Hello world! This is a test of Supertonic." \
  --lang en \
  --voice-style ../assets/voice_styles/M1.json \
  --total-step 8 \
  --speed 1.05

```

### Python Batch Example

```bash
python py/example_onnx.py \
  --onnx-dir ../assets/onnx \
  --batch \
  --voice-style ../assets/voice_styles/M1.json ../assets/voice_styles/F1.json \
  --text "The sun sets." "夕日が沈む。" \
  --lang en ko \
  --total-step 10 \
  --speed 1.0

```

Both examples load the same ONNX models from `assets/onnx/`, preprocess input text, execute the full denoising loop, and write resulting WAV files to the `results/` directory.

## Summary

- Supertonic's ONNX inference pipeline operates through four specialized models loaded via `load_text_to_speech` in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)
- Text processing combines NFKD normalization with [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) mapping to create integer token sequences
- Duration prediction establishes temporal structure before the text encoder generates contextual embeddings
- An iterative denoising loop (default 8 steps) refines Gaussian latents through the `vector_estimator` ONNX model
- The `vocoder` produces final audio waveforms, with optional chunking and silence insertion for long-form content
- Both Rust and Python implementations share identical logic, differing only in runtime bindings (ORT vs. onnxruntime)

## Frequently Asked Questions

### How does Supertonic handle multilingual text input?

The pipeline normalizes input using NFKD Unicode normalization and strips emojis through `preprocess_text` (lines 89-112), then maps characters to IDs via [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json). Language tags are injected during preprocessing, allowing the same four ONNX models to process multiple languages without separate model loading.

### What is the purpose of the vector estimator in the pipeline?

The `vector_estimator` (denoiser) performs iterative refinement of the audio latent representation. During each of the `total-step` iterations (default 8), it executes `vector_est_ort.run` (lines 442-466) to denoise the latent tensor using text embeddings and style conditioning, similar to a diffusion model's sampling process.

### Can I adjust the speech speed without retraining the models?

Yes. The pipeline applies a user-specified speed factor to rescale the raw durations output by the `duration_predictor`. This occurs in `apply speed` (lines 611-614) before the latent sampling phase, enabling real-time speed adjustment between 0.5x and 2.0x without model modification.

### Where does the actual audio generation happen in the codebase?

The final waveform generation occurs in `vocoder_ort.run` (lines 670-674) within [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) (or the equivalent Python helper). This ONNX session takes the denoised latent tensor and outputs raw audio samples, which are then written to WAV format via `write_wav_file` using the `hound` crate.