How Supertonic's ONNX Inference Pipeline Works Internally

Supertonic's ONNX inference pipeline processes text-to-speech synthesis by chaining four ONNX models—duration predictor, text encoder, vector estimator, and vocoder—through a Rust core with Python bindings, converting normalized Unicode text into raw audio via an iterative denoising loop.

Supertonic is an open-source text-to-speech engine developed by Supertone Inc. that leverages ONNX Runtime for high-performance inference. The pipeline is implemented primarily in rust/src/helper.rs with a mirrored Python implementation in py/helper.py, enabling language-agnostic deployment while maintaining consistent behavior across both environments.

Configuration and Model Loading

The pipeline initialization begins by reading tts.json to extract sample-rate and chunk-size parameters through the load_cfgs function (lines 46-52). Subsequently, the load_text_to_speech function (lines 805-819) loads four ONNX model files into ort::Session objects:

  • duration_predictor.onnx
  • text_encoder.onnx
  • vector_estimator.onnx (denoiser)
  • vocoder.onnx

These sessions remain cached throughout the synthesis process to avoid reloading overhead.

Unicode Text Processing

Input strings undergo normalization through the UnicodeProcessor::call method, which chains preprocess_text (lines 89-112) to apply NFKD normalization, strip emojis and punctuation, and inject language tags. Each character is then mapped to integer IDs using text_to_unicode_values (lines 115-117), which references the pre-computed unicode_indexer.json dictionary.

A binary mask indicating valid token positions is generated via get_text_mask (lines 32-35), ensuring the model ignores padding positions during attention computations.

Duration Prediction

The duration_predictor receives three inputs: the text_ids tensor, a style-specific DP (duration predictor) tensor, and the text mask. The dp_ort.run execution (lines 602-607) returns raw duration values for each token. These durations are then rescaled by a user-specified speed factor in the apply speed logic (lines 611-614), allowing temporal stretching or compression of the output speech.

Text Encoding

With durations established, the text_encoder processes the same text_ids, a style-specific TTL tensor, and the mask through text_enc_ort.run (lines 618-622). This produces a dense text_emb tensor that captures phonetic and linguistic context for the subsequent denoising phase.

Latent Sampling and Denoising Loop

The pipeline generates a zero-mean Gaussian latent tensor sized to match the predicted total duration via sample_noisy_latent (lines 37-88). This latent representation undergoes refinement through an iterative denoising process (defaulting to 8 steps).

For each step, the vector_estimator executes vector_est_ort.run (lines 442-466) within the loop, consuming the current latent tensor, step index, text embedding, style tensors, and both text and latent masks. This diffusion-inspired process gradually denoises the latent representation into a structured audio embedding.

Vocoder and Waveform Generation

The final denoised latent tensor feeds directly into the vocoder via vocoder_ort.run (lines 670-674). Unlike intermediate stages that manipulate latent representations, the vocoder outputs a raw floating-point waveform representing the actual audio signal.

Chunking and Silence Insertion

For long inputs, chunk_text (lines 30-49) segments text at paragraph or sentence boundaries while respecting an abbreviation list to prevent spurious splits. In single-speaker mode, the pipeline optionally inserts short silence intervals (default 0.3 seconds) between chunks through the call method's silence handling logic (lines 93-115), creating natural pauses in extended utterances.

Batch Processing vs Single-Speaker Mode

Supertonic supports two execution modes. When the --batch flag is active, TextToSpeech::batch (lines 220-226) bypasses chunking and processes each (text, style, lang) triple independently, generating parallel outputs. In standard single-speaker mode, the pipeline iterates over chunks sequentially, concatenating results with optional silence padding.

Output Writing

The final floating-point waveform is persisted to disk using write_wav_file (lines 94-108), which utilizes the hound crate to encode 16-bit PCM WAV files. This occurs after all processing stages complete, ensuring the entire pipeline executes before I/O operations begin.

Running the Inference Pipeline

The following commands demonstrate complete inference workflows for both Rust and Python implementations.

Rust CLI Example

cargo run --release --bin example_onnx -- \
  --onnx-dir ../assets/onnx \
  --text "Hello world! This is a test of Supertonic." \
  --lang en \
  --voice-style ../assets/voice_styles/M1.json \
  --total-step 8 \
  --speed 1.05

Python Batch Example

python py/example_onnx.py \
  --onnx-dir ../assets/onnx \
  --batch \
  --voice-style ../assets/voice_styles/M1.json ../assets/voice_styles/F1.json \
  --text "The sun sets." "夕日が沈む。" \
  --lang en ko \
  --total-step 10 \
  --speed 1.0

Both examples load the same ONNX models from assets/onnx/, preprocess input text, execute the full denoising loop, and write resulting WAV files to the results/ directory.

Summary

  • Supertonic's ONNX inference pipeline operates through four specialized models loaded via load_text_to_speech in rust/src/helper.rs
  • Text processing combines NFKD normalization with unicode_indexer.json mapping to create integer token sequences
  • Duration prediction establishes temporal structure before the text encoder generates contextual embeddings
  • An iterative denoising loop (default 8 steps) refines Gaussian latents through the vector_estimator ONNX model
  • The vocoder produces final audio waveforms, with optional chunking and silence insertion for long-form content
  • Both Rust and Python implementations share identical logic, differing only in runtime bindings (ORT vs. onnxruntime)

Frequently Asked Questions

How does Supertonic handle multilingual text input?

The pipeline normalizes input using NFKD Unicode normalization and strips emojis through preprocess_text (lines 89-112), then maps characters to IDs via unicode_indexer.json. Language tags are injected during preprocessing, allowing the same four ONNX models to process multiple languages without separate model loading.

What is the purpose of the vector estimator in the pipeline?

The vector_estimator (denoiser) performs iterative refinement of the audio latent representation. During each of the total-step iterations (default 8), it executes vector_est_ort.run (lines 442-466) to denoise the latent tensor using text embeddings and style conditioning, similar to a diffusion model's sampling process.

Can I adjust the speech speed without retraining the models?

Yes. The pipeline applies a user-specified speed factor to rescale the raw durations output by the duration_predictor. This occurs in apply speed (lines 611-614) before the latent sampling phase, enabling real-time speed adjustment between 0.5x and 2.0x without model modification.

Where does the actual audio generation happen in the codebase?

The final waveform generation occurs in vocoder_ort.run (lines 670-674) within rust/src/helper.rs (or the equivalent Python helper). This ONNX session takes the denoised latent tensor and outputs raw audio samples, which are then written to WAV format via write_wav_file using the hound crate.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →