Supertonic Internal ONNX Pipeline Architecture: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder Explained

Supertonic’s text-to-speech engine processes text through a fixed four-stage ONNX pipeline—Duration Predictor, Text Encoder, Vector Estimator (diffusion denoiser), and Vocoder—to generate raw audio waveforms from tokenized input.

The supertone-inc/supertonic repository implements a neural TTS system using ONNX Runtime. Its internal ONNX pipeline architecture sequences four specialized models to convert Unicode text into audible speech, handling everything from phoneme duration prediction to diffusion-based latent generation and final waveform synthesis.

The Four-Stage ONNX Pipeline

Each model in the pipeline is loaded as a separate ONNX session via the loadTextToSpeech function in web/helper.js (JavaScript) or load_text_to_speech in rust/src/helper.rs (Rust). The components execute in a strict order defined in the _infer method.

Duration Predictor

The Duration Predictor accepts tokenized text IDs (text_ids), a style tensor (style_dp), and a text mask (text_mask). It outputs a duration tensor specifying how many audio samples each phoneme should occupy, effectively controlling speech tempo. According to the source code in web/helper.js lines 44-47, the model path is referenced as dpPath, while the Rust implementation in rust/src/helper.rs lines 13-16 defines it as dp_path. The predicted durations are modified in-place by a speed factor (default 1.05) before proceeding to the next stage.

Text Encoder

The Text Encoder consumes the same text_ids and text_mask alongside a style-TTL tensor (style_ttl) to produce a high-dimensional text_emb embedding. This embedding serves as the primary conditioning signal for the diffusion process. In web/helper.js, this corresponds to textEncPath (lines 45-48), and in rust/src/helper.rs to text_enc_path (lines 14-17).

Vector Estimator (Diffusion Denoiser)

The Vector Estimator implements the core diffusion model, iteratively denoising a randomly-initialized latent tensor across multiple time steps. For each step, it receives the text_emb, style_ttl, a latent mask derived from the predicted durations, the current diffusion step (current_step), and the total step count (total_step). It returns a denoised_latent that gradually refines toward a clean speech representation. The model is referenced as vectorEstPath in web/helper.js (lines 46-48) and vector_est_path in rust/src/helper.rs (lines 15-18).

Vocoder

The Vocoder performs the final conversion from latent space to audio waveform. It takes the fully denoised latent tensor and generates wav_tts, a raw PCM waveform at the sample rate defined in cfgs.ae.sample_rate. The model path is vocoderPath in web/helper.js (lines 47-49) and vocoder_path in rust/src/helper.rs (lines 16-19).

End-to-End Inference Flow

The _infer method orchestrates these four models in a deterministic sequence. Both JavaScript (web/helper.js lines 162-268) and Rust (rust/src/helper.rs lines 75-150) implementations follow identical logic:

  1. Configuration Loading: loadCfgs reads tts.json while loadTextProcessor loads unicode_indexer.json for token mapping.
  2. Session Initialization: loadTextToSpeech (or load_text_to_speech) creates ONNX Runtime sessions for all four models.
  3. Text Preprocessing: UnicodeProcessor.call tokenizes input strings into textIds and binary textMask tensors.
  4. Duration Prediction: dpOrt.run (JS) or the equivalent Rust session call produces the duration tensor.
  5. Text Encoding: textEncOrt.run generates the conditioning embeddings.
  6. Latent Initialization: sampleNoisyLatent (JS) or sample_noisy_latent (Rust) creates a Gaussian noise tensor sized according to the predicted durations and model chunk settings.
  7. Diffusion Loop: For totalStep iterations (default approximately 30), the code builds a current_step tensor and invokes vectorEstOrt.run (or the Rust equivalent), replacing the latent with the denoised output each iteration.
  8. Vocoding: The final latent feeds into vocoderOrt.run to produce wav_tts.
  9. Post-Processing: Optional silence padding and chunk concatenation occur in the call or batch methods.

Code Examples

JavaScript Implementation (onnxruntime-web)

The following example demonstrates loading the pipeline and running inference using the JavaScript bindings:

import * as ort from 'onnxruntime-web';
import { loadTextToSpeech, loadVoiceStyle } from './helper.js';

async function synthesize(text, language = 'en') {
  // 1️⃣ Load the four ONNX models
  const onnxDir = '/assets/onnx';
  const { textToSpeech, cfgs } = await loadTextToSpeech(onnxDir);

  // 2️⃣ Load voice style (speaker characteristics)
  const style = await loadVoiceStyle([`${onnxDir}/voice_style_neutral.json`]);

  // 3️⃣ Run inference with default totalStep from config
  const { wav } = await textToSpeech.call(text, language, style[0], cfgs.ttl.total_step);

  // 4️⃣ Play the generated Float32 PCM audio
  const audioCtx = new AudioContext({ sampleRate: textToSpeech.sampleRate });
  const buffer = audioCtx.createBuffer(1, wav.length, audioCtx.sampleRate);
  buffer.copyToChannel(new Float32Array(wav), 0);
  const source = audioCtx.createBufferSource();
  source.buffer = buffer;
  source.connect(audioCtx.destination);
  source.start();
}

synthesize('Hello world, this demonstrates the Supertonic ONNX TTS pipeline.');

Key implementation details: The loadTextToSpeech function (lines 44-58 of web/helper.js) initializes the four sessions, while the _infer method (lines 162-235) implements the diffusion loop and vocoder call (lines 262-266).

Rust Implementation (native onnxruntime)

For server-side or desktop applications, the Rust API provides equivalent functionality:

use supertonic::rust::helper::{load_text_to_speech, Style, Config};

fn main() -> anyhow::Result<()> {
    // 1️⃣ Initialize the pipeline with CPU inference
    let onnx_dir = "../assets/onnx";
    let mut tts = load_text_to_speech(onnx_dir, false)?;

    // 2️⃣ Load voice style from JSON
    let voice_style = std::fs::read_to_string("../assets/onnx/voice_style_neutral.json")?;
    let style: Style = serde_json::from_str(&voice_style)?;

    // 3️⃣ Run inference with default speed factor 1.05
    let text = "Hello Rust world, demonstrating the ONNX TTS pipeline.";
    let (wav, _duration) = tts._infer(
        &[text.to_string()],
        &["en".to_string()],
        &style,
        tts.cfgs.ttl.total_step as usize,
        1.05,
    )?;

    // 4️⃣ Write to 16-bit PCM WAV file
    let spec = hound::WavSpec {
        channels: 1,
        sample_rate: tts.sample_rate as u32,
        bits_per_sample: 16,
        sample_format: hound::SampleFormat::Int,
    };
    let mut writer = hound::WavWriter::create("out.wav", spec)?;
    for s in wav {
        writer.write_sample((s * i16::MAX as f32) as i16)?;
    }
    writer.finalize()?;
    Ok(())
}

Key implementation details: The load_text_to_speech function (rust/src/helper.rs lines 13-18) creates the four ONNX sessions, while TextToSpeech::_infer (lines 75-150) contains the diffusion loop (lines 155-185).

Summary

  • Four specialized models form the Supertonic ONNX pipeline: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder.
  • Fixed execution order: Duration prediction → Text encoding → Diffusion denoising (Vector Estimator) → Vocoding.
  • Dual implementation: Identical logic is implemented in both web/helper.js (JavaScript/WebAssembly) and rust/src/helper.rs (native Rust).
  • Diffusion-based generation: The Vector Estimator iteratively refines a noisy latent over approximately 30 steps (totalStep) before the Vocoder converts it to audio.
  • Style conditioning: Both style_ttl (for text encoder and diffusion) and style_dp (for duration predictor) tensors encode speaker-specific characteristics loaded from JSON voice style files.

Frequently Asked Questions

What is the execution order of models in the Supertonic internal ONNX pipeline architecture?

The pipeline executes in a strict sequence: first the Duration Predictor determines phoneme lengths, then the Text Encoder generates conditioning embeddings, followed by the Vector Estimator which runs iteratively to denoise the latent representation, and finally the Vocoder converts the clean latent into a raw waveform.

How does the Vector Estimator differ from the Vocoder in the Supertonic TTS system?

The Vector Estimator is a diffusion model that operates in the latent space, gradually refining a random noise tensor into a structured speech representation over multiple steps (default ~30). The Vocoder is a deterministic decode model that takes the final denoised latent and directly synthesizes the audible PCM waveform at the target sample rate (cfgs.ae.sample_rate).

Where are the ONNX model paths defined in the Supertonic codebase?

Model paths are defined in web/helper.js within the loadTextToSpeech function (lines 44-49) as dpPath, textEncPath, vectorEstPath, and vocoderPath. In the Rust implementation (rust/src/helper.rs), these correspond to dp_path, text_enc_path, vector_est_path, and vocoder_path in the load_text_to_speech function (lines 13-19).

What controls the speed and duration of the generated speech in the Supertonic pipeline?

The Duration Predictor outputs the base phoneme durations, which are multiplied by a speed factor (default 1.05) to adjust speech tempo. Additionally, the totalStep parameter controls the number of diffusion iterations in the Vector Estimator, though this primarily affects voice quality rather than duration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →