# Supertonic Internal ONNX Pipeline Architecture: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder Explained

> Explore Supertonic's ONNX pipeline for text-to-speech: learn about the Duration Predictor, Text Encoder, Vector Estimator, and Vocoder. Generate raw audio from text tokens.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: architecture
- Published: 2026-06-14

---

**Supertonic’s text-to-speech engine processes text through a fixed four-stage ONNX pipeline—Duration Predictor, Text Encoder, Vector Estimator (diffusion denoiser), and Vocoder—to generate raw audio waveforms from tokenized input.**

The supertone-inc/supertonic repository implements a neural TTS system using ONNX Runtime. Its **internal ONNX pipeline architecture** sequences four specialized models to convert Unicode text into audible speech, handling everything from phoneme duration prediction to diffusion-based latent generation and final waveform synthesis.

## The Four-Stage ONNX Pipeline

Each model in the pipeline is loaded as a separate ONNX session via the `loadTextToSpeech` function in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) (JavaScript) or `load_text_to_speech` in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) (Rust). The components execute in a strict order defined in the `_infer` method.

### Duration Predictor

The **Duration Predictor** accepts tokenized text IDs (`text_ids`), a style tensor (`style_dp`), and a text mask (`text_mask`). It outputs a `duration` tensor specifying how many audio samples each phoneme should occupy, effectively controlling speech tempo. According to the source code in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) lines 44-47, the model path is referenced as `dpPath`, while the Rust implementation in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) lines 13-16 defines it as `dp_path`. The predicted durations are modified in-place by a speed factor (default **1.05**) before proceeding to the next stage.

### Text Encoder

The **Text Encoder** consumes the same `text_ids` and `text_mask` alongside a style-TTL tensor (`style_ttl`) to produce a high-dimensional `text_emb` embedding. This embedding serves as the primary conditioning signal for the diffusion process. In [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js), this corresponds to `textEncPath` (lines 45-48), and in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) to `text_enc_path` (lines 14-17).

### Vector Estimator (Diffusion Denoiser)

The **Vector Estimator** implements the core diffusion model, iteratively denoising a randomly-initialized latent tensor across multiple time steps. For each step, it receives the `text_emb`, `style_ttl`, a latent mask derived from the predicted durations, the current diffusion step (`current_step`), and the total step count (`total_step`). It returns a `denoised_latent` that gradually refines toward a clean speech representation. The model is referenced as `vectorEstPath` in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) (lines 46-48) and `vector_est_path` in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) (lines 15-18).

### Vocoder

The **Vocoder** performs the final conversion from latent space to audio waveform. It takes the fully denoised latent tensor and generates `wav_tts`, a raw PCM waveform at the sample rate defined in `cfgs.ae.sample_rate`. The model path is `vocoderPath` in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) (lines 47-49) and `vocoder_path` in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) (lines 16-19).

## End-to-End Inference Flow

The `_infer` method orchestrates these four models in a deterministic sequence. Both JavaScript ([`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) lines 162-268) and Rust ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) lines 75-150) implementations follow identical logic:

1. **Configuration Loading**: `loadCfgs` reads [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json) while `loadTextProcessor` loads [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) for token mapping.
2. **Session Initialization**: `loadTextToSpeech` (or `load_text_to_speech`) creates ONNX Runtime sessions for all four models.
3. **Text Preprocessing**: `UnicodeProcessor.call` tokenizes input strings into `textIds` and binary `textMask` tensors.
4. **Duration Prediction**: `dpOrt.run` (JS) or the equivalent Rust session call produces the duration tensor.
5. **Text Encoding**: `textEncOrt.run` generates the conditioning embeddings.
6. **Latent Initialization**: `sampleNoisyLatent` (JS) or `sample_noisy_latent` (Rust) creates a Gaussian noise tensor sized according to the predicted durations and model chunk settings.
7. **Diffusion Loop**: For `totalStep` iterations (default approximately 30), the code builds a `current_step` tensor and invokes `vectorEstOrt.run` (or the Rust equivalent), replacing the latent with the denoised output each iteration.
8. **Vocoding**: The final latent feeds into `vocoderOrt.run` to produce `wav_tts`.
9. **Post-Processing**: Optional silence padding and chunk concatenation occur in the `call` or `batch` methods.

## Code Examples

### JavaScript Implementation (onnxruntime-web)

The following example demonstrates loading the pipeline and running inference using the JavaScript bindings:

```javascript
import * as ort from 'onnxruntime-web';
import { loadTextToSpeech, loadVoiceStyle } from './helper.js';

async function synthesize(text, language = 'en') {
  // 1️⃣ Load the four ONNX models
  const onnxDir = '/assets/onnx';
  const { textToSpeech, cfgs } = await loadTextToSpeech(onnxDir);

  // 2️⃣ Load voice style (speaker characteristics)
  const style = await loadVoiceStyle([`${onnxDir}/voice_style_neutral.json`]);

  // 3️⃣ Run inference with default totalStep from config
  const { wav } = await textToSpeech.call(text, language, style[0], cfgs.ttl.total_step);

  // 4️⃣ Play the generated Float32 PCM audio
  const audioCtx = new AudioContext({ sampleRate: textToSpeech.sampleRate });
  const buffer = audioCtx.createBuffer(1, wav.length, audioCtx.sampleRate);
  buffer.copyToChannel(new Float32Array(wav), 0);
  const source = audioCtx.createBufferSource();
  source.buffer = buffer;
  source.connect(audioCtx.destination);
  source.start();
}

synthesize('Hello world, this demonstrates the Supertonic ONNX TTS pipeline.');

```

*Key implementation details*: The `loadTextToSpeech` function (lines 44-58 of [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js)) initializes the four sessions, while the `_infer` method (lines 162-235) implements the diffusion loop and vocoder call (lines 262-266).

### Rust Implementation (native onnxruntime)

For server-side or desktop applications, the Rust API provides equivalent functionality:

```rust
use supertonic::rust::helper::{load_text_to_speech, Style, Config};

fn main() -> anyhow::Result<()> {
    // 1️⃣ Initialize the pipeline with CPU inference
    let onnx_dir = "../assets/onnx";
    let mut tts = load_text_to_speech(onnx_dir, false)?;

    // 2️⃣ Load voice style from JSON
    let voice_style = std::fs::read_to_string("../assets/onnx/voice_style_neutral.json")?;
    let style: Style = serde_json::from_str(&voice_style)?;

    // 3️⃣ Run inference with default speed factor 1.05
    let text = "Hello Rust world, demonstrating the ONNX TTS pipeline.";
    let (wav, _duration) = tts._infer(
        &[text.to_string()],
        &["en".to_string()],
        &style,
        tts.cfgs.ttl.total_step as usize,
        1.05,
    )?;

    // 4️⃣ Write to 16-bit PCM WAV file
    let spec = hound::WavSpec {
        channels: 1,
        sample_rate: tts.sample_rate as u32,
        bits_per_sample: 16,
        sample_format: hound::SampleFormat::Int,
    };
    let mut writer = hound::WavWriter::create("out.wav", spec)?;
    for s in wav {
        writer.write_sample((s * i16::MAX as f32) as i16)?;
    }
    writer.finalize()?;
    Ok(())
}

```

*Key implementation details*: The `load_text_to_speech` function ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) lines 13-18) creates the four ONNX sessions, while `TextToSpeech::_infer` (lines 75-150) contains the diffusion loop (lines 155-185).

## Summary

- **Four specialized models** form the Supertonic ONNX pipeline: Duration Predictor, Text Encoder, Vector Estimator, and Vocoder.
- **Fixed execution order**: Duration prediction → Text encoding → Diffusion denoising (Vector Estimator) → Vocoding.
- **Dual implementation**: Identical logic is implemented in both [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) (JavaScript/WebAssembly) and [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) (native Rust).
- **Diffusion-based generation**: The Vector Estimator iteratively refines a noisy latent over approximately 30 steps (`totalStep`) before the Vocoder converts it to audio.
- **Style conditioning**: Both `style_ttl` (for text encoder and diffusion) and `style_dp` (for duration predictor) tensors encode speaker-specific characteristics loaded from JSON voice style files.

## Frequently Asked Questions

### What is the execution order of models in the Supertonic internal ONNX pipeline architecture?

The pipeline executes in a strict sequence: first the **Duration Predictor** determines phoneme lengths, then the **Text Encoder** generates conditioning embeddings, followed by the **Vector Estimator** which runs iteratively to denoise the latent representation, and finally the **Vocoder** converts the clean latent into a raw waveform.

### How does the Vector Estimator differ from the Vocoder in the Supertonic TTS system?

The **Vector Estimator** is a diffusion model that operates in the latent space, gradually refining a random noise tensor into a structured speech representation over multiple steps (default ~30). The **Vocoder** is a deterministic decode model that takes the final denoised latent and directly synthesizes the audible PCM waveform at the target sample rate (`cfgs.ae.sample_rate`).

### Where are the ONNX model paths defined in the Supertonic codebase?

Model paths are defined in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) within the `loadTextToSpeech` function (lines 44-49) as `dpPath`, `textEncPath`, `vectorEstPath`, and `vocoderPath`. In the Rust implementation ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)), these correspond to `dp_path`, `text_enc_path`, `vector_est_path`, and `vocoder_path` in the `load_text_to_speech` function (lines 13-19).

### What controls the speed and duration of the generated speech in the Supertonic pipeline?

The **Duration Predictor** outputs the base phoneme durations, which are multiplied by a **speed factor** (default 1.05) to adjust speech tempo. Additionally, the `totalStep` parameter controls the number of diffusion iterations in the Vector Estimator, though this primarily affects voice quality rather than duration.