# What Kind of Models Does Supertonic Use? Inside the 99M Parameter ONNX Architecture

> Discover Supertonic's 99M parameter ONNX models for text-to-speech. Learn about their compact, privacy-preserving architecture for on-device inference.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: internals
- Published: 2026-06-13

---

**Supertonic uses four compact, open-weight ONNX models—totaling approximately 99 million parameters—that implement a complete text-to-speech pipeline including a duration predictor, text encoder, flow-matching vector estimator, and vocoder, designed for privacy-preserving on-device inference.**

The supertone-inc/supertonic repository ships a lightweight, open-weight text-to-speech system built entirely on ONNX Runtime. Unlike cloud-dependent TTS services, Supertonic bundles its intelligence into four specialized neural networks that run locally across Python, Node.js, and browser environments. Understanding what kind of models Supertonic uses reveals why it can deliver high-quality speech synthesis on devices as modest as a Raspberry Pi while maintaining complete data privacy.

## The Four-Core ONNX Models

Supertonic’s architecture divides the TTS pipeline into four distinct **ONNX inference graphs** that are downloaded automatically from Hugging Face (e.g., `Supertone/supertonic-3`) on first use. These models are loaded at runtime by the `load_onnx_all` utility in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) and shared across all language bindings.

### Duration Predictor (`duration_predictor.onnx`)

The **Duration Predictor** estimates the temporal length of the input utterance in seconds. This lightweight regression network analyzes the text encoding to forecast the total audio duration before any waveform generation begins, enabling efficient memory allocation and streaming preparation.

### Text Encoder (`text_encoder.onnx`)

The **Text Encoder** converts pre-processed Unicode tokens into a dense latent text embedding. Input text is first normalized, wrapped with language tags (e.g., `<en>…</en>`), and tokenized via a shared Unicode indexer ([`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json)) that supports **31 languages**. The encoder then projects these tokens into a latent space that conditions the subsequent audio generation steps.

### Vector Estimator (`vector_estimator.onnx`)

The **Vector Estimator** implements a **flow-matching latent diffusion** process. This denoising model refines a random latent tensor into a realistic spectrogram representation, conditioned on both the text embedding and the voice style tensors. The model accepts parameters like `total_step` (defaulting to 8) and a `speed` factor to trade off between inference speed and audio quality.

### Vocoder (`vocoder.onnx`)

The **Vocoder** performs the final audio synthesis, converting the refined latent representation into a **44.1 kHz 16-bit PCM waveform**. This last stage produces the actual audio output that can be saved to disk or streamed to an audio device.

## How Supertonic Loads Its Models

At runtime, the framework initializes the pipeline through a unified loading mechanism. In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), the `load_onnx_all` function instantiates the four ONNX sessions, while [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js) provides an equivalent JavaScript implementation. Both utilities manage model caching, memory layout, and cross-platform ONNX Runtime configuration (including WebGPU support for browser deployments).

## Architectural Highlights

### Flow-Matching Latent Diffusion

The `vector_estimator.onnx` model employs a **flow-matching** denoising approach rather than traditional discrete steps. This allows the system to generate high-fidelity audio with as few as 8 inference steps (`total_step=8`), drastically reducing latency while maintaining naturalness. The denoising process is guided by the text embedding and voice style parameters passed from the encoder.

### Language-Agnostic Tokenisation

Supertonic handles **31 languages** through a unified tokenization pipeline. Rather than maintaining separate models per language, the system uses a single [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json) vocabulary shared across all supported locales. Text normalization and language tag injection (e.g., `<ko>`, `<jp>`) happen before the Text Encoder processes the sequence, enabling multilingual synthesis without switching model weights.

### Voice-Style Conditioning

Voice characteristics are controlled via two small tensors—`style_ttl` and `style_dp`—stored in JSON files (typically a few kilobytes each). These tensors are loaded via `load_voice_style` and batched into the inference context. Because the style data is separate from the base 99M-parameter model, users can swap speaker identities on-device instantly by loading different style JSONs from `assets/voice_styles/` without reloading the heavy ONNX weights.

## Cross-Platform Inference Examples

### Python – Synthesising Speech

The Python SDK demonstrates how the four ONNX models work together through the high-level `TTS` class:

```python
from supertonic import TTS

# First run automatically downloads the ONNX assets from Hugging Face

tts = TTS(auto_download=True)

# Load a preset voice style (e.g. “M1”)

style = tts.get_voice_style(voice_name="M1")

# Generate audio

wav, duration = tts.synthesize(
    text="Supertonic runs fast on‑device with no cloud.",
    lang="en",                     # Use any supported language code

    voice_style=style,
    total_steps=8,                 # Quality vs speed trade‑off

    speed=1.05,                    # Faster speech = lower duration

)

# Save the result

tts.save_audio(wav, "output.wav")

```

*The Python SDK internally calls `load_text_to_speech` → `load_onnx_all` to instantiate the four ONNX modules.*

### Node.js – Batch Synthesis

The JavaScript implementation mirrors the Python architecture using the same underlying ONNX files:

```javascript
import { loadTextToSpeech, loadVoiceStyle } from "./helper.js";

async function main() {
  const onnxDir = "./assets/onnx";
  const tts = await loadTextToSpeech(onnxDir);            // loads the four ONNX files
  const style = loadVoiceStyle([ "./assets/voice_styles/M1.json" ]);

  const texts = ["Hello world!", "Supertonic is on‑device."];
  const langs = ["en", "en"];
  const { wav, duration } = await tts.batch(texts, langs, style, 8, 1.05);
  // `wav` is a Float32Array; you can write it to a .wav file with `writeWavFile`
}
main();

```

### Browser (WebGPU) – Client-Side Inference

For web deployments, Supertonic uses `onnxruntime-web` to execute the same four models entirely within the browser:

```javascript
import { loadTextToSpeech, loadVoiceStyle } from "./helper.js";

(async () => {
  const onnxDir = "/models/onnx";       // served from your static site
  const tts = await loadTextToSpeech(onnxDir);
  const style = loadVoiceStyle([ "/styles/M1.json" ]);

  const { wav } = await tts.call(
    "Supertonic runs entirely in your browser.",
    "en",
    style,
    8,
    1.05
  );

  // Convert `wav` to an AudioBuffer and play via Web Audio API
})();

```

*The browser version preserves complete privacy by performing all inference client-side without cloud round-trips.*

## Model Configuration and Key Files

The repository organizes model metadata and inference utilities in specific locations:

- **[`assets/tts.json`](https://github.com/supertone-inc/supertonic/blob/main/assets/tts.json)** – Contains model configuration including `ae.sample_rate`, `ttl.chunk_compress_factor`, and latent dimensions used by all runtimes.
- **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)** – Implements `load_onnx_all`, text normalization, and the Unicode tokenization pipeline.
- **[`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js)** – JavaScript port of the core utilities, handling ONNX Runtime Web initialization.
- **`assets/voice_styles/*.json`** – Stores the `style_ttl` and `style_dp` tensors for each voice preset.

## Summary

- **Supertonic uses four specialized ONNX models** (Duration Predictor, Text Encoder, Vector Estimator, Vocoder) totaling approximately 99 million parameters.
- **Models are loaded via `load_onnx_all`** in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) (and equivalents in other languages), enabling zero-dependency, on-device inference.
- **Flow-matching architecture** in the Vector Estimator allows quality synthesis with only 8 denoising steps.
- **31 languages are supported** through a unified Unicode tokenizer and language-tag conditioning.
- **Voice styles are separate from base weights**, allowing instant speaker switching via small JSON files without reloading the 99M-parameter backbone.

## Frequently Asked Questions

### What is the total parameter count of Supertonic models?

The Supertonic 3 release contains approximately **99 million parameters** distributed across the four ONNX files. This compact size enables the model to run on resource-constrained devices like Raspberry Pi while maintaining competitive naturalness metrics.

### How does Supertonic handle multiple languages?

Supertonic supports **31 languages** through a language-agnostic tokenization approach. Text is normalized and wrapped with language tags (e.g., `<en>`, `<ko>`) before being processed by the single shared Text Encoder, eliminating the need for separate models per language.

### Can Supertonic models run in a web browser?

Yes. Supertonic uses `onnxruntime-web` with optional WebGPU acceleration to execute the same four ONNX models entirely client-side. The browser implementation loads the models from static assets and performs inference locally, ensuring complete privacy and zero server costs.

### What file format does Supertonic use for voice styles?

Voice styles are stored as **JSON files** containing two tensors: `style_ttl` and `style_dp`. These files are typically only a few kilobytes and are loaded separately from the base model weights, allowing users to swap speakers instantly without reloading the 99M-parameter ONNX files.