What Kind of Models Does Supertonic Use? Inside the 99M Parameter ONNX Architecture

Supertonic uses four compact, open-weight ONNX models—totaling approximately 99 million parameters—that implement a complete text-to-speech pipeline including a duration predictor, text encoder, flow-matching vector estimator, and vocoder, designed for privacy-preserving on-device inference.

The supertone-inc/supertonic repository ships a lightweight, open-weight text-to-speech system built entirely on ONNX Runtime. Unlike cloud-dependent TTS services, Supertonic bundles its intelligence into four specialized neural networks that run locally across Python, Node.js, and browser environments. Understanding what kind of models Supertonic uses reveals why it can deliver high-quality speech synthesis on devices as modest as a Raspberry Pi while maintaining complete data privacy.

The Four-Core ONNX Models

Supertonic’s architecture divides the TTS pipeline into four distinct ONNX inference graphs that are downloaded automatically from Hugging Face (e.g., Supertone/supertonic-3) on first use. These models are loaded at runtime by the load_onnx_all utility in py/helper.py and shared across all language bindings.

Duration Predictor (duration_predictor.onnx)

The Duration Predictor estimates the temporal length of the input utterance in seconds. This lightweight regression network analyzes the text encoding to forecast the total audio duration before any waveform generation begins, enabling efficient memory allocation and streaming preparation.

Text Encoder (text_encoder.onnx)

The Text Encoder converts pre-processed Unicode tokens into a dense latent text embedding. Input text is first normalized, wrapped with language tags (e.g., <en>…</en>), and tokenized via a shared Unicode indexer (unicode_indexer.json) that supports 31 languages. The encoder then projects these tokens into a latent space that conditions the subsequent audio generation steps.

Vector Estimator (vector_estimator.onnx)

The Vector Estimator implements a flow-matching latent diffusion process. This denoising model refines a random latent tensor into a realistic spectrogram representation, conditioned on both the text embedding and the voice style tensors. The model accepts parameters like total_step (defaulting to 8) and a speed factor to trade off between inference speed and audio quality.

Vocoder (vocoder.onnx)

The Vocoder performs the final audio synthesis, converting the refined latent representation into a 44.1 kHz 16-bit PCM waveform. This last stage produces the actual audio output that can be saved to disk or streamed to an audio device.

How Supertonic Loads Its Models

At runtime, the framework initializes the pipeline through a unified loading mechanism. In py/helper.py, the load_onnx_all function instantiates the four ONNX sessions, while nodejs/helper.js provides an equivalent JavaScript implementation. Both utilities manage model caching, memory layout, and cross-platform ONNX Runtime configuration (including WebGPU support for browser deployments).

Architectural Highlights

Flow-Matching Latent Diffusion

The vector_estimator.onnx model employs a flow-matching denoising approach rather than traditional discrete steps. This allows the system to generate high-fidelity audio with as few as 8 inference steps (total_step=8), drastically reducing latency while maintaining naturalness. The denoising process is guided by the text embedding and voice style parameters passed from the encoder.

Language-Agnostic Tokenisation

Supertonic handles 31 languages through a unified tokenization pipeline. Rather than maintaining separate models per language, the system uses a single unicode_indexer.json vocabulary shared across all supported locales. Text normalization and language tag injection (e.g., <ko>, <jp>) happen before the Text Encoder processes the sequence, enabling multilingual synthesis without switching model weights.

Voice-Style Conditioning

Voice characteristics are controlled via two small tensors—style_ttl and style_dp—stored in JSON files (typically a few kilobytes each). These tensors are loaded via load_voice_style and batched into the inference context. Because the style data is separate from the base 99M-parameter model, users can swap speaker identities on-device instantly by loading different style JSONs from assets/voice_styles/ without reloading the heavy ONNX weights.

Cross-Platform Inference Examples

Python – Synthesising Speech

The Python SDK demonstrates how the four ONNX models work together through the high-level TTS class:

from supertonic import TTS

# First run automatically downloads the ONNX assets from Hugging Face

tts = TTS(auto_download=True)

# Load a preset voice style (e.g. “M1”)

style = tts.get_voice_style(voice_name="M1")

# Generate audio

wav, duration = tts.synthesize(
    text="Supertonic runs fast on‑device with no cloud.",
    lang="en",                     # Use any supported language code

    voice_style=style,
    total_steps=8,                 # Quality vs speed trade‑off

    speed=1.05,                    # Faster speech = lower duration

)

# Save the result

tts.save_audio(wav, "output.wav")

The Python SDK internally calls load_text_to_speech → load_onnx_all to instantiate the four ONNX modules.

Node.js – Batch Synthesis

The JavaScript implementation mirrors the Python architecture using the same underlying ONNX files:

import { loadTextToSpeech, loadVoiceStyle } from "./helper.js";

async function main() {
  const onnxDir = "./assets/onnx";
  const tts = await loadTextToSpeech(onnxDir);            // loads the four ONNX files
  const style = loadVoiceStyle([ "./assets/voice_styles/M1.json" ]);

  const texts = ["Hello world!", "Supertonic is on‑device."];
  const langs = ["en", "en"];
  const { wav, duration } = await tts.batch(texts, langs, style, 8, 1.05);
  // `wav` is a Float32Array; you can write it to a .wav file with `writeWavFile`
}
main();

Browser (WebGPU) – Client-Side Inference

For web deployments, Supertonic uses onnxruntime-web to execute the same four models entirely within the browser:

import { loadTextToSpeech, loadVoiceStyle } from "./helper.js";

(async () => {
  const onnxDir = "/models/onnx";       // served from your static site
  const tts = await loadTextToSpeech(onnxDir);
  const style = loadVoiceStyle([ "/styles/M1.json" ]);

  const { wav } = await tts.call(
    "Supertonic runs entirely in your browser.",
    "en",
    style,
    8,
    1.05
  );

  // Convert `wav` to an AudioBuffer and play via Web Audio API
})();

The browser version preserves complete privacy by performing all inference client-side without cloud round-trips.

Model Configuration and Key Files

The repository organizes model metadata and inference utilities in specific locations:

  • assets/tts.json – Contains model configuration including ae.sample_rate, ttl.chunk_compress_factor, and latent dimensions used by all runtimes.
  • py/helper.py – Implements load_onnx_all, text normalization, and the Unicode tokenization pipeline.
  • nodejs/helper.js – JavaScript port of the core utilities, handling ONNX Runtime Web initialization.
  • assets/voice_styles/*.json – Stores the style_ttl and style_dp tensors for each voice preset.

Summary

  • Supertonic uses four specialized ONNX models (Duration Predictor, Text Encoder, Vector Estimator, Vocoder) totaling approximately 99 million parameters.
  • Models are loaded via load_onnx_all in py/helper.py (and equivalents in other languages), enabling zero-dependency, on-device inference.
  • Flow-matching architecture in the Vector Estimator allows quality synthesis with only 8 denoising steps.
  • 31 languages are supported through a unified Unicode tokenizer and language-tag conditioning.
  • Voice styles are separate from base weights, allowing instant speaker switching via small JSON files without reloading the 99M-parameter backbone.

Frequently Asked Questions

What is the total parameter count of Supertonic models?

The Supertonic 3 release contains approximately 99 million parameters distributed across the four ONNX files. This compact size enables the model to run on resource-constrained devices like Raspberry Pi while maintaining competitive naturalness metrics.

How does Supertonic handle multiple languages?

Supertonic supports 31 languages through a language-agnostic tokenization approach. Text is normalized and wrapped with language tags (e.g., <en>, <ko>) before being processed by the single shared Text Encoder, eliminating the need for separate models per language.

Can Supertonic models run in a web browser?

Yes. Supertonic uses onnxruntime-web with optional WebGPU acceleration to execute the same four ONNX models entirely client-side. The browser implementation loads the models from static assets and performs inference locally, ensuring complete privacy and zero server costs.

What file format does Supertonic use for voice styles?

Voice styles are stored as JSON files containing two tensors: style_ttl and style_dp. These files are typically only a few kilobytes and are loaded separately from the base model weights, allowing users to swap speakers instantly without reloading the 99M-parameter ONNX files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →