What is Supertonic? The Open-Source On-Device Multilingual TTS System Explained

Supertonic is a lightning-fast, on-device multilingual text-to-speech (TTS) system that runs entirely locally using ONNX Runtime, supporting 31 languages with high-quality 44.1 kHz audio and zero cloud dependencies.

Supertonic is an open-weight TTS engine developed by Supertone Inc. that delivers studio-quality speech synthesis without internet connectivity or API keys. Unlike billion-parameter cloud models, this repository (supertone-inc/supertonic) packages a compact 99 million-parameter ONNX model capable of running on edge devices like Raspberry Pi and e-readers while maintaining competitive audio quality.

Architecture Overview

Supertonic’s inference pipeline consists of three distinct components exported as a single ONNX graph. According to the source code in README.md, the system processes text through a normalization stage, generates latent representations via diffusion-style modeling, and decodes waveforms through a lightweight neural vocoder.

Text Normalizer & Tokenizer

The first stage parses input strings, expands abbreviations, and inserts expressive tags (e.g., <laugh>, <breath>) to produce natural-sounding prosody. This component handles the initial text processing before tokenization, ensuring that numbers, dates, and punctuation are properly normalized for the acoustic model.

Latent Flow-Matching Module

A diffusion-inspired latent flow-matching module maps token sequences into a compressed latent space. This 99M-parameter neural network operates as the core "brain" of the system, transforming linguistic features into audio representations that capture speaker characteristics and intonation patterns.

Speech Decoder (Vocoder)

The final stage employs a lightweight neural vocoder to generate 44.1 kHz PCM waveforms from latent representations. This ONNX-exported decoder produces broadcast-quality audio while maintaining the computational efficiency required for CPU-only inference.

Multilingual Support

Supertonic includes language-specific token vocabularies alongside a shared acoustic model capable of synthesizing speech in 31 languages. Users can specify a target language using ISO codes (e.g., lang="en" for English) or enable automatic language detection with lang="na" (language-agnostic mode). The system automatically handles phoneme conversion and prosody modeling appropriate for each language family.

Runtime Flexibility

Because the model is stored as a standardized ONNX graph, the same inference assets can execute across diverse runtimes without modification. The repository provides SDKs and working examples for the following platforms:

Each runtime implementation follows an identical pattern: load ONNX assets from the assets/ directory, instantiate a helper class, and call synthesize(text, lang, voice_style, …).

Extensibility Features

Voice Builder

Supertonic includes a Voice Builder web interface that converts short reference recordings into JSON voice-style files. Users place these generated files into assets/voice_styles/ to create custom speaker profiles without retraining the base model. The TTS class in py/helper.py loads these styles via get_voice_style(voice_name="CustomName").

Local HTTP Server

The Python SDK can operate as a self-hosted HTTP service exposing two endpoints:

  • Native POST /v1/tts for direct synthesis requests
  • OpenAI-compatible POST /v1/audio/speech for drop-in replacement of cloud TTS APIs

This server implementation allows existing applications to migrate from paid cloud services to local inference without code changes.

Performance Characteristics

Benchmarks demonstrate that Supertonic achieves real-time factors (RTF) well below 1.0 on standard CPUs, making it significantly faster than GPU-dependent models like VoxCPM2. Memory consumption remains under 500 MiB during inference, enabling deployment on memory-constrained devices including embedded systems and mobile phones.

Code Examples

Python Implementation

The canonical example resides in py/example_onnx.py. This implementation demonstrates the complete workflow from model initialization to audio export:

from supertonic import TTS

# Auto-download the model on first run from Hugging Face.

tts = TTS(auto_download=True)

# Load a voice style (e.g., the default "M1" preset).

style = tts.get_voice_style(voice_name="M1")

# Synthesize speech with configurable quality steps.

wav, duration = tts.synthesize(
    text="Supertonic is a lightning fast, on-device TTS system.",
    lang="en",            # Use "na" for automatic language detection

    voice_style=style,
    total_steps=8,        # 5=low quality, 12=high quality

    speed=1.05,           # 0.7=slow, 2.0=fast

)

# Export to WAV format.

tts.save_audio(wav, "output.wav")
print(f"Generated {duration[0]:.2f}s of audio")

Execution:

cd py
uv sync                    # Install dependencies

uv run example_onnx.py     # Downloads models and generates audio

Cross-Platform Snippets

Node.js (nodejs/example_onnx.js):

const { TTS } = require("./helper");
const tts = new TTS();
tts.synthesize("Hello world", "en").then(wav => {
  // Write wav buffer to file
});

C++ (cpp/example_onnx.cpp):

#include "helper.h"
int main() {
  auto tts = TTS();
  auto wav = tts.synthesize("Hello", "en");
  // Write PCM data to WAV file
  return 0;
}

Go (go/example_onnx.go):

tts, _ := helper.NewTTS()
wav, _ := tts.Synthesize("Hello", "en")
// Write wav to disk

Java (java/ExampleONNX.java):

TTS tts = new TTS();
float[] wav = tts.synthesize("Hello", "en");
// Convert float array to audio file

Rust (rust/example_onnx.rs):

let tts = TTS::new()?;
let wav = tts.synthesize("Hello", "en")?;
// Handle audio output

Summary

  • Supertonic is a 99M-parameter ONNX-based TTS system that runs entirely on-device without cloud dependencies.
  • The architecture combines a text normalizer, latent flow-matching diffusion model, and 44.1 kHz neural vocoder into a single ONNX graph.
  • Supports 31 languages with automatic language detection capabilities and expressive markup tags.
  • Provides SDKs for 11 runtimes including Python, Node.js, Java, C++, C#, Go, Swift, Rust, and Flutter.
  • Features a Voice Builder for custom speaker creation and an OpenAI-compatible HTTP server for API migration.
  • Achieves sub-1.0 real-time factors on CPU with under 500 MiB memory usage.

Frequently Asked Questions

Is Supertonic free for commercial use?

Yes, Supertonic is released as an open-weight model under the repository supertone-inc/supertonic. The ONNX model and all SDK code are available for commercial and personal use without API fees or usage limits, as all inference occurs locally on your hardware.

How does Supertonic compare to cloud TTS services like OpenAI or ElevenLabs?

Supertonic offers competitive audio quality with significantly lower latency and zero network dependencies. While cloud services may offer larger parameter counts, Supertonic's 99M-parameter model achieves real-time synthesis on CPUs without requiring GPUs, making it cost-effective for high-volume applications. The built-in POST /v1/audio/speech endpoint provides OpenAI API compatibility for easy migration.

What hardware is required to run Supertonic?

Supertonic runs on any hardware supporting ONNX Runtime, including Raspberry Pi devices, e-readers, and standard laptops. The system requires less than 500 MiB of RAM and operates efficiently on CPU-only environments, though WebGPU acceleration is available for browser deployments.

Can I create custom voices with Supertonic?

Yes, the Voice Builder tool allows you to generate custom voice styles from short reference recordings. The resulting JSON files are placed in assets/voice_styles/ and loaded via the get_voice_style() method in the Python SDK or equivalent APIs in other language bindings. No model retraining or fine-tuning is required to use custom voices.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →