What is Supertonic? The Open-Source On-Device Multilingual TTS System Explained
Supertonic is a lightning-fast, on-device multilingual text-to-speech (TTS) system that runs entirely locally using ONNX Runtime, supporting 31 languages with high-quality 44.1 kHz audio and zero cloud dependencies.
Supertonic is an open-weight TTS engine developed by Supertone Inc. that delivers studio-quality speech synthesis without internet connectivity or API keys. Unlike billion-parameter cloud models, this repository (supertone-inc/supertonic) packages a compact 99 million-parameter ONNX model capable of running on edge devices like Raspberry Pi and e-readers while maintaining competitive audio quality.
Architecture Overview
Supertonic’s inference pipeline consists of three distinct components exported as a single ONNX graph. According to the source code in README.md, the system processes text through a normalization stage, generates latent representations via diffusion-style modeling, and decodes waveforms through a lightweight neural vocoder.
Text Normalizer & Tokenizer
The first stage parses input strings, expands abbreviations, and inserts expressive tags (e.g., <laugh>, <breath>) to produce natural-sounding prosody. This component handles the initial text processing before tokenization, ensuring that numbers, dates, and punctuation are properly normalized for the acoustic model.
Latent Flow-Matching Module
A diffusion-inspired latent flow-matching module maps token sequences into a compressed latent space. This 99M-parameter neural network operates as the core "brain" of the system, transforming linguistic features into audio representations that capture speaker characteristics and intonation patterns.
Speech Decoder (Vocoder)
The final stage employs a lightweight neural vocoder to generate 44.1 kHz PCM waveforms from latent representations. This ONNX-exported decoder produces broadcast-quality audio while maintaining the computational efficiency required for CPU-only inference.
Multilingual Support
Supertonic includes language-specific token vocabularies alongside a shared acoustic model capable of synthesizing speech in 31 languages. Users can specify a target language using ISO codes (e.g., lang="en" for English) or enable automatic language detection with lang="na" (language-agnostic mode). The system automatically handles phoneme conversion and prosody modeling appropriate for each language family.
Runtime Flexibility
Because the model is stored as a standardized ONNX graph, the same inference assets can execute across diverse runtimes without modification. The repository provides SDKs and working examples for the following platforms:
- Python – Primary SDK via
pip install supertonic, with examples inpy/example_onnx.py - Node.js – JavaScript bindings demonstrated in
nodejs/example_onnx.js - Web/Browser – WebGPU-accelerated inference in
web/main.js - Java – Maven-based project with entry point
java/ExampleONNX.java - C++ – CMake build system with
cpp/example_onnx.cpp - C# – .NET 9 compatible implementation in
csharp/ExampleONNX.cs - Go – Native modules with example
go/example_onnx.go - Swift – Swift Package Manager support in
swift/example_onnx - iOS – Xcode project template at
ios/ExampleiOSApp - Rust – Cargo crate using the
ortONNX runtime inrust/example_onnx.rs - Flutter – Dart plugin with cross-platform mobile support in
flutter/
Each runtime implementation follows an identical pattern: load ONNX assets from the assets/ directory, instantiate a helper class, and call synthesize(text, lang, voice_style, …).
Extensibility Features
Voice Builder
Supertonic includes a Voice Builder web interface that converts short reference recordings into JSON voice-style files. Users place these generated files into assets/voice_styles/ to create custom speaker profiles without retraining the base model. The TTS class in py/helper.py loads these styles via get_voice_style(voice_name="CustomName").
Local HTTP Server
The Python SDK can operate as a self-hosted HTTP service exposing two endpoints:
- Native
POST /v1/ttsfor direct synthesis requests - OpenAI-compatible
POST /v1/audio/speechfor drop-in replacement of cloud TTS APIs
This server implementation allows existing applications to migrate from paid cloud services to local inference without code changes.
Performance Characteristics
Benchmarks demonstrate that Supertonic achieves real-time factors (RTF) well below 1.0 on standard CPUs, making it significantly faster than GPU-dependent models like VoxCPM2. Memory consumption remains under 500 MiB during inference, enabling deployment on memory-constrained devices including embedded systems and mobile phones.
Code Examples
Python Implementation
The canonical example resides in py/example_onnx.py. This implementation demonstrates the complete workflow from model initialization to audio export:
from supertonic import TTS
# Auto-download the model on first run from Hugging Face.
tts = TTS(auto_download=True)
# Load a voice style (e.g., the default "M1" preset).
style = tts.get_voice_style(voice_name="M1")
# Synthesize speech with configurable quality steps.
wav, duration = tts.synthesize(
text="Supertonic is a lightning fast, on-device TTS system.",
lang="en", # Use "na" for automatic language detection
voice_style=style,
total_steps=8, # 5=low quality, 12=high quality
speed=1.05, # 0.7=slow, 2.0=fast
)
# Export to WAV format.
tts.save_audio(wav, "output.wav")
print(f"Generated {duration[0]:.2f}s of audio")
Execution:
cd py
uv sync # Install dependencies
uv run example_onnx.py # Downloads models and generates audio
Cross-Platform Snippets
Node.js (nodejs/example_onnx.js):
const { TTS } = require("./helper");
const tts = new TTS();
tts.synthesize("Hello world", "en").then(wav => {
// Write wav buffer to file
});
C++ (cpp/example_onnx.cpp):
#include "helper.h"
int main() {
auto tts = TTS();
auto wav = tts.synthesize("Hello", "en");
// Write PCM data to WAV file
return 0;
}
Go (go/example_onnx.go):
tts, _ := helper.NewTTS()
wav, _ := tts.Synthesize("Hello", "en")
// Write wav to disk
Java (java/ExampleONNX.java):
TTS tts = new TTS();
float[] wav = tts.synthesize("Hello", "en");
// Convert float array to audio file
Rust (rust/example_onnx.rs):
let tts = TTS::new()?;
let wav = tts.synthesize("Hello", "en")?;
// Handle audio output
Summary
- Supertonic is a 99M-parameter ONNX-based TTS system that runs entirely on-device without cloud dependencies.
- The architecture combines a text normalizer, latent flow-matching diffusion model, and 44.1 kHz neural vocoder into a single ONNX graph.
- Supports 31 languages with automatic language detection capabilities and expressive markup tags.
- Provides SDKs for 11 runtimes including Python, Node.js, Java, C++, C#, Go, Swift, Rust, and Flutter.
- Features a Voice Builder for custom speaker creation and an OpenAI-compatible HTTP server for API migration.
- Achieves sub-1.0 real-time factors on CPU with under 500 MiB memory usage.
Frequently Asked Questions
Is Supertonic free for commercial use?
Yes, Supertonic is released as an open-weight model under the repository supertone-inc/supertonic. The ONNX model and all SDK code are available for commercial and personal use without API fees or usage limits, as all inference occurs locally on your hardware.
How does Supertonic compare to cloud TTS services like OpenAI or ElevenLabs?
Supertonic offers competitive audio quality with significantly lower latency and zero network dependencies. While cloud services may offer larger parameter counts, Supertonic's 99M-parameter model achieves real-time synthesis on CPUs without requiring GPUs, making it cost-effective for high-volume applications. The built-in POST /v1/audio/speech endpoint provides OpenAI API compatibility for easy migration.
What hardware is required to run Supertonic?
Supertonic runs on any hardware supporting ONNX Runtime, including Raspberry Pi devices, e-readers, and standard laptops. The system requires less than 500 MiB of RAM and operates efficiently on CPU-only environments, though WebGPU acceleration is available for browser deployments.
Can I create custom voices with Supertonic?
Yes, the Voice Builder tool allows you to generate custom voice styles from short reference recordings. The resulting JSON files are placed in assets/voice_styles/ and loaded via the get_voice_style() method in the Python SDK or equivalent APIs in other language bindings. No model retraining or fine-tuning is required to use custom voices.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →