# What is Supertonic? The Open-Source On-Device Multilingual TTS System Explained

> Discover Supertonic, a fast open-source on-device multilingual TTS system. Get high-quality 44.1 kHz audio in 31 languages with zero cloud dependencies. Runs locally via ONNX Runtime.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: getting-started
- Published: 2026-06-13

---

**Supertonic is a lightning-fast, on-device multilingual text-to-speech (TTS) system that runs entirely locally using ONNX Runtime, supporting 31 languages with high-quality 44.1 kHz audio and zero cloud dependencies.**

Supertonic is an open-weight TTS engine developed by Supertone Inc. that delivers studio-quality speech synthesis without internet connectivity or API keys. Unlike billion-parameter cloud models, this repository (`supertone-inc/supertonic`) packages a compact 99 million-parameter ONNX model capable of running on edge devices like Raspberry Pi and e-readers while maintaining competitive audio quality.

## Architecture Overview

Supertonic’s inference pipeline consists of three distinct components exported as a single ONNX graph. According to the source code in [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md), the system processes text through a normalization stage, generates latent representations via diffusion-style modeling, and decodes waveforms through a lightweight neural vocoder.

### Text Normalizer & Tokenizer

The first stage parses input strings, expands abbreviations, and inserts expressive tags (e.g., `<laugh>`, `<breath>`) to produce natural-sounding prosody. This component handles the initial text processing before tokenization, ensuring that numbers, dates, and punctuation are properly normalized for the acoustic model.

### Latent Flow-Matching Module

A diffusion-inspired **latent flow-matching module** maps token sequences into a compressed latent space. This 99M-parameter neural network operates as the core "brain" of the system, transforming linguistic features into audio representations that capture speaker characteristics and intonation patterns.

### Speech Decoder (Vocoder)

The final stage employs a lightweight neural vocoder to generate **44.1 kHz PCM waveforms** from latent representations. This ONNX-exported decoder produces broadcast-quality audio while maintaining the computational efficiency required for CPU-only inference.

## Multilingual Support

Supertonic includes language-specific token vocabularies alongside a shared acoustic model capable of synthesizing speech in **31 languages**. Users can specify a target language using ISO codes (e.g., `lang="en"` for English) or enable automatic language detection with `lang="na"` (language-agnostic mode). The system automatically handles phoneme conversion and prosody modeling appropriate for each language family.

## Runtime Flexibility

Because the model is stored as a standardized ONNX graph, the same inference assets can execute across diverse runtimes without modification. The repository provides SDKs and working examples for the following platforms:

- **Python** – Primary SDK via `pip install supertonic`, with examples in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py)
- **Node.js** – JavaScript bindings demonstrated in [`nodejs/example_onnx.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/example_onnx.js)
- **Web/Browser** – WebGPU-accelerated inference in [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js)
- **Java** – Maven-based project with entry point [`java/ExampleONNX.java`](https://github.com/supertone-inc/supertonic/blob/main/java/ExampleONNX.java)
- **C++** – CMake build system with [`cpp/example_onnx.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/example_onnx.cpp)
- **C#** – .NET 9 compatible implementation in [`csharp/ExampleONNX.cs`](https://github.com/supertone-inc/supertonic/blob/main/csharp/ExampleONNX.cs)
- **Go** – Native modules with example [`go/example_onnx.go`](https://github.com/supertone-inc/supertonic/blob/main/go/example_onnx.go)
- **Swift** – Swift Package Manager support in `swift/example_onnx`
- **iOS** – Xcode project template at `ios/ExampleiOSApp`
- **Rust** – Cargo crate using the `ort` ONNX runtime in [`rust/example_onnx.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/example_onnx.rs)
- **Flutter** – Dart plugin with cross-platform mobile support in `flutter/`

Each runtime implementation follows an identical pattern: load ONNX assets from the `assets/` directory, instantiate a helper class, and call `synthesize(text, lang, voice_style, …)`.

## Extensibility Features

### Voice Builder

Supertonic includes a **Voice Builder** web interface that converts short reference recordings into JSON voice-style files. Users place these generated files into `assets/voice_styles/` to create custom speaker profiles without retraining the base model. The `TTS` class in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) loads these styles via `get_voice_style(voice_name="CustomName")`.

### Local HTTP Server

The Python SDK can operate as a self-hosted HTTP service exposing two endpoints:
- Native `POST /v1/tts` for direct synthesis requests
- OpenAI-compatible `POST /v1/audio/speech` for drop-in replacement of cloud TTS APIs

This server implementation allows existing applications to migrate from paid cloud services to local inference without code changes.

## Performance Characteristics

Benchmarks demonstrate that Supertonic achieves **real-time factors (RTF) well below 1.0** on standard CPUs, making it significantly faster than GPU-dependent models like VoxCPM2. Memory consumption remains under **500 MiB** during inference, enabling deployment on memory-constrained devices including embedded systems and mobile phones.

## Code Examples

### Python Implementation

The canonical example resides in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py). This implementation demonstrates the complete workflow from model initialization to audio export:

```python
from supertonic import TTS

# Auto-download the model on first run from Hugging Face.

tts = TTS(auto_download=True)

# Load a voice style (e.g., the default "M1" preset).

style = tts.get_voice_style(voice_name="M1")

# Synthesize speech with configurable quality steps.

wav, duration = tts.synthesize(
    text="Supertonic is a lightning fast, on-device TTS system.",
    lang="en",            # Use "na" for automatic language detection

    voice_style=style,
    total_steps=8,        # 5=low quality, 12=high quality

    speed=1.05,           # 0.7=slow, 2.0=fast

)

# Export to WAV format.

tts.save_audio(wav, "output.wav")
print(f"Generated {duration[0]:.2f}s of audio")

```

**Execution:**

```bash
cd py
uv sync                    # Install dependencies

uv run example_onnx.py     # Downloads models and generates audio

```

### Cross-Platform Snippets

**Node.js** ([`nodejs/example_onnx.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/example_onnx.js)):

```javascript
const { TTS } = require("./helper");
const tts = new TTS();
tts.synthesize("Hello world", "en").then(wav => {
  // Write wav buffer to file
});

```

**C++** ([`cpp/example_onnx.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/example_onnx.cpp)):

```cpp
#include "helper.h"
int main() {
  auto tts = TTS();
  auto wav = tts.synthesize("Hello", "en");
  // Write PCM data to WAV file
  return 0;
}

```

**Go** ([`go/example_onnx.go`](https://github.com/supertone-inc/supertonic/blob/main/go/example_onnx.go)):

```go
tts, _ := helper.NewTTS()
wav, _ := tts.Synthesize("Hello", "en")
// Write wav to disk

```

**Java** ([`java/ExampleONNX.java`](https://github.com/supertone-inc/supertonic/blob/main/java/ExampleONNX.java)):

```java
TTS tts = new TTS();
float[] wav = tts.synthesize("Hello", "en");
// Convert float array to audio file

```

**Rust** ([`rust/example_onnx.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/example_onnx.rs)):

```rust
let tts = TTS::new()?;
let wav = tts.synthesize("Hello", "en")?;
// Handle audio output

```

## Summary

- **Supertonic** is a 99M-parameter ONNX-based TTS system that runs entirely on-device without cloud dependencies.
- The architecture combines a text normalizer, latent flow-matching diffusion model, and 44.1 kHz neural vocoder into a single ONNX graph.
- Supports **31 languages** with automatic language detection capabilities and expressive markup tags.
- Provides SDKs for **11 runtimes** including Python, Node.js, Java, C++, C#, Go, Swift, Rust, and Flutter.
- Features a **Voice Builder** for custom speaker creation and an **OpenAI-compatible HTTP server** for API migration.
- Achieves sub-1.0 real-time factors on CPU with under 500 MiB memory usage.

## Frequently Asked Questions

### Is Supertonic free for commercial use?

Yes, Supertonic is released as an open-weight model under the repository `supertone-inc/supertonic`. The ONNX model and all SDK code are available for commercial and personal use without API fees or usage limits, as all inference occurs locally on your hardware.

### How does Supertonic compare to cloud TTS services like OpenAI or ElevenLabs?

Supertonic offers competitive audio quality with significantly lower latency and zero network dependencies. While cloud services may offer larger parameter counts, Supertonic's 99M-parameter model achieves real-time synthesis on CPUs without requiring GPUs, making it cost-effective for high-volume applications. The built-in `POST /v1/audio/speech` endpoint provides OpenAI API compatibility for easy migration.

### What hardware is required to run Supertonic?

Supertonic runs on any hardware supporting ONNX Runtime, including Raspberry Pi devices, e-readers, and standard laptops. The system requires less than 500 MiB of RAM and operates efficiently on CPU-only environments, though WebGPU acceleration is available for browser deployments.

### Can I create custom voices with Supertonic?

Yes, the **Voice Builder** tool allows you to generate custom voice styles from short reference recordings. The resulting JSON files are placed in `assets/voice_styles/` and loaded via the `get_voice_style()` method in the Python SDK or equivalent APIs in other language bindings. No model retraining or fine-tuning is required to use custom voices.