How to Compare Supertonic's Performance Against Other TTS Engines
Compare Supertonic's performance against other TTS engines by evaluating Word Error Rate (WER) for accuracy, Real-Time Factor (RTF) for inference speed, and parameter count for model efficiency, using the MiniMax-MLS-Test dataset and the reference scripts provided in the py/example_onnx.py file.
Supertonic 3 is a 99-million-parameter on-device text-to-speech engine developed by supertone-inc. When evaluating how it stacks up against cloud-based alternatives like ElevenLabs or OpenAI TTS-1, you need a standardized methodology that accounts for accuracy, latency, and resource consumption. The repository provides reference implementations and benchmark data that make cross-engine comparisons straightforward and reproducible.
The Three Dimensions of TTS Performance
Supertonic evaluates performance along three orthogonal axes that serve as the standard comparison points for any TTS system. These dimensions allow you to benchmark against both open-weight models like VoxCPM2, OmniVoice, and Qwen3-TTS, as well as proprietary cloud services.
Accuracy (WER and CER)
Measure transcription accuracy using Word Error Rate (WER) for languages with clear lexical matches, and Character Error Rate (CER) for language-agnostic scripts. The README.md contains a side-by-side comparison table in the Reading Accuracy section showing Supertonic 3 versus large open-weight models. To calculate this for any engine, align generated transcripts against ground truth using an off-the-shelf ASR like Whisper, then compute the edit distance.
Speed and Latency (RTF)
Track the Real-Time Factor (RTF) on CPU versus GPU baselines, plus the measured latency in milliseconds per utterance. The Runtime Footprint section in README.md provides a visual chart demonstrating that Supertonic 3 achieves sub-100ms latency on laptop CPUs while beating A100-GPU baselines. When comparing against other engines, record wall-clock time including model-loading overhead to ensure fair comparison.
Model Size and Memory Footprint
Document the parameter count and disk footprint of the ONNX checkpoint. Supertonic 3 uses approximately 99 million parameters, making it roughly 15× smaller than 0.7B-class open TTS models according to the Model Size section in README.md. Measure peak memory usage during inference to account for runtime overhead beyond the static model size.
Setting Up a Reproducible Benchmark
To compare Supertonic against any other engine, follow this reproducible pipeline using the publicly available MiniMax-MLS-Test dataset:
-
Select the Test Corpus – Download the MiniMax-MLS-Test dataset from Hugging Face (
MiniMaxAI/TTS-MLS-Test), which provides the multilingual benchmark used by Supertonic 3. -
Run Inference – Invoke each engine with identical text strings from the corpus, capture the raw audio output, and record wall-clock time including any model-loading overhead.
-
Normalize Audio – Convert all outputs to a common sampling rate of 44.1 kHz 16-bit PCM. Supertonic handles this natively, but other engines may require resampling.
-
Compute Accuracy – Transcribe generated audio using Whisper or another ASR, align predictions against ground-truth text, and calculate WER/CER.
-
Report Metrics – Summarize per-language WER/CER, overall RTF, peak memory usage, and model size in a comparison table.
Measuring Latency with the Python SDK
The Python SDK includes a reference implementation in py/example_onnx.py that demonstrates basic latency measurement. This script loads the 99M ONNX checkpoint and prints elapsed time per utterance.
from supertonic import TTS
import time
tts = TTS(auto_download=True) # Loads the 99M ONNX checkpoint
text = "Supertonic is lightning fast and runs locally."
start = time.perf_counter()
wav, _ = tts.synthesize(text=text, lang="en")
elapsed = time.perf_counter() - start
print(f"Latency: {elapsed * 1000:.1f} ms")
Source: py/example_onnx.py
Calculating Word Error Rate (WER)
Adapt the reference script to loop over the MiniMax-MLS test set and calculate WER using an external ASR. This example uses Whisper to generate transcripts and compute edit distance against ground truth.
import whisper
from supertonic import TTS
from datasets import load_dataset
# Load benchmark corpus
ds = load_dataset("MiniMaxAI/TTS-MLS-Test", split="test")
tts = TTS(auto_download=True)
asr = whisper.load_model("base")
def synth_and_score(example):
wav, _ = tts.synthesize(text=example["text"], lang=example["lang"])
pred = asr.transcribe(wav.squeeze(), language="en")["text"]
# Simple WER calculation
wer = whisper.utils.edit_distance(
example["text"].split(),
pred.split()
) / len(example["text"].split())
return {"wer": wer}
results = ds.map(synth_and_score, batched=False)
print("Average WER:", sum(r["wer"] for r in results) / len(results))
Source: py/example_onnx.py (adapted)
Cross-Platform Benchmarking in Rust
For low-level performance testing, the Rust implementation in rust/src/example_onnx.rs provides the same inference contract. This allows you to measure latency without Python overhead.
use supertonic::helper::Supertonic;
use std::time::Instant;
fn main() {
let mut st = Supertonic::new().unwrap(); // Loads the ONNX model
let text = "Supertonic 通过本地推理实现高速语音合成。";
let start = Instant::now();
let wav = st.synthesize(text, "zh").unwrap();
let elapsed = start.elapsed().as_millis();
println!("Latency (zh): {} ms", elapsed);
}
Source: rust/src/example_onnx.rs
Why Supertonic Outperforms Larger Models
Supertonic 3 maintains competitive performance against multi-billion-parameter cloud services through specific architectural optimizations:
-
Compact Architecture – The 99M-parameter flow-matching model uses a minimal ONNX runtime graph, eliminating the overhead associated with larger transformer architectures.
-
Optimized ONNX Runtime – The inference engine compiles for target hardware (CPU, GPU, or WebGPU) with batch-processing enabled, reducing per-utterance overhead compared to dynamic graph execution.
-
Expression Tags – Ten built-in tags (
<laugh>,<breath>, etc.) handle complex prose internally, reducing error propagation that typically inflates WER in systems requiring external text normalization.
Summary
- Compare Supertonic's performance using three metrics: WER/CER for accuracy, RTF for speed, and parameter count for efficiency.
- Use the MiniMax-MLS-Test dataset and the reference scripts in
py/example_onnx.pyto ensure reproducible benchmarks. - Supertonic 3 achieves sub-100ms latency on laptop CPUs while maintaining competitive accuracy with models 15× larger.
- The ONNX-based architecture enables consistent benchmarking across Python, Rust, and Go implementations using the same 99M-parameter checkpoint.
Frequently Asked Questions
What dataset should I use to benchmark Supertonic?
Use the MiniMax-MLS-Test dataset available on Hugging Face (MiniMaxAI/TTS-MLS-Test). This multilingual benchmark is the same dataset referenced in the README.md accuracy tables, ensuring your results align with the published WER/CER metrics for Supertonic 3.
How does Supertonic achieve lower latency than GPU-based models?
Supertonic uses a 99M-parameter flow-matching architecture with an optimized ONNX runtime that compiles specifically for the target hardware. By avoiding the memory transfer overhead and dynamic graph execution of large cloud models, it achieves sub-100ms inference on laptop CPUs, outperforming A100 GPU baselines for certain workloads.
Can I benchmark Supertonic against proprietary APIs like ElevenLabs?
Yes. The benchmarking pipeline is engine-agnostic. Send identical text strings to any proprietary API (ElevenLabs, OpenAI TTS-1, Gemini 2.5), capture the audio output, normalize to 44.1 kHz 16-bit PCM, and compute WER using Whisper. The workflow in py/example_onnx.py can be extended to include any service with a /v1/audio/speech endpoint.
What hardware specifications are needed to reproduce the reported benchmarks?
The reported CPU benchmarks run on standard laptop hardware without GPU acceleration. For GPU comparisons, the repository provides ONNX Runtime backends for CUDA and WebGPU. Ensure you have sufficient RAM to load the 99M-parameter checkpoint (approximately 400MB disk footprint) plus overhead for the audio processing pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →