Supertonic vs VoxCPM2 WER/CER: How the 99M Parameter Model Matches 2B Parameter Accuracy

Supertonic achieves comparable or lower word error rates (WER) and character error rates (CER) than VoxCPM2 across the majority of benchmarked languages, despite using 20× fewer parameters (99M vs 2B) and running inference on CPU-only devices via ONNX Runtime.

Supertonic is an open-weight text-to-speech (TTS) system developed by Supertone Inc. that prioritizes on-device deployment without sacrificing pronunciation quality. According to the supertone-inc/supertonic source code, the model was evaluated against VoxCPM2—a 2 billion parameter open-weight TTS model—using the MiniMax-MLS-test benchmark to measure transcription accuracy across 31 languages.

Architectural Innovations in Supertonic

Supertonic’s efficiency stems from two core architectural contributions implemented in the ONNX Runtime inference pipeline.

Length-Aware Rotary Position Embedding (LARoPE)

As documented in README.md, LARoPE improves text-speech alignment within cross-attention layers by incorporating positional information that adapts to sequence length. This mechanism ensures precise phoneme

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →