Supertonic v2 vs v3: Key Differences and Migration Guide

Supertonic v3 is a drop-in replacement for v2 that triples language coverage to 31 languages, adds expressive control tags, improves accuracy by up to 38% WER reduction, and reduces model size by 25%—all while maintaining the same ONNX inference contract.

The supertone-inc/supertonic repository provides state-of-the-art neural text-to-speech models with open-source ONNX runtimes. Understanding the differences between Supertonic v2 and v3 helps developers choose the right version for multilingual synthesis workloads and migration strategies.

Architecture and Scale Improvements

Supertonic v3 substantially scales the model capacity while optimizing for deployment efficiency.

Parameter Count and Model Size

According to the project README.md, Supertonic v2 utilizes approximately 66 million parameters (line 50), whereas Supertonic v3 expands to roughly 99 million parameters (line 81). Despite the 50% parameter increase, v3 achieves a smaller disk footprint through aggressive optimization.

The v3 release applies OnnxSlim techniques to compress the model assets, resulting in approximately 150 MB downloads compared to v2's ~200 MB (line 47). This reduction eliminates bloat while preserving output quality.

Runtime Performance

Benchmarks documented in the repository (lines 73-78) demonstrate that v3 delivers faster CPU latency and lower memory usage than v2, performing comparably to much larger GPU-bound baselines. The architectural improvements in v3 eliminate the synthesis bottlenecks present in v2 without requiring hardware acceleration.

Language Support and Accuracy

The most significant functional difference lies in multilingual capabilities and reading precision.

Expanded Language Coverage

Supertonic v2 supports 5 languages (English plus 4 others) as noted at line 44. Supertonic v3 expands this to 31 languages with full multilingual support (line 33), enabling single-model polyglot synthesis without voice switching.

Word Error Rate Improvements

Per-language accuracy metrics in README.md (lines 41-56) show consistent reductions in Word Error Rate (WER) and Character Error Rate (CER):

  • English: WER improves from 2.52 to 2.06
  • Spanish: WER improves from 1.81 to 1.13

These improvements represent up to 38% error reduction, often outperforming larger open-source alternatives like VoxCPM2. Additionally, v3 significantly reduces repeat and skip failures during synthesis (line 270), producing smoother, more natural speech patterns than v2.

Speaker Similarity

When switching between voice styles, v3 maintains improved speaker similarity across the shared language set compared to v2 (line 270), ensuring consistent vocal characteristics across multilingual content.

Expression Tags and Voice Control

Supertonic v3 introduces 10 inline expression tags that v2 lacks entirely (line 28). These markup annotations allow fine-grained prosodic control without preprocessing:

  • <laugh> - adds laughter intonation
  • <breath> - inserts natural breathing pauses
  • <sigh> - conveys exasperation or relief

These tags process natively within the v3 model architecture, enabling nuanced emotional expression that v2 cannot replicate.

Backward Compatibility and ONNX Interface

Despite the architectural upgrades, v3 maintains full backward compatibility with v2 integrations.

Drop-in Replacement Architecture

The v3 ONNX assets remain v2-compatible (line 42), utilizing the identical inference contract. Existing implementations can upgrade by simply changing the model_dir parameter from the v2-specific branch (release/supertonic-2) to the default v3 assets (assets/onnx/).

File Structure Differences

  • v2 assets: Located in release/supertonic-2/ branch with legacy formatting
  • v3 assets: Stored in assets/onnx/ with optimized slimming

Both versions use the same Python SDK interface defined in py/example_onnx.py and py/helper.py, ensuring helper utilities for loading models and voice styles function identically across versions.

Code Examples

Basic v2 Inference

Load the legacy checkpoint from the v2 branch:

from supertonic import TTS

tts_v2 = TTS(
    model_dir="assets/supertonic-2/onnx",
    auto_download=False
)

style = tts_v2.get_voice_style(voice_name="M1")
wav, _ = tts_v2.synthesize(
    text="Hello, this is Supertonic version 2.",
    lang="en",
    voice_style=style,
    total_steps=8,
    speed=1.0,
)
tts_v2.save_audio(wav, "v2_output.wav")

Migration to v3

Upgrade by changing only the model directory path:

from supertonic import TTS

tts_v3 = TTS(
    model_dir="assets/onnx",  # v3 default path

    auto_download=False
)

style = tts_v3.get_voice_style(voice_name="M1")
wav, _ = tts_v3.synthesize(
    text="Hello, this is Supertonic version 3 with 31 languages.",
    lang="en",
    voice_style=style,
    total_steps=8,
    speed=1.0,
)
tts_v3.save_audio(wav, "v3_output.wav")

Expression Tags (v3 Only)

Leverage v3-exclusive prosodic controls:

from supertonic import TTS

tts = TTS(auto_download=True)
style = tts.get_voice_style(voice_name="M1")

text = "The meeting was great <laugh>! Let's celebrate <sigh>."
wav, _ = tts.synthesize(
    text=text,
    lang="en",
    voice_style=style,
    total_steps=10,
    speed=1.1,
)
tts.save_audio(wav, "v3_expression.wav")

Summary

  • Supertonic v3 scales from 66M to 99M parameters while reducing disk size from ~200MB to ~150MB through OnnxSlim optimization
  • Language support expands from 5 to 31 languages with significant WER improvements (e.g., English 2.52→2.06, Spanish 1.81→1.13)
  • Expression tags (<laugh>, <breath>, etc.) provide v3-exclusive prosodic control unavailable in v2
  • Backward compatibility allows drop-in replacement via the same ONNX inference contract—only the model_dir path changes
  • v3 eliminates repeat/skip failures and improves speaker similarity across voice styles

Frequently Asked Questions

Can I use Supertonic v3 with existing v2 integration code?

Yes. Supertonic v3 maintains v2-compatible ONNX assets using the same inference contract. You can upgrade by simply changing the model_dir parameter from assets/supertonic-2/onnx to assets/onnx/. The Python SDK in py/example_onnx.py processes both versions identically without API modifications.

What are expression tags and how do they work in v3?

Expression tags are inline markup annotations (such as <laugh>, <breath>, and <sigh>) that inject natural prosodic nuance into synthesized speech. These 10 tags process natively within the v3 model architecture, allowing real-time emotional variation without text preprocessing or external audio editing. This feature is exclusive to v3 and unavailable in v2.

Which languages does v3 support compared to v2?

Supertonic v2 supports 5 languages (English plus 4 others), while v3 expands coverage to 31 languages with full multilingual capabilities. The v3 model achieves lower Word Error Rates across all supported languages, including significant improvements in English and Spanish accuracy.

Is v3 faster than v2 on CPU-only deployments?

Yes. Despite the increased parameter count, v3 delivers faster CPU inference latency and lower memory usage than v2. The optimized architecture in v3, combined with OnnxSlim compression, produces performance comparable to much larger GPU-bound baselines while maintaining the smaller v2 hardware requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →