# Supertonic v2 vs v3: Key Differences and Migration Guide

> Discover Supertonic v3 key differences from v2. Get 3x language coverage, improved accuracy, smaller size, and expressive control tags. Migrate easily with the same ONNX contract.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: migration-guide
- Published: 2026-06-12

---

**Supertonic v3 is a drop-in replacement for v2 that triples language coverage to 31 languages, adds expressive control tags, improves accuracy by up to 38% WER reduction, and reduces model size by 25%—all while maintaining the same ONNX inference contract.**

The `supertone-inc/supertonic` repository provides state-of-the-art neural text-to-speech models with open-source ONNX runtimes. Understanding the differences between Supertonic v2 and v3 helps developers choose the right version for multilingual synthesis workloads and migration strategies.

## Architecture and Scale Improvements

Supertonic v3 substantially scales the model capacity while optimizing for deployment efficiency.

### Parameter Count and Model Size

According to the project [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md), Supertonic v2 utilizes approximately **66 million parameters** (line 50), whereas Supertonic v3 expands to roughly **99 million parameters** (line 81). Despite the 50% parameter increase, v3 achieves a smaller disk footprint through aggressive optimization.

The v3 release applies **OnnxSlim** techniques to compress the model assets, resulting in approximately **150 MB** downloads compared to v2's **~200 MB** (line 47). This reduction eliminates bloat while preserving output quality.

### Runtime Performance

Benchmarks documented in the repository (lines 73-78) demonstrate that v3 delivers **faster CPU latency** and **lower memory usage** than v2, performing comparably to much larger GPU-bound baselines. The architectural improvements in v3 eliminate the synthesis bottlenecks present in v2 without requiring hardware acceleration.

## Language Support and Accuracy

The most significant functional difference lies in multilingual capabilities and reading precision.

### Expanded Language Coverage

Supertonic v2 supports **5 languages** (English plus 4 others) as noted at line 44. Supertonic v3 expands this to **31 languages** with full multilingual support (line 33), enabling single-model polyglot synthesis without voice switching.

### Word Error Rate Improvements

Per-language accuracy metrics in [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md) (lines 41-56) show consistent reductions in Word Error Rate (WER) and Character Error Rate (CER):

- **English**: WER improves from 2.52 to 2.06
- **Spanish**: WER improves from 1.81 to 1.13

These improvements represent up to 38% error reduction, often outperforming larger open-source alternatives like VoxCPM2. Additionally, v3 significantly reduces **repeat and skip failures** during synthesis (line 270), producing smoother, more natural speech patterns than v2.

### Speaker Similarity

When switching between voice styles, v3 maintains **improved speaker similarity** across the shared language set compared to v2 (line 270), ensuring consistent vocal characteristics across multilingual content.

## Expression Tags and Voice Control

Supertonic v3 introduces **10 inline expression tags** that v2 lacks entirely (line 28). These markup annotations allow fine-grained prosodic control without preprocessing:

- `<laugh>` - adds laughter intonation
- `<breath>` - inserts natural breathing pauses
- `<sigh>` - conveys exasperation or relief

These tags process natively within the v3 model architecture, enabling nuanced emotional expression that v2 cannot replicate.

## Backward Compatibility and ONNX Interface

Despite the architectural upgrades, v3 maintains **full backward compatibility** with v2 integrations.

### Drop-in Replacement Architecture

The v3 ONNX assets remain **v2-compatible** (line 42), utilizing the identical inference contract. Existing implementations can upgrade by simply changing the `model_dir` parameter from the v2-specific branch (`release/supertonic-2`) to the default v3 assets (`assets/onnx/`).

### File Structure Differences

- **v2 assets**: Located in `release/supertonic-2/` branch with legacy formatting
- **v3 assets**: Stored in `assets/onnx/` with optimized slimming

Both versions use the same Python SDK interface defined in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) and [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), ensuring helper utilities for loading models and voice styles function identically across versions.

## Code Examples

### Basic v2 Inference

Load the legacy checkpoint from the v2 branch:

```python
from supertonic import TTS

tts_v2 = TTS(
    model_dir="assets/supertonic-2/onnx",
    auto_download=False
)

style = tts_v2.get_voice_style(voice_name="M1")
wav, _ = tts_v2.synthesize(
    text="Hello, this is Supertonic version 2.",
    lang="en",
    voice_style=style,
    total_steps=8,
    speed=1.0,
)
tts_v2.save_audio(wav, "v2_output.wav")

```

### Migration to v3

Upgrade by changing only the model directory path:

```python
from supertonic import TTS

tts_v3 = TTS(
    model_dir="assets/onnx",  # v3 default path

    auto_download=False
)

style = tts_v3.get_voice_style(voice_name="M1")
wav, _ = tts_v3.synthesize(
    text="Hello, this is Supertonic version 3 with 31 languages.",
    lang="en",
    voice_style=style,
    total_steps=8,
    speed=1.0,
)
tts_v3.save_audio(wav, "v3_output.wav")

```

### Expression Tags (v3 Only)

Leverage v3-exclusive prosodic controls:

```python
from supertonic import TTS

tts = TTS(auto_download=True)
style = tts.get_voice_style(voice_name="M1")

text = "The meeting was great <laugh>! Let's celebrate <sigh>."
wav, _ = tts.synthesize(
    text=text,
    lang="en",
    voice_style=style,
    total_steps=10,
    speed=1.1,
)
tts.save_audio(wav, "v3_expression.wav")

```

## Summary

- **Supertonic v3** scales from 66M to 99M parameters while reducing disk size from ~200MB to ~150MB through OnnxSlim optimization
- Language support expands from 5 to 31 languages with significant WER improvements (e.g., English 2.52→2.06, Spanish 1.81→1.13)
- **Expression tags** (`<laugh>`, `<breath>`, etc.) provide v3-exclusive prosodic control unavailable in v2
- **Backward compatibility** allows drop-in replacement via the same ONNX inference contract—only the `model_dir` path changes
- v3 eliminates repeat/skip failures and improves speaker similarity across voice styles

## Frequently Asked Questions

### Can I use Supertonic v3 with existing v2 integration code?

Yes. Supertonic v3 maintains **v2-compatible ONNX assets** using the same inference contract. You can upgrade by simply changing the `model_dir` parameter from `assets/supertonic-2/onnx` to `assets/onnx/`. The Python SDK in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) processes both versions identically without API modifications.

### What are expression tags and how do they work in v3?

Expression tags are inline markup annotations (such as `<laugh>`, `<breath>`, and `<sigh>`) that inject natural prosodic nuance into synthesized speech. These 10 tags process natively within the v3 model architecture, allowing real-time emotional variation without text preprocessing or external audio editing. This feature is exclusive to v3 and unavailable in v2.

### Which languages does v3 support compared to v2?

Supertonic v2 supports 5 languages (English plus 4 others), while v3 expands coverage to **31 languages** with full multilingual capabilities. The v3 model achieves lower Word Error Rates across all supported languages, including significant improvements in English and Spanish accuracy.

### Is v3 faster than v2 on CPU-only deployments?

Yes. Despite the increased parameter count, v3 delivers faster CPU inference latency and lower memory usage than v2. The optimized architecture in v3, combined with OnnxSlim compression, produces performance comparable to much larger GPU-bound baselines while maintaining the smaller v2 hardware requirements.