How to Migrate from Supertonic v2 to v3 While Maintaining ONNX Interface Compatibility

You can migrate from Supertonic v2 to v3 without changing any inference code by simply replacing the four ONNX model files and the tts.json configuration file, because v3 preserves the exact same ONNX input/output signatures as v2.

Supertonic v3 expands support to 31 languages and improves inference speed, yet deliberately maintains the public ONNX inference contract established in v2. This guide walks you through the exact steps to upgrade the supertone-inc/supertonic assets while keeping your existing integration code intact.

Understanding the ONNX-Based Architecture

Supertonic uses a pipeline of four ONNX models that remain structurally identical between v2 and v3. The load_onnx_all function in py/helper.py (lines 82-97) initializes these models, and the TextToSpeech class orchestrates inference.

The four model files are:

  • duration_predictor.onnx – Predicts speech duration per input token (loaded as dp_ort)
  • text_encoder.onnx – Converts token IDs and style embeddings into text latents (loaded as text_enc_ort)
  • vector_estimator.onnx – Refines latents through flow-matching steps (loaded as vector_est_ort)
  • vocoder.onnx – Generates the final 44.1 kHz waveform (loaded as vocoder_ort)

The inference pipeline in TextToSpeech._infer feeds the same tensor names to each model regardless of version. These inputs include text_ids, style_ttl, style_dp, text_mask, and latent, ensuring that as long as the model files are compatible, your calling code requires no modifications.

Downloading the v3 ONNX Assets

The v3 models are hosted on Hugging Face and use the same filenames as v2, allowing for a simple file swap.

  1. Clone the v3 repository using Git LFS:
git lfs install
git clone https://huggingface.co/Supertone/supertonic-3 assets
  1. Replace your existing assets by overwriting the contents of your current assets directory with the new files:
    • duration_predictor.onnx
    • text_encoder.onnx
    • vector_estimator.onnx
    • vocoder.onnx
    • tts.json

Because the filenames are identical, the load_onnx_all helper automatically picks up the v3 versions when your application restarts.

Verifying Compatibility with Existing Code

You can verify the migration using the provided example_onnx.py script without any source changes:

from supertonic import TTS

# Initialize with local assets

tts = TTS(auto_download=False)

# Load voice style from JSON

style = tts.get_voice_style(voice_name="M1")

# Synthesize audio using the same API as v2

wav, dur = tts.synthesize(
    text="Supertonic v3 maintains full ONNX backward compatibility.",
    lang="en",
    voice_style=style,
    total_steps=8,
    speed=1.05,
)

tts.save_audio(wav, "v3_output.wav")

Running this script should produce audio output immediately. If you encounter no shape mismatches or input errors, your migration is complete.

Expanding to New Languages (Optional)

While not required for migration, v3 adds significant language coverage. The AVAILABLE_LANGS constant in py/helper.py (lines 13-14) now supports 31 languages compared to v2's smaller set.

You can now pass new ISO-639-1 codes such as vi (Vietnamese), uk (Ukrainian), or tr (Turkish) to tts.synthesize without changing your inference logic.

Cross-Language Migration Checklist

All official SDKs follow the same pattern: they load the four ONNX files from a directory you specify. To migrate from Supertonic v2 to v3 across different languages, simply replace the assets folder:

Verifying On-Device Performance

After migration, run your existing benchmark scripts to confirm improvements. The v3 models offer faster real-time factors and reduced memory usage compared to v2, though the ONNX input/output interfaces remain identical. You can use the test_all.sh script to validate performance across all supported language bindings.

Summary

  • Supertonic v3 preserves the exact ONNX signatures (text_ids, style_ttl, style_dp, text_mask, etc.) used in v2, ensuring zero-code-change migration.
  • Replace five files (duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, vocoder.onnx, and tts.json) from the Hugging Face repository to upgrade.
  • load_onnx_all in py/helper.py and its equivalents in other SDKs handle the new assets automatically.
  • 31 languages are now available using the same synthesis API.
  • All language SDKs (Python, Node.js, Go, Java, C++, C#, Swift, Flutter, iOS) follow the same migration pattern: swap assets, keep code.

Frequently Asked Questions

Do I need to modify my Python inference code when upgrading from v2 to v3?

No. According to the Supertonic source code, the TextToSpeech class and load_onnx_all function accept the same input tensors and model filenames. Simply replace the ONNX files and tts.json in your assets directory; your existing tts.synthesize() calls will work unchanged.

Where can I download the official v3 ONNX model files?

The official v3 assets are hosted at https://huggingface.co/Supertone/supertonic-3. You must use Git LFS to clone the repository, which contains the four .onnx files and the configuration JSON required for inference.

Will my v2 integrations break if I accidentally mix v2 and v3 model files?

Yes, you must replace all four ONNX files (duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, vocoder.onnx) simultaneously with their v3 counterparts. The tts.json file should also be updated to match the v3 release. Mixing versions across the pipeline will cause inference errors or garbled output.

Does v3 support the same voice styles and model configurations as v2?

Yes, the voice style JSON files and the tts.json configuration schema remain compatible. However, v3 adds support for 31 languages (expanded from v2's set) and improves accuracy, so you may benefit from re-exporting styles if you need the new linguistic capabilities.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →