API Differences Between Supertonic 2 and Supertonic 3: What Changed in the TTS Engine
Supertonic 2 and Supertonic 3 expose identical HTTP endpoints at /v1/tts and /v1/audio/speech, meaning you can upgrade from v2 to v3 without changing client code—the only differences lie in the underlying model architecture, language support, and inference performance.
The supertone-inc/supertonic repository maintains both versions under different release branches, but contrary to typical major version bumps, the API differences between Supertonic 2 and Supertonic 3 are confined to the server-side model rather than the request contract. Understanding these distinctions helps you optimize deployment strategy while maintaining backward compatibility with existing integrations.
API Surface: No Breaking Changes
Both versions share the same HTTP API surface. According to the repository documentation, the server exposes two endpoints regardless of which model is loaded:
- Native endpoint:
POST /v1/tts - OpenAI-compatible endpoint:
POST /v1/audio/speech
As noted in README.md, the request and response contracts remain unchanged between versions. The distinction is not in the REST interface but in the model that backs the service and the capabilities that model provides.
Model Architecture and Performance Differences
While the API endpoints stay constant, the underlying text-to-speech engine differs significantly between versions.
Language Coverage Expansion
Supertonic 2 supports 5 languages, as maintained in the release/supertonic-2 branch. Supertonic 3, located on the main branch, expands this to 31 languages. The lang field in your API requests accepts the expanded code list only when running the v3 model; attempts to use unsupported language codes with the v2 model will return validation errors.
Model Size and Inference Speed
The parameter count dropped dramatically between versions:
- Supertonic 2: Larger architecture ranging from approximately 0.7B to 2B parameters, requiring more memory and GPU resources for optimal performance
- Supertonic 3: Streamlined to approximately 99M parameters, enabling fast CPU inference with reduced memory footprint
As documented in the repository's performance notes, this reduction allows v3 to run efficiently on CPU hardware where v2 required more substantial compute resources.
Audio Quality Improvements
Despite the smaller model size, Supertonic 3 delivers superior acoustic output. The v3 model exhibits reduced repeat/skip failures and improved speaker similarity across the shared language set compared to v2's baseline performance. The output is louder, clearer, and contains fewer artifacts, though these quality improvements are transparent to the API client.
Server Configuration and Deployment
Selecting between versions occurs at server startup rather than through API versioning headers.
Branch Selection
The code paths are separated by Git branches:
- Supertonic 2:
release/supertonic-2branch - Supertonic 3:
mainbranch (default)
Runtime Model Selection
When using the Python SDK, specify the model version via command-line arguments:
# Install the SDK with server capabilities
pip install "supertonic[serve]"
# Launch Supertonic 2
supertonic serve --model v2 --host 127.0.0.1 --port 7788
# Launch Supertonic 3 (default on main branch)
supertonic serve --model v3 --host 127.0.0.1 --port 7788
The --model parameter determines which binary assets the server loads, while the HTTP interface remains identical regardless of selection.
Voice Builder and ONNX Compatibility
The ONNX interface remains v2-compatible in Supertonic 3, ensuring that existing integrations can point to the new model without code changes. As implemented in the go/ directory and documented in the main README, the inference contract stays constant.
Voice Builder JSON files work across both versions. While Supertonic 2 only supported JSON files specific to v2, Supertonic 3 introduced version-specific JSON files for both versions, allowing the same voice definition to be used with either model backend.
Practical Implementation Examples
Starting the Server for Each Version
# Supertonic 2 - checkout the v2 release branch
git checkout release/supertonic-2
supertonic serve --host 127.0.0.1 --port 7788
# Supertonic 3 - use the main branch (default)
git checkout main
supertonic serve --host 127.0.0.1 --port 7788
Making API Requests
The request structure is identical for both versions. Only the available voice/language codes differ:
curl -X POST http://127.0.0.1:7788/v1/tts \
-H "Content-Type: application/json" \
-d '{
"text": "Hello, world!",
"voice": "en-US-female",
"output_format": "wav"
}' --output hello.wav
When using Supertonic 3, you can substitute "voice" with any of the 31 supported language codes (e.g., "ko-KR-female", "de-DE-male"), whereas Supertonic 2 limits you to the original 5-language set.
OpenAI-Compatible Endpoint Usage
curl -X POST http://127.0.0.1:7788/v1/audio/speech \
-H "Authorization: Bearer dummy" \
-H "Content-Type: application/json" \
-d '{
"model": "supertonic",
"input": "Bonjour le monde",
"voice": "fr-FR-male"
}' --output bonjour.wav
The OpenAI-compatible endpoint at /v1/audio/speech accepts identical payloads for both versions, returning audio generated by whichever model the server instance has loaded.
Summary
- Supertonic 2 and 3 share identical HTTP API endpoints (
/v1/ttsand/v1/audio/speech) with no request contract changes between versions. - Language support expanded from 5 languages in v2 to 31 languages in v3, though the JSON request structure remains unchanged.
- Model architecture shifted from 0.7B–2B parameters in v2 to approximately 99M parameters in v3, enabling CPU-efficient inference while improving audio quality.
- ONNX and Voice Builder compatibility is maintained—v3 uses the v2-compatible ONNX interface, and voice JSON files work across both versions.
- Version selection happens at deployment via Git branches (
release/supertonic-2vs.main) or the--modelCLI flag, not through API versioning.
Frequently Asked Questions
Do I need to modify my client code when upgrading from Supertonic 2 to 3?
No. The HTTP API surface remains identical between versions. Your existing POST /v1/tts or POST /v1/audio/speech requests will work without modification. You only need to ensure that the voice parameter uses language codes supported by the specific model version running on your server.
Can I use the same voice JSON files with both Supertonic versions?
Yes. While Supertonic 2 only supported v2-specific JSON files, Supertonic 3 introduced version-specific JSON files that work for both versions. The same voice definition can be used with either model backend, as the ONNX interface maintains backward compatibility.
Why does Supertonic 3 support more languages with a smaller model?
Supertonic 3 utilizes a more efficient architecture (~99M parameters compared to v2's 0.7B–2B) that achieves better multilingual coverage through improved training techniques and data efficiency. The smaller size enables faster CPU inference and lower memory usage while delivering higher quality output with fewer artifacts than the larger v2 model.
How do I check which model version my server is currently running?
Check which Git branch your deployment uses (release/supertonic-2 for v2, main for v3) or review the server startup logs when using the --model flag. If you attempt to use language codes unsupported by the loaded model, the server will return a validation error indicating the discrepancy between your request and the active model's capabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →