# API Differences Between Supertonic 2 and Supertonic 3: What Changed in the TTS Engine

> Understand the API differences between Supertonic 2 and Supertonic 3. Upgrade seamlessly as HTTP endpoints remain identical, focusing on model architecture and performance enhancements.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: api-reference
- Published: 2026-06-14

---

**Supertonic 2 and Supertonic 3 expose identical HTTP endpoints at `/v1/tts` and `/v1/audio/speech`, meaning you can upgrade from v2 to v3 without changing client code—the only differences lie in the underlying model architecture, language support, and inference performance.**

The `supertone-inc/supertonic` repository maintains both versions under different release branches, but contrary to typical major version bumps, the API differences between Supertonic 2 and Supertonic 3 are confined to the server-side model rather than the request contract. Understanding these distinctions helps you optimize deployment strategy while maintaining backward compatibility with existing integrations.

## API Surface: No Breaking Changes

Both versions share the **same HTTP API surface**. According to the repository documentation, the server exposes two endpoints regardless of which model is loaded:

- **Native endpoint**: `POST /v1/tts`
- **OpenAI-compatible endpoint**: `POST /v1/audio/speech`

As noted in [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md), the request and response contracts remain unchanged between versions. The distinction is not in the REST interface but in the **model that backs the service** and the capabilities that model provides.

## Model Architecture and Performance Differences

While the API endpoints stay constant, the underlying text-to-speech engine differs significantly between versions.

### Language Coverage Expansion

**Supertonic 2** supports **5 languages**, as maintained in the `release/supertonic-2` branch. **Supertonic 3**, located on the `main` branch, expands this to **31 languages**. The `lang` field in your API requests accepts the expanded code list only when running the v3 model; attempts to use unsupported language codes with the v2 model will return validation errors.

### Model Size and Inference Speed

The parameter count dropped dramatically between versions:

- **Supertonic 2**: Larger architecture ranging from approximately **0.7B to 2B parameters**, requiring more memory and GPU resources for optimal performance
- **Supertonic 3**: Streamlined to approximately **99M parameters**, enabling fast CPU inference with reduced memory footprint

As documented in the repository's performance notes, this reduction allows v3 to run efficiently on CPU hardware where v2 required more substantial compute resources.

### Audio Quality Improvements

Despite the smaller model size, Supertonic 3 delivers **superior acoustic output**. The v3 model exhibits reduced repeat/skip failures and improved speaker similarity across the shared language set compared to v2's baseline performance. The output is louder, clearer, and contains fewer artifacts, though these quality improvements are transparent to the API client.

## Server Configuration and Deployment

Selecting between versions occurs at server startup rather than through API versioning headers.

### Branch Selection

The code paths are separated by Git branches:

- **Supertonic 2**: `release/supertonic-2` branch
- **Supertonic 3**: `main` branch (default)

### Runtime Model Selection

When using the Python SDK, specify the model version via command-line arguments:

```bash

# Install the SDK with server capabilities

pip install "supertonic[serve]"

# Launch Supertonic 2

supertonic serve --model v2 --host 127.0.0.1 --port 7788

# Launch Supertonic 3 (default on main branch)

supertonic serve --model v3 --host 127.0.0.1 --port 7788

```

The `--model` parameter determines which binary assets the server loads, while the HTTP interface remains identical regardless of selection.

## Voice Builder and ONNX Compatibility

The **ONNX interface remains v2-compatible** in Supertonic 3, ensuring that existing integrations can point to the new model without code changes. As implemented in the `go/` directory and documented in the main README, the inference contract stays constant.

Voice Builder JSON files work across both versions. While Supertonic 2 only supported JSON files specific to v2, Supertonic 3 introduced version-specific JSON files for both versions, allowing the same voice definition to be used with either model backend.

## Practical Implementation Examples

### Starting the Server for Each Version

```bash

# Supertonic 2 - checkout the v2 release branch

git checkout release/supertonic-2
supertonic serve --host 127.0.0.1 --port 7788

# Supertonic 3 - use the main branch (default)

git checkout main
supertonic serve --host 127.0.0.1 --port 7788

```

### Making API Requests

The request structure is identical for both versions. Only the available voice/language codes differ:

```bash
curl -X POST http://127.0.0.1:7788/v1/tts \
     -H "Content-Type: application/json" \
     -d '{
           "text": "Hello, world!",
           "voice": "en-US-female",
           "output_format": "wav"
         }' --output hello.wav

```

When using Supertonic 3, you can substitute `"voice"` with any of the 31 supported language codes (e.g., `"ko-KR-female"`, `"de-DE-male"`), whereas Supertonic 2 limits you to the original 5-language set.

### OpenAI-Compatible Endpoint Usage

```bash
curl -X POST http://127.0.0.1:7788/v1/audio/speech \
     -H "Authorization: Bearer dummy" \
     -H "Content-Type: application/json" \
     -d '{
           "model": "supertonic",
           "input": "Bonjour le monde",
           "voice": "fr-FR-male"
         }' --output bonjour.wav

```

The OpenAI-compatible endpoint at `/v1/audio/speech` accepts identical payloads for both versions, returning audio generated by whichever model the server instance has loaded.

## Summary

- **Supertonic 2 and 3 share identical HTTP API endpoints** (`/v1/tts` and `/v1/audio/speech`) with no request contract changes between versions.
- **Language support expanded** from 5 languages in v2 to 31 languages in v3, though the JSON request structure remains unchanged.
- **Model architecture shifted** from 0.7B–2B parameters in v2 to approximately 99M parameters in v3, enabling CPU-efficient inference while improving audio quality.
- **ONNX and Voice Builder compatibility** is maintained—v3 uses the v2-compatible ONNX interface, and voice JSON files work across both versions.
- **Version selection happens at deployment** via Git branches (`release/supertonic-2` vs. `main`) or the `--model` CLI flag, not through API versioning.

## Frequently Asked Questions

### Do I need to modify my client code when upgrading from Supertonic 2 to 3?

No. The HTTP API surface remains identical between versions. Your existing `POST /v1/tts` or `POST /v1/audio/speech` requests will work without modification. You only need to ensure that the `voice` parameter uses language codes supported by the specific model version running on your server.

### Can I use the same voice JSON files with both Supertonic versions?

Yes. While Supertonic 2 only supported v2-specific JSON files, Supertonic 3 introduced version-specific JSON files that work for both versions. The same voice definition can be used with either model backend, as the ONNX interface maintains backward compatibility.

### Why does Supertonic 3 support more languages with a smaller model?

Supertonic 3 utilizes a more efficient architecture (~99M parameters compared to v2's 0.7B–2B) that achieves better multilingual coverage through improved training techniques and data efficiency. The smaller size enables faster CPU inference and lower memory usage while delivering higher quality output with fewer artifacts than the larger v2 model.

### How do I check which model version my server is currently running?

Check which Git branch your deployment uses (`release/supertonic-2` for v2, `main` for v3) or review the server startup logs when using the `--model` flag. If you attempt to use language codes unsupported by the loaded model, the server will return a validation error indicating the discrepancy between your request and the active model's capabilities.