# How to Fine-Tune Supertonic Models: Voice Builder vs. Training Pipeline

> Discover how to fine-tune Supertonic models. Customize voice output with Voice Builder or rebuild the training pipeline using research papers from the repository.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Supertonic is an inference-only library that does not support traditional fine-tuning; instead, you customize voice output using the Voice Builder tool to generate style JSON files, or rebuild the training pipeline from scratch using the research papers referenced in the repository.**

The `supertone-inc/supertonic` repository provides pre-trained ONNX models for text-to-speech synthesis across multiple platforms, but it does not include a training or fine-tuning pipeline. If you need to adapt the model to a specific voice or domain, you must use the Voice Builder workflow or implement the research-grade training stack yourself.

## Why Supertonic Does Not Support Traditional Fine-Tuning

According to the source code in [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md) (lines 19-30), Supertonic is explicitly designed as an **inference-only** library. The repository contains no training scripts, optimization loops, or dataset loaders—only optimized ONNX runtimes for Python, Node.js, WebGPU, Java, C++, C#, Go, Swift, iOS, Rust, and Flutter.

The codebase references advanced training techniques like self-purification flow-matching and Length-Aware RoPE (lines 492-504), but these describe the research methodology used to create the released models, not user-accessible fine-tuning interfaces. Because the heavy-weight training infrastructure—including large multilingual speech corpora and GPU-intensive training loops—is not part of the open-source distribution, you cannot directly fine-tune an existing Supertonic model within this repository.

## Using Voice Builder for Voice Customization

Since traditional fine-tuning is unavailable, Supertonic provides **Voice Builder**, an official tool that converts a short reference recording into a version-specific JSON file. This creates a new voice style that the inference engine can load without modifying the underlying ONNX weights (lines 88-91).

### Step-by-Step Voice Builder Workflow

1. **Obtain the base model.** Download a pre-trained model such as `supertonic-3` from the Hugging Face hub (lines 111-120).
2. **Record a reference.** Use the Voice Builder web application to process a short audio sample of your target voice.
3. **Export the JSON.** Save the generated voice-style file (e.g., [`my_voice.json`](https://github.com/supertone-inc/supertonic/blob/main/my_voice.json)).
4. **Load via SDK.** Use the Python or Node.js SDK to load the JSON as a voice style during inference.

### Python Implementation

The [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) file provides the `TTS` class that handles model loading and voice style application.

```python
from supertonic import TTS

# Auto-download the public model on first run

tts = TTS(auto_download=True)

# Load a custom voice style JSON produced by Voice Builder

voice_style = tts.load_voice_style("my_voice.json")

# Synthesize speech with the custom style

wav, duration = tts.synthesize(
    text="Welcome to the fine-tuned Supertonic demo.",
    lang="en",
    voice_style=voice_style,
    total_steps=8,
    speed=1.0,
)

# Save the output WAV file

tts.save_audio(wav, "output.wav")
print(f"Generated {duration[0]:.2f}s of audio")

```

### Node.js Implementation

For server-side applications, the [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js) wrapper exposes an identical API.

```javascript
const { TTS } = require("./helper.js");

// Load model and optional custom voice style
const tts = new TTS({ autoDownload: true });
tts.loadVoiceStyle("my_voice.json");

// Generate audio (returns Float32Array + duration)
const { wav, duration } = tts.synthesize({
  text: "Fine-tuning via Voice Builder.",
  lang: "en",
  totalSteps: 8,
  speed: 1.0,
});

```

### Command-Line Interface

You can also serve models via HTTP for integration testing:

```bash
pip install "supertonic[serve]"
supertonic serve --host 127.0.0.1 --port 7788

```

## Implementing a Full Training Pipeline

If you require capabilities beyond Voice Builder—such as adding new languages or drastically altering acoustic characteristics—you must **rebuild the training pipeline yourself**. The research papers cited in [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md) (lines 490-502) describe the architecture, but implementing them requires:

- PyTorch or similar deep-learning frameworks
- Large-scale GPU clusters
- Access to the original training scripts (not included in the open-source release)
- Multilingual speech corpora for training

This path effectively means creating a new model from scratch rather than fine-tuning the existing ONNX files.

## Summary

- Supertonic is **inference-only**; no fine-tuning scripts exist in the repository.
- Customize voices using **Voice Builder** to generate JSON style files compatible with the existing ONNX models.
- Load custom styles via `load_voice_style()` in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) or `loadVoiceStyle()` in [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js).
- Full retraining requires implementing the research pipeline described in the cited papers, which is outside the scope of the open-source distribution.
- Reference examples are available in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) and `flutter/lib/helper.dart`.

## Frequently Asked Questions

### Can I fine-tune Supertonic models on my own dataset?

No. The `supertone-inc/supertonic` repository does not contain training scripts or fine-tuning interfaces. It ships only pre-trained ONNX models optimized for inference. To adapt the model to new data, you must either use Voice Builder for style transfer or rebuild the training infrastructure from the research papers.

### What is the Voice Builder tool?

Voice Builder is the official web application that converts a short reference recording into a JSON voice-style file. This file instructs the inference engine in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) to adjust synthesis parameters without modifying the underlying model weights, effectively creating a customized voice.

### How do I load a custom voice style in Python?

Use the `TTS` class from the Supertonic SDK. After initialization with `auto_download=True`, call `tts.load_voice_style("path/to/voice.json")` to load your Voice Builder output. Pass the resulting object to the `voice_style` parameter in `tts.synthesize()`.

### Where can I find the training code for Supertonic?

The training code is not included in the open-source release. The [`README.md`](https://github.com/supertone-inc/supertonic/blob/main/README.md) references research papers describing self-purification flow-matching and Length-Aware RoPE techniques (lines 492-504), but the actual implementation requires proprietary datasets and GPU-intensive training loops that are not part of the public repository.