How to Fine-Tune Supertonic Models: Voice Builder vs. Training Pipeline
Supertonic is an inference-only library that does not support traditional fine-tuning; instead, you customize voice output using the Voice Builder tool to generate style JSON files, or rebuild the training pipeline from scratch using the research papers referenced in the repository.
The supertone-inc/supertonic repository provides pre-trained ONNX models for text-to-speech synthesis across multiple platforms, but it does not include a training or fine-tuning pipeline. If you need to adapt the model to a specific voice or domain, you must use the Voice Builder workflow or implement the research-grade training stack yourself.
Why Supertonic Does Not Support Traditional Fine-Tuning
According to the source code in README.md (lines 19-30), Supertonic is explicitly designed as an inference-only library. The repository contains no training scripts, optimization loops, or dataset loaders—only optimized ONNX runtimes for Python, Node.js, WebGPU, Java, C++, C#, Go, Swift, iOS, Rust, and Flutter.
The codebase references advanced training techniques like self-purification flow-matching and Length-Aware RoPE (lines 492-504), but these describe the research methodology used to create the released models, not user-accessible fine-tuning interfaces. Because the heavy-weight training infrastructure—including large multilingual speech corpora and GPU-intensive training loops—is not part of the open-source distribution, you cannot directly fine-tune an existing Supertonic model within this repository.
Using Voice Builder for Voice Customization
Since traditional fine-tuning is unavailable, Supertonic provides Voice Builder, an official tool that converts a short reference recording into a version-specific JSON file. This creates a new voice style that the inference engine can load without modifying the underlying ONNX weights (lines 88-91).
Step-by-Step Voice Builder Workflow
- Obtain the base model. Download a pre-trained model such as
supertonic-3from the Hugging Face hub (lines 111-120). - Record a reference. Use the Voice Builder web application to process a short audio sample of your target voice.
- Export the JSON. Save the generated voice-style file (e.g.,
my_voice.json). - Load via SDK. Use the Python or Node.js SDK to load the JSON as a voice style during inference.
Python Implementation
The py/helper.py file provides the TTS class that handles model loading and voice style application.
from supertonic import TTS
# Auto-download the public model on first run
tts = TTS(auto_download=True)
# Load a custom voice style JSON produced by Voice Builder
voice_style = tts.load_voice_style("my_voice.json")
# Synthesize speech with the custom style
wav, duration = tts.synthesize(
text="Welcome to the fine-tuned Supertonic demo.",
lang="en",
voice_style=voice_style,
total_steps=8,
speed=1.0,
)
# Save the output WAV file
tts.save_audio(wav, "output.wav")
print(f"Generated {duration[0]:.2f}s of audio")
Node.js Implementation
For server-side applications, the nodejs/helper.js wrapper exposes an identical API.
const { TTS } = require("./helper.js");
// Load model and optional custom voice style
const tts = new TTS({ autoDownload: true });
tts.loadVoiceStyle("my_voice.json");
// Generate audio (returns Float32Array + duration)
const { wav, duration } = tts.synthesize({
text: "Fine-tuning via Voice Builder.",
lang: "en",
totalSteps: 8,
speed: 1.0,
});
Command-Line Interface
You can also serve models via HTTP for integration testing:
pip install "supertonic[serve]"
supertonic serve --host 127.0.0.1 --port 7788
Implementing a Full Training Pipeline
If you require capabilities beyond Voice Builder—such as adding new languages or drastically altering acoustic characteristics—you must rebuild the training pipeline yourself. The research papers cited in README.md (lines 490-502) describe the architecture, but implementing them requires:
- PyTorch or similar deep-learning frameworks
- Large-scale GPU clusters
- Access to the original training scripts (not included in the open-source release)
- Multilingual speech corpora for training
This path effectively means creating a new model from scratch rather than fine-tuning the existing ONNX files.
Summary
- Supertonic is inference-only; no fine-tuning scripts exist in the repository.
- Customize voices using Voice Builder to generate JSON style files compatible with the existing ONNX models.
- Load custom styles via
load_voice_style()inpy/helper.pyorloadVoiceStyle()innodejs/helper.js. - Full retraining requires implementing the research pipeline described in the cited papers, which is outside the scope of the open-source distribution.
- Reference examples are available in
py/example_onnx.pyandflutter/lib/helper.dart.
Frequently Asked Questions
Can I fine-tune Supertonic models on my own dataset?
No. The supertone-inc/supertonic repository does not contain training scripts or fine-tuning interfaces. It ships only pre-trained ONNX models optimized for inference. To adapt the model to new data, you must either use Voice Builder for style transfer or rebuild the training infrastructure from the research papers.
What is the Voice Builder tool?
Voice Builder is the official web application that converts a short reference recording into a JSON voice-style file. This file instructs the inference engine in py/helper.py to adjust synthesis parameters without modifying the underlying model weights, effectively creating a customized voice.
How do I load a custom voice style in Python?
Use the TTS class from the Supertonic SDK. After initialization with auto_download=True, call tts.load_voice_style("path/to/voice.json") to load your Voice Builder output. Pass the resulting object to the voice_style parameter in tts.synthesize().
Where can I find the training code for Supertonic?
The training code is not included in the open-source release. The README.md references research papers describing self-purification flow-matching and Length-Aware RoPE techniques (lines 492-504), but the actual implementation requires proprietary datasets and GPU-intensive training loops that are not part of the public repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →