How to Implement Custom Voice Styles for Supertonic: A Complete Guide

Create a JSON file containing style_ttl and style_dp tensors, place it in assets/voice_styles/, and load it via the --voice-style CLI flag or loadVoiceStyle() function to apply custom acoustic characteristics to your speech synthesis.

Supertonic is an open-source neural text-to-speech library that generates speech by applying voice-style JSON files describing speaker characteristics like timing, pitch, and intensity. The repository includes pre-extracted styles (M1-M5, F1-F5) in assets/voice_styles/, but you can implement custom voice styles by creating compatible JSON files and loading them through any of the supported language bindings including Python, Rust, Go, Swift, or JavaScript.

Understanding the Voice Style Architecture

Supertonic uses a unified JSON schema across all platforms. Each voice style file contains two flattened tensors that control different aspects of the neural network: style_ttl (timbre-level latent vectors for the text-to-latent network) and style_dp (diffusion-process latent vectors for the decoder network). The loader concatenates these tensors when multiple files are provided, enabling batch generation with different styles per text.

How the Loader Works Across Bindings

Each language binding implements the same loading logic in its respective helper file:

  • JavaScript: In web/helper.js, the loadVoiceStyle() function fetches JSON files, parses them, and flattens the tensors into a single voiceStyle object.
  • Rust: The load_voice_style() function in rust/src/helper.rs deserializes files into VoiceStyleData and concatenates the tensors.
  • Swift: swift/Sources/Helper.swift uses JSONDecoder to read files into Data objects and process the arrays.
  • Python: py/helper.py provides load_voice_style() for the Python binding.
  • Go: go/helper.go implements LoadVoiceStyle() with JSON unmarshaling.

Creating a Custom Voice Style Step-by-Step

1. Copy an Existing Style Template

Start with the minimal example in assets/voice_styles/M1.json. Copy this file to create your base, for example MyCustom.json.

2. Modify the Tensors

Edit the floating-point arrays in style_ttl.data and style_dp.data. These values represent the acoustic characteristics of your speaker. For production use, generate these tensors by running the Supertonic training pipeline on your target speaker's recordings. For experimentation, manually tweak the existing values.

3. Validate and Place the File

Validate your JSON syntax using a tool like jq. Store the file in assets/voice_styles/ relative to your runtime root, or any location if using absolute paths via CLI.

4. Reference the Style in Your Application

For web applications, add an option to the dropdown in web/index.html:

<option value="assets/voice_styles/MyCustom.json">My Custom (MC)</option>

For CLI usage across all languages, use the --voice-style flag.

Implementation Examples by Language

Web UI Integration

In web/main.js, the selection handler automatically calls loadVoiceStyle() when users select a style from the dropdown. Ensure your custom style is added to web/index.html as shown above.

Python CLI


# Install dependencies first

uv pip install -r py/requirements.txt

# Run with custom style

uv run py/example_onnx.py \
    --voice-style assets/voice_styles/MyCustom.json \
    --text "Welcome to Supertonic" \
    --lang en

Rust CLI

cargo run --release -- \
    --voice-style assets/voice_styles/MyCustom.json \
    --text "Welcome to Supertonic" \
    --lang en

Go CLI

go run go/example_onnx.go go/helper.go \
    --voice-style assets/voice_styles/MyCustom.json \
    --text "Welcome to Supertonic" \
    --lang en

Swift iOS Implementation

let style = try Helper.loadVoiceStyle(
    ["../assets/voice_styles/MyCustom.json"], 
    verbose: true
)
let audio = try synthesizer.synthesize(
    text: "Welcome to Supertonic", 
    style: style
)

Key Source Files for Reference

Summary

Frequently Asked Questions

What is the exact format of the voice style JSON file?

The voice style JSON must contain two top-level keys: style_ttl and style_dp. Each contains a data field with 2-D or 3-D float arrays representing latent vectors. The loader flattens these arrays during parsing, so you can structure them as nested lists compatible with your training pipeline's output.

Can I use multiple custom voice styles in a single batch?

Yes. The loadVoiceStyle() functions in all language bindings accept an array of file paths and concatenate the tensors. This allows you to provide a different style for each text segment in batch generation workflows, as demonstrated in the Go example at go/example_onnx.go.

Do I need to retrain the Supertonic model to create a custom voice style?

No. You can create custom styles by manually editing the tensor values in the JSON file, though for optimal results you should extract these values using the Supertonic training pipeline on your target speaker's recordings. The model itself remains unchanged; only the style parameters are customized.

Why does my custom style fail to load with a JSON error?

Ensure your JSON syntax is valid and the file paths are correct relative to your execution directory. Validate the file using jq or an online validator. The tensors must be properly formatted as nested arrays of floats. If using the web UI, verify that the path in the <option> value matches the actual file location relative to the server root.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →