How to Create Custom Voice Styles for Supertonic: A Step-by-Step Guide
To create custom voice styles for Supertonic, author a JSON file containing style_ttl and style_dp tensor arrays, place it in assets/voice_styles/, and reference the file path via the web UI dropdown or the --voice-style CLI flag.
Supertonic is an open-source text-to-speech library that generates speech by applying voice-style JSON files describing acoustic characteristics like timing, pitch, and intensity. While the repository ships with pre-extracted styles (M1-M5, F1-F5) located in assets/voice_styles/, you can extend the system with custom speaker profiles. This guide explains the loader architecture, the required JSON schema, and implementation steps across Python, Rust, Swift, Go, and JavaScript bindings.
Understanding the Voice Style Architecture
Supertonic loads voice styles through a consistent architecture shared across all language bindings. The system deserializes JSON files containing latent tensors that control the timbre and diffusion process of the neural TTS model.
The JSON Schema for Voice Styles
Every voice style file must follow this structure:
{
"style_ttl": {
"data": [
[ ... ]
]
},
"style_dp": {
"data": [
[ ... ]
]
}
}
The style_ttl field contains timbre-level latent vectors for the text-to-latent network, while style_dp contains diffusion-process latent vectors for the decoder network. Both contain 2-D or 3-D float tensors that the loader flattens and concatenates when processing multiple style files.
How Loaders Work Across Language Bindings
Each language binding implements a loadVoiceStyle function that performs identical operations:
web/helper.js: TheloadVoiceStyle()function fetches JSON files, parses them, and flattens thestyle_ttlandstyle_dptensors into a singlevoiceStyleobject.rust/src/helper.rs: Theload_voice_style()function opens files for validation, deserializes each intoVoiceStyleData, and concatenates the tensors.swift/Sources/Helper.swift: TheloadVoiceStyle()method reads files intoData, decodes them withJSONDecoder, and merges the arrays.go/helper.go: TheLoadVoiceStyle()function unmarshals JSON intoVoiceStyleDataand appends the tensors.py/helper.py: The Python implementation mirrors this logic for the Python binding.
All loaders support batch generation by concatenating the data arrays of multiple supplied files.
Step-by-Step Guide to Creating Custom Voice Styles
Step 1: Copy and Modify an Existing Style
Start by copying an existing style from assets/voice_styles/M1.json to a new file such as MyCustom.json. Edit the floating-point values in the style_ttl.data and style_dp.data arrays to sculpt the acoustic characteristics.
For production-quality styles, generate new tensors by running the Supertonic training script on a target speaker's recordings. For quick experimentation, manually adjusting values in an existing JSON file creates variation in speaking style and timbre.
Step 2: Validate and Place the JSON File
Validate your JSON syntax using a tool like jq or an online validator to prevent deserialization errors. Store the validated file under assets/voice_styles/ relative to the runtime root. The web UI expects this path structure, while CLI tools accept absolute or relative paths via the --voice-style flag.
Step 3: Reference Your Custom Style
For the Web UI, add an option to the dropdown in web/index.html:
<select id="voiceStyleSelect">
<option value="assets/voice_styles/MyCustom.json">My Custom (MC)</option>
</select>
The selection logic in web/main.js automatically calls loadVoiceStyle() when the dropdown changes.
For CLI usage, pass the file path using the --voice-style argument:
# Python example
uv run py/example_onnx.py \
--voice-style assets/voice_styles/MyCustom.json \
--text "Hello world" \
--lang en
Code Examples by Language
Python CLI Implementation
Install dependencies and run synthesis with your custom style:
uv pip install -r py/requirements.txt
uv run py/example_onnx.py \
--voice-style assets/voice_styles/MyCustom.json \
--text "Welcome to Supertonic" \
--lang en
The py/helper.py file handles parsing the JSON and building the Style object for the ONNX runtime.
Web UI Integration
In the browser environment, the loader fetches the JSON asynchronously:
// The helper.js loadVoiceStyle function handles the fetch
const style = await loadVoiceStyle(['assets/voice_styles/MyCustom.json']);
Ensure your web server serves the assets/ directory so the fetch request in web/helper.js resolves correctly.
Rust, Go, and Swift Examples
Rust:
cargo run --release -- \
--voice-style assets/voice_styles/MyCustom.json \
--text "Welcome to Supertonic" \
--lang en
Go:
go run go/example_onnx.go go/helper.go \
--voice-style assets/voice_styles/MyCustom.json \
--text "Welcome to Supertonic" \
--lang en
Swift (iOS):
let style = try Helper.loadVoiceStyle(
["../assets/voice_styles/MyCustom.json"],
verbose: true
)
let audio = try synthesizer.synthesize(
text: "Welcome to Supertonic",
style: style
)
All examples invoke the respective loadVoiceStyle routine, which returns a Style object ready for the synthesis engine.
Summary
- Voice styles are JSON tensor files containing
style_ttlandstyle_dparrays that control acoustic characteristics in the Supertonic TTS pipeline. - All language bindings share identical loading logic implemented in
helperfiles (Rust:rust/src/helper.rs, Python:py/helper.py, Swift:swift/Sources/Helper.swift, Go:go/helper.go, Web:web/helper.js). - Create custom styles by copying an existing JSON from
assets/voice_styles/, modifying the tensor values, validating the JSON, and storing the file in the assets directory. - Reference custom styles either by adding a dropdown option in
web/index.htmlfor the web UI or by passing the--voice-styleflag to CLI examples in Python, Rust, or Go.
Frequently Asked Questions
What is the exact format required for the voice style JSON files?
Voice style JSON files must contain two top-level keys: style_ttl and style_dp. Each key maps to an object with a data field containing nested float arrays representing latent tensors. The style_ttl tensor controls the timbre-level characteristics through the text-to-latent network, while style_dp controls the diffusion process in the decoder network. Both tensors must maintain matching shapes for the concatenation logic in loadVoiceStyle to function correctly.
Can I use multiple voice style files at once?
Yes. All Supertonic language bindings support loading multiple voice style files simultaneously. The loaders concatenate the data arrays from each file, enabling batch generation where each text input can utilize a different speaker style. Pass an array of file paths to loadVoiceStyle in Swift or JavaScript, or invoke the CLI with multiple style files to process them in sequence.
How do I generate completely new voice styles instead of modifying existing ones?
To generate original voice styles rather than tweaking existing JSON values, you must run the Supertonic training pipeline on a new speaker's audio recordings. The training process extracts the style_ttl and style_dp tensors from acoustic features of the target voice. Refer to the original Supertonic research paper and training scripts for the specific pipeline configuration and data requirements needed to extract these latent representations.
Why does my custom voice style fail to load in the web interface?
The web interface typically fails to load custom styles due to path resolution issues or CORS policy restrictions. Ensure the JSON file resides under the assets/voice_styles/ directory relative to the web server root, and verify that your server configuration permits fetching local JSON files. Check the browser console for fetch errors in web/helper.js, and confirm the JSON syntax is valid using a linter before testing in the UI.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →