How to Integrate Custom Voice Builder JSON Files for Voice Cloning with Supertonic

Supertonic's TTS engine consumes Voice Builder JSON files containing style_ttl and style_dp parameters directly through the load_voice_style helper or --voice-style CLI flags.

The Supertonic repository (supertone-inc/supertonic) provides a cross-platform text-to-speech engine that supports custom voice cloning via JSON configuration files. These files are generated by the Voice Builder service and contain the neural style tensors required to replicate a specific voice. Integrating these custom voices requires placing the JSON files in the correct asset directories and loading them through standardized helper functions available across Python, Rust, Node.js, Java, Swift, Flutter, and Web implementations.

Generate Your Custom Voice Builder JSON

Before integration, you must create the voice style file using the Voice Builder interface. Navigate to https://supertonic.supertone.ai/voice-builder and record a short reference clip. After processing, the service returns a voice-style JSON file containing the style_ttl and style_dp parameters that define the acoustic characteristics of your cloned voice.

Download the resulting file (e.g., my_custom_voice.json) to your local machine. This file is portable and can be used across all Supertonic language bindings without modification.

Place the JSON File in the Assets Directory

Supertonic expects voice style files to reside in specific asset directories depending on your deployment target.

For command-line tools and local development:

  • Place the JSON file under assets/voice_styles/ in the repository root.

For mobile and cross-platform deployments:

  • iOS/Swift: Copy to ios/ExampleiOSApp/voice_styles/
  • Android/Java: Copy to android/app/src/main/assets/voice_styles/
  • Flutter: Copy to flutter/assets/voice_styles/

The consistent directory naming convention ensures that the helper libraries can locate and load the tensors regardless of the underlying platform.

Load the Voice Style via CLI or Programmatically

Supertonic exposes two primary methods for loading custom voice styles: CLI flags for rapid testing and programmatic loading for application integration.

Command-Line Interface (CLI)

All language bindings support the --voice-style (or -voice-style) flag, which accepts one or more comma-separated JSON paths.

Python:

python -m supertonic.example_onnx \
    --text "Hello, this is my custom voice." \
    --voice-style assets/voice_styles/my_custom_voice.json

Rust:

cargo run --example example_onnx \
    -- --text "Hello, this is my custom voice." \
    --voice-style ../assets/voice_styles/my_custom_voice.json

Node.js:

node nodejs/example_onnx.js \
    --text "Hello, this is my custom voice." \
    --voice-style ../assets/voice_styles/my_custom_voice.json

Java:

mvn exec:java -Dexec.args="--text \"Hello, this is my custom voice.\" \
    --voice-style ../assets/voice_styles/my_custom_voice.json"

Library Integration (Python)

For production applications, use the load_voice_style function implemented in py/helper.py. This function pre-allocates tensors and returns a Style object ready for inference.

from py.helper import load_voice_style, load_text_to_speech

# Load the custom style (batch size = 1)

style = load_voice_style(["assets/voice_styles/my_custom_voice.json"], verbose=True)

# Initialise the TTS engine

tts = load_text_to_speech(onnx_dir="assets/onnx")

# The style object is now ready for synthesis

Other Language Bindings

The helper pattern is consistent across all supported languages:

Rust (rust/src/helper.rs):

let style = load_voice_style(&["../assets/voice_styles/my_custom_voice.json".to_string()]);

Swift (swift/ExampleONNX.swift):

let style = try loadVoiceStyle(paths: ["../assets/voice_styles/my_custom_voice.json"])

Flutter (flutter/lib/main.dart):

final style = await loadVoiceStyle(['assets/voice_styles/my_custom_voice.json']);

Web (web/main.js):

import { loadVoiceStyle } from './helper.js';
const style = await loadVoiceStyle(['assets/voice_styles/my_custom_voice.json']);

Synthesize Speech with the Custom Voice

Once loaded, pass the Style object (or JSON path) to the synthesize method. The inference pipeline automatically substitutes the built-in voice tensors with your custom parameters.

Python example:

wav, _ = tts.synthesize("Hello, this is my custom voice.", voice_style=style)

The synthesize call returns raw audio data (typically WAV format) that reflects the acoustic characteristics defined in your Voice Builder JSON file. The default voice in web implementations loads from assets/voice_styles/M1.json, but any custom path overrides this behavior.

Key Implementation Files

The following source files handle voice style loading across the codebase:

  • py/helper.py: Python implementation of load_voice_style that parses JSON and builds the Style tensor batch.
  • rust/src/helper.rs: Rust implementation mirroring the Python loader; powers the Rust CLI and core library.
  • go/example_onnx.go: Demonstrates -voice-style flag handling in Go applications.
  • nodejs/example_onnx.js: Shows JSON path passing in JavaScript/Node.js environments.
  • java/ExampleONNX.java: Implements --voice-style CLI argument parsing for Java.
  • swift/ExampleONNX.swift: Handles bundle-based JSON loading for iOS applications.
  • flutter/lib/main.dart: Dart implementation for mobile cross-platform deployment.
  • web/main.js: Default voice style handling for browser-based demonstrations.

Summary

  • Voice Builder JSON files contain style_ttl and style_dp parameters that define custom voice characteristics.
  • Place downloaded JSON files in assets/voice_styles/ (or platform-specific asset bundles for mobile).
  • Use the --voice-style CLI flag for command-line testing across all languages.
  • Use load_voice_style (implemented in py/helper.py and rust/src/helper.rs) for programmatic loading in applications.
  • Pass the resulting Style object to synthesize to generate audio with your cloned voice.

Frequently Asked Questions

What parameters are contained in the Voice Builder JSON file?

The JSON file contains style_ttl and style_dp parameters. These are neural network tensors that encode the prosodic and acoustic features of the recorded voice, allowing the TTS engine to replicate the specific vocal characteristics during inference.

Can I load multiple custom voice styles simultaneously?

Yes. The load_voice_style function and --voice-style CLI flag both accept arrays or comma-separated paths of JSON files. This enables batch processing where you can switch between multiple cloned voices within the same application session by passing different indices to the synthesis call.

Where should I place the JSON file for mobile deployments?

For iOS, place the file in ios/ExampleiOSApp/voice_styles/. For Android, use android/app/src/main/assets/voice_styles/. For Flutter projects, use flutter/assets/voice_styles/. Ensure the files are included in your platform's asset bundle configuration so they are available at runtime.

Is the Voice Builder JSON format compatible across all Supertonic language bindings?

Yes. The JSON format is standardized across Python, Rust, Node.js, Java, Swift, Flutter, and Web implementations. A single JSON file generated by the Voice Builder service works identically across all platforms without conversion or modification, as implemented in the respective helper files (py/helper.py, rust/src/helper.rs, etc.).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →