How to Load and Use Voice Styles in Supertonic TTS: A Complete Guide

Voice styles in Supertonic TTS are JSON files containing pre-computed style_ttl and style_dp tensors that you load via language-specific helpers like loadVoiceStyle in Swift or loadStyleFromJSON in JavaScript, then pass to the TextToSpeech.call method for speaker-specific synthesis.

Supertonic, the open-source text-to-speech engine by Supertone, separates acoustic modeling from speaker identity through modular voice styles. This guide shows you how to load these styles from JSON files and integrate them into your TTS pipeline using the official C++ core and language bindings.

What Is a Voice Style in Supertonic TTS?

In the Supertonic architecture, a voice style is a self-contained JSON file that stores speaker-specific embeddings as two pre-computed tensors. According to the supertone-inc/supertonic source code, each file contains:

  • style_ttl: Tensor for the text-to-latent model component with shape [1, dim1, dim2]
  • style_dp: Tensor for the duration-predictor component with shape [1, dim1, dim2]

These tensors are serialized as nested arrays of Float values. At runtime, the language bindings flatten these arrays and wrap them in ONNX Runtime ORTValue objects, which the acoustic model consumes during inference.

How to Load Voice Styles in Supertonic TTS

The repository provides language-specific helpers to handle JSON parsing, tensor flattening, and ONNX tensor creation. All implementations support batch loading of multiple styles for multi-speaker synthesis.

Swift (iOS) Implementation

In swift/Sources/Helper.swift, the loadVoiceStyle function (lines 741-806) provides the primary loading mechanism:

func loadVoiceStyle(_ voiceStylePaths: [String], verbose: Bool) throws -> Style {
    // Reads JSON files, flattens TTL & DP tensors, creates ORTValues
    // Returns Style struct wrapping the ONNX tensors
    return Style(ttl: ortValueTTL, dp: ortValueDP)
}

To load a single voice style:

let stylePath = Bundle.main.path(forResource: "M1", ofType: "json",
                                 inDirectory: "assets/voice_styles")!
let style = try loadVoiceStyle([stylePath], verbose: true)

Web (JavaScript) Implementation

For browser applications, web/main.js exposes the loadStyleFromJSON wrapper (lines 60-66) that calls the core loadVoiceStyle utility:

async function loadStyleFromJSON(stylePath) {
    const style = await loadVoiceStyle([stylePath], true);
    return style;
}

The underlying tensor flattening logic resides in web/helper.js, which converts the nested JSON arrays into flat Float32 buffers required by ONNX Runtime Web.

Python and Command-Line Usage

All language bindings support the --voice-style flag for CLI usage. As shown in test_all.sh and the Python examples, you can load styles from the assets/voice_styles/ directory:

uv run example_onnx.py \
  --voice-style assets/voice_styles/M1.json \
  --text "Hello world" \
  --lang en

For batch processing with multiple speakers, pass comma-separated paths:

uv run example_onnx.py \
  --batch \
  --voice-style assets/voice_styles/M1.json,assets/voice_styles/F1.json \
  --text "Good morning" "How are you?" \
  --lang en en

Using Voice Styles for Speech Synthesis

Once loaded, the Style object wraps the ONNX tensors and feeds directly into the synthesis pipeline via the TextToSpeech.call or TextToSpeech.batch methods.

Single Speaker Inference

In Swift, pass the style to the call method as implemented in swift/Sources/ExampleONNX.swift (lines 108-121):

let tts = try loadTextToSpeech(onnxDir, false, env)
let style = try loadVoiceStyle([stylePath], verbose: true)
let (wav, duration) = try tts.call("Welcome to Supertonic.", "en", style, 100)

In JavaScript web applications:

const style = await loadStyleFromJSON("assets/voice_styles/F2.json");
const audioBuffer = await synthesizeText("Welcome to Supertonic", "en", style);

Batch Processing with Multiple Styles

Supertonic supports loading multiple voice styles simultaneously for parallel inference. In Swift, pass an array of paths to loadVoiceStyle:

let styles = try loadVoiceStyle([
    "path/to/M1.json",
    "path/to/F1.json"
], verbose: true)
let (wavs, durations) = try tts.batch(texts, langs, styles, totalSteps)

This creates a batched Style tensor where the first dimension equals the number of voice styles, enabling the ONNX Runtime to synthesize different speakers in a single session.

Voice Style File Locations and Format

Pre-extracted voice styles ship with the repository under assets/voice_styles/. The SDK includes ten default speakers:

Native SDKs reference these via relative paths (../assets/voice_styles/), while web builds access them from the repository root or CDN deployment. Each JSON file follows the schema:

{
  "style_ttl": [[[0.01, 0.02, ...]]],
  "style_dp": [[[0.03, 0.04, ...]]]
}

The nested array structure represents the [1, dim1, dim2] tensor shape expected by the acoustic model.

Summary

  • Voice styles in Supertonic TTS are JSON files containing style_ttl and style_dp tensors that define speaker characteristics separate from the acoustic model
  • Load styles using loadVoiceStyle in Swift (swift/Sources/Helper.swift lines 741-806) or loadStyleFromJSON in JavaScript (web/main.js lines 60-66)
  • Pass the resulting Style object to TextToSpeech.call for inference; supply multiple paths for batch multi-speaker synthesis
  • Use the --voice-style flag in Python, Node.js, Go, Rust, C#, Java, and C++ CLI examples to specify JSON file paths
  • Pre-trained styles reside in assets/voice_styles/ and include M1-M5 and F1-F5 variants

Frequently Asked Questions

Where are the voice style files located in the Supertonic repository?

Pre-extracted voice styles are stored in the assets/voice_styles/ directory at the repository root. This location contains JSON files for ten default speakers (M1-M5 and F1-F5), which are referenced by native SDKs using relative paths (../assets/voice_styles/) and by web implementations from the root or deployed CDN.

Can I load multiple voice styles at once for batch processing?

Yes. The loadVoiceStyle function accepts an array of file paths and creates batched tensors where the first dimension equals the number of styles provided. This enables synthesizing speech for multiple speakers in a single inference call, as demonstrated in the CLI examples using comma-separated --voice-style arguments across Python, JavaScript, and Swift implementations.

What is the internal structure of voice style JSON files?

Each JSON file contains two top-level fields: style_ttl (text-to-latent tensor) and style_dp (duration-predictor tensor). Both store data as nested arrays of Float values with shape [1, dim1, dim2]. The language bindings flatten these arrays and wrap them in ONNX Runtime ORTValue objects before the inference session begins.

How do I switch voice styles in the Supertonic web demo?

In the browser interface defined in web/index.html, select your desired style from the dropdown menu. This triggers the loadStyleFromJSON function in web/main.js, which asynchronously loads the JSON from assets/voice_styles/, flattens the tensors using utilities in web/helper.js, and returns a Style object ready for the synthesis pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →