# How to Create Custom Voice Styles for Supertonic: A Step-by-Step Guide

> Learn how to create custom voice styles for Supertonic with this step by step guide. Author a JSON file place it in assets voice_styles and reference it via the web UI or CLI.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**To create custom voice styles for Supertonic, author a JSON file containing `style_ttl` and `style_dp` tensor arrays, place it in `assets/voice_styles/`, and reference the file path via the web UI dropdown or the `--voice-style` CLI flag.**

Supertonic is an open-source text-to-speech library that generates speech by applying **voice-style JSON files** describing acoustic characteristics like timing, pitch, and intensity. While the repository ships with pre-extracted styles (M1-M5, F1-F5) located in `assets/voice_styles/`, you can extend the system with custom speaker profiles. This guide explains the loader architecture, the required JSON schema, and implementation steps across Python, Rust, Swift, Go, and JavaScript bindings.

## Understanding the Voice Style Architecture

Supertonic loads voice styles through a consistent architecture shared across all language bindings. The system deserializes JSON files containing latent tensors that control the timbre and diffusion process of the neural TTS model.

### The JSON Schema for Voice Styles

Every voice style file must follow this structure:

```json
{
  "style_ttl": {
    "data": [
      [ ... ]
    ]
  },
  "style_dp": {
    "data": [
      [ ... ]
    ]
  }
}

```

The `style_ttl` field contains timbre-level latent vectors for the text-to-latent network, while `style_dp` contains diffusion-process latent vectors for the decoder network. Both contain 2-D or 3-D float tensors that the loader flattens and concatenates when processing multiple style files.

### How Loaders Work Across Language Bindings

Each language binding implements a `loadVoiceStyle` function that performs identical operations:

- **[`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js)**: The `loadVoiceStyle()` function fetches JSON files, parses them, and flattens the `style_ttl` and `style_dp` tensors into a single `voiceStyle` object.
- **[`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)**: The `load_voice_style()` function opens files for validation, deserializes each into `VoiceStyleData`, and concatenates the tensors.
- **[`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)**: The `loadVoiceStyle()` method reads files into `Data`, decodes them with `JSONDecoder`, and merges the arrays.
- **[`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go)**: The `LoadVoiceStyle()` function unmarshals JSON into `VoiceStyleData` and appends the tensors.
- **[`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)**: The Python implementation mirrors this logic for the Python binding.

All loaders support batch generation by concatenating the `data` arrays of multiple supplied files.

## Step-by-Step Guide to Creating Custom Voice Styles

### Step 1: Copy and Modify an Existing Style

Start by copying an existing style from [`assets/voice_styles/M1.json`](https://github.com/supertone-inc/supertonic/blob/main/assets/voice_styles/M1.json) to a new file such as [`MyCustom.json`](https://github.com/supertone-inc/supertonic/blob/main/MyCustom.json). Edit the floating-point values in the `style_ttl.data` and `style_dp.data` arrays to sculpt the acoustic characteristics.

For production-quality styles, generate new tensors by running the Supertonic training script on a target speaker's recordings. For quick experimentation, manually adjusting values in an existing JSON file creates variation in speaking style and timbre.

### Step 2: Validate and Place the JSON File

Validate your JSON syntax using a tool like `jq` or an online validator to prevent deserialization errors. Store the validated file under `assets/voice_styles/` relative to the runtime root. The web UI expects this path structure, while CLI tools accept absolute or relative paths via the `--voice-style` flag.

### Step 3: Reference Your Custom Style

**For the Web UI**, add an option to the dropdown in [`web/index.html`](https://github.com/supertone-inc/supertonic/blob/main/web/index.html):

```html
<select id="voiceStyleSelect">
  <option value="assets/voice_styles/MyCustom.json">My Custom (MC)</option>
</select>

```

The selection logic in [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js) automatically calls `loadVoiceStyle()` when the dropdown changes.

**For CLI usage**, pass the file path using the `--voice-style` argument:

```bash

# Python example

uv run py/example_onnx.py \
    --voice-style assets/voice_styles/MyCustom.json \
    --text "Hello world" \
    --lang en

```

## Code Examples by Language

### Python CLI Implementation

Install dependencies and run synthesis with your custom style:

```bash
uv pip install -r py/requirements.txt
uv run py/example_onnx.py \
    --voice-style assets/voice_styles/MyCustom.json \
    --text "Welcome to Supertonic" \
    --lang en

```

The [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) file handles parsing the JSON and building the `Style` object for the ONNX runtime.

### Web UI Integration

In the browser environment, the loader fetches the JSON asynchronously:

```javascript
// The helper.js loadVoiceStyle function handles the fetch
const style = await loadVoiceStyle(['assets/voice_styles/MyCustom.json']);

```

Ensure your web server serves the `assets/` directory so the fetch request in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) resolves correctly.

### Rust, Go, and Swift Examples

**Rust:**

```bash
cargo run --release -- \
    --voice-style assets/voice_styles/MyCustom.json \
    --text "Welcome to Supertonic" \
    --lang en

```

**Go:**

```bash
go run go/example_onnx.go go/helper.go \
    --voice-style assets/voice_styles/MyCustom.json \
    --text "Welcome to Supertonic" \
    --lang en

```

**Swift (iOS):**

```swift
let style = try Helper.loadVoiceStyle(
    ["../assets/voice_styles/MyCustom.json"], 
    verbose: true
)
let audio = try synthesizer.synthesize(
    text: "Welcome to Supertonic", 
    style: style
)

```

All examples invoke the respective `loadVoiceStyle` routine, which returns a `Style` object ready for the synthesis engine.

## Summary

- **Voice styles are JSON tensor files** containing `style_ttl` and `style_dp` arrays that control acoustic characteristics in the Supertonic TTS pipeline.
- **All language bindings share identical loading logic** implemented in `helper` files (Rust: [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs), Python: [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), Swift: [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift), Go: [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go), Web: [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js)).
- **Create custom styles** by copying an existing JSON from `assets/voice_styles/`, modifying the tensor values, validating the JSON, and storing the file in the assets directory.
- **Reference custom styles** either by adding a dropdown option in [`web/index.html`](https://github.com/supertone-inc/supertonic/blob/main/web/index.html) for the web UI or by passing the `--voice-style` flag to CLI examples in Python, Rust, or Go.

## Frequently Asked Questions

### What is the exact format required for the voice style JSON files?

Voice style JSON files must contain two top-level keys: `style_ttl` and `style_dp`. Each key maps to an object with a `data` field containing nested float arrays representing latent tensors. The `style_ttl` tensor controls the timbre-level characteristics through the text-to-latent network, while `style_dp` controls the diffusion process in the decoder network. Both tensors must maintain matching shapes for the concatenation logic in `loadVoiceStyle` to function correctly.

### Can I use multiple voice style files at once?

Yes. All Supertonic language bindings support loading multiple voice style files simultaneously. The loaders concatenate the `data` arrays from each file, enabling batch generation where each text input can utilize a different speaker style. Pass an array of file paths to `loadVoiceStyle` in Swift or JavaScript, or invoke the CLI with multiple style files to process them in sequence.

### How do I generate completely new voice styles instead of modifying existing ones?

To generate original voice styles rather than tweaking existing JSON values, you must run the Supertonic training pipeline on a new speaker's audio recordings. The training process extracts the `style_ttl` and `style_dp` tensors from acoustic features of the target voice. Refer to the original Supertonic research paper and training scripts for the specific pipeline configuration and data requirements needed to extract these latent representations.

### Why does my custom voice style fail to load in the web interface?

The web interface typically fails to load custom styles due to path resolution issues or CORS policy restrictions. Ensure the JSON file resides under the `assets/voice_styles/` directory relative to the web server root, and verify that your server configuration permits fetching local JSON files. Check the browser console for fetch errors in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js), and confirm the JSON syntax is valid using a linter before testing in the UI.