# Supertonic Voice Style JSON Format: How to Create Custom Voices

> Learn the Supertonic voice style JSON format for custom voices. Discover how to create voice presets with speaker embeddings and runtime parameters using the Voice Builder.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Supertonic stores voice presets as flat UTF-8 JSON files containing a speaker embedding vector and runtime parameters like `speaking_rate` and `pitch_shift`, which you can generate using the Voice Builder web service and load into any Supertonic SDK.**

The supertone-inc/supertonic repository implements a cross-platform text-to-speech engine that uses JSON-based voice style definitions to condition neural speech synthesis. Each voice style is a self-contained configuration file that the Python, Node.js, Swift, Rust, and other SDKs parse at runtime to reproduce specific speaker characteristics. Understanding this JSON format is essential for creating custom voices that work consistently across all Supertonic implementations.

## Understanding the Voice Style JSON Schema

The voice style JSON is a strictly flat object (no nested structures) that must be UTF-8 encoded. According to the source code, every SDK parses this file using standard JSON libraries and passes the resulting dictionary to the TTS inference engine.

### Core Identity Fields

These fields define the speaker's identity and model alignment:

- **voice_name**: Human-readable identifier used by SDK selectors (e.g., `"M1"`).
- **gender**: `"male"` or `"female"` for UI categorization.
- **language**: ISO-639-1 code (e.g., `"en"`) indicating the primary training language.
- **speaker_id**: Integer matching the speaker embedding in the ONNX model (e.g., `0`).
- **embedding**: Float32 array containing 256 or 512 values that constitute the actual voice style vector conditioning the latent-space flow.

### Runtime Modification Fields

These parameters allow on-the-fly adjustments without recomputing the embedding:

- **speaking_rate**: Speed multiplier applied after synthesis (default `1.0`).
- **pitch_shift**: Pitch adjustment in semitones (default `0.0`).
- **audio_norm**: Target RMS for output waveform normalization (default `1.0`).

## How to Create Custom Voices in Supertonic

Supertonic does not include a built-in voice-cloning pipeline in the open-source repository. Instead, you generate custom voice styles using the **Voice Builder** web service at `https://supertonic.supertone.ai/voice-builder`.

1. **Record a reference sample** – Capture 10–30 seconds of clean, noiseless speech in any language.
2. **Upload to Voice Builder** – The service extracts the speaker embedding, normalizes audio levels, and generates a schema-compliant JSON file.
3. **Download the JSON** – Files are version-specific (e.g., [`M1.json`](https://github.com/supertone-inc/supertonic/blob/main/M1.json) for Supertonic 2, [`M1_v3.json`](https://github.com/supertone-inc/supertonic/blob/main/M1_v3.json) for Supertonic 3).
4. **Place in assets** – Move the file to `assets/voice_styles/` in your project directory, or any location you will reference via the `--voice-style` CLI flag.
5. **Reference in code** – Load the file path and pass the parsed object to `tts.synthesize()`.

## Loading and Modifying Voice Styles Across SDKs

Each SDK implementation follows the same pattern: parse the JSON, optionally adjust runtime fields, and pass the style object to the synthesizer.

### Python Implementation

In [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), the SDK reads the JSON, validates required keys, and forwards the dictionary to the inference engine.

```python
from supertonic import TTS
import json
import pathlib

# Load the voice style JSON

voice_path = pathlib.Path("../assets/voice_styles/M1.json")
voice_style = json.loads(voice_path.read_text())

# Adjust runtime parameters

voice_style["speaking_rate"] = 1.2  # 20% faster

voice_style["pitch_shift"] = -0.3   # Slightly lower pitch

# Initialize and synthesize

tts = TTS(auto_download=True)
wav, duration = tts.synthesize(
    text="Hello, custom voice!",
    lang="en",
    voice_style=voice_style,
    total_steps=8,
    speed=voice_style["speaking_rate"],
)
tts.save_audio(wav, "custom.wav")

```

### Other Language SDKs

The same JSON schema works across all implementations:

- **Web**: [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js) loads `assets/voice_styles/*.json` and injects the dictionary into the ONNX Runtime session.
- **Swift**: [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) parses JSON into a `VoiceStyle` struct for the TTS class.
- **Rust**: [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) implements `serde::Deserialize` for the schema and supports the `--voice-style` CLI flag.
- **Node.js**: [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js) uses `fs.readFileSync` and `JSON.parse` before passing to the model.
- **Java**: [`java/Helper.java`](https://github.com/supertone-inc/supertonic/blob/main/java/Helper.java) maps JSON to a `VoiceStyle` POJO using `ObjectMapper`.
- **C#**: [`csharp/Helper.cs`](https://github.com/supertone-inc/supertonic/blob/main/csharp/Helper.cs) deserializes via `System.Text.Json.JsonSerializer`.
- **Go**: [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) parses into a `VoiceStyle` struct.
- **C++**: [`cpp/example_onnx.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/example_onnx.cpp) loads the file into `nlohmann::json` and extracts the embedding vector.

## Summary

- Supertonic voice styles are flat UTF-8 JSON files containing a speaker embedding vector and runtime parameters.
- The `embedding` field (256 or 512 Float32 values) determines the core voice characteristics, while `speaking_rate`, `pitch_shift`, and `audio_norm` allow post-processing adjustments.
- Create custom voices using the Voice Builder web service, then place the generated JSON in `assets/voice_styles/`.
- All SDKs—including Python, Node.js, Swift, Rust, Java, C#, Go, and C++—parse the same schema, enabling cross-platform voice consistency.

## Frequently Asked Questions

### Can I manually edit the embedding values in the voice style JSON?

While you can technically modify the float array in the `embedding` field, manual changes are only useful if you understand the model's latent space structure. The Voice Builder service is the recommended method for generating valid embeddings from audio samples.

### What is the minimum audio length required for Voice Builder?

The Voice Builder service requires 10–30 seconds of clean, noiseless speech to extract a reliable speaker embedding. Longer recordings do not necessarily improve quality beyond this threshold.

### Do voice style JSON files work across different Supertonic versions?

JSON files are version-specific. A file named [`M1.json`](https://github.com/supertone-inc/supertonic/blob/main/M1.json) targets Supertonic 2, while [`M1_v3.json`](https://github.com/supertone-inc/supertonic/blob/main/M1_v3.json) targets Supertonic 3. Always use the version-specific file that matches your SDK release.

### Where should I store custom voice JSON files in my project?

Store custom voices in `assets/voice_styles/` relative to your application root, or any directory path you pass to the `--voice-style` CLI argument. The Python, Node.js, and Rust examples all demonstrate loading from this standard location.