Supertonic Voice Style JSON Format: How to Create Custom Voices

Supertonic stores voice presets as flat UTF-8 JSON files containing a speaker embedding vector and runtime parameters like speaking_rate and pitch_shift, which you can generate using the Voice Builder web service and load into any Supertonic SDK.

The supertone-inc/supertonic repository implements a cross-platform text-to-speech engine that uses JSON-based voice style definitions to condition neural speech synthesis. Each voice style is a self-contained configuration file that the Python, Node.js, Swift, Rust, and other SDKs parse at runtime to reproduce specific speaker characteristics. Understanding this JSON format is essential for creating custom voices that work consistently across all Supertonic implementations.

Understanding the Voice Style JSON Schema

The voice style JSON is a strictly flat object (no nested structures) that must be UTF-8 encoded. According to the source code, every SDK parses this file using standard JSON libraries and passes the resulting dictionary to the TTS inference engine.

Core Identity Fields

These fields define the speaker's identity and model alignment:

  • voice_name: Human-readable identifier used by SDK selectors (e.g., "M1").
  • gender: "male" or "female" for UI categorization.
  • language: ISO-639-1 code (e.g., "en") indicating the primary training language.
  • speaker_id: Integer matching the speaker embedding in the ONNX model (e.g., 0).
  • embedding: Float32 array containing 256 or 512 values that constitute the actual voice style vector conditioning the latent-space flow.

Runtime Modification Fields

These parameters allow on-the-fly adjustments without recomputing the embedding:

  • speaking_rate: Speed multiplier applied after synthesis (default 1.0).
  • pitch_shift: Pitch adjustment in semitones (default 0.0).
  • audio_norm: Target RMS for output waveform normalization (default 1.0).

How to Create Custom Voices in Supertonic

Supertonic does not include a built-in voice-cloning pipeline in the open-source repository. Instead, you generate custom voice styles using the Voice Builder web service at https://supertonic.supertone.ai/voice-builder.

  1. Record a reference sample – Capture 10–30 seconds of clean, noiseless speech in any language.
  2. Upload to Voice Builder – The service extracts the speaker embedding, normalizes audio levels, and generates a schema-compliant JSON file.
  3. Download the JSON – Files are version-specific (e.g., M1.json for Supertonic 2, M1_v3.json for Supertonic 3).
  4. Place in assets – Move the file to assets/voice_styles/ in your project directory, or any location you will reference via the --voice-style CLI flag.
  5. Reference in code – Load the file path and pass the parsed object to tts.synthesize().

Loading and Modifying Voice Styles Across SDKs

Each SDK implementation follows the same pattern: parse the JSON, optionally adjust runtime fields, and pass the style object to the synthesizer.

Python Implementation

In py/helper.py, the SDK reads the JSON, validates required keys, and forwards the dictionary to the inference engine.

from supertonic import TTS
import json
import pathlib

# Load the voice style JSON

voice_path = pathlib.Path("../assets/voice_styles/M1.json")
voice_style = json.loads(voice_path.read_text())

# Adjust runtime parameters

voice_style["speaking_rate"] = 1.2  # 20% faster

voice_style["pitch_shift"] = -0.3   # Slightly lower pitch

# Initialize and synthesize

tts = TTS(auto_download=True)
wav, duration = tts.synthesize(
    text="Hello, custom voice!",
    lang="en",
    voice_style=voice_style,
    total_steps=8,
    speed=voice_style["speaking_rate"],
)
tts.save_audio(wav, "custom.wav")

Other Language SDKs

The same JSON schema works across all implementations:

  • Web: web/main.js loads assets/voice_styles/*.json and injects the dictionary into the ONNX Runtime session.
  • Swift: swift/Sources/Helper.swift parses JSON into a VoiceStyle struct for the TTS class.
  • Rust: rust/src/helper.rs implements serde::Deserialize for the schema and supports the --voice-style CLI flag.
  • Node.js: nodejs/helper.js uses fs.readFileSync and JSON.parse before passing to the model.
  • Java: java/Helper.java maps JSON to a VoiceStyle POJO using ObjectMapper.
  • C#: csharp/Helper.cs deserializes via System.Text.Json.JsonSerializer.
  • Go: go/helper.go parses into a VoiceStyle struct.
  • C++: cpp/example_onnx.cpp loads the file into nlohmann::json and extracts the embedding vector.

Summary

  • Supertonic voice styles are flat UTF-8 JSON files containing a speaker embedding vector and runtime parameters.
  • The embedding field (256 or 512 Float32 values) determines the core voice characteristics, while speaking_rate, pitch_shift, and audio_norm allow post-processing adjustments.
  • Create custom voices using the Voice Builder web service, then place the generated JSON in assets/voice_styles/.
  • All SDKs—including Python, Node.js, Swift, Rust, Java, C#, Go, and C++—parse the same schema, enabling cross-platform voice consistency.

Frequently Asked Questions

Can I manually edit the embedding values in the voice style JSON?

While you can technically modify the float array in the embedding field, manual changes are only useful if you understand the model's latent space structure. The Voice Builder service is the recommended method for generating valid embeddings from audio samples.

What is the minimum audio length required for Voice Builder?

The Voice Builder service requires 10–30 seconds of clean, noiseless speech to extract a reliable speaker embedding. Longer recordings do not necessarily improve quality beyond this threshold.

Do voice style JSON files work across different Supertonic versions?

JSON files are version-specific. A file named M1.json targets Supertonic 2, while M1_v3.json targets Supertonic 3. Always use the version-specific file that matches your SDK release.

Where should I store custom voice JSON files in my project?

Store custom voices in assets/voice_styles/ relative to your application root, or any directory path you pass to the --voice-style CLI argument. The Python, Node.js, and Rust examples all demonstrate loading from this standard location.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →