How to Load and Switch Between Voice Style Presets (M1‑M5, F1‑F5) in Supertonic
Supertonic provides ten pre‑extracted voice‑style presets (M1‑M5 and F1‑F5) stored as JSON files in assets/voice_styles/, which you can load using loadVoiceStyle (JavaScript) or load_voice_style (Python) and switch between at runtime by updating the file path or batch index.
The Supertonic text‑to‑speech engine ships with five male and five female voice style presets that capture distinct speaker characteristics. Each preset is a JSON file containing latent tensors that steer the acoustic model and duration predictor. Whether you are building a web interface or a Python application, you can load these voice style presets and switch between them programmatically to change the speaker voice instantly without reloading the underlying TTS model.
Understanding Voice Style Presets
Supertonic’s voice style presets are pre‑extracted latent representations stored in assets/voice_styles/. Each JSON file contains two tensors:
style_ttl– The tone‑level latent that steers the acoustic model timbre.style_dp– The duration‑predictor latent that influences speech timing and rhythm.
These files are named M1.json through M5.json (male voices) and F1.json through F5.json (female voices). When loaded, the tensors are flattened, packed into batch‑wise arrays of shape [batch, dim1, dim2], and wrapped in a Style object that the TextToSpeech pipeline consumes during inference.
How to Load Voice Style Presets
The loading pipeline works identically across language bindings. The loader reads the JSON files, allocates batch tensors, and copies the flattened data into ONNX Runtime tensors (ort.Tensor in JavaScript, np.ndarray in Python).
Loading a Single Preset in Python
To load a single voice style preset, pass a one‑element list containing the JSON path to load_voice_style in py/helper.py:
from supertonic.py.helper import load_voice_style, TextToSpeech
# Load Female 3 preset
preset_path = "assets/voice_styles/F3.json"
style = load_voice_style([preset_path])
# Initialize TTS and synthesize
tts = TextToSpeech.load_text_to_speech("assets/onnx", use_gpu=False)
wav, duration = tts("Hello world!", "en", style, total_step=8)
Loading a Single Preset in JavaScript
In the web implementation, use loadVoiceStyle from web/helper.js:
import { loadVoiceStyle } from './helper.js';
const style = await loadVoiceStyle(['assets/voice_styles/M3.json']);
Loading Multiple Presets in a Batch
You can load multiple presets simultaneously to enable instant switching without file I/O delays. Pass a list of paths to create a Style object with a batch dimension equal to the list length:
preset_paths = [
"assets/voice_styles/M1.json",
"assets/voice_styles/M2.json",
"assets/voice_styles/M3.json",
"assets/voice_styles/M4.json",
"assets/voice_styles/M5.json",
]
# Returns a Style with batch dimension = 5
style_batch = load_voice_style(preset_paths)
When invoking the model, slice the specific preset by indexing the batch dimension:
# Use the third male preset (index 2)
wav, dur = tts("Testing batch loading", "en", style_batch.ttl[2:3], total_step=8)
Switching Between Presets at Runtime
Switching presets is accomplished either by selecting a different index from an already‑loaded batch or by reloading the style JSON with a new file path.
Web Interface Implementation
The web demo in web/index.html demonstrates real‑time switching using a <select> element. The dropdown lists all ten presets and fires a change event that reloads the style via loadVoiceStyle:
<label for="voiceStyleSelect">Voice Style:</label>
<select id="voiceStyleSelect">
<option value="assets/voice_styles/M1.json">Male 1 (M1)</option>
<option value="assets/voice_styles/M2.json">Male 2 (M2)</option>
<option value="assets/voice_styles/M3.json">Male 3 (M3)</option>
<option value="assets/voice_styles/M4.json">Male 4 (M4)</option>
<option value="assets/voice_styles/M5.json">Male 5 (M5)</option>
<option value="assets/voice_styles/F1.json">Female 1 (F1)</option>
<option value="assets/voice_styles/F2.json">Female 2 (F2)</option>
<option value="assets/voice_styles/F3.json">Female 3 (F3)</option>
<option value="assets/voice_styles/F4.json">Female 4 (F4)</option>
<option value="assets/voice_styles/F5.json">Female 5 (F5)</option>
</select>
The event handler in web/main.js updates the current style path and reloads:
import { loadVoiceStyle } from './helper.js';
const voiceStyleSelect = document.getElementById('voiceStyleSelect');
let currentStyle;
// Load default on startup
(async () => {
currentStyle = await loadVoiceStyle(['assets/voice_styles/M1.json']);
})();
// Switch on user selection
voiceStyleSelect.addEventListener('change', async (e) => {
const newPath = e.target.value;
currentStyle = await loadVoiceStyle([newPath]);
});
Python Programmatic Switching
In Python, if you pre‑loaded multiple presets into a batch, switch between them by slicing the batch dimension without reloading files:
# Switch to Female 4 (index 3 in a batch loaded with F1-F5)
current_style = style_batch.ttl[3:4], style_batch.dp[3:4]
wav, dur = tts("Switched voice", "en", current_style, total_step=8)
Alternatively, load a new preset entirely by calling load_voice_style with a different path.
Node‑JS Implementation
The Node‑JS binding in nodejs/helper.js mirrors the web implementation:
import { loadVoiceStyle, loadTextToSpeech } from './helper.js';
(async () => {
const style = await loadVoiceStyle(['assets/voice_styles/F4.json']);
const tts = await loadTextToSpeech('assets/onnx');
const { wav, duration } = await tts.synthesize("Node example", "en", style, 8);
})();
Summary
- Voice style presets are JSON files (
M1.json–M5.json,F1.json–F5.json) stored inassets/voice_styles/containingstyle_ttlandstyle_dptensors. - Load presets using
load_voice_style(Python) orloadVoiceStyle(JavaScript) from the respectivehelpermodules. - Batch loading allows you to load all ten presets at once into a single
Styleobject and switch by indexing[n:n+1]. - Runtime switching in web apps uses a
<select>element to reload JSON paths; in Python, slice pre‑loaded batches or reload specific files. - The
Styleobject wraps ONNX Runtime tensors and is passed directly to theTextToSpeechpipeline for inference.
Frequently Asked Questions
What is the difference between the M1‑M5 and F1‑F5 voice style presets?
The M1‑M5 presets represent five distinct male speaker characteristics, while F1‑F5 represent five distinct female speaker characteristics. According to the Supertonic source code, each preset contains unique latent tensors (style_ttl and style_dp) that shape the timbre, pitch, and speaking rhythm of the generated speech.
Can I load multiple voice style presets simultaneously?
Yes. Pass a list of JSON file paths to load_voice_style (Python) or loadVoiceStyle (JavaScript). The function inspects the first preset to obtain tensor dimensionality, allocates buffers of shape [batch, dim1, dim2], and returns a Style object with a batch dimension equal to the number of paths provided.
How do I switch presets without reloading the entire TTS model?
Pre‑load all desired presets into a single batch Style object. Then, instead of reloading JSON files, slice the batch dimension (e.g., style_batch.ttl[2:3]) to select the specific preset index. This keeps the ONNX session in memory and only swaps the style tensors, enabling instant voice switching.
Where are the preset files located in the Supertonic repository?
The preset JSON files are located in the assets/voice_styles/ directory at the repository root. The web interface references them at assets/voice_styles/M1.json through F5.json, while Python examples use the relative path from the script location to assets/voice_styles/.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →