How to integrate Supertonic into a web application
Supertonic provides a browser-based text-to-speech pipeline using ONNX Runtime Web that enables client-side synthesis through the loadTextToSpeech and TextToSpeech.call methods defined in web/helper.js.
Supertonic is an open-source text-to-speech (TTS) engine developed by Supertone Inc. that runs entirely in the browser using ONNX Runtime Web. Learning how to integrate Supertonic into a web application allows you to generate speech with zero network latency while keeping user data private, as implemented in the supertone-inc/supertonic repository.
Prerequisites and Model Assets
Before initializing the pipeline, download the ONNX model files from the Hugging Face hub and place them in web/assets/onnx/. The required files include duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx, along with the tts.json configuration file.
You also need a voice style JSON file (e.g., web/assets/voice_styles/M1.json) which defines the speaker characteristics. You can create additional styles using the Voice Builder tool and export them to this directory.
Initializing the TTS Pipeline
The loadTextToSpeech function in web/helper.js handles model initialization and runtime setup. It attempts to create a WebGPU session first for GPU acceleration, automatically falling back to WebAssembly (WASM) if WebGPU is unavailable.
import { loadTextToSpeech } from './helper.js';
const { textToSpeech, cfgs } = await loadTextToSpeech(
'assets/onnx',
{ executionProviders: ['webgpu'] }
);
This function returns a TextToSpeech instance and configuration objects. The executionProviders option accepts webgpu as the primary target, with the helper managing the WASM fallback internally.
Loading Voice Styles
Voice styles control the speaker characteristics and are loaded using loadVoiceStyle from web/helper.js. This function processes the JSON style file and converts the TTL and DP tensors into ONNX ort.Tensor objects required by the inference engine.
import { loadVoiceStyle } from './helper.js';
const style = await loadVoiceStyle(
['assets/voice_styles/M1.json'],
true
);
Passing true as the second argument enables style processing. The function returns a Style instance that you pass to the synthesis method.
Generating Speech
The TextToSpeech class orchestrates the full synthesis pipeline including text preprocessing, duration prediction, latent noise sampling, denoising loops, and vocoder inference. Call the call method with your text, language code, style object, and synthesis parameters.
const { wav, duration } = await textToSpeech.call(
text,
lang,
style,
totalStep, // Quality steps: 5 (low) to 12 (high)
speed, // Speech rate multiplier (1.0 = normal)
silenceDuration // Optional silence padding
);
The method returns a Float32Array containing the raw waveform and a duration array. The totalStep parameter controls quality versus speed—values between 5 and 12 are typical, with higher values producing better quality but requiring more computation.
Converting to Playable Audio
The raw waveform requires conversion to a standard audio format. The writeWavFile utility in web/helper.js encodes the Float32Array into a WAV ArrayBuffer.
import { writeWavFile } from './helper.js';
const wavBuffer = writeWavFile(wav, textToSpeech.sampleRate);
const blob = new Blob([wavBuffer], { type: 'audio/wav' });
const url = URL.createObjectURL(blob);
// Attach to audio element
audioElement.src = url;
This creates a downloadable Blob URL that can be attached to an <audio> element for playback or saved to disk.
Complete Integration Example
The following example combines all components into a minimal HTML page based on the reference implementation in web/main.js and web/index.html.
<!-- index.html -->
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Supertonic Web TTS</title>
<script type="module" src="main.js"></script>
</head>
<body>
<h1>Supertonic Web TTS</h1>
<textarea id="text" rows="4" placeholder="Enter text…"></textarea><br>
<label>Language:
<select id="langSelect">
<option value="en">English</option>
<option value="ko">Korean</option>
</select>
</label><br>
<button id="generateBtn" disabled>Generate</button>
<div id="statusBox"><span id="statusText"></span></div>
<div id="results"></div>
</body>
</html>
// main.js
import {
loadTextToSpeech,
loadVoiceStyle,
writeWavFile
} from './helper.js';
let tts = null;
let style = null;
async function init() {
// Load models with WebGPU → WASM fallback
const { textToSpeech } = await loadTextToSpeech(
'assets/onnx',
{ executionProviders: ['webgpu'] }
);
tts = textToSpeech;
// Load voice style
style = await loadVoiceStyle(['assets/voice_styles/M1.json'], true);
document.getElementById('generateBtn').disabled = false;
document.getElementById('statusText').textContent = 'Ready';
}
init();
document.getElementById('generateBtn').addEventListener('click', async () => {
const text = document.getElementById('text').value.trim();
const lang = document.getElementById('langSelect').value;
if (!text) return alert('Please enter text');
// Synthesize with quality step 8
const { wav, duration } = await tts.call(
text,
lang,
style,
8, // totalStep
1.0 // speed
);
// Create WAV blob
const wavBuffer = writeWavFile(wav, tts.sampleRate);
const blob = new Blob([wavBuffer], { type: 'audio/wav' });
const url = URL.createObjectURL(blob);
// Display results
const resultsDiv = document.getElementById('results');
resultsDiv.innerHTML = `
<audio controls src="${url}"></audio>
<p>Duration: ${duration[0].toFixed(2)} s</p>
<a href="${url}" download="speech.wav">Download</a>
`;
});
Summary
- Place ONNX assets in
web/assets/onnx/including the four model files (duration_predictor.onnx,text_encoder.onnx,vector_estimator.onnx,vocoder.onnx) andtts.json. - Initialize the pipeline using
loadTextToSpeechfromweb/helper.js, which automatically handles WebGPU and WASM fallback. - Load voice styles with
loadVoiceStyleto createStyleinstances from JSON files inweb/assets/voice_styles/. - Synthesize speech by calling
textToSpeech.callwith text, language, style, and quality parameters. - Render audio using
writeWavFileto convert the Float32 waveform into a downloadable WAV Blob.
Frequently Asked Questions
Does Supertonic require a server backend for web integration?
No. Supertonic runs entirely client-side using ONNX Runtime Web according to the source code in supertone-inc/supertonic. All model inference happens in the browser using WebGPU or WebAssembly, eliminating server costs and keeping user data private.
What hardware acceleration does Supertonic support in browsers?
Supertonic attempts to use WebGPU first for GPU acceleration, falling back automatically to WebAssembly (WASM) if the browser or device lacks WebGPU support. This behavior is handled internally by the loadTextToSpeech function in web/helper.js.
How do I customize the voice style in Supertonic?
Create custom voice styles using the Voice Builder tool and export them as JSON files. Place these files in web/assets/voice_styles/ and load them using loadVoiceStyle(['path/to/style.json'], true). The boolean flag enables style tensor processing required by the TTS engine.
What is the difference between totalStep values in the synthesis call?
The totalStep parameter controls the number of denoising steps during inference. Values range from 5 (fast, lower quality) to 12 (slow, higher quality). The default examples typically use 8 steps as a balance between synthesis speed and audio quality.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →