How to Implement Real-Time Text-to-Speech with Supertonic in a Web Application
Supertonic runs real-time TTS entirely in the browser using ONNX Runtime Web with WebGPU acceleration, loading models once via loadTextToSpeech and synthesizing speech on demand through the TextToSpeech inference pipeline.
The supertone-inc/supertonic repository provides an on-device, multilingual TTS system that ships inference graphs as ONNX models. The web demo in the /web directory demonstrates how to run Supertonic directly in a browser without backend dependencies, enabling sub-second latency for real-time applications.
Architecture Overview
Supertonic’s web implementation relies on ONNX Runtime Web (onnxruntime-web) with a WebGPU-first execution strategy that automatically falls back to WebAssembly when WebGPU is unavailable. The system maintains a strict separation between model management, text preprocessing, and inference orchestration.
The four ONNX model files required for inference are stored in assets/onnx/:
duration_predictor.onnxtext_encoder.onnxvector_estimator.onnxvocoder.onnx
Voice styles are pre-extracted tensor pairs (style_ttl and style_dp) stored as JSON files in assets/voice_styles/. These represent speaker characteristics for presets like M1 through M5 and F1 through F5.
Setting Up the Environment
Loading ONNX Models
The loadTextToSpeech function in web/helper.js initializes the inference environment. It accepts a path to the model assets and configuration options for execution providers.
import { loadTextToSpeech } from './helper.js';
const { textToSpeech, cfgs } = await loadTextToSpeech('assets/onnx', {
executionProviders: ['webgpu'],
graphOptimizationLevel: 'all'
});
This function returns a TextToSpeech instance and configuration objects. The WebGPU provider delivers GPU-accelerated inference, while the graph optimization ensures minimal latency during the denoising steps.
Loading Voice Styles
Voice styles are loaded via loadVoiceStyle, also defined in web/helper.js. This function accepts an array of JSON file paths and returns a batched Style object containing tensors for the voice characteristics.
import { loadVoiceStyle } from './helper.js';
const currentStyle = await loadVoiceStyle(['assets/voice_styles/M1.json']);
You can load multiple styles simultaneously to enable voice switching without reinitializing the model sessions.
Preprocessing and Inference
Text Normalization with UnicodeProcessor
The UnicodeProcessor class in web/helper.js handles all text preprocessing before tokenization. It performs Unicode normalization, strips emojis, replaces punctuation, and wraps the input in language tags (e.g., <en>…</en>) as required by the model.
const processor = new UnicodeProcessor();
const normalizedText = processor.process("Hello, world!", "en");
// Result: "<en>Hello, world!</en>"
This step is mandatory because the TTS model expects language-tagged token streams rather than raw user input.
The Four-Stage Pipeline
The TextToSpeech class orchestrates inference through four sequential stages:
- Duration Prediction: Predicts phoneme durations using
duration_predictor.onnx - Text Encoding: Converts token IDs to latent representations via
text_encoder.onnx - Latent Diffusion: Iteratively denoises the latent representation using the vector estimator. The
totalStepparameter controls quality versus speed (fewer steps = faster inference) - Vocoding: Converts the final latent to raw waveform audio using
vocoder.onnx
The call method executes this pipeline asynchronously, accepting a progress callback to track denoising steps:
const { wav, duration } = await textToSpeech.call(
text, // normalized input
lang, // language code (e.g., "en")
currentStyle, // Style object from loadVoiceStyle
steps, // number of denoising iterations (e.g., 8)
speed, // speaking rate multiplier (e.g., 1.05)
0.3, // guidance scale
(step, total) => console.log(`Progress: ${step}/${total}`)
);
Real-Time Implementation Pattern
To achieve real-time interaction, initiate the models once during page load, then invoke synthesis on user input events with debouncing. Because the pipeline runs entirely in the browser’s UI thread via ONNX Runtime Web, you can trigger synthesis on every keystroke without blocking the interface.
let debounceTimer;
const textarea = document.getElementById('input');
const audioElement = document.getElementById('player');
textarea.addEventListener('keyup', () => {
clearTimeout(debounceTimer);
debounceTimer = setTimeout(async () => {
const txt = textarea.value;
const lang = 'en';
const speed = 1.0;
const steps = 8;
const { wav } = await textToSpeech.call(
txt, lang, currentStyle, steps, speed, 0.3,
(step, total) => console.log(`denoise ${step}/${total}`)
);
// Convert to WAV and play immediately
const wavBuf = writeWavFile(wav, textToSpeech.sampleRate);
const url = URL.createObjectURL(new Blob([wavBuf], { type: 'audio/wav' }));
audioElement.src = url;
audioElement.play();
}, 300);
});
This pattern mirrors the generateSpeech() implementation in web/main.js but isolates the synthesis logic for reuse in chat interfaces or live captioning systems.
Complete Integration Example
Here is a complete implementation combining initialization, voice loading, and real-time synthesis:
import { loadTextToSpeech, loadVoiceStyle, UnicodeProcessor, writeWavFile } from './helper.js';
class SupertonicTTS {
constructor() {
this.tts = null;
this.style = null;
this.processor = new UnicodeProcessor();
}
async initialize(modelPath = 'assets/onnx', stylePath = 'assets/voice_styles/M1.json') {
const { textToSpeech } = await loadTextToSpeech(modelPath, {
executionProviders: ['webgpu'],
graphOptimizationLevel: 'all'
});
this.tts = textToSpeech;
this.style = await loadVoiceStyle([stylePath]);
}
async speak(text, lang = 'en', steps = 8, speed = 1.0) {
if (!this.tts) throw new Error('TTS not initialized');
const normalized = this.processor.process(text, lang);
const { wav } = await this.tts.call(
normalized, lang, this.style, steps, speed, 0.3
);
const wavBuf = writeWavFile(wav, this.tts.sampleRate);
const blob = new Blob([wavBuf], { type: 'audio/wav' });
const url = URL.createObjectURL(blob);
const audio = new Audio(url);
audio.play();
return audio;
}
}
// Usage
const tts = new SupertonicTTS();
await tts.initialize();
// Real-time usage with debouncing
document.getElementById('chatInput').addEventListener('input', (e) => {
clearTimeout(window.ttsTimer);
window.ttsTimer = setTimeout(() => {
tts.speak(e.target.value, 'en', 8, 1.05);
}, 250);
});
Summary
- Model Loading: Use
loadTextToSpeechfromweb/helper.jsto initialize ONNX Runtime Web with WebGPU priority - Voice Styles: Load pre-extracted speaker tensors via
loadVoiceStylefrom JSON files inassets/voice_styles/ - Text Prep: Pass all input through
UnicodeProcessorto add language tags and normalize Unicode - Inference: Call
textToSpeech.call()with the normalized text, voice style, step count, and speed parameters - Audio Output: Convert raw waveforms to playable WAV files using
writeWavFileand stream via Blob URLs - Real-Time: Initialize once on page load, then debounce synthesis calls on user input events for sub-second response times
Frequently Asked Questions
What browsers support Supertonic’s real-time TTS?
Supertonic requires ONNX Runtime Web, which supports WebGPU in Chrome 113+, Edge 113+, and Firefox Nightly with WebGPU enabled. If WebGPU is unavailable, the library automatically falls back to WebAssembly, which works in all modern browsers but with higher latency. Check the web/README.md for specific version requirements.
How do I reduce latency for real-time applications?
Reduce the steps parameter in textToSpeech.call() from the default 20 to 8 or fewer. Each denoising step adds approximately 50-100ms of computation time. According to the source code in web/helper.js, the totalStep value directly controls the number of iterations in the _infer method’s denoising loop.
Can I switch voices without reloading the page?
Yes. Call loadVoiceStyle with a different JSON file path (e.g., assets/voice_styles/F3.json) and pass the returned Style object to subsequent textToSpeech.call() invocations. The underlying ONNX sessions remain cached, so voice switching adds only the JSON parsing overhead.
What file formats does the audio output use?
The writeWavFile function in web/helper.js generates standard 16-bit PCM WAV files with the sample rate defined by textToSpeech.sampleRate (typically 22050Hz or 24000Hz). The output is returned as a Uint8Array suitable for creating Blob URLs or uploading to servers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →