# How to Implement Real-Time Text-to-Speech with Supertonic in a Web Application

> Implement real-time TTS in your web app using Supertonic. Leverage ONNX Runtime Web and WebGPU for fast, in-browser speech synthesis. Learn how to load models and synthesize speech efficiently.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Supertonic runs real-time TTS entirely in the browser using ONNX Runtime Web with WebGPU acceleration, loading models once via `loadTextToSpeech` and synthesizing speech on demand through the `TextToSpeech` inference pipeline.**

The `supertone-inc/supertonic` repository provides an on-device, multilingual TTS system that ships inference graphs as ONNX models. The web demo in the `/web` directory demonstrates how to run Supertonic directly in a browser without backend dependencies, enabling sub-second latency for real-time applications.

## Architecture Overview

Supertonic’s web implementation relies on **ONNX Runtime Web** (`onnxruntime-web`) with a WebGPU-first execution strategy that automatically falls back to WebAssembly when WebGPU is unavailable. The system maintains a strict separation between model management, text preprocessing, and inference orchestration.

The four ONNX model files required for inference are stored in `assets/onnx/`:

- `duration_predictor.onnx`
- `text_encoder.onnx`
- `vector_estimator.onnx`
- `vocoder.onnx`

Voice styles are pre-extracted tensor pairs (`style_ttl` and `style_dp`) stored as JSON files in `assets/voice_styles/`. These represent speaker characteristics for presets like M1 through M5 and F1 through F5.

## Setting Up the Environment

### Loading ONNX Models

The `loadTextToSpeech` function in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) initializes the inference environment. It accepts a path to the model assets and configuration options for execution providers.

```javascript
import { loadTextToSpeech } from './helper.js';

const { textToSpeech, cfgs } = await loadTextToSpeech('assets/onnx', {
  executionProviders: ['webgpu'],
  graphOptimizationLevel: 'all'
});

```

This function returns a `TextToSpeech` instance and configuration objects. The WebGPU provider delivers GPU-accelerated inference, while the graph optimization ensures minimal latency during the denoising steps.

### Loading Voice Styles

Voice styles are loaded via `loadVoiceStyle`, also defined in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js). This function accepts an array of JSON file paths and returns a batched `Style` object containing tensors for the voice characteristics.

```javascript
import { loadVoiceStyle } from './helper.js';

const currentStyle = await loadVoiceStyle(['assets/voice_styles/M1.json']);

```

You can load multiple styles simultaneously to enable voice switching without reinitializing the model sessions.

## Preprocessing and Inference

### Text Normalization with UnicodeProcessor

The **UnicodeProcessor** class in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) handles all text preprocessing before tokenization. It performs Unicode normalization, strips emojis, replaces punctuation, and wraps the input in language tags (e.g., `<en>…</en>`) as required by the model.

```javascript
const processor = new UnicodeProcessor();
const normalizedText = processor.process("Hello, world!", "en");
// Result: "<en>Hello, world!</en>"

```

This step is mandatory because the TTS model expects language-tagged token streams rather than raw user input.

### The Four-Stage Pipeline

The `TextToSpeech` class orchestrates inference through four sequential stages:

1. **Duration Prediction**: Predicts phoneme durations using `duration_predictor.onnx`
2. **Text Encoding**: Converts token IDs to latent representations via `text_encoder.onnx`
3. **Latent Diffusion**: Iteratively denoises the latent representation using the vector estimator. The `totalStep` parameter controls quality versus speed (fewer steps = faster inference)
4. **Vocoding**: Converts the final latent to raw waveform audio using `vocoder.onnx`

The `call` method executes this pipeline asynchronously, accepting a progress callback to track denoising steps:

```javascript
const { wav, duration } = await textToSpeech.call(
  text,           // normalized input
  lang,           // language code (e.g., "en")
  currentStyle,   // Style object from loadVoiceStyle
  steps,          // number of denoising iterations (e.g., 8)
  speed,          // speaking rate multiplier (e.g., 1.05)
  0.3,            // guidance scale
  (step, total) => console.log(`Progress: ${step}/${total}`)
);

```

## Real-Time Implementation Pattern

To achieve real-time interaction, initiate the models once during page load, then invoke synthesis on user input events with debouncing. Because the pipeline runs entirely in the browser’s UI thread via ONNX Runtime Web, you can trigger synthesis on every keystroke without blocking the interface.

```javascript
let debounceTimer;
const textarea = document.getElementById('input');
const audioElement = document.getElementById('player');

textarea.addEventListener('keyup', () => {
  clearTimeout(debounceTimer);
  debounceTimer = setTimeout(async () => {
    const txt = textarea.value;
    const lang = 'en';
    const speed = 1.0;
    const steps = 8;

    const { wav } = await textToSpeech.call(
      txt, lang, currentStyle, steps, speed, 0.3,
      (step, total) => console.log(`denoise ${step}/${total}`)
    );

    // Convert to WAV and play immediately
    const wavBuf = writeWavFile(wav, textToSpeech.sampleRate);
    const url = URL.createObjectURL(new Blob([wavBuf], { type: 'audio/wav' }));
    audioElement.src = url;
    audioElement.play();
  }, 300);
});

```

This pattern mirrors the `generateSpeech()` implementation in [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js) but isolates the synthesis logic for reuse in chat interfaces or live captioning systems.

## Complete Integration Example

Here is a complete implementation combining initialization, voice loading, and real-time synthesis:

```javascript
import { loadTextToSpeech, loadVoiceStyle, UnicodeProcessor, writeWavFile } from './helper.js';

class SupertonicTTS {
  constructor() {
    this.tts = null;
    this.style = null;
    this.processor = new UnicodeProcessor();
  }

  async initialize(modelPath = 'assets/onnx', stylePath = 'assets/voice_styles/M1.json') {
    const { textToSpeech } = await loadTextToSpeech(modelPath, {
      executionProviders: ['webgpu'],
      graphOptimizationLevel: 'all'
    });
    this.tts = textToSpeech;
    this.style = await loadVoiceStyle([stylePath]);
  }

  async speak(text, lang = 'en', steps = 8, speed = 1.0) {
    if (!this.tts) throw new Error('TTS not initialized');
    
    const normalized = this.processor.process(text, lang);
    const { wav } = await this.tts.call(
      normalized, lang, this.style, steps, speed, 0.3
    );
    
    const wavBuf = writeWavFile(wav, this.tts.sampleRate);
    const blob = new Blob([wavBuf], { type: 'audio/wav' });
    const url = URL.createObjectURL(blob);
    
    const audio = new Audio(url);
    audio.play();
    return audio;
  }
}

// Usage
const tts = new SupertonicTTS();
await tts.initialize();

// Real-time usage with debouncing
document.getElementById('chatInput').addEventListener('input', (e) => {
  clearTimeout(window.ttsTimer);
  window.ttsTimer = setTimeout(() => {
    tts.speak(e.target.value, 'en', 8, 1.05);
  }, 250);
});

```

## Summary

- **Model Loading**: Use `loadTextToSpeech` from [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) to initialize ONNX Runtime Web with WebGPU priority
- **Voice Styles**: Load pre-extracted speaker tensors via `loadVoiceStyle` from JSON files in `assets/voice_styles/`
- **Text Prep**: Pass all input through `UnicodeProcessor` to add language tags and normalize Unicode
- **Inference**: Call `textToSpeech.call()` with the normalized text, voice style, step count, and speed parameters
- **Audio Output**: Convert raw waveforms to playable WAV files using `writeWavFile` and stream via Blob URLs
- **Real-Time**: Initialize once on page load, then debounce synthesis calls on user input events for sub-second response times

## Frequently Asked Questions

### What browsers support Supertonic’s real-time TTS?

Supertonic requires ONNX Runtime Web, which supports WebGPU in Chrome 113+, Edge 113+, and Firefox Nightly with WebGPU enabled. If WebGPU is unavailable, the library automatically falls back to WebAssembly, which works in all modern browsers but with higher latency. Check the [`web/README.md`](https://github.com/supertone-inc/supertonic/blob/main/web/README.md) for specific version requirements.

### How do I reduce latency for real-time applications?

Reduce the `steps` parameter in `textToSpeech.call()` from the default 20 to 8 or fewer. Each denoising step adds approximately 50-100ms of computation time. According to the source code in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js), the `totalStep` value directly controls the number of iterations in the `_infer` method’s denoising loop.

### Can I switch voices without reloading the page?

Yes. Call `loadVoiceStyle` with a different JSON file path (e.g., [`assets/voice_styles/F3.json`](https://github.com/supertone-inc/supertonic/blob/main/assets/voice_styles/F3.json)) and pass the returned `Style` object to subsequent `textToSpeech.call()` invocations. The underlying ONNX sessions remain cached, so voice switching adds only the JSON parsing overhead.

### What file formats does the audio output use?

The `writeWavFile` function in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) generates standard 16-bit PCM WAV files with the sample rate defined by `textToSpeech.sampleRate` (typically 22050Hz or 24000Hz). The output is returned as a Uint8Array suitable for creating Blob URLs or uploading to servers.