# How to integrate Supertonic into a web application

> Integrate Supertonic into your web app for client-side text-to-speech synthesis. Discover how to use ONNX Runtime Web and leverage quick setup with our helper methods.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Supertonic provides a browser-based text-to-speech pipeline using ONNX Runtime Web that enables client-side synthesis through the `loadTextToSpeech` and `TextToSpeech.call` methods defined in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js).**

Supertonic is an open-source text-to-speech (TTS) engine developed by Supertone Inc. that runs entirely in the browser using ONNX Runtime Web. Learning how to integrate Supertonic into a web application allows you to generate speech with zero network latency while keeping user data private, as implemented in the `supertone-inc/supertonic` repository.

## Prerequisites and Model Assets

Before initializing the pipeline, download the ONNX model files from the Hugging Face hub and place them in `web/assets/onnx/`. The required files include `duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, and `vocoder.onnx`, along with the [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json) configuration file.

You also need a voice style JSON file (e.g., [`web/assets/voice_styles/M1.json`](https://github.com/supertone-inc/supertonic/blob/main/web/assets/voice_styles/M1.json)) which defines the speaker characteristics. You can create additional styles using the Voice Builder tool and export them to this directory.

## Initializing the TTS Pipeline

The `loadTextToSpeech` function in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) handles model initialization and runtime setup. It attempts to create a WebGPU session first for GPU acceleration, automatically falling back to WebAssembly (WASM) if WebGPU is unavailable.

```javascript
import { loadTextToSpeech } from './helper.js';

const { textToSpeech, cfgs } = await loadTextToSpeech(
  'assets/onnx',
  { executionProviders: ['webgpu'] }
);

```

This function returns a `TextToSpeech` instance and configuration objects. The `executionProviders` option accepts `webgpu` as the primary target, with the helper managing the WASM fallback internally.

## Loading Voice Styles

Voice styles control the speaker characteristics and are loaded using `loadVoiceStyle` from [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js). This function processes the JSON style file and converts the TTL and DP tensors into ONNX `ort.Tensor` objects required by the inference engine.

```javascript
import { loadVoiceStyle } from './helper.js';

const style = await loadVoiceStyle(
  ['assets/voice_styles/M1.json'],
  true
);

```

Passing `true` as the second argument enables style processing. The function returns a `Style` instance that you pass to the synthesis method.

## Generating Speech

The `TextToSpeech` class orchestrates the full synthesis pipeline including text preprocessing, duration prediction, latent noise sampling, denoising loops, and vocoder inference. Call the `call` method with your text, language code, style object, and synthesis parameters.

```javascript
const { wav, duration } = await textToSpeech.call(
  text,
  lang,
  style,
  totalStep,      // Quality steps: 5 (low) to 12 (high)
  speed,          // Speech rate multiplier (1.0 = normal)
  silenceDuration // Optional silence padding
);

```

The method returns a Float32Array containing the raw waveform and a duration array. The `totalStep` parameter controls quality versus speed—values between 5 and 12 are typical, with higher values producing better quality but requiring more computation.

## Converting to Playable Audio

The raw waveform requires conversion to a standard audio format. The `writeWavFile` utility in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) encodes the Float32Array into a WAV ArrayBuffer.

```javascript
import { writeWavFile } from './helper.js';

const wavBuffer = writeWavFile(wav, textToSpeech.sampleRate);
const blob = new Blob([wavBuffer], { type: 'audio/wav' });
const url = URL.createObjectURL(blob);

// Attach to audio element
audioElement.src = url;

```

This creates a downloadable Blob URL that can be attached to an `<audio>` element for playback or saved to disk.

## Complete Integration Example

The following example combines all components into a minimal HTML page based on the reference implementation in [`web/main.js`](https://github.com/supertone-inc/supertonic/blob/main/web/main.js) and [`web/index.html`](https://github.com/supertone-inc/supertonic/blob/main/web/index.html).

```html
<!-- index.html -->
<!DOCTYPE html>
<html lang="en">
<head>
  <meta charset="UTF-8">
  <title>Supertonic Web TTS</title>
  <script type="module" src="main.js"></script>
</head>
<body>
  <h1>Supertonic Web TTS</h1>
  <textarea id="text" rows="4" placeholder="Enter text…"></textarea><br>
  <label>Language:
    <select id="langSelect">
      <option value="en">English</option>
      <option value="ko">Korean</option>
    </select>
  </label><br>
  <button id="generateBtn" disabled>Generate</button>
  <div id="statusBox"><span id="statusText"></span></div>
  <div id="results"></div>
</body>
</html>

```

```javascript
// main.js
import {
  loadTextToSpeech,
  loadVoiceStyle,
  writeWavFile
} from './helper.js';

let tts = null;
let style = null;

async function init() {
  // Load models with WebGPU → WASM fallback
  const { textToSpeech } = await loadTextToSpeech(
    'assets/onnx',
    { executionProviders: ['webgpu'] }
  );
  tts = textToSpeech;

  // Load voice style
  style = await loadVoiceStyle(['assets/voice_styles/M1.json'], true);

  document.getElementById('generateBtn').disabled = false;
  document.getElementById('statusText').textContent = 'Ready';
}
init();

document.getElementById('generateBtn').addEventListener('click', async () => {
  const text = document.getElementById('text').value.trim();
  const lang = document.getElementById('langSelect').value;
  
  if (!text) return alert('Please enter text');

  // Synthesize with quality step 8
  const { wav, duration } = await tts.call(
    text, 
    lang, 
    style, 
    8,      // totalStep
    1.0     // speed
  );

  // Create WAV blob
  const wavBuffer = writeWavFile(wav, tts.sampleRate);
  const blob = new Blob([wavBuffer], { type: 'audio/wav' });
  const url = URL.createObjectURL(blob);

  // Display results
  const resultsDiv = document.getElementById('results');
  resultsDiv.innerHTML = `
    <audio controls src="${url}"></audio>
    <p>Duration: ${duration[0].toFixed(2)} s</p>
    <a href="${url}" download="speech.wav">Download</a>
  `;
});

```

## Summary

- **Place ONNX assets** in `web/assets/onnx/` including the four model files (`duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, `vocoder.onnx`) and [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json).
- **Initialize the pipeline** using `loadTextToSpeech` from [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js), which automatically handles WebGPU and WASM fallback.
- **Load voice styles** with `loadVoiceStyle` to create `Style` instances from JSON files in `web/assets/voice_styles/`.
- **Synthesize speech** by calling `textToSpeech.call` with text, language, style, and quality parameters.
- **Render audio** using `writeWavFile` to convert the Float32 waveform into a downloadable WAV Blob.

## Frequently Asked Questions

### Does Supertonic require a server backend for web integration?

No. Supertonic runs entirely client-side using ONNX Runtime Web according to the source code in `supertone-inc/supertonic`. All model inference happens in the browser using WebGPU or WebAssembly, eliminating server costs and keeping user data private.

### What hardware acceleration does Supertonic support in browsers?

Supertonic attempts to use **WebGPU** first for GPU acceleration, falling back automatically to **WebAssembly (WASM)** if the browser or device lacks WebGPU support. This behavior is handled internally by the `loadTextToSpeech` function in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js).

### How do I customize the voice style in Supertonic?

Create custom voice styles using the Voice Builder tool and export them as JSON files. Place these files in `web/assets/voice_styles/` and load them using `loadVoiceStyle(['path/to/style.json'], true)`. The boolean flag enables style tensor processing required by the TTS engine.

### What is the difference between totalStep values in the synthesis call?

The `totalStep` parameter controls the number of denoising steps during inference. Values range from **5 (fast, lower quality)** to **12 (slow, higher quality)**. The default examples typically use 8 steps as a balance between synthesis speed and audio quality.