Integrating Supertonic with Flutter for Cross-Platform TTS Apps

Supertonic’s Flutter example provides a complete on-device TTS integration using ONNX Runtime to synthesize speech locally across macOS, Windows, Linux, Android, and iOS without network dependencies.

The supertone-inc/supertonic repository delivers a production-ready Flutter implementation that embeds a neural text-to-speech engine directly into mobile and desktop applications. By leveraging the flutter_onnxruntime package and bundling pre-converted ONNX models, you can generate high-quality speech in 31 languages while processing all data on-device for maximum privacy and offline functionality.

Architecture Overview

The Supertonic Flutter integration follows a clean three-layer architecture that separates UI concerns from heavy inference logic.

UI Layer (Flutter Widgets)

Located in flutter/lib/main.dart, the SupertonicApp widget constructs the interface using standard Flutter components. It manages a text input field, a language dropdown supporting 31 language codes, sliders for denoising steps and speech speed, and action buttons for generation, playback, and downloads. The UI listens to model loading states via _loadModels and disables controls during inference to prevent conflicting operations.

TTS Service Layer (helper.dart)

All neural processing resides in flutter/lib/helper.dart. The loadTextToSpeech function initializes four ONNX sessions from assets/onnx/ for the duration predictor, text encoder, vector estimator, and vocoder models. The loadVoiceStyle function parses JSON configurations from assets/voice_styles/ to apply specific speaker characteristics.

When synthesizing, TextToSpeech.call executes the full inference pipeline: tokenizing input through preprocessText (which applies NFKD-like Unicode decomposition and emoji removal), predicting durations via dpOrt, encoding text through textEncOrt, performing iterative denoising using vectorEstOrt for the configured number of steps, and finally generating audio waveforms through vocoderOrt.

Asset Layer

Pre-trained models live in flutter/assets/onnx/ and voice configurations in flutter/assets/voice_styles/. The flutter/pubspec.yaml declares these directories under the flutter assets section, ensuring ONNX Runtime can access them at runtime. The helper methods use copyModelToFile to stream bundled resources into temporary files accessible by the native runtime.

Implementation Walkthrough

Initializing the Engine

During app startup, call _loadModels to prepare the inference environment:

final tts = await loadTextToSpeech('assets/onnx', useGpu: false);
final style = await loadVoiceStyle(['assets/voice_styles/M1.json']);

This loads the four required ONNX models—duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx—and initializes the voice style parameters. The useGpu parameter controls whether inference runs on GPU or CPU hardware.

Generating Speech

When the user triggers synthesis, the app executes _generateSpeech:

  1. Preprocess input: preprocessText normalizes punctuation, removes emojis, and wraps the string in language tags.
  2. Run inference: tts.call chunks the text and executes the ONNX sessions sequentially, sampling a noisy latent and denoising it for the configured number of steps.
  3. Export audio: writeWavFile converts the float waveform to 16-bit PCM WAV format.
  4. Playback: The UI streams the temporary file using just_audio for immediate audio playback.

Handling File Operations

The _downloadFile method copies the generated WAV from temporary storage to the user's Downloads folder, enabling permanent offline access to synthesized content.

Code Examples

Minimal TTS Integration

For quick integration into existing Flutter widgets:

import 'package:just_audio/just_audio.dart';
import 'package:path_provider/path_provider.dart';
import 'helper.dart';

Future<void> synthesizeSpeech(String text) async {
  // Load models (typically done at app startup)
  final tts = await loadTextToSpeech('assets/onnx', useGpu: false);
  final style = await loadVoiceStyle(['assets/voice_styles/M1.json']);
  
  // Process and generate
  final processed = preprocessText(text, 'en');
  final result = await tts.call(processed, 'en', style, 8); // 8 denoising steps
  
  // Extract waveform data
  final wav = result['wav'] as List<double>;
  final sampleRate = tts.sampleRate;
  
  // Write to temporary file and play
  final dir = await getTemporaryDirectory();
  final path = '${dir.path}/demo.wav';
  writeWavFile(path, wav, sampleRate);
  
  await AudioPlayer().setAudioSource(AudioSource.uri(Uri.file(path)));
  await AudioPlayer().play();
}

UI Orchestration

The complete implementation in main.dart wires these components together:

  • App initialization: _loadModels validates asset availability before enabling UI controls.
  • Generation trigger: _generateSpeech coordinates preprocessing, inference, and audio routing.
  • Download handler: _downloadFile persists temporary outputs to permanent storage.
  • Parameter binding: Slider widgets update _totalSteps (denoising iterations) and _speed (speech rate) that get passed to tts.call.

Key Configuration Files

flutter/pubspec.yaml declares the ONNX dependencies and asset bundles:

dependencies:
  flutter_onnxruntime: ^[version]
  just_audio: ^[version]
  
flutter:
  assets:
    - assets/onnx/
    - assets/voice_styles/

The assets/onnx/ directory must contain the four model files exported from the Supertonic training pipeline, while assets/voice_styles/ holds JSON definitions such as M1.json that specify speaker characteristics for the vector estimator.

Summary

  • Supertonic enables fully offline TTS in Flutter apps by running ONNX Runtime inference on-device across all supported platforms.
  • Three-layer architecture separates UI (main.dart), inference logic (helper.dart), and model assets (assets/).
  • Key integration points include loadTextToSpeech for initialization, preprocessText for input normalization, and TextToSpeech.call for the neural synthesis pipeline.
  • Audio output converts float waveforms to standard WAV files via writeWavFile, playable through just_audio or similar plugins.
  • No network requirements make this suitable for privacy-sensitive applications requiring 31-language support.

Frequently Asked Questions

Does Supertonic support GPU acceleration on Flutter?

Yes. Pass useGpu: true to loadTextToSpeech in helper.dart to enable GPU inference via ONNX Runtime, though availability depends on the target platform's GPU drivers and the flutter_onnxruntime package implementation.

How many languages does the Flutter implementation support?

The Supertonic Flutter example supports 31 languages through ISO language codes passed to preprocessText and tts.call, with built-in text normalization for Hangul Jamo decomposition and Latin accent handling.

What are the system requirements for running Supertonic in Flutter?

The app requires sufficient RAM to load the four ONNX models (approximately several hundred MB) and temporary disk space for asset extraction. The flutter_onnxruntime package handles platform-specific bindings for macOS, Windows, Linux, Android, and iOS.

Can I use custom voice styles with the Flutter integration?

Yes. Place your custom JSON voice style files in assets/voice_styles/ and reference them via loadVoiceStyle. The M1.json format defines speaker characteristics that the vector estimator uses during the denoising diffusion process.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →