How to Integrate Supertonic into an iOS App Using Swift: Complete Implementation Guide

Integrate Supertonic by adding the ONNX Runtime Swift package, bundling the ONNX models and voice styles, initializing TTSService in a @MainActor view model, and calling synthesize(text:nfe:voice:language:) to generate local WAV files.

Supertonic is an open-source, on-device text-to-speech (TTS) engine developed by supertone-inc that leverages ONNX Runtime for local inference without network dependencies. Unlike cloud-based TTS solutions, Supertonic processes all audio generation directly on the device using optimized ONNX models. This guide walks through the complete Swift integration based on the reference implementation in the supertone-inc/supertonic repository.

Architecture Overview

The iOS integration follows a three-layer architecture that separates model inference from UI concerns.

Model & Runtime Layer handles ONNX model loading and inference. Located in [TTSService.swift](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSService.swift), this layer loads duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx via the ONNX Runtime Swift bindings.

Business Logic Layer exposes a simple async API through the TTSService class. The synthesize(text:nfe:voice:language:) method prepares voice-style JSON, assembles ONNX inputs, runs the inference loop, and writes the resulting WAV file to disk.

UI Layer manages state and presentation. [TTSViewModel.swift](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSViewModel.swift) holds reactive state for text input, NFE sliders, and playback status, while [ContentView.swift](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/ContentView.swift) renders the SwiftUI interface.

Step 1: Add the ONNX Runtime Dependency

Supertonic requires Microsoft's ONNX Runtime Swift package to execute the trained models locally.

Add the dependency to your Package.swift or Xcode project:

// Package.swift
.package(
    url: "https://github.com/microsoft/onnxruntime-swift-package-manager.git",
    from: "1.16.0"
)

The exact dependency declaration is available in the repository's [swift/Package.swift](https://github.com/supertone-inc/supertonic/blob/main/swift/Package.swift).

Step 2: Bundle Supertonic Assets

Copy the required model files and voice configurations into your iOS target bundle. You need the onnx/ directory containing the four model files and the voice_styles/ directory containing JSON configuration files.

Execute these commands from the repository root:

cd ios/ExampleiOSApp
mkdir -p onnx voice_styles
rsync -a ../../assets/onnx/ onnx/
rsync -a ../../assets/voice_styles/ voice_styles/

In Xcode, add these folders as Folder References (blue folders) and ensure they appear under Copy Bundle Resources in your target's build phases. The TTSService class uses locateOnnxDirInBundle() and locateVoiceStyleURL() to discover these resources at runtime.

Step 3: Initialize TTSService in Your View Model

Instantiate TTSService within a @MainActor view model to handle ONNX setup and state management.

import SwiftUI

@MainActor
final class TTSViewModel: ObservableObject {
    @Published var audioURL: URL?
    private var service: TTSService?

    func startup() {
        do {
            service = try TTSService()
        } catch {
            print("Failed to initialize TTS: \(error)")
        }
    }
}

This pattern matches the implementation in [TTSViewModel.swift](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSViewModel.swift). The TTSService initializer throws if assets are missing or ONNX runtime fails to load.

Step 4: Synthesize Speech with Async/Await

Call the synthesize method to convert text into speech. This method runs the full inference pipeline including duration prediction, text encoding, and vocoding.

func generateSpeech(
    for text: String,
    nfe: Int = 8,
    voice: TTSService.Voice = .male,
    language: TTSService.Language = .en
) async throws -> URL {
    guard let service = service else {
        throw NSError(domain: "TTS", code: -1, userInfo: nil)
    }
    
    return try await service.synthesize(
        text: text,
        nfe: nfe,
        voice: voice,
        language: language
    )
}

Internally, synthesize performs the following operations:

  • Loads the appropriate voice-style JSON via locateVoiceStyleURL()
  • Chunks input text using chunkText() to respect model sequence limits
  • Executes the ONNX inference loop across duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx
  • Writes the output to a temporary WAV file via writeWavFile() and returns the URL

The full inference flow is implemented in [TTSService.swift](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSService.swift).

Step 5: Play Generated Audio with AVAudioPlayer

Use AVAudioPlayer to play the synthesized WAV file. The example project includes a wrapper class to manage playback state.

import AVFoundation

final class AudioPlayer {
    private var player: AVAudioPlayer?

    func play(url: URL, finished: @escaping () -> Void) {
        player = try? AVAudioPlayer(contentsOf: url)
        player?.delegate = AudioDelegate(finished)
        player?.play()
    }

    func stop() {
        player?.stop()
        player = nil
    }
}

Wire this player into your view model's togglePlay() method to update UI state while audio plays.

Step 6: Wire the UI to SwiftUI Components

Connect the view model to your SwiftUI interface to create a complete TTS experience.

struct ContentView: View {
    @StateObject private var viewModel = TTSViewModel()

    var body: some View {
        VStack {
            TextEditor(text: $viewModel.text)
                .frame(height: 150)
            
            Slider(value: $viewModel.nfe, in: 2...15, step: 1)
            
            Picker("Voice", selection: $viewModel.voice) {
                Text("Male").tag(TTSService.Voice.male)
                Text("Female").tag(TTSService.Voice.female)
            }
            
            Picker("Language", selection: $viewModel.language) {
                ForEach(TTSService.Language.allCases, id: \.self) { lang in
                    Text(lang.displayName).tag(lang)
                }
            }
            
            Button(viewModel.isGenerating ? "Generating…" : "Generate") {
                Task { await viewModel.generate() }
            }
            .disabled(viewModel.isGenerating)
            
            Button(viewModel.isPlaying ? "Stop" : "Play") {
                viewModel.togglePlay()
            }
            .disabled(viewModel.audioURL == nil)
        }
        .onAppear { viewModel.startup() }
    }
}

This implementation mirrors [ContentView.swift](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/ContentView.swift), binding all TTS parameters to the UI.

Summary

Integrating Supertonic into an iOS Swift app requires four main components:

  • Dependency Management: Add the ONNX Runtime Swift package via SPM to enable local inference.
  • Asset Bundling: Include the onnx/ and voice_styles/ directories as folder references in your Xcode target.
  • Service Initialization: Create TTSService in a @MainActor view model to handle model loading and synthesis.
  • Audio Pipeline: Call synthesize() to generate WAV files, then play them using AVAudioPlayer wrapped in a custom player class.

Frequently Asked Questions

What are the system requirements for running Supertonic on iOS?

Supertonic runs on any iOS device capable of executing ONNX Runtime models, which includes all modern 64-bit iPhones and iPads. The inference occurs entirely on-device using CPU acceleration, meaning no network connection is required after the initial app installation and asset bundling.

How do I customize voice styles and languages in Supertonic?

Voice styles and languages are controlled via the voice and language parameters in the synthesize(text:nfe:voice:language:) method. The TTSService.Voice enum includes options like .male and .female, while TTSService.Language supports multiple locales. The service loads corresponding JSON configurations from the voice_styles/ bundle directory using locateVoiceStyleURL() to apply the correct acoustic parameters for the selected voice.

Can I adjust the quality or speed of the synthesized speech?

Yes, use the nfe (Number of Flow Executions) parameter to balance quality against inference speed. The reference UI implements a slider ranging from 2 to 15, where lower values generate faster but lower-quality audio, and higher values produce higher-fidelity output with increased latency. This integer is passed directly to the ONNX inference session in TTSService.swift.

Where are the generated audio files stored?

The synthesize() method writes temporary WAV files to the app's cache directory and returns the file URL. These files persist until the system cleans temporary storage or the app terminates. The example implementation plays audio directly from the returned URL using AVAudioPlayer as shown in the AudioPlayer.swift wrapper, and you can move the file to permanent storage if persistence is required.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →