# How to Implement Streaming Audio Playback for Supertonic TTS Output

> Implement streaming audio playback for Supertonic TTS. Feed PCM samples directly to AVAudioEngine for efficient audio generation without temporary files.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Implement streaming audio playback for Supertonic TTS by feeding PCM samples directly to an `AVAudioEngine` as they are generated, eliminating the temporary file write and `AVAudioPlayer` dependency used in the standard iOS example.**

The Supertonic text-to-speech engine from supertone-inc/supertonic generates high-quality speech using ONNX models, but the stock iOS implementation writes complete WAV files before playback begins. This creates unnecessary latency for real-time applications. By modifying the [`TTSService.swift`](https://github.com/supertone-inc/supertonic/blob/main/TTSService.swift) architecture to stream chunks directly to `AVAudioPlayerNode`, you can achieve low-latency audio output that starts playing while synthesis continues.

## Understanding the Batch Playback Bottleneck

The current implementation in [`ios/ExampleiOSApp/TTSService.swift`](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSService.swift) follows a batch processing pattern. The `synthesize` method loads the ONNX model, runs full-length inference through `textToSpeech.call`, and writes the resulting float-PCM array to a temporary *.wav* file using `writeWavFile` logic. Only then does [`ios/ExampleiOSApp/AudioPlayer.swift`](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/AudioPlayer.swift) load that file into an `AVAudioPlayer` and begin playback.

This approach forces three sequential delays: complete inference computation, filesystem write operations, and audio file memory mapping. For utterances longer than a few seconds, this blocks the UI thread and creates audible gaps between the synthesis request and audio output.

## Streaming Architecture Overview

A streaming pipeline replaces the file-based `AudioPlayer` with an `AVAudioEngine` containing an `AVAudioPlayerNode`. Instead of waiting for the complete `wav` array, you split the PCM data into small buffers (typically 20 ms chunks) and schedule them with `playerNode.scheduleBuffer`. The audio engine pulls these buffers as needed, allowing playback to commence as soon as the first inference chunk is ready.

This architecture eliminates filesystem I/O entirely and reduces latency from hundreds of milliseconds to the duration of a single buffer plus initial inference time.

## Step-by-Step Implementation

### Model Inference and PCM Chunking

The `textToSpeech.call` method in [`TTSService.swift`](https://github.com/supertone-inc/supertonic/blob/main/TTSService.swift) returns a complete `(wav, duration)` tuple where `wav` is a `[Float]` array. For streaming, split this array into chunks of 882 samples (20 ms at 44.1 kHz). Convert each chunk from float (-1.0 to 1.0) to Int16 (clamped and multiplied by 32767) using the same scaling logic found in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js).

### Audio Engine Configuration

Initialize an `AVAudioEngine` and attach an `AVAudioPlayerNode`. Configure the format as 44.1 kHz, mono, 16-bit PCM (`AVAudioCommonFormat.pcmFormatInt16`). Connect the player node to the main mixer node and start the engine with `engine.start()`. All `AVAudioEngine` operations must occur on the main thread.

### Buffer Scheduling and Playback

Wrap each converted chunk in an `AVAudioPCMBuffer` with the player node's output format. Schedule buffers using `playerNode.scheduleBuffer(_:completionHandler:)`, chaining completion handlers to detect when the final buffer finishes. Call `playerNode.play()` immediately after scheduling the first buffer to begin streaming.

## Complete Swift Implementation

Below is a drop-in replacement for the file-based `TTSService` that streams directly to the audio hardware:

```swift
import AVFoundation
import onnxruntime

final class StreamingTTSService {
    private let env: ORTEnv
    private let textToSpeech: TextToSpeech
    private let sampleRate: Int
    private let engine = AVAudioEngine()
    private let playerNode = AVAudioPlayerNode()
    private let chunkSize = 882  // 20ms at 44.1kHz

    init() throws {
        let bundleOnnxDir = try TTSService.locateOnnxDirInBundle()
        env = try ORTEnv(loggingLevel: .warning)
        textToSpeech = try loadTextToSpeech(bundleOnnxDir, false, env)
        sampleRate = textToSpeech.sampleRate
        
        let format = AVAudioFormat(commonFormat: .pcmFormatInt16,
                                   sampleRate: Double(sampleRate),
                                   channels: 1,
                                   interleaved: true)!
        engine.attach(playerNode)
        engine.connect(playerNode, to: engine.mainMixerNode, format: format)
        try engine.start()
    }

    private func makeBuffers(from wav: [Float]) -> [AVAudioPCMBuffer] {
        var buffers: [AVAudioPCMBuffer] = []
        var index = 0
        
        while index < wav.count {
            let remaining = wav.count - index
            let count = min(chunkSize, remaining)
            
            let int16Data = wav[index..<index+count].map { sample -> Int16 in
                let clamped = max(-1.0, min(1.0, sample))
                return Int16(clamped * 32767)
            }
            
            let buffer = AVAudioPCMBuffer(pcmFormat: playerNode.outputFormat(forBus: 0),
                                          frameCapacity: AVAudioFrameCount(count))!
            buffer.frameLength = buffer.frameCapacity
            buffer.int16ChannelData?.pointee.assign(from: int16Data, count: count)
            
            buffers.append(buffer)
            index += count
        }
        return buffers
    }

    func synthesizeAndStream(text: String,
                             nfe: Int,
                             voice: TTSService.Voice,
                             language: TTSService.Language,
                             onFinish: (() -> Void)? = nil) async throws {
        let styleURL = try TTSService.locateVoiceStyleURL(voice: voice)
        let style = try loadVoiceStyle([styleURL.path], verbose: false)
        
        let (wav, _) = try textToSpeech.call(text, language.rawValue, style, nfe)
        let buffers = makeBuffers(from: wav)
        
        for (i, buffer) in buffers.enumerated() {
            let isLast = i == buffers.count - 1
            playerNode.scheduleBuffer(buffer) {
                if isLast { onFinish?() }
            }
        }
        
        playerNode.play()
    }
}

```

## UI Integration

Replace the `TTSService.synthesize` call in [`ios/ExampleiOSApp/TTSViewModel.swift`](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSViewModel.swift) with the streaming service. The completion handler maintains compatibility with existing UI callbacks:

```swift
@StateObject private var viewModel = TTSViewModel()

Button("Speak") {
    Task {
        do {
            let service = try StreamingTTSService()
            try await service.synthesizeAndStream(
                text: viewModel.inputText,
                nfe: 8,
                voice: viewModel.selectedVoice,
                language: viewModel.selectedLanguage,
                onFinish: { print("Playback complete") }
            )
        } catch {
            viewModel.errorMessage = error.localizedDescription
        }
    }
}

```

## Cross-Platform Parallels

The streaming concept applies across the Supertonic ecosystem. The web demo in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) uses `writeWavFile` to generate an `ArrayBuffer` that can be streamed to the Web Audio API `AudioBufferSourceNode` instead of downloading. Similarly, the Flutter implementation in `flutter/lib/helper.dart` can feed PCM chunks to a `StreamController` rather than writing temporary files, maintaining architectural consistency with the iOS approach.

## Summary

- **Eliminate file I/O** by removing `writeWavFile` and `AVAudioPlayer` dependencies from [`TTSService.swift`](https://github.com/supertone-inc/supertonic/blob/main/TTSService.swift).
- **Use `AVAudioEngine`** with `AVAudioPlayerNode` to push PCM buffers directly to hardware.
- **Chunk at 20ms intervals** (882 samples at 44.1 kHz) to balance latency with scheduling overhead.
- **Maintain thread safety** by keeping `AVAudioEngine` operations on the main thread while running ONNX inference on background tasks.
- **Scale floats to Int16** using the same clamping logic (`sample * 32767`) found in the web helper utilities.

## Frequently Asked Questions

### How does streaming audio playback reduce latency compared to the standard Supertonic iOS example?

The standard implementation in [`ios/ExampleiOSApp/TTSService.swift`](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSService.swift) waits for complete inference, writes a temporary WAV file, and then loads it into `AVAudioPlayer`. Streaming bypasses the filesystem entirely by feeding PCM chunks directly to `AVAudioPlayerNode`, reducing startup delay from hundreds of milliseconds to approximately 20-40ms (the first buffer duration).

### Can I pause and resume playback if synthesis is slower than real-time?

Yes. Call `playerNode.pause()` when buffer underruns are detected and `playerNode.play()` when additional chunks are ready. This creates back-pressure handling ensuring smooth audio even when `textToSpeech.call` produces output slower than playback consumes it.

### Is the streaming approach compatible with the Web and Flutter Supertonic examples?

Absolutely. The architecture mirrors the web implementation in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) where PCM data can be streamed to Web Audio API nodes instead of file downloads, and the Flutter code in `flutter/lib/helper.dart` can use platform channels to stream buffers rather than file paths.

### What audio format should I use for the AVAudioEngine configuration?

Configure the format as 44.1 kHz sample rate, 1 channel (mono), and 16-bit integer PCM (`AVAudioCommonFormat.pcmFormatInt16`). This matches the bit depth used in `writeWavFile` and ensures compatibility with the Int16 conversion logic while minimizing memory bandwidth compared to float buffers.