How to Implement Streaming Audio Playback for Supertonic TTS Output
Implement streaming audio playback for Supertonic TTS by feeding PCM samples directly to an AVAudioEngine as they are generated, eliminating the temporary file write and AVAudioPlayer dependency used in the standard iOS example.
The Supertonic text-to-speech engine from supertone-inc/supertonic generates high-quality speech using ONNX models, but the stock iOS implementation writes complete WAV files before playback begins. This creates unnecessary latency for real-time applications. By modifying the TTSService.swift architecture to stream chunks directly to AVAudioPlayerNode, you can achieve low-latency audio output that starts playing while synthesis continues.
Understanding the Batch Playback Bottleneck
The current implementation in ios/ExampleiOSApp/TTSService.swift follows a batch processing pattern. The synthesize method loads the ONNX model, runs full-length inference through textToSpeech.call, and writes the resulting float-PCM array to a temporary .wav file using writeWavFile logic. Only then does ios/ExampleiOSApp/AudioPlayer.swift load that file into an AVAudioPlayer and begin playback.
This approach forces three sequential delays: complete inference computation, filesystem write operations, and audio file memory mapping. For utterances longer than a few seconds, this blocks the UI thread and creates audible gaps between the synthesis request and audio output.
Streaming Architecture Overview
A streaming pipeline replaces the file-based AudioPlayer with an AVAudioEngine containing an AVAudioPlayerNode. Instead of waiting for the complete wav array, you split the PCM data into small buffers (typically 20 ms chunks) and schedule them with playerNode.scheduleBuffer. The audio engine pulls these buffers as needed, allowing playback to commence as soon as the first inference chunk is ready.
This architecture eliminates filesystem I/O entirely and reduces latency from hundreds of milliseconds to the duration of a single buffer plus initial inference time.
Step-by-Step Implementation
Model Inference and PCM Chunking
The textToSpeech.call method in TTSService.swift returns a complete (wav, duration) tuple where wav is a [Float] array. For streaming, split this array into chunks of 882 samples (20 ms at 44.1 kHz). Convert each chunk from float (-1.0 to 1.0) to Int16 (clamped and multiplied by 32767) using the same scaling logic found in web/helper.js.
Audio Engine Configuration
Initialize an AVAudioEngine and attach an AVAudioPlayerNode. Configure the format as 44.1 kHz, mono, 16-bit PCM (AVAudioCommonFormat.pcmFormatInt16). Connect the player node to the main mixer node and start the engine with engine.start(). All AVAudioEngine operations must occur on the main thread.
Buffer Scheduling and Playback
Wrap each converted chunk in an AVAudioPCMBuffer with the player node's output format. Schedule buffers using playerNode.scheduleBuffer(_:completionHandler:), chaining completion handlers to detect when the final buffer finishes. Call playerNode.play() immediately after scheduling the first buffer to begin streaming.
Complete Swift Implementation
Below is a drop-in replacement for the file-based TTSService that streams directly to the audio hardware:
import AVFoundation
import onnxruntime
final class StreamingTTSService {
private let env: ORTEnv
private let textToSpeech: TextToSpeech
private let sampleRate: Int
private let engine = AVAudioEngine()
private let playerNode = AVAudioPlayerNode()
private let chunkSize = 882 // 20ms at 44.1kHz
init() throws {
let bundleOnnxDir = try TTSService.locateOnnxDirInBundle()
env = try ORTEnv(loggingLevel: .warning)
textToSpeech = try loadTextToSpeech(bundleOnnxDir, false, env)
sampleRate = textToSpeech.sampleRate
let format = AVAudioFormat(commonFormat: .pcmFormatInt16,
sampleRate: Double(sampleRate),
channels: 1,
interleaved: true)!
engine.attach(playerNode)
engine.connect(playerNode, to: engine.mainMixerNode, format: format)
try engine.start()
}
private func makeBuffers(from wav: [Float]) -> [AVAudioPCMBuffer] {
var buffers: [AVAudioPCMBuffer] = []
var index = 0
while index < wav.count {
let remaining = wav.count - index
let count = min(chunkSize, remaining)
let int16Data = wav[index..<index+count].map { sample -> Int16 in
let clamped = max(-1.0, min(1.0, sample))
return Int16(clamped * 32767)
}
let buffer = AVAudioPCMBuffer(pcmFormat: playerNode.outputFormat(forBus: 0),
frameCapacity: AVAudioFrameCount(count))!
buffer.frameLength = buffer.frameCapacity
buffer.int16ChannelData?.pointee.assign(from: int16Data, count: count)
buffers.append(buffer)
index += count
}
return buffers
}
func synthesizeAndStream(text: String,
nfe: Int,
voice: TTSService.Voice,
language: TTSService.Language,
onFinish: (() -> Void)? = nil) async throws {
let styleURL = try TTSService.locateVoiceStyleURL(voice: voice)
let style = try loadVoiceStyle([styleURL.path], verbose: false)
let (wav, _) = try textToSpeech.call(text, language.rawValue, style, nfe)
let buffers = makeBuffers(from: wav)
for (i, buffer) in buffers.enumerated() {
let isLast = i == buffers.count - 1
playerNode.scheduleBuffer(buffer) {
if isLast { onFinish?() }
}
}
playerNode.play()
}
}
UI Integration
Replace the TTSService.synthesize call in ios/ExampleiOSApp/TTSViewModel.swift with the streaming service. The completion handler maintains compatibility with existing UI callbacks:
@StateObject private var viewModel = TTSViewModel()
Button("Speak") {
Task {
do {
let service = try StreamingTTSService()
try await service.synthesizeAndStream(
text: viewModel.inputText,
nfe: 8,
voice: viewModel.selectedVoice,
language: viewModel.selectedLanguage,
onFinish: { print("Playback complete") }
)
} catch {
viewModel.errorMessage = error.localizedDescription
}
}
}
Cross-Platform Parallels
The streaming concept applies across the Supertonic ecosystem. The web demo in web/helper.js uses writeWavFile to generate an ArrayBuffer that can be streamed to the Web Audio API AudioBufferSourceNode instead of downloading. Similarly, the Flutter implementation in flutter/lib/helper.dart can feed PCM chunks to a StreamController rather than writing temporary files, maintaining architectural consistency with the iOS approach.
Summary
- Eliminate file I/O by removing
writeWavFileandAVAudioPlayerdependencies fromTTSService.swift. - Use
AVAudioEnginewithAVAudioPlayerNodeto push PCM buffers directly to hardware. - Chunk at 20ms intervals (882 samples at 44.1 kHz) to balance latency with scheduling overhead.
- Maintain thread safety by keeping
AVAudioEngineoperations on the main thread while running ONNX inference on background tasks. - Scale floats to Int16 using the same clamping logic (
sample * 32767) found in the web helper utilities.
Frequently Asked Questions
How does streaming audio playback reduce latency compared to the standard Supertonic iOS example?
The standard implementation in ios/ExampleiOSApp/TTSService.swift waits for complete inference, writes a temporary WAV file, and then loads it into AVAudioPlayer. Streaming bypasses the filesystem entirely by feeding PCM chunks directly to AVAudioPlayerNode, reducing startup delay from hundreds of milliseconds to approximately 20-40ms (the first buffer duration).
Can I pause and resume playback if synthesis is slower than real-time?
Yes. Call playerNode.pause() when buffer underruns are detected and playerNode.play() when additional chunks are ready. This creates back-pressure handling ensuring smooth audio even when textToSpeech.call produces output slower than playback consumes it.
Is the streaming approach compatible with the Web and Flutter Supertonic examples?
Absolutely. The architecture mirrors the web implementation in web/helper.js where PCM data can be streamed to Web Audio API nodes instead of file downloads, and the Flutter code in flutter/lib/helper.dart can use platform channels to stream buffers rather than file paths.
What audio format should I use for the AVAudioEngine configuration?
Configure the format as 44.1 kHz sample rate, 1 channel (mono), and 16-bit integer PCM (AVAudioCommonFormat.pcmFormatInt16). This matches the bit depth used in writeWavFile and ensures compatibility with the Int16 conversion logic while minimizing memory bandwidth compared to float buffers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →