Implementing Streaming Synthesis with Supertonic: Real-Time ONNX TTS

You can implement streaming synthesis with Supertonic by wrapping the private _infer method to process text chunks incrementally, yielding PCM buffers through Go channels or Python generators before the full utterance completes.

Supertonic is a multi-language, ONNX-based text-to-speech engine developed by supertone-inc that ships with a unified Go core and thin language-specific wrappers. While the repository provides high-quality offline synthesis via blocking APIs, implementing streaming synthesis requires extending the inference pipeline to emit audio frames as soon as individual text chunks finish denoising.

Understanding Supertonic’s Blocking Architecture

The current TextToSpeech implementation in go/helper.go processes speech through three deterministic stages that aggregate results before returning.

Text Preprocessing and Chunking

The pipeline begins with preprocessText (lines 44-73), which performs Unicode normalization, emoji removal, and language-tag wrapping. Long inputs are automatically segmented by chunkText, using a maximum length of 300 characters for most languages and 120 for Korean or Japanese text.

The Monolithic _infer Pipeline

Both the Call and Batch methods internally invoke the private _infer routine. This function runs the full ONNX cascade—duration predictor, text encoder, latent diffusion, and vocoder—for the entire input, returning a complete []float32 waveform. Because _infer aggregates all chunk results into a single slice, callers must wait for the entire utterance to finish before receiving any audio data.

Designing a Streaming Interface for Supertonic

A streaming API can be built on top of the existing _infer logic by processing chunks separately and converting latent tensors to PCM incrementally.

Step Existing Function Streaming Modification
1️⃣ chunkText (splits text) Process each chunk separately instead of batching.
2️⃣ _infer (full inference) Expose InferChunk to run the ONNX pipeline for one chunk and return raw latent tensors immediately.
3️⃣ writeWavFile Convert latent to PCM incrementally and push buffers through a Go channel or Python generator.
4️⃣ Call (concatenates) Replace concatenation with a streaming loop that forwards data to the consumer.

Because heavy resources like ONNX sessions and Unicode indexers are cached in the TextToSpeech instance, per-chunk overhead is limited to inference and latent-to-waveform conversion.

Go Implementation: Channel-Based Streaming

The Go core can be extended with a thin wrapper that returns a receive-only channel. Create a new file (e.g., streaming.go) in the package:

package supertonic

import (
	"context"
	"fmt"
)

// StreamChunk runs inference for ONE chunk and sends PCM frames
// on the provided channel. The channel is closed when the chunk
// (plus optional silence) is fully processed.
func (tts *TextToSpeech) StreamChunk(
	ctx context.Context,
	chunk string,
	lang string,
	style *Style,
	totalStep int,
	speed float32,
	silenceDuration float32,
	out chan<- []float32,
) error {
	// 1️⃣ Run the private inference for the single chunk.
	wav, dur, err := tts._infer([]string{chunk}, []string{lang}, style, totalStep, speed)
	if err != nil {
		return err
	}
	// 2️⃣ Trim to the exact duration.
	frames := int(float32(tts.SampleRate) * dur[0])
	if frames > len(wav) {
		frames = len(wav)
	}
	chunkPCM := wav[:frames]

	// 3️⃣ Send the PCM data (non-blocking, abort on ctx Done).
	select {
	case out <- chunkPCM:
	case <-ctx.Done():
		return ctx.Err()
	}

	// 4️⃣ Append optional silence.
	if silenceDuration > 0 {
		silence := make([]float32, int(silenceDuration*float32(tts.SampleRate)))
		select {
		case out <- silence:
		case <-ctx.Done():
			return ctx.Err()
		}
	}
	return nil
}

// Stream synthesises a full text stream by iterating over chunks.
func (tts *TextToSpeech) Stream(
	ctx context.Context,
	text string,
	lang string,
	style *Style,
	totalStep int,
	speed float32,
	silenceDuration float32,
) (<-chan []float32, error) {
	ch := make(chan []float32, 2) // buffered to avoid blocking the pipeline
	maxLen := 300
	if lang == "ko" || lang == "ja" {
		maxLen = 120
	}
	chunks := chunkText(text, maxLen)

	// Launch a goroutine that streams each chunk sequentially.
	go func() {
		defer close(ch)
		for _, c := range chunks {
			if err := tts.StreamChunk(ctx, c, lang, style, totalStep, speed, silenceDuration, ch); err != nil {
				fmt.Printf("stream error: %v\n", err)
				return
			}
		}
	}()
	return ch, nil
}

The Stream method returns a receive-only channel (<-chan []float32), allowing consumers to read PCM buffers as they become available. The context.Context argument enables cancellation mid-stream without leaking goroutines.

Python Implementation: Generator-Based Streaming

For the Python wrapper in py/helper.py, implement a generator function that yields NumPy arrays incrementally:


# streaming_example.py

import argparse
import queue
import threading

from helper import load_text_to_speech, load_voice_style, sanitize_filename
import numpy as np

def stream_generator(tts, text, lang, style, total_step, speed, silence):
    """Yield PCM chunks for each text segment."""
    max_len = 300 if lang not in ("ko", "ja") else 120
    # reuse the same chunking routine as the Go core (mirrored in helper.py)

    chunks = tts.text_processor.chunk_text(text, max_len)

    for i, chunk in enumerate(chunks):
        wav, duration = tts._infer([chunk], [lang], style, total_step, speed)
        wav_len = int(tts.sample_rate * duration[0])
        yield wav[:wav_len]                     # raw PCM as a NumPy array

        if i != len(chunks) - 1 and silence > 0:
            yield np.zeros(int(silence * tts.sample_rate), dtype=np.float32)

if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("--text", required=True)
    parser.add_argument("--lang", default="en")
    parser.add_argument("--onnx-dir", default="../assets/onnx")
    args = parser.parse_args()

    tts = load_text_to_speech(args.onnx_dir, use_gpu=False)
    style = load_voice_style(["../assets/voice_styles/M1.json"], verbose=False)

    # Producer thread pushes PCM chunks into a queue.

    q = queue.Queue(maxsize=4)

    def producer():
        for pcm in stream_generator(tts, args.text, args.lang, style,
                                   total_step=8, speed=1.05, silence=0.2):
            q.put(pcm)
        q.put(None)  # sentinel

    threading.Thread(target=producer, daemon=True).start()

    # Consumer reads from the queue and streams to sounddevice.

    import sounddevice as sd
    while True:
        chunk = q.get()
        if chunk is None:
            break
        sd.play(chunk, samplerate=tts.sample_rate)
        sd.wait()

This producer-consumer pattern uses a bounded queue.Queue to decouple the ONNX inference thread from the audio playback callback, ensuring the model can prefetch while the sound device consumes data at its own pace.

Integration with Existing Model Architecture

Streaming extensions leverage the same underlying ONNX models without modification:

  • duration_predictor.onnx – Estimates phoneme durations per chunk.
  • text_encoder.onnx – Converts tokens to latent representations.
  • vector_estimator.onnx – Handles style conditioning.
  • vocoder.onnx – Converts mel spectrograms to PCM waveforms.

The heavy lifting occurs via ort.DynamicAdvancedSession bindings in LoadTextToSpeech (lines 58-90 of go/helper.go). Because sessions persist across calls, you can allocate a TextToSpeech instance per client or share a singleton when language and style are identical.

Key implementation files:

File Role
go/helper.go Core TTS logic, LoadTextToSpeech, _infer pipeline.
go/example_onnx.go CLI demo exercising the blocking API.
py/helper.py Python mirror with load_text_to_speech and text processing.
rust/src/example_onnx.rs Rust reference for native streaming services.
nodejs/example_onnx.js Node.js wrapper extensible to async generators.
ios/ExampleiOSApp/TTSService.swift iOS side suitable for Combine-based streams.

Summary

  • Supertonic’s core aggregates all chunks before returning audio, requiring a custom wrapper for streaming synthesis.
  • Go implementations should use buffered channels and goroutines to stream []float32 slices from StreamChunk.
  • Python implementations can mirror the logic with generators and queue.Queue for thread-safe audio sinks.
  • The same ONNX Runtime sessions are reused, keeping per-chunk overhead minimal.
  • Source files in go/helper.go and py/helper.py provide the _infer and text processing primitives needed for real-time extensions.

Frequently Asked Questions

Does Supertonic include a built-in streaming API?

No. The repository provides blocking Call() and Batch() methods that return complete waveforms. You must implement streaming synthesis by wrapping the private _infer routine to process chunks incrementally and yield audio buffers through channels or generators.

What is the latency improvement when implementing streaming?

Streaming reduces time-to-first-audio from the full utterance duration (seconds) to the processing time of a single chunk (typically 50-200ms), depending on the totalStep denoising iterations and CPU/GPU speed.

Can I reuse the same ONNX models for streaming?

Yes. The same model files—duration_predictor.onnx, text_encoder.onnx, vector_estimator.onnx, and vocoder.onnx—are used without modification. The only change is the application-level logic that iterates over chunks rather than processing them as a batch.

How do I handle user cancellation during streaming?

Pass a context.Context to the Stream method (Go) or check an event flag in the generator loop (Python). In the Go example above, ctx.Done() is checked before sending on the channel, allowing immediate abortion without leaving orphaned ONNX sessions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →