# Implementing Streaming Synthesis with Supertonic: Real-Time ONNX TTS

> Implement streaming synthesis with Supertonic for real-time ONNX TTS. Process text chunks incrementally and get PCM buffers before the utterance completes.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**You can implement streaming synthesis with Supertonic by wrapping the private `_infer` method to process text chunks incrementally, yielding PCM buffers through Go channels or Python generators before the full utterance completes.**

Supertonic is a multi-language, ONNX-based text-to-speech engine developed by supertone-inc that ships with a unified Go core and thin language-specific wrappers. While the repository provides high-quality offline synthesis via blocking APIs, implementing **streaming synthesis** requires extending the inference pipeline to emit audio frames as soon as individual text chunks finish denoising.

## Understanding Supertonic’s Blocking Architecture

The current `TextToSpeech` implementation in [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) processes speech through three deterministic stages that aggregate results before returning.

### Text Preprocessing and Chunking

The pipeline begins with `preprocessText` (lines 44-73), which performs Unicode normalization, emoji removal, and language-tag wrapping. Long inputs are automatically segmented by `chunkText`, using a maximum length of 300 characters for most languages and 120 for Korean or Japanese text.

### The Monolithic _infer Pipeline

Both the `Call` and `Batch` methods internally invoke the private `_infer` routine. This function runs the full ONNX cascade—duration predictor, text encoder, latent diffusion, and vocoder—for the entire input, returning a complete `[]float32` waveform. Because `_infer` aggregates all chunk results into a single slice, callers must wait for the entire utterance to finish before receiving any audio data.

## Designing a Streaming Interface for Supertonic

A streaming API can be built on top of the existing `_infer` logic by processing chunks separately and converting latent tensors to PCM incrementally.

| Step | Existing Function | Streaming Modification |
|------|-------------------|----------------------|
| 1️⃣ | `chunkText` (splits text) | Process each chunk separately instead of batching. |
| 2️⃣ | `_infer` (full inference) | Expose `InferChunk` to run the ONNX pipeline for one chunk and return raw latent tensors immediately. |
| 3️⃣ | `writeWavFile` | Convert latent to PCM incrementally and push buffers through a **Go channel** or **Python generator**. |
| 4️⃣ | `Call` (concatenates) | Replace concatenation with a streaming loop that forwards data to the consumer. |

Because heavy resources like ONNX sessions and Unicode indexers are cached in the `TextToSpeech` instance, per-chunk overhead is limited to inference and latent-to-waveform conversion.

## Go Implementation: Channel-Based Streaming

The Go core can be extended with a thin wrapper that returns a receive-only channel. Create a new file (e.g., [`streaming.go`](https://github.com/supertone-inc/supertonic/blob/main/streaming.go)) in the package:

```go
package supertonic

import (
	"context"
	"fmt"
)

// StreamChunk runs inference for ONE chunk and sends PCM frames
// on the provided channel. The channel is closed when the chunk
// (plus optional silence) is fully processed.
func (tts *TextToSpeech) StreamChunk(
	ctx context.Context,
	chunk string,
	lang string,
	style *Style,
	totalStep int,
	speed float32,
	silenceDuration float32,
	out chan<- []float32,
) error {
	// 1️⃣ Run the private inference for the single chunk.
	wav, dur, err := tts._infer([]string{chunk}, []string{lang}, style, totalStep, speed)
	if err != nil {
		return err
	}
	// 2️⃣ Trim to the exact duration.
	frames := int(float32(tts.SampleRate) * dur[0])
	if frames > len(wav) {
		frames = len(wav)
	}
	chunkPCM := wav[:frames]

	// 3️⃣ Send the PCM data (non-blocking, abort on ctx Done).
	select {
	case out <- chunkPCM:
	case <-ctx.Done():
		return ctx.Err()
	}

	// 4️⃣ Append optional silence.
	if silenceDuration > 0 {
		silence := make([]float32, int(silenceDuration*float32(tts.SampleRate)))
		select {
		case out <- silence:
		case <-ctx.Done():
			return ctx.Err()
		}
	}
	return nil
}

// Stream synthesises a full text stream by iterating over chunks.
func (tts *TextToSpeech) Stream(
	ctx context.Context,
	text string,
	lang string,
	style *Style,
	totalStep int,
	speed float32,
	silenceDuration float32,
) (<-chan []float32, error) {
	ch := make(chan []float32, 2) // buffered to avoid blocking the pipeline
	maxLen := 300
	if lang == "ko" || lang == "ja" {
		maxLen = 120
	}
	chunks := chunkText(text, maxLen)

	// Launch a goroutine that streams each chunk sequentially.
	go func() {
		defer close(ch)
		for _, c := range chunks {
			if err := tts.StreamChunk(ctx, c, lang, style, totalStep, speed, silenceDuration, ch); err != nil {
				fmt.Printf("stream error: %v\n", err)
				return
			}
		}
	}()
	return ch, nil
}

```

The `Stream` method returns a receive-only channel (`<-chan []float32`), allowing consumers to read PCM buffers as they become available. The `context.Context` argument enables cancellation mid-stream without leaking goroutines.

## Python Implementation: Generator-Based Streaming

For the Python wrapper in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), implement a generator function that yields NumPy arrays incrementally:

```python

# streaming_example.py

import argparse
import queue
import threading

from helper import load_text_to_speech, load_voice_style, sanitize_filename
import numpy as np

def stream_generator(tts, text, lang, style, total_step, speed, silence):
    """Yield PCM chunks for each text segment."""
    max_len = 300 if lang not in ("ko", "ja") else 120
    # reuse the same chunking routine as the Go core (mirrored in helper.py)

    chunks = tts.text_processor.chunk_text(text, max_len)

    for i, chunk in enumerate(chunks):
        wav, duration = tts._infer([chunk], [lang], style, total_step, speed)
        wav_len = int(tts.sample_rate * duration[0])
        yield wav[:wav_len]                     # raw PCM as a NumPy array

        if i != len(chunks) - 1 and silence > 0:
            yield np.zeros(int(silence * tts.sample_rate), dtype=np.float32)

if __name__ == "__main__":
    parser = argparse.ArgumentParser()
    parser.add_argument("--text", required=True)
    parser.add_argument("--lang", default="en")
    parser.add_argument("--onnx-dir", default="../assets/onnx")
    args = parser.parse_args()

    tts = load_text_to_speech(args.onnx_dir, use_gpu=False)
    style = load_voice_style(["../assets/voice_styles/M1.json"], verbose=False)

    # Producer thread pushes PCM chunks into a queue.

    q = queue.Queue(maxsize=4)

    def producer():
        for pcm in stream_generator(tts, args.text, args.lang, style,
                                   total_step=8, speed=1.05, silence=0.2):
            q.put(pcm)
        q.put(None)  # sentinel

    threading.Thread(target=producer, daemon=True).start()

    # Consumer reads from the queue and streams to sounddevice.

    import sounddevice as sd
    while True:
        chunk = q.get()
        if chunk is None:
            break
        sd.play(chunk, samplerate=tts.sample_rate)
        sd.wait()

```

This producer-consumer pattern uses a bounded `queue.Queue` to decouple the ONNX inference thread from the audio playback callback, ensuring the model can prefetch while the sound device consumes data at its own pace.

## Integration with Existing Model Architecture

Streaming extensions leverage the same underlying ONNX models without modification:

- **duration_predictor.onnx** – Estimates phoneme durations per chunk.
- **text_encoder.onnx** – Converts tokens to latent representations.
- **vector_estimator.onnx** – Handles style conditioning.
- **vocoder.onnx** – Converts mel spectrograms to PCM waveforms.

The heavy lifting occurs via `ort.DynamicAdvancedSession` bindings in `LoadTextToSpeech` (lines 58-90 of [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go)). Because sessions persist across calls, you can allocate a `TextToSpeech` instance per client or share a singleton when language and style are identical.

**Key implementation files:**

| File | Role |
|------|------|
| [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) | Core TTS logic, `LoadTextToSpeech`, `_infer` pipeline. |
| [`go/example_onnx.go`](https://github.com/supertone-inc/supertonic/blob/main/go/example_onnx.go) | CLI demo exercising the blocking API. |
| [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) | Python mirror with `load_text_to_speech` and text processing. |
| [`rust/src/example_onnx.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/example_onnx.rs) | Rust reference for native streaming services. |
| [`nodejs/example_onnx.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/example_onnx.js) | Node.js wrapper extensible to async generators. |
| [`ios/ExampleiOSApp/TTSService.swift`](https://github.com/supertone-inc/supertonic/blob/main/ios/ExampleiOSApp/TTSService.swift) | iOS side suitable for Combine-based streams. |

## Summary

- **Supertonic’s core** aggregates all chunks before returning audio, requiring a custom wrapper for streaming synthesis.
- **Go implementations** should use buffered channels and goroutines to stream `[]float32` slices from `StreamChunk`.
- **Python implementations** can mirror the logic with generators and `queue.Queue` for thread-safe audio sinks.
- The same **ONNX Runtime** sessions are reused, keeping per-chunk overhead minimal.
- Source files in [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) and [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) provide the `_infer` and text processing primitives needed for real-time extensions.

## Frequently Asked Questions

### Does Supertonic include a built-in streaming API?

No. The repository provides blocking `Call()` and `Batch()` methods that return complete waveforms. You must implement streaming synthesis by wrapping the private `_infer` routine to process chunks incrementally and yield audio buffers through channels or generators.

### What is the latency improvement when implementing streaming?

Streaming reduces time-to-first-audio from the full utterance duration (seconds) to the processing time of a single chunk (typically 50-200ms), depending on the `totalStep` denoising iterations and CPU/GPU speed.

### Can I reuse the same ONNX models for streaming?

Yes. The same model files—`duration_predictor.onnx`, `text_encoder.onnx`, `vector_estimator.onnx`, and `vocoder.onnx`—are used without modification. The only change is the application-level logic that iterates over chunks rather than processing them as a batch.

### How do I handle user cancellation during streaming?

Pass a `context.Context` to the `Stream` method (Go) or check an event flag in the generator loop (Python). In the Go example above, `ctx.Done()` is checked before sending on the channel, allowing immediate abortion without leaving orphaned ONNX sessions.