# How to Optimize Supertonic for Low‑Latency Real‑Time Synthesis on CPU

> Optimize Supertonic for low-latency CPU synthesis. Configure ONNX Runtime, reduce steps to 6-8, boost speed to 1.2-1.5, and use singleton sessions for minimal overhead.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: performance
- Published: 2026-06-15

---

**Optimizing Supertonic for low‑latency CPU synthesis requires configuring the ONNX Runtime CPU execution provider, reducing the diffusion `total_step` parameter to 6–8 iterations, increasing the `speed` factor to 1.2–1.5, and maintaining singleton model sessions to eliminate initialization overhead.**

Supertonic is an open‑source text‑to‑speech engine developed by supertone‑inc that performs all inference on CPU via ONNX Runtime. Achieving sub‑200 ms latency for real‑time applications demands careful tuning of the chunking strategy, denoising iterations, and session reuse rather than raw hardware acceleration.

## Core Architecture for CPU Performance

Supertonic’s real‑time efficiency on CPU rests on three architectural pillars implemented in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) and [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs):

- **ONNX Runtime with CPUExecutionProvider** – All heavy computation (duration prediction, text encoding, latent denoising, and vocoding) executes through ONNX Runtime using the CPU provider. In [`py/helper.py:284‑289`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L284), the session is created with `providers = ["CPUExecutionProvider"]`, and the same pattern appears in the Rust bindings at [`rust/src/helper.rs:30‑33`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs#L30).

- **Chunk‑based text processing** – Long utterances are split into language‑specific chunks (120 tokens for CJK, 300 for others) to cap maximum tensor dimensions. This prevents latency spikes from memory allocation. The logic resides in [`py/helper.py:388‑429`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L388) and the Rust equivalent at [`rust/src/helper.rs:30‑50`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs#L30).

- **Controlled denoising steps** – The flow‑matching decoder iterates for a configurable `total_step` count. Each step invokes a full forward pass of `vector_estimator.onnx`, so reducing steps directly cuts CPU time. The loop is implemented in [`py/helper.py:200‑215`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L200) and [`rust/src/helper.rs:74‑84`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs#L74).

## Pipeline Stages on CPU

Understanding the data flow helps identify optimization bottlenecks. The pipeline executes sequentially within the ONNX Runtime CPU provider:

| Stage | Implementation | Key Code Location |
|-------|----------------|-------------------|
| **Unicode preprocessing** | Normalizes input, strips emojis, adds language tags (e.g., `<en>…</en>`) | [`py/helper.py:21‑105`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L21) |
| **Text → ID conversion** | Maps characters to integers via [`unicode_indexer.json`](https://github.com/supertone-inc/supertonic/blob/main/unicode_indexer.json); produces `text_ids` and `text_mask` tensors | [`py/helper.py:111‑130`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L111) |
| **Duration prediction** | `duration_predictor.onnx` outputs per‑token durations scaled by the `speed` factor | [`py/helper.py:190‑192`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L190) |
| **Text encoding** | `text_encoder.onnx` generates dense embeddings `text_emb` | [`py/helper.py:194‑197`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L194) |
| **Latent sampling** | Gaussian tensor of shape `(bsz, latent_dim, latent_len)` is sampled and masked to predicted waveform length | [`py/helper.py:164‑176`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L164) |
| **Denoising (flow‑matching)** | Iteratively calls `vector_estimator.onnx` for `total_step` iterations to refine the latent representation | [`py/helper.py:200‑214`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L200) |
| **Vocoding** | `vocoder.onnx` converts final latent to 44.1 kHz waveform | [`py/helper.py:214‑215`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L214) |
| **Post‑processing** | Applies optional silence padding between chunks and concatenates outputs | [`py/helper.py:327‑344`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L327) |

## Low‑Latency Optimization Strategies

Apply these specific strategies to minimize wall‑clock time on CPU:

**Reduce `total_step` to 6–8 iterations** – Halving the denoising steps roughly halves inference time. Quality degrades gracefully; 6 steps are often sufficient for voice assistants. Pass this to `tts.synthesize()` (Python) or `tts.call()` (Rust).

**Increase the `speed` factor** – Values of 1.2–1.5 reduce the predicted duration per token, shortening the latent length and reducing work for the denoiser and vocoder. This is applied at [`py/helper.py:190`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L190).

**Pre‑load and reuse model sessions** – Session creation incurs one‑time graph optimization costs. Keep a singleton `TextToSpeech` instance across requests rather than instantiating per utterance. In Python, cache the `TTS` object; in Rust, keep the `load_text_to_speech` result in a long‑running service.

**Use batch inference for concurrent requests** – ONNX Runtime processes batches efficiently via cache‑friendly memory access. Call `tts.batch(texts, langs, style, total_step, speed)` instead of looping over single synthesis calls.

**Minimize chunk size for short inputs** – While the default `max_len` (120 tokens for CJK, 300 for others) is optimized, manually splitting very short inputs and calling the API per‑sentence avoids internal chunking overhead in [`py/helper.py:388‑429`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L388).

**Disable unnecessary post‑processing** – Eliminate silence padding by setting `silence_duration=0.0` and skip preprocessing steps by calling `tts._infer()` directly after preparing `text_ids` yourself.

**Leverage SIMD‑enabled ONNX Runtime** – Ensure you install the standard `onnxruntime` package, which ships with AVX2/AVX‑512 kernels for 1.5–2× speedup on supported CPUs without code changes.

## Python Implementation Example

This daemon keeps ONNX sessions alive, uses reduced diffusion steps, and streams audio without padding:

```python
from supertonic import TTS, load_voice_style
import sounddevice as sd

# Load ONNX assets once (CPU only)

tts = TTS(auto_download=False)  # loads config + models

style = load_voice_style(
    ["assets/voice/style_m1.json"], verbose=False
)

# Low‑latency settings

TOTAL_STEPS = 6          # fewer diffusion steps

SPEED_FACTOR = 1.3       # slightly faster speech

SILENCE = 0.0            # no extra padding

def synthesize(text: str, lang: str = "en"):
    wav, _ = tts.synthesize(
        text=text,
        lang=lang,
        voice_style=style,
        total_steps=TOTAL_STEPS,
        speed=SPEED_FACTOR,
        silence_duration=SILENCE,
    )
    return wav.squeeze()

# Real‑time streaming loop

if __name__ == "__main__":
    while True:
        cmd = input(">> ")
        if not cmd:
            continue
        audio = synthesize(cmd)
        sd.play(audio, 44100)
        sd.wait()

```

The example achieves sub‑200 ms latency for short commands by reusing the session and limiting iterations.

## Rust Implementation Example

The Rust implementation mirrors the Python optimization strategy using the same ONNX Runtime parameters:

```rust
use supertonic::helper::{
    load_text_to_speech, load_voice_style, TextToSpeech,
};

fn main() -> anyhow::Result<()> {
    // Load models (CPU only) once
    let mut tts = load_text_to_speech("assets", false)?;
    let style = load_voice_style(
        &["assets/voice/style_m1.json".to_string()], 
        false
    )?;

    // Low‑latency parameters
    let total_steps = 6usize;
    let speed = 1.3f32;
    let silence = 0.0f32;

    // Synthesize
    let (wav, _duration) = tts.call(
        "Supertonic runs fast on CPU.",
        "en",
        &style,
        total_steps,
        speed,
        silence,
    )?;

    // Stream or write output
    supertonic::helper::write_wav_file(
        "output.wav", 
        &wav, 
        tts.sample_rate
    )?;
    Ok(())
}

```

This approach leverages the same session reuse and step reduction as the Python implementation, suitable for embedded systems.

## Summary

- **Use CPU execution provider** by setting `providers = ["CPUExecutionProvider"]` in `load_text_to_speech()` at [`py/helper.py:284`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L284) to eliminate GPU initialization overhead.
- **Reduce `total_step`** to 6–8 iterations to cut diffusion time by 50% with acceptable quality trade‑offs.
- **Increase `speed`** to 1.2–1.5 to shorten latent lengths and reduce vocoder workload.
- **Maintain singleton sessions** by reusing the `TTS` (Python) or `TextToSpeech` (Rust) instance across requests to avoid repeated graph optimization costs.
- **Batch requests** when possible via `tts.batch()` to maximize cache efficiency and throughput.
- **Strip post‑processing** by setting `silence_duration=0.0` or calling `_infer()` directly to remove padding overhead.

## Frequently Asked Questions

### How does reducing `total_step` affect audio quality?

Reducing `total_step` from the default to 6–8 steps decreases the number of flow‑matching iterations performed in [`py/helper.py:200‑215`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L200). Quality degrades gracefully but remains natural for most voice assistant applications; fewer steps produce slightly rougher transitions but maintain intelligibility while cutting latency by roughly half.

### Can I run Supertonic on a Raspberry Pi or other embedded CPU?

Yes. By using the CPU execution provider enforced in [`rust/src/helper.rs:30‑33`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs#L30) and applying the low‑latency settings (reduced steps, increased speed, no silence padding), Supertonic achieves real‑time synthesis on modest ARM and x86 CPUs. The Rust implementation is particularly suited for resource‑constrained environments due to lower memory overhead than the Python runtime.

### What is the difference between `tts.synthesize()` and `tts.batch()`?

`synthesize()` processes a single text string through the full pipeline including Unicode preprocessing at [`py/helper.py:21‑105`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L21). `batch()` accepts multiple texts and processes them in a single ONNX Runtime forward pass, improving cache locality and throughput when synthesizing multiple utterances concurrently. Both methods respect the `total_step` and `speed` parameters defined in the inference loop.

### Where does the chunking logic reside and how can I adjust it?

The chunking logic that splits text into 120‑token (CJK) or 300‑token (other languages) segments is implemented in [`py/helper.py:388‑429`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py#L388) and [`rust/src/helper.rs:30‑50`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs#L30). These values are loaded from [`assets/tts.json`](https://github.com/supertone-inc/supertonic/blob/main/assets/tts.json). You can manually split input strings before calling the API to bypass internal chunking overhead, or modify the configuration JSON to adjust `max_len` for your specific latency requirements.