How to Optimize Supertonic for Low‑Latency Real‑Time Synthesis on CPU
Optimizing Supertonic for low‑latency CPU synthesis requires configuring the ONNX Runtime CPU execution provider, reducing the diffusion total_step parameter to 6–8 iterations, increasing the speed factor to 1.2–1.5, and maintaining singleton model sessions to eliminate initialization overhead.
Supertonic is an open‑source text‑to‑speech engine developed by supertone‑inc that performs all inference on CPU via ONNX Runtime. Achieving sub‑200 ms latency for real‑time applications demands careful tuning of the chunking strategy, denoising iterations, and session reuse rather than raw hardware acceleration.
Core Architecture for CPU Performance
Supertonic’s real‑time efficiency on CPU rests on three architectural pillars implemented in py/helper.py and rust/src/helper.rs:
-
ONNX Runtime with CPUExecutionProvider – All heavy computation (duration prediction, text encoding, latent denoising, and vocoding) executes through ONNX Runtime using the CPU provider. In
py/helper.py:284‑289, the session is created withproviders = ["CPUExecutionProvider"], and the same pattern appears in the Rust bindings atrust/src/helper.rs:30‑33. -
Chunk‑based text processing – Long utterances are split into language‑specific chunks (120 tokens for CJK, 300 for others) to cap maximum tensor dimensions. This prevents latency spikes from memory allocation. The logic resides in
py/helper.py:388‑429and the Rust equivalent atrust/src/helper.rs:30‑50. -
Controlled denoising steps – The flow‑matching decoder iterates for a configurable
total_stepcount. Each step invokes a full forward pass ofvector_estimator.onnx, so reducing steps directly cuts CPU time. The loop is implemented inpy/helper.py:200‑215andrust/src/helper.rs:74‑84.
Pipeline Stages on CPU
Understanding the data flow helps identify optimization bottlenecks. The pipeline executes sequentially within the ONNX Runtime CPU provider:
| Stage | Implementation | Key Code Location |
|---|---|---|
| Unicode preprocessing | Normalizes input, strips emojis, adds language tags (e.g., <en>…</en>) |
py/helper.py:21‑105 |
| Text → ID conversion | Maps characters to integers via unicode_indexer.json; produces text_ids and text_mask tensors |
py/helper.py:111‑130 |
| Duration prediction | duration_predictor.onnx outputs per‑token durations scaled by the speed factor |
py/helper.py:190‑192 |
| Text encoding | text_encoder.onnx generates dense embeddings text_emb |
py/helper.py:194‑197 |
| Latent sampling | Gaussian tensor of shape (bsz, latent_dim, latent_len) is sampled and masked to predicted waveform length |
py/helper.py:164‑176 |
| Denoising (flow‑matching) | Iteratively calls vector_estimator.onnx for total_step iterations to refine the latent representation |
py/helper.py:200‑214 |
| Vocoding | vocoder.onnx converts final latent to 44.1 kHz waveform |
py/helper.py:214‑215 |
| Post‑processing | Applies optional silence padding between chunks and concatenates outputs | py/helper.py:327‑344 |
Low‑Latency Optimization Strategies
Apply these specific strategies to minimize wall‑clock time on CPU:
Reduce total_step to 6–8 iterations – Halving the denoising steps roughly halves inference time. Quality degrades gracefully; 6 steps are often sufficient for voice assistants. Pass this to tts.synthesize() (Python) or tts.call() (Rust).
Increase the speed factor – Values of 1.2–1.5 reduce the predicted duration per token, shortening the latent length and reducing work for the denoiser and vocoder. This is applied at py/helper.py:190.
Pre‑load and reuse model sessions – Session creation incurs one‑time graph optimization costs. Keep a singleton TextToSpeech instance across requests rather than instantiating per utterance. In Python, cache the TTS object; in Rust, keep the load_text_to_speech result in a long‑running service.
Use batch inference for concurrent requests – ONNX Runtime processes batches efficiently via cache‑friendly memory access. Call tts.batch(texts, langs, style, total_step, speed) instead of looping over single synthesis calls.
Minimize chunk size for short inputs – While the default max_len (120 tokens for CJK, 300 for others) is optimized, manually splitting very short inputs and calling the API per‑sentence avoids internal chunking overhead in py/helper.py:388‑429.
Disable unnecessary post‑processing – Eliminate silence padding by setting silence_duration=0.0 and skip preprocessing steps by calling tts._infer() directly after preparing text_ids yourself.
Leverage SIMD‑enabled ONNX Runtime – Ensure you install the standard onnxruntime package, which ships with AVX2/AVX‑512 kernels for 1.5–2× speedup on supported CPUs without code changes.
Python Implementation Example
This daemon keeps ONNX sessions alive, uses reduced diffusion steps, and streams audio without padding:
from supertonic import TTS, load_voice_style
import sounddevice as sd
# Load ONNX assets once (CPU only)
tts = TTS(auto_download=False) # loads config + models
style = load_voice_style(
["assets/voice/style_m1.json"], verbose=False
)
# Low‑latency settings
TOTAL_STEPS = 6 # fewer diffusion steps
SPEED_FACTOR = 1.3 # slightly faster speech
SILENCE = 0.0 # no extra padding
def synthesize(text: str, lang: str = "en"):
wav, _ = tts.synthesize(
text=text,
lang=lang,
voice_style=style,
total_steps=TOTAL_STEPS,
speed=SPEED_FACTOR,
silence_duration=SILENCE,
)
return wav.squeeze()
# Real‑time streaming loop
if __name__ == "__main__":
while True:
cmd = input(">> ")
if not cmd:
continue
audio = synthesize(cmd)
sd.play(audio, 44100)
sd.wait()
The example achieves sub‑200 ms latency for short commands by reusing the session and limiting iterations.
Rust Implementation Example
The Rust implementation mirrors the Python optimization strategy using the same ONNX Runtime parameters:
use supertonic::helper::{
load_text_to_speech, load_voice_style, TextToSpeech,
};
fn main() -> anyhow::Result<()> {
// Load models (CPU only) once
let mut tts = load_text_to_speech("assets", false)?;
let style = load_voice_style(
&["assets/voice/style_m1.json".to_string()],
false
)?;
// Low‑latency parameters
let total_steps = 6usize;
let speed = 1.3f32;
let silence = 0.0f32;
// Synthesize
let (wav, _duration) = tts.call(
"Supertonic runs fast on CPU.",
"en",
&style,
total_steps,
speed,
silence,
)?;
// Stream or write output
supertonic::helper::write_wav_file(
"output.wav",
&wav,
tts.sample_rate
)?;
Ok(())
}
This approach leverages the same session reuse and step reduction as the Python implementation, suitable for embedded systems.
Summary
- Use CPU execution provider by setting
providers = ["CPUExecutionProvider"]inload_text_to_speech()atpy/helper.py:284to eliminate GPU initialization overhead. - Reduce
total_stepto 6–8 iterations to cut diffusion time by 50% with acceptable quality trade‑offs. - Increase
speedto 1.2–1.5 to shorten latent lengths and reduce vocoder workload. - Maintain singleton sessions by reusing the
TTS(Python) orTextToSpeech(Rust) instance across requests to avoid repeated graph optimization costs. - Batch requests when possible via
tts.batch()to maximize cache efficiency and throughput. - Strip post‑processing by setting
silence_duration=0.0or calling_infer()directly to remove padding overhead.
Frequently Asked Questions
How does reducing total_step affect audio quality?
Reducing total_step from the default to 6–8 steps decreases the number of flow‑matching iterations performed in py/helper.py:200‑215. Quality degrades gracefully but remains natural for most voice assistant applications; fewer steps produce slightly rougher transitions but maintain intelligibility while cutting latency by roughly half.
Can I run Supertonic on a Raspberry Pi or other embedded CPU?
Yes. By using the CPU execution provider enforced in rust/src/helper.rs:30‑33 and applying the low‑latency settings (reduced steps, increased speed, no silence padding), Supertonic achieves real‑time synthesis on modest ARM and x86 CPUs. The Rust implementation is particularly suited for resource‑constrained environments due to lower memory overhead than the Python runtime.
What is the difference between tts.synthesize() and tts.batch()?
synthesize() processes a single text string through the full pipeline including Unicode preprocessing at py/helper.py:21‑105. batch() accepts multiple texts and processes them in a single ONNX Runtime forward pass, improving cache locality and throughput when synthesizing multiple utterances concurrently. Both methods respect the total_step and speed parameters defined in the inference loop.
Where does the chunking logic reside and how can I adjust it?
The chunking logic that splits text into 120‑token (CJK) or 300‑token (other languages) segments is implemented in py/helper.py:388‑429 and rust/src/helper.rs:30‑50. These values are loaded from assets/tts.json. You can manually split input strings before calling the API to bypass internal chunking overhead, or modify the configuration JSON to adjust max_len for your specific latency requirements.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →