What is Supertonic's Batch Processing Mode and When Should You Use It?
Supertonic batch processing mode disables automatic text chunking and processes parallel lists of short texts in a single ONNX inference call, returning separate waveforms for each input rather than one concatenated audio file.
Supertonic by supertone-inc is an open-source text-to-speech library that supports multiple programming languages through ONNX-based inference. The library operates in two distinct synthesis modes: a default long-form mode that automatically splits and concatenates text segments, and a specialized batch processing mode designed for high-throughput generation of multiple short utterances.
How Supertonic Batch Processing Mode Works
Unlike the default mode—which automatically splits texts exceeding the model's maximum chunk size and inserts ≈0.3 s pauses between segments—batch processing mode disables this chunking behavior entirely. When invoked via the --batch flag or the batch method, the system expects parallel lists of texts, voice-style files, and language codes, processing each entry as-is through a single inference step.
Input Validation and Tensor Pre-allocation
In py/helper.py at line 246, the TextToSpeech.batch method first validates that the three input lists—texts, voice-styles, and languages—contain identical lengths. It then pre-allocates tensors for style vectors and text IDs once for the full batch size (e.g., np.zeros([bsz, …])), eliminating the overhead of repeated memory allocation during inference.
Single Inference Pass Architecture
The core optimization occurs in the ONNX runtime execution. According to the source code in py/helper.py, the batch method runs the four model components—duration_predictor, text_encoder, vector_estimator, and vocoder—exactly once with a batch dimension. This architecture is mirrored across language bindings, including nodejs/helper.js at line 300, swift/Sources/Helper.swift at line 734, and rust/src/helper.rs at line 720, where the batch function builds a batch of style vectors and executes the ONNX runtime in a single call.
Batch Processing Mode vs Default Long-Form Mode
| Feature | Default Mode | Batch Processing Mode |
|---|---|---|
| Text handling | Automatic chunking for long texts | No automatic splitting; inputs processed as-is |
| Output format | Single concatenated audio stream | List of separate waveforms (one per input) |
| Timing | Adds ≈0.3 s pause between chunks | No automatic pauses; caller controls timing |
| Input structure | Single text string | Parallel lists of texts, styles, and languages |
Implementation Across Language Bindings
The batch processing API is consistent across all Supertonic language implementations, sharing the same validation logic and single-inference architecture.
Python Implementation
In the Python bindings, the TextToSpeech.batch method in py/helper.py handles the --batch CLI flag passed through example_onnx.py. The method validates input lengths, pre-allocates NumPy arrays, and returns a tuple of (wavs, durations) where wavs is a list of waveform arrays.
Node.js, Swift, Rust, and Go
The Node.js implementation in nodejs/helper.js at line 300 pre-allocates tensors and calls the ONNX runtime once. Similarly, Swift's batch function in swift/Sources/Helper.swift at line 734, Rust's pub fn batch in rust/src/helper.rs at line 720, and Go's batch routine in go/helper.go all implement the same parallel-list processing logic. Java and C# bindings access batch functionality through JNI and helper classes respectively, following the same validation and inference pattern.
When to Use Supertonic Batch Processing Mode
Use batch processing mode when synthesizing multiple short utterances where per-call overhead would otherwise dominate latency:
- Generating UI prompts or voice commands: Process hundreds of short labels ("Hello", "Good morning", "Confirm") in one inference call.
- CPU-only deployments: Reduce costly context switches and model loading overhead by batching inputs into a single ONNX execution.
- Deterministic timing requirements: When you need exact control over audio duration without automatic 0.3 s pauses between segments.
- Bulk style-transfer experiments: Supply different voice-style files for each text in the batch to generate varied samples efficiently.
Avoid batch mode for long paragraphs or articles, as the model's maximum input length will be exceeded, causing truncation or errors. For long-form content, use the default mode's automatic chunking.
Code Examples
Python Command Line
python example_onnx.py \
--batch \
--voice-style ../style1.onnx,../style2.onnx \
--text 'Hello world' 'Good morning' \
--lang en en
Node.js Programmatic API
const { TextToSpeech, loadTextToSpeech, loadVoiceStyle } = require('./helper');
const tts = await loadTextToSpeech('model_dir');
const style = await loadVoiceStyle(['style1.onnx', 'style2.onnx']);
const texts = ['Hello world', 'Good morning'];
const langs = ['en', 'en'];
const [wavs, durations] = await tts.batch(texts, langs, style, totalStep = 20);
Swift Command Line
./ExampleONNX \
--batch \
--voice-style ../style1.onnx,../style2.onnx \
--text 'Hello world' 'Good morning' \
--lang en en
Summary
- Supertonic batch processing mode processes parallel lists of short texts in a single ONNX inference pass, returning separate waveforms rather than concatenated audio.
- The mode disables automatic text chunking and the 0.3 s inter-chunk pauses present in default mode.
- Implementations across Python (
py/helper.py), Node.js (nodejs/helper.js), Swift (swift/Sources/Helper.swift), Rust (rust/src/helper.rs), and Go (go/helper.go) share identical validation and pre-allocation logic. - Use batch mode for high-throughput synthesis of short prompts, CPU-only environments, and scenarios requiring deterministic timing control.
- Do not use batch mode for long-form content; rely on default mode's automatic chunking for paragraphs or articles.
Frequently Asked Questions
What is the difference between Supertonic's batch processing mode and default mode?
Default mode automatically splits long texts into chunks, synthesizes each separately, and concatenates results with ≈0.3 s pauses. Batch processing mode disables chunking, expects parallel lists of inputs, and runs a single inference for the entire batch, returning separate waveforms without automatic pauses.
Can I use batch processing mode for long articles or paragraphs?
No. Batch mode does not perform automatic text chunking, so inputs exceeding the model's maximum length will truncate or error. For long-form content, use the default mode which handles segmentation automatically.
Which programming languages support Supertonic batch processing?
All official Supertonic bindings support batch processing: Python (TextToSpeech.batch in py/helper.py), Node.js (batch in nodejs/helper.js), Swift (batch in swift/Sources/Helper.swift), Rust (pub fn batch in rust/src/helper.rs), Go (go/helper.go), Java (via JNI), and C# (via helper classes).
How does batch processing improve performance?
Batch processing reduces per-call overhead by validating inputs once, pre-allocating tensors for the full batch size, and executing the ONNX models (duration_predictor, text_encoder, vector_estimator, vocoder) in a single inference pass. This eliminates redundant model loads and context switches, significantly reducing latency when processing multiple utterances.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →