How to Implement Batch Processing with Supertonic's Python SDK batch() Method
Use the TextToSpeech.batch() method in py/helper.py to process parallel lists of texts, language codes, and voice styles in a single ONNX forward pass, returning a batched waveform tensor and duration vector.
Supertonic's open-source Python SDK provides a high-performance TextToSpeech class for neural speech synthesis. While the __call__ method handles single utterances, the batch() method enables efficient bulk processing by running multiple inputs through the ONNX pipeline simultaneously. This guide demonstrates the exact implementation patterns found in the supertone-inc/supertonic repository.
Prerequisites: Loading Models and Voice Styles
Before calling batch(), you must initialize the inference engine and prepare voice styles with matching batch dimensions.
Load the ONNX runtime using load_text_to_speech from py/helper.py. This creates a TextToSpeech instance that wraps the underlying ONNX models.
Load voice styles using load_voice_style, which returns a Style object. When you pass a list of style file paths, the resulting tensors include a leading batch dimension equal to the number of styles.
from supertonic.py.helper import load_text_to_speech, load_voice_style
# Initialize the TTS engine (CPU by default)
tts = load_text_to_speech("../assets/onnx", use_gpu=False)
# Load multiple voice styles simultaneously
style = load_voice_style([
"../assets/voice_styles/M1.json", # English male
"../assets/voice_styles/F1.json", # Korean female
], verbose=True)
Understanding the batch() Method Signature
The batch() method is defined in py/helper.py at lines 46-55. It forwards arguments to the private _infer routine and returns a tuple of (waveform, duration).
Input requirements (all lists must have identical length):
text_list: List of strings to synthesizelang_list: List of language codes (e.g.,"en","ko")style: AStyleobject with batch dimension matching the list lengthtotal_step: Number of denoising steps (higher values improve quality)speed: Speech rate multiplier (default 1.05)
Output specifications:
waveform: Tensor of shape[B, T]where B is batch size and T is audio samplesduration: Vector of shape[B]containing the valid sample count for each utterance
Critical constraint: Batch mode skips automatic text chunking. Each text string must fit within the model's maximum length (approximately 300 characters for most languages).
Implementing Batch Processing
Minimal Batch Processing Script
The following pattern demonstrates the complete workflow from model loading to audio export. This matches the implementation found in the repository's example files.
from supertonic.py.helper import load_text_to_speech, load_voice_style
import soundfile as sf
# 1. Load inference engine
tts = load_text_to_speech("../assets/onnx", use_gpu=False)
# 2. Load voice styles (creates batch dimension = 2)
style = load_voice_style([
"../assets/voice_styles/M1.json",
"../assets/voice_styles/F1.json",
])
# 3. Prepare parallel input lists (length must match style batch size)
texts = [
"The sunrise painted the sky in orange.",
"오늘 아침에 공원을 산책했는데, 새소리와 바람 소리가 너무 좋아서 한참을 멈춰 서서 들었어요."
]
langs = ["en", "ko"]
# 4. Run batch inference
wav, dur = tts.batch(texts, langs, style, total_step=8, speed=1.05)
# 5. Save individual files using duration to trim padding
for i, (w, d) in enumerate(zip(wav, dur)):
samples = int(tts.sample_rate * d.item())
sf.write(f"output_{i+1}.wav", w[:samples], tts.sample_rate)
Using the Bundled CLI Example
The repository includes py/example_onnx.py, which demonstrates batch processing via command-line interface. The batch call occurs at lines 102-104.
uv run py/example_onnx.py \
--voice-style ../assets/voice_styles/M1.json ../assets/voice_styles/F1.json \
--text "The sun sets behind the mountains." "오늘 저녁에 별을 보며 산책했어요." \
--lang en ko \
--batch \
--total-step 10 \
--speed 1.2
The --batch flag triggers the parallel list parsing logic, automatically wiring the inputs to TextToSpeech.batch() with the specified inference parameters.
Integration into Production Applications
For server-side implementations, wrap the batch inference in a function that returns audio bytes rather than writing files directly.
def synthesize_batch(
onnx_dir: str,
style_paths: list[str],
texts: list[str],
langs: list[str],
total_step: int = 8,
speed: float = 1.05,
) -> list[bytes]:
"""Return WAV bytes for each (text, style) pair."""
import io
tts = load_text_to_speech(onnx_dir, use_gpu=False)
style = load_voice_style(style_paths)
wav, dur = tts.batch(texts, langs, style, total_step, speed)
wav_bytes = []
for w, d in zip(wav, dur):
buf = io.BytesIO()
valid_samples = int(tts.sample_rate * d.item())
sf.write(buf, w[:valid_samples], tts.sample_rate, format="WAV")
wav_bytes.append(buf.getvalue())
return wav_bytes
This pattern enables integration with web frameworks (FastAPI, Flask) or task queues (Celery) without file system dependencies.
Key Constraints and Performance Considerations
Batch dimension alignment is strictly enforced. If you provide three text strings, you must provide exactly three language codes and load exactly three voice styles. Mismatched dimensions will raise runtime errors during the ONNX forward pass.
Memory scaling follows the batch size. Each entry in the batch requires sufficient GPU or CPU memory for the full attention computation. For large batches, monitor memory usage or process in chunks.
No automatic chunking distinguishes batch mode from single inference. While __call__ automatically splits long texts, batch() processes each element in a single forward pass. Pre-validate text lengths to avoid truncation errors.
Performance optimization: Because _infer runs the ONNX pipeline once for the entire batch rather than looping in Python, latency overhead per utterance decreases significantly compared to sequential __call__ invocations.
Summary
- The
batch()method inpy/helper.pyenables parallel speech synthesis by accepting aligned lists of texts, languages, and batched voice styles. - The method returns a stacked waveform tensor
[B, T]and duration vector[B]that must be unpacked manually. - Batch processing skips automatic text chunking, requiring manual validation of input lengths (approximately 300 characters maximum).
- The repository provides working examples in
py/example_onnx.py(CLI) andpy/helper.py(SDK implementation).
Frequently Asked Questions
What is the maximum batch size for the Supertonic Python SDK?
The SDK does not enforce a hardcoded batch size limit. Maximum batch size depends on available system memory (RAM for CPU, VRAM for GPU) and the length of your input texts. Since batch mode processes all inputs simultaneously in the ONNX runtime, monitor memory usage and reduce batch size if you encounter out-of-memory errors.
Why does batch processing fail when my texts have different lengths?
The batch() method requires all input lists (texts, languages, and style files) to have identical lengths because it creates a one-to-one mapping between these parallel arrays. If your texts vary in count from your language codes or loaded voice styles, the method will raise an error during the _infer call. Ensure your preprocessing logic aligns all input arrays to the same length before invocation.
Does batch processing support automatic text chunking for long inputs?
No. Unlike the single-inference __call__ method, batch() processes each text element in a single forward pass without automatic chunking. This design choice reduces computational overhead but requires you to manually segment texts longer than approximately 300 characters before adding them to the batch. Pre-process long documents by splitting them at sentence boundaries or character limits.
How do I use different total_step values for individual items in a batch?
The batch() method applies a single total_step value to all items in the batch. If you require different denoising steps for different utterances, you must either run separate batches grouped by step count, or process items individually using the __call__ method. The internal _infer routine in py/helper.py uses broadcast operations that assume uniform hyperparameters across the batch dimension.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →