# How to Implement Batch Processing with Supertonic's Python SDK batch() Method

> Implement batch processing efficiently with Supertonic's Python SDK batch() method. Process parallel lists of texts in a single ONNX forward pass for faster results.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-14

---

**Use the `TextToSpeech.batch()` method in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) to process parallel lists of texts, language codes, and voice styles in a single ONNX forward pass, returning a batched waveform tensor and duration vector.**

Supertonic's open-source Python SDK provides a high-performance `TextToSpeech` class for neural speech synthesis. While the `__call__` method handles single utterances, the `batch()` method enables efficient bulk processing by running multiple inputs through the ONNX pipeline simultaneously. This guide demonstrates the exact implementation patterns found in the `supertone-inc/supertonic` repository.

## Prerequisites: Loading Models and Voice Styles

Before calling `batch()`, you must initialize the inference engine and prepare voice styles with matching batch dimensions.

**Load the ONNX runtime** using `load_text_to_speech` from [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py). This creates a `TextToSpeech` instance that wraps the underlying ONNX models.

**Load voice styles** using `load_voice_style`, which returns a `Style` object. When you pass a list of style file paths, the resulting tensors include a leading batch dimension equal to the number of styles.

```python
from supertonic.py.helper import load_text_to_speech, load_voice_style

# Initialize the TTS engine (CPU by default)

tts = load_text_to_speech("../assets/onnx", use_gpu=False)

# Load multiple voice styles simultaneously

style = load_voice_style([
    "../assets/voice_styles/M1.json",   # English male

    "../assets/voice_styles/F1.json",   # Korean female

], verbose=True)

```

## Understanding the batch() Method Signature

The `batch()` method is defined in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) at lines 46-55. It forwards arguments to the private `_infer` routine and returns a tuple of `(waveform, duration)`.

**Input requirements** (all lists must have identical length):
- `text_list`: List of strings to synthesize
- `lang_list`: List of language codes (e.g., `"en"`, `"ko"`)
- `style`: A `Style` object with batch dimension matching the list length
- `total_step`: Number of denoising steps (higher values improve quality)
- `speed`: Speech rate multiplier (default 1.05)

**Output specifications**:
- `waveform`: Tensor of shape `[B, T]` where B is batch size and T is audio samples
- `duration`: Vector of shape `[B]` containing the valid sample count for each utterance

**Critical constraint**: Batch mode skips automatic text chunking. Each text string must fit within the model's maximum length (approximately 300 characters for most languages).

## Implementing Batch Processing

### Minimal Batch Processing Script

The following pattern demonstrates the complete workflow from model loading to audio export. This matches the implementation found in the repository's example files.

```python
from supertonic.py.helper import load_text_to_speech, load_voice_style
import soundfile as sf

# 1. Load inference engine

tts = load_text_to_speech("../assets/onnx", use_gpu=False)

# 2. Load voice styles (creates batch dimension = 2)

style = load_voice_style([
    "../assets/voice_styles/M1.json",
    "../assets/voice_styles/F1.json",
])

# 3. Prepare parallel input lists (length must match style batch size)

texts = [
    "The sunrise painted the sky in orange.",
    "오늘 아침에 공원을 산책했는데, 새소리와 바람 소리가 너무 좋아서 한참을 멈춰 서서 들었어요."
]
langs = ["en", "ko"]

# 4. Run batch inference

wav, dur = tts.batch(texts, langs, style, total_step=8, speed=1.05)

# 5. Save individual files using duration to trim padding

for i, (w, d) in enumerate(zip(wav, dur)):
    samples = int(tts.sample_rate * d.item())
    sf.write(f"output_{i+1}.wav", w[:samples], tts.sample_rate)

```

### Using the Bundled CLI Example

The repository includes [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py), which demonstrates batch processing via command-line interface. The batch call occurs at lines 102-104.

```bash
uv run py/example_onnx.py \
  --voice-style ../assets/voice_styles/M1.json ../assets/voice_styles/F1.json \
  --text "The sun sets behind the mountains." "오늘 저녁에 별을 보며 산책했어요." \
  --lang en ko \
  --batch \
  --total-step 10 \
  --speed 1.2

```

The `--batch` flag triggers the parallel list parsing logic, automatically wiring the inputs to `TextToSpeech.batch()` with the specified inference parameters.

### Integration into Production Applications

For server-side implementations, wrap the batch inference in a function that returns audio bytes rather than writing files directly.

```python
def synthesize_batch(
    onnx_dir: str,
    style_paths: list[str],
    texts: list[str],
    langs: list[str],
    total_step: int = 8,
    speed: float = 1.05,
) -> list[bytes]:
    """Return WAV bytes for each (text, style) pair."""
    import io
    
    tts = load_text_to_speech(onnx_dir, use_gpu=False)
    style = load_voice_style(style_paths)
    
    wav, dur = tts.batch(texts, langs, style, total_step, speed)
    
    wav_bytes = []
    for w, d in zip(wav, dur):
        buf = io.BytesIO()
        valid_samples = int(tts.sample_rate * d.item())
        sf.write(buf, w[:valid_samples], tts.sample_rate, format="WAV")
        wav_bytes.append(buf.getvalue())
    return wav_bytes

```

This pattern enables integration with web frameworks (FastAPI, Flask) or task queues (Celery) without file system dependencies.

## Key Constraints and Performance Considerations

**Batch dimension alignment** is strictly enforced. If you provide three text strings, you must provide exactly three language codes and load exactly three voice styles. Mismatched dimensions will raise runtime errors during the ONNX forward pass.

**Memory scaling** follows the batch size. Each entry in the batch requires sufficient GPU or CPU memory for the full attention computation. For large batches, monitor memory usage or process in chunks.

**No automatic chunking** distinguishes batch mode from single inference. While `__call__` automatically splits long texts, `batch()` processes each element in a single forward pass. Pre-validate text lengths to avoid truncation errors.

**Performance optimization**: Because `_infer` runs the ONNX pipeline once for the entire batch rather than looping in Python, latency overhead per utterance decreases significantly compared to sequential `__call__` invocations.

## Summary

- The `batch()` method in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) enables parallel speech synthesis by accepting aligned lists of texts, languages, and batched voice styles.
- The method returns a stacked waveform tensor `[B, T]` and duration vector `[B]` that must be unpacked manually.
- Batch processing skips automatic text chunking, requiring manual validation of input lengths (approximately 300 characters maximum).
- The repository provides working examples in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) (CLI) and [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) (SDK implementation).

## Frequently Asked Questions

### What is the maximum batch size for the Supertonic Python SDK?

The SDK does not enforce a hardcoded batch size limit. Maximum batch size depends on available system memory (RAM for CPU, VRAM for GPU) and the length of your input texts. Since batch mode processes all inputs simultaneously in the ONNX runtime, monitor memory usage and reduce batch size if you encounter out-of-memory errors.

### Why does batch processing fail when my texts have different lengths?

The `batch()` method requires all input lists (texts, languages, and style files) to have identical lengths because it creates a one-to-one mapping between these parallel arrays. If your texts vary in count from your language codes or loaded voice styles, the method will raise an error during the `_infer` call. Ensure your preprocessing logic aligns all input arrays to the same length before invocation.

### Does batch processing support automatic text chunking for long inputs?

No. Unlike the single-inference `__call__` method, `batch()` processes each text element in a single forward pass without automatic chunking. This design choice reduces computational overhead but requires you to manually segment texts longer than approximately 300 characters before adding them to the batch. Pre-process long documents by splitting them at sentence boundaries or character limits.

### How do I use different total_step values for individual items in a batch?

The `batch()` method applies a single `total_step` value to all items in the batch. If you require different denoising steps for different utterances, you must either run separate batches grouped by step count, or process items individually using the `__call__` method. The internal `_infer` routine in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) uses broadcast operations that assume uniform hyperparameters across the batch dimension.