# How to Implement Batch Processing for Multiple Texts in Supertonic: A Complete Guide

> Learn how to implement batch processing for multiple texts in Supertonic using the batch method. Efficiently process texts with parallel lists and concatenate waveforms with silence padding.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**Supertonic processes multiple texts efficiently by accepting parallel lists of strings, languages, and voice-style files through the `batch` method, concatenating the resulting waveforms with silence padding while skipping automatic chunking logic.**

Batch processing for multiple texts in Supertonic enables you to synthesize multiple utterances in a single inference call, returning one concatenated audio buffer. The `TextToSpeech` class implements this capability identically across all supported languages—Python, Node.js, Rust, Swift, and C++—by forwarding parallel input lists directly to the underlying ONNX inference pipeline in the supertone-inc/supertonic repository.

## Understanding the Batch Processing Architecture

When you invoke **batch mode**, the library bypasses the automatic text chunking that normally splits long passages into smaller segments. Instead, each entry in your input lists is processed independently through the full ONNX pipeline—including tokenization, duration prediction, vector estimation, and vocoder inference—before the resulting waveforms are concatenated.

### How the Batch Method Works

The internal flow follows a consistent pattern across all language implementations:

1. **Input Validation** – The implementation verifies that `text_list`, `lang_list`, and voice-style files contain matching entry counts.

2. **Sequential Inference** – The `_infer` routine (implemented in each language's helper file) loops through the input lists, running the full ONNX pipeline for each text-language pair.

3. **Waveform Concatenation** – Successfully generated waveforms are joined with a default silence segment of approximately 0.3 seconds between utterances.

4. **Duration Aggregation** – Individual durations are summed to reflect the total length of the concatenated output.

### Cross-Language API Consistency

According to the supertone-inc/supertonic source code, the **batch API contract** remains identical across all language implementations. Whether you are calling from Python ([`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py)), Node.js ([`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js)), Rust ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)), Swift ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)), or C++ ([`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp)), the method accepts the same parameter structure and returns a tuple containing the waveform array and duration metadata.

## Batch Method Signature and Parameters

The `batch` method signature follows this pattern across all languages:

```text
batch(text_list: List[str],
      lang_list: List[str],
      style: Style,
      total_step: int,
      speed: float = 1.05) → (wav: np.ndarray, duration: np.ndarray)

```

**Parameter specifications:**

- **`text_list`** – Array or list of strings to synthesize. Each string represents one utterance.
- **`lang_list`** – Parallel array of ISO language codes (e.g., "en", "es") corresponding to each text entry.
- **`style`** – Style configuration object containing voice parameters and path references.
- **`total_step`** – Integer specifying the inference steps for the diffusion model.
- **`speed`** – Optional float controlling speech tempo (default 1.05).

## Implementing Batch Processing in Python

In the Python implementation located in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), the `batch` method (lines 46-55) delegates directly to `_infer` after validating input dimensions.

### CLI Approach

Use the `--batch` flag in [`py/example_onnx.py`](https://github.com/supertone-inc/supertonic/blob/main/py/example_onnx.py) to process multiple texts from the command line:

```bash
python example_onnx.py \
  --batch \
  --voice-style ../style1.json ../style2.json \
  --text "Hello world!" "¡Hola mundo!" \
  --lang en es \
  --total-step 200

```

### Programmatic Implementation

Import the helper classes and invoke `batch` with parallel lists:

```python
from py.helper import TextToSpeech, load_voice_style, load_text_to_speech

# Initialize the TTS engine

tts = load_text_to_speech("./onnx", use_gpu=False)
style = load_voice_style(["./style1.json", "./style2.json"])

# Process multiple texts in one call

wav, duration = tts.batch(
    ["Hello world!", "¡Hola mundo!"],
    ["en", "es"],
    style,
    total_step=200,
    speed=1.05,
)

```

The method returns `wav` as a NumPy array containing the concatenated audio samples and `duration` as an array reflecting total length.

## Implementing Batch Processing in Node.js

The Node.js wrapper in [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js) exposes an asynchronous `batch` method (lines 300-302) that mirrors the Python signature.

### CLI Approach

Execute batch processing via the example script:

```bash
node example_onnx.js \
  --batch \
  --voice-style ../style1.json,../style2.json \
  --text "Hello world!|¡Hola mundo!" \
  --lang en,es \
  --total-step 200

```

### Programmatic Implementation

Load the modules and await the batch result:

```javascript
const { TextToSpeech, loadVoiceStyle, loadTextToSpeech } = require("./helper.js");

(async () => {
  const tts = await loadTextToSpeech("./onnx", false);
  const style = await loadVoiceStyle(["./style1.json", "./style2.json"]);

  const { wav, duration } = await tts.batch(
    ["Hello world!", "¡Hola mundo!"],
    ["en", "es"],
    style,
    200,
    1.05,
  );
  // wav contains Float32 samples; duration is total length
})();

```

## Implementing Batch Processing in Rust

The Rust implementation in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) provides a synchronous `pub fn batch` (lines 720-728) that forwards vectors to the inference engine.

### CLI Usage

Run the compiled example with batch flags:

```bash
cargo run --release -- \
  --batch \
  --voice-style ../style1.json,../style2.json \
  --text "Hello world!|¡Hola mundo!" \
  --lang en,es \
  --total-step 200

```

### Code Example

```rust
// Inside your Rust application
let (wav, duration) = text_to_speech.batch(
    vec!["Hello world!".to_string(), "¡Hola mundo!".to_string()],
    vec!["en".to_string(), "es".to_string()],
    &style,
    total_step,
    1.05,
)?;

```

## Implementing Batch Processing in Swift

Swift developers can access batch functionality through [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) (lines 734-740).

### CLI Usage

```bash
.build/release/example_onnx \
  --batch \
  --voice-style ../style1.json,../style2.json \
  --text "Hello world!|¡Hola mundo!" \
  --lang en,es \
  --total-step 200

```

### Code Example

```swift
let (wav, duration) = try helper.batch(
    ["Hello world!", "¡Hola mundo!"],
    ["en", "es"],
    style,
    totalStep,
    speed: 1.05
)

```

## Input Validation and Error Handling

The Supertonic implementation enforces strict validation before processing begins. If the entry counts in `text_list`, `lang_list`, and the voice-style file array do not match, the library raises an error immediately. This prevents partial processing and ensures that every text entry has a corresponding language code and voice configuration.

The concatenation logic in `_infer` handles the waveform joining and silence insertion automatically. When multiple texts are provided, the implementation inserts approximately 0.3 seconds of silence between successive waveforms before returning the final buffer.

## Summary

- **Batch processing** in Supertonic accepts parallel lists of texts, languages, and styles through the `batch` method available in Python, Node.js, Rust, Swift, and C++.
- The implementation resides in [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py), [`nodejs/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/nodejs/helper.js), [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs), and [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift), each forwarding to a shared internal inference pipeline.
- Input lists must contain matching entry counts; otherwise, validation errors occur before inference begins.
- Output waveforms are automatically concatenated with ~0.3s silence padding between utterances.
- Both CLI (`--batch` flag) and programmatic APIs are available across all supported languages.

## Frequently Asked Questions

### What is the maximum number of texts I can process in one batch?

The Supertonic source code does not enforce a hard limit on batch size; however, you are constrained by available system memory when concatenating waveforms. Since each text is processed sequentially through the ONNX pipeline, very large batches increase total processing time linearly while accumulating results in memory before returning the final buffer.

### Can I mix different voice styles in a single batch call?

Yes. You must provide a voice-style configuration for each text entry in your batch, passed as the `style` parameter. The implementation validates that the number of style configurations matches the number of text entries, allowing you to synthesize different voices or speaking styles in a single concatenated output.

### Does batch processing affect audio quality compared to single-text synthesis?

No. According to the `TextToSpeech` implementation in the Supertonic repository, batch processing uses the identical ONNX inference pipeline (`_infer`) as single-text synthesis. The only difference is the bypassing of automatic chunking logic and the subsequent concatenation of results; the underlying duration predictor, vector estimator, and vocoder models process each utterance with the same parameters.

### How do I adjust the silence between concatenated utterances in batch mode?

The default silence padding of approximately 0.3 seconds is hardcoded in the internal `_infer` logic across language implementations. To customize silence duration, process texts individually and manually concatenate the resulting `wav` arrays, inserting your preferred amount of zero-padding or silence between segments.