How to Implement Batch Processing for Multiple Texts in Supertonic: A Complete Guide

Supertonic processes multiple texts efficiently by accepting parallel lists of strings, languages, and voice-style files through the batch method, concatenating the resulting waveforms with silence padding while skipping automatic chunking logic.

Batch processing for multiple texts in Supertonic enables you to synthesize multiple utterances in a single inference call, returning one concatenated audio buffer. The TextToSpeech class implements this capability identically across all supported languages—Python, Node.js, Rust, Swift, and C++—by forwarding parallel input lists directly to the underlying ONNX inference pipeline in the supertone-inc/supertonic repository.

Understanding the Batch Processing Architecture

When you invoke batch mode, the library bypasses the automatic text chunking that normally splits long passages into smaller segments. Instead, each entry in your input lists is processed independently through the full ONNX pipeline—including tokenization, duration prediction, vector estimation, and vocoder inference—before the resulting waveforms are concatenated.

How the Batch Method Works

The internal flow follows a consistent pattern across all language implementations:

  1. Input Validation – The implementation verifies that text_list, lang_list, and voice-style files contain matching entry counts.

  2. Sequential Inference – The _infer routine (implemented in each language's helper file) loops through the input lists, running the full ONNX pipeline for each text-language pair.

  3. Waveform Concatenation – Successfully generated waveforms are joined with a default silence segment of approximately 0.3 seconds between utterances.

  4. Duration Aggregation – Individual durations are summed to reflect the total length of the concatenated output.

Cross-Language API Consistency

According to the supertone-inc/supertonic source code, the batch API contract remains identical across all language implementations. Whether you are calling from Python (py/helper.py), Node.js (nodejs/helper.js), Rust (rust/src/helper.rs), Swift (swift/Sources/Helper.swift), or C++ (cpp/helper.cpp), the method accepts the same parameter structure and returns a tuple containing the waveform array and duration metadata.

Batch Method Signature and Parameters

The batch method signature follows this pattern across all languages:

batch(text_list: List[str],
      lang_list: List[str],
      style: Style,
      total_step: int,
      speed: float = 1.05) → (wav: np.ndarray, duration: np.ndarray)

Parameter specifications:

  • text_list – Array or list of strings to synthesize. Each string represents one utterance.
  • lang_list – Parallel array of ISO language codes (e.g., "en", "es") corresponding to each text entry.
  • style – Style configuration object containing voice parameters and path references.
  • total_step – Integer specifying the inference steps for the diffusion model.
  • speed – Optional float controlling speech tempo (default 1.05).

Implementing Batch Processing in Python

In the Python implementation located in py/helper.py, the batch method (lines 46-55) delegates directly to _infer after validating input dimensions.

CLI Approach

Use the --batch flag in py/example_onnx.py to process multiple texts from the command line:

python example_onnx.py \
  --batch \
  --voice-style ../style1.json ../style2.json \
  --text "Hello world!" "¡Hola mundo!" \
  --lang en es \
  --total-step 200

Programmatic Implementation

Import the helper classes and invoke batch with parallel lists:

from py.helper import TextToSpeech, load_voice_style, load_text_to_speech

# Initialize the TTS engine

tts = load_text_to_speech("./onnx", use_gpu=False)
style = load_voice_style(["./style1.json", "./style2.json"])

# Process multiple texts in one call

wav, duration = tts.batch(
    ["Hello world!", "¡Hola mundo!"],
    ["en", "es"],
    style,
    total_step=200,
    speed=1.05,
)

The method returns wav as a NumPy array containing the concatenated audio samples and duration as an array reflecting total length.

Implementing Batch Processing in Node.js

The Node.js wrapper in nodejs/helper.js exposes an asynchronous batch method (lines 300-302) that mirrors the Python signature.

CLI Approach

Execute batch processing via the example script:

node example_onnx.js \
  --batch \
  --voice-style ../style1.json,../style2.json \
  --text "Hello world!|¡Hola mundo!" \
  --lang en,es \
  --total-step 200

Programmatic Implementation

Load the modules and await the batch result:

const { TextToSpeech, loadVoiceStyle, loadTextToSpeech } = require("./helper.js");

(async () => {
  const tts = await loadTextToSpeech("./onnx", false);
  const style = await loadVoiceStyle(["./style1.json", "./style2.json"]);

  const { wav, duration } = await tts.batch(
    ["Hello world!", "¡Hola mundo!"],
    ["en", "es"],
    style,
    200,
    1.05,
  );
  // wav contains Float32 samples; duration is total length
})();

Implementing Batch Processing in Rust

The Rust implementation in rust/src/helper.rs provides a synchronous pub fn batch (lines 720-728) that forwards vectors to the inference engine.

CLI Usage

Run the compiled example with batch flags:

cargo run --release -- \
  --batch \
  --voice-style ../style1.json,../style2.json \
  --text "Hello world!|¡Hola mundo!" \
  --lang en,es \
  --total-step 200

Code Example

// Inside your Rust application
let (wav, duration) = text_to_speech.batch(
    vec!["Hello world!".to_string(), "¡Hola mundo!".to_string()],
    vec!["en".to_string(), "es".to_string()],
    &style,
    total_step,
    1.05,
)?;

Implementing Batch Processing in Swift

Swift developers can access batch functionality through swift/Sources/Helper.swift (lines 734-740).

CLI Usage

.build/release/example_onnx \
  --batch \
  --voice-style ../style1.json,../style2.json \
  --text "Hello world!|¡Hola mundo!" \
  --lang en,es \
  --total-step 200

Code Example

let (wav, duration) = try helper.batch(
    ["Hello world!", "¡Hola mundo!"],
    ["en", "es"],
    style,
    totalStep,
    speed: 1.05
)

Input Validation and Error Handling

The Supertonic implementation enforces strict validation before processing begins. If the entry counts in text_list, lang_list, and the voice-style file array do not match, the library raises an error immediately. This prevents partial processing and ensures that every text entry has a corresponding language code and voice configuration.

The concatenation logic in _infer handles the waveform joining and silence insertion automatically. When multiple texts are provided, the implementation inserts approximately 0.3 seconds of silence between successive waveforms before returning the final buffer.

Summary

  • Batch processing in Supertonic accepts parallel lists of texts, languages, and styles through the batch method available in Python, Node.js, Rust, Swift, and C++.
  • The implementation resides in py/helper.py, nodejs/helper.js, rust/src/helper.rs, and swift/Sources/Helper.swift, each forwarding to a shared internal inference pipeline.
  • Input lists must contain matching entry counts; otherwise, validation errors occur before inference begins.
  • Output waveforms are automatically concatenated with ~0.3s silence padding between utterances.
  • Both CLI (--batch flag) and programmatic APIs are available across all supported languages.

Frequently Asked Questions

What is the maximum number of texts I can process in one batch?

The Supertonic source code does not enforce a hard limit on batch size; however, you are constrained by available system memory when concatenating waveforms. Since each text is processed sequentially through the ONNX pipeline, very large batches increase total processing time linearly while accumulating results in memory before returning the final buffer.

Can I mix different voice styles in a single batch call?

Yes. You must provide a voice-style configuration for each text entry in your batch, passed as the style parameter. The implementation validates that the number of style configurations matches the number of text entries, allowing you to synthesize different voices or speaking styles in a single concatenated output.

Does batch processing affect audio quality compared to single-text synthesis?

No. According to the TextToSpeech implementation in the Supertonic repository, batch processing uses the identical ONNX inference pipeline (_infer) as single-text synthesis. The only difference is the bypassing of automatic chunking logic and the subsequent concatenation of results; the underlying duration predictor, vector estimator, and vocoder models process each utterance with the same parameters.

How do I adjust the silence between concatenated utterances in batch mode?

The default silence padding of approximately 0.3 seconds is hardcoded in the internal _infer logic across language implementations. To customize silence duration, process texts individually and manually concatenate the resulting wav arrays, inserting your preferred amount of zero-padding or silence between segments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →