# How to Handle Long Text Synthesis with Text Chunking in Supertonic

> Learn how Supertonic expertly handles long text synthesis using automatic, intelligent text chunking. Get seamless audio from your lengthy documents.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-05-14

---

**Supertonic automatically splits long input text into manageable chunks using a hierarchical algorithm that respects paragraph and sentence boundaries, then synthesizes each chunk separately and stitches them together with configurable silence intervals to produce seamless audio.**

The `supertone-inc/supertonic` repository provides a neural text-to-speech engine capable of handling arbitrarily long passages without memory overruns. By implementing intelligent text chunking consistently across JavaScript, Rust, and Swift bindings, the library preserves prosodic continuity while respecting the acoustic encoder's token limits.

## Understanding the Chunking Architecture

Supertonic's chunking system operates through a deterministic pipeline that balances semantic coherence with technical constraints. The architecture relies on configuration parameters that define hardware limits and runtime behavior.

### Chunk Size Configuration

The maximum token length for a single inference is governed by two configuration values located in [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json). The **`ae.base_chunk_size`** parameter defines the acoustic encoder's hard limit, while **`ttl.chunk_compress_factor`** acts as a multiplier to determine the effective chunk size used during synthesis. According to the source code in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) [lines 73-77](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js#L73-L77), the system calculates the final threshold by multiplying these values.

For most languages, the effective maximum character length defaults to **300 characters**, while Korean and Japanese inputs use a reduced limit of **120 characters** to account for tokenization differences.

### Hierarchical Text Splitting Algorithm

The chunking engine implements a multi-stage fallback strategy to preserve natural speech boundaries. As implemented in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) [lines 476-513](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js#L476-L513), the `chunkText` function executes the following steps:

1. **Paragraph segmentation** – Splits input on double-newline patterns (`\n\s*\n`) to preserve structural breaks.
2. **Sentence segmentation** – Divides paragraphs on sentence-ending punctuation (`.!?`) while protecting common abbreviations like *Dr.* and *e.g.*
3. **Chunk assembly** – Accumulates sentences until adding another would exceed the maximum length.
4. **Recursive fallback** – When a single sentence exceeds the limit, splits on commas; if still too long, splits on spaces as a final resort.

This algorithm guarantees that no chunk exceeds the configured `max_len`, ensuring the acoustic encoder receives only valid inputs.

## Language-Specific Implementations

Supertonic maintains consistent chunking behavior across all bindings through parallel implementations of the core algorithm.

### JavaScript (Web)

The Web implementation in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) exposes `chunkText` as a utility function. The `TextToSpeech.call` method [lines 71-94](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js#L71-L94) orchestrates the synthesis loop by first invoking `chunkText`, then processing each segment through the private `_infer` routine.

```javascript
import { loadTextToSpeech } from './helper.js';

// Load models and configuration
const { textToSpeech, cfgs } = await loadTextToSpeech('assets/onnx');

const longText = `Your very long text content here...`;

const style = await loadVoiceStyle(['assets/voice_styles/M1.json']);
const result = await textToSpeech.call(
  longText,
  'en',
  style,
  8,        // total steps
  1.05,     // speed factor
  0.3,      // silence between chunks (seconds)
);

```

### Rust

The Rust library implements identical logic in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) [lines 30-49](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs#L30-L49) through the `chunk_text` function. The `TextToSpeech::synthesize` method handles chunk iteration internally.

```rust
use supertonic::helper::{load_cfgs, load_voice_style, TextToSpeech};

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let cfgs = load_cfgs("assets/onnx").await?;
    let style = load_voice_style(vec!["assets/voice_styles/M1.json"]).await?;
    let tts = TextToSpeech::new(cfgs, style);

    let long_text = "Your very long text content here...";
    let wav = tts.synthesize(long_text, "en", 8, 1.05, 0.3).await?;
    
    Ok(())
}

```

### Swift

In the Swift binding, [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) [lines 34-66](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift#L34-L66) implements `chunkText` with the same hierarchical logic. The `TextToSpeech` class manages the synthesis loop and audio concatenation.

```swift
import Supertonic

let helper = Helper()
let cfgs = try await helper.loadCfgs(from: "onnx")
let style = try await helper.loadVoiceStyle(paths: ["voice_styles/M1.json"])
let tts = TextToSpeech(cfgs: cfgs, style: style)

let longText = "Your very long text content here..."
let wavData = try await tts.synthesize(
    text: longText,
    lang: "en",
    totalStep: 8,
    speed: 1.05,
    silenceDuration: 0.3
)

```

## The Synthesis Loop and Audio Stitching

After chunking, Supertonic processes each text segment through the complete inference pipeline: duration prediction, encoding, latent estimation, and vocoding. The system inserts a configurable silence interval (default **0.3 seconds**) between chunks to simulate natural pauses and prevent abrupt concatenation artifacts.

The `TextToSpeech.call` method in JavaScript demonstrates this workflow: it iterates over the chunk array, invokes `_infer` for each segment, prepends silence to all chunks after the first, and accumulates the waveforms into a single output buffer.

## Configuration Flags and CLI Options

Automatic chunking is enabled by default in all Supertonic interfaces. However, users can disable this behavior when processing pre-segmented text using the **`--batch`** flag in the CLI, which treats each supplied text string as a single synthesis unit regardless of length.

The silence duration between chunks is configurable via the `silenceDuration` parameter in the `call` or `synthesize` methods across all language bindings.

## Summary

- **Hierarchical chunking** preserves semantic boundaries by splitting on paragraphs, then sentences, then commas, then spaces.
- **Configuration-driven limits** derive from `base_chunk_size` and `chunk_compress_factor` in [`tts.json`](https://github.com/supertone-inc/supertonic/blob/main/tts.json), yielding effective limits of 300 characters (or 120 for Korean/Japanese).
- **Cross-language consistency** ensures identical chunk boundaries whether using JavaScript ([`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js)), Rust ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs)), or Swift ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)).
- **Automatic stitching** inserts 0.3 seconds of silence between chunks by default, with configurable intervals available.
- **CLI control** via the `--batch` flag allows disabling chunking for batch processing of pre-segmented inputs.

## Frequently Asked Questions

### How does Supertonic prevent memory overruns when synthesizing long texts?

Supertonic prevents memory overruns by enforcing a maximum character limit per inference based on the acoustic encoder's capacity. The `chunkText` algorithm in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) [lines 476-513](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js#L476-L513) splits input hierarchically before synthesis, ensuring each chunk fits within the configured `max_len` (300 characters for most languages, 120 for Korean/Japanese). This segmentation occurs before the neural network processes the text, keeping memory usage bounded regardless of input length.

### Can I customize the silence duration between chunks?

Yes, the silence duration between chunks is fully configurable through the `silenceDuration` parameter in the synthesis method. In JavaScript, pass the value to `textToSpeech.call()` as the sixth argument; in Rust, provide it to `tts.synthesize()`; in Swift, use the `silenceDuration` parameter of `tts.synthesize()`. The default value is **0.3 seconds**, but you can adjust this to create longer pauses between paragraphs or shorter gaps for faster speech.

### What happens if a single sentence exceeds the maximum chunk size?

When a single sentence exceeds the maximum chunk size, Supertonic implements recursive fallback splitting. First, it attempts to split the sentence on **comma boundaries** to preserve grammatical phrasing. If segments remain too long, it falls back to **space-separated words** as the final granularity. This ensures no text is rejected due to length while maintaining the largest possible semantic units, as implemented in the `chunkText` functions across [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js), [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs), and [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift).

### How do I disable automatic chunking for batch processing?

To disable automatic chunking, use the **`--batch`** flag when running Supertonic from the command line. This flag instructs the engine to treat each input text as a single synthesis unit, bypassing the `chunkText` algorithm entirely. Use this option only when you have pre-segmented your text externally or when processing individual sentences that are guaranteed to fit within the acoustic encoder's limits.