How to Handle Long Text Synthesis with Text Chunking in Supertonic

Supertonic automatically splits long input text into manageable chunks using a hierarchical algorithm that respects paragraph and sentence boundaries, then synthesizes each chunk separately and stitches them together with configurable silence intervals to produce seamless audio.

The supertone-inc/supertonic repository provides a neural text-to-speech engine capable of handling arbitrarily long passages without memory overruns. By implementing intelligent text chunking consistently across JavaScript, Rust, and Swift bindings, the library preserves prosodic continuity while respecting the acoustic encoder's token limits.

Understanding the Chunking Architecture

Supertonic's chunking system operates through a deterministic pipeline that balances semantic coherence with technical constraints. The architecture relies on configuration parameters that define hardware limits and runtime behavior.

Chunk Size Configuration

The maximum token length for a single inference is governed by two configuration values located in tts.json. The ae.base_chunk_size parameter defines the acoustic encoder's hard limit, while ttl.chunk_compress_factor acts as a multiplier to determine the effective chunk size used during synthesis. According to the source code in web/helper.js lines 73-77, the system calculates the final threshold by multiplying these values.

For most languages, the effective maximum character length defaults to 300 characters, while Korean and Japanese inputs use a reduced limit of 120 characters to account for tokenization differences.

Hierarchical Text Splitting Algorithm

The chunking engine implements a multi-stage fallback strategy to preserve natural speech boundaries. As implemented in web/helper.js lines 476-513, the chunkText function executes the following steps:

  1. Paragraph segmentation – Splits input on double-newline patterns (\n\s*\n) to preserve structural breaks.
  2. Sentence segmentation – Divides paragraphs on sentence-ending punctuation (.!?) while protecting common abbreviations like Dr. and e.g.
  3. Chunk assembly – Accumulates sentences until adding another would exceed the maximum length.
  4. Recursive fallback – When a single sentence exceeds the limit, splits on commas; if still too long, splits on spaces as a final resort.

This algorithm guarantees that no chunk exceeds the configured max_len, ensuring the acoustic encoder receives only valid inputs.

Language-Specific Implementations

Supertonic maintains consistent chunking behavior across all bindings through parallel implementations of the core algorithm.

JavaScript (Web)

The Web implementation in web/helper.js exposes chunkText as a utility function. The TextToSpeech.call method lines 71-94 orchestrates the synthesis loop by first invoking chunkText, then processing each segment through the private _infer routine.

import { loadTextToSpeech } from './helper.js';

// Load models and configuration
const { textToSpeech, cfgs } = await loadTextToSpeech('assets/onnx');

const longText = `Your very long text content here...`;

const style = await loadVoiceStyle(['assets/voice_styles/M1.json']);
const result = await textToSpeech.call(
  longText,
  'en',
  style,
  8,        // total steps
  1.05,     // speed factor
  0.3,      // silence between chunks (seconds)
);

Rust

The Rust library implements identical logic in rust/src/helper.rs lines 30-49 through the chunk_text function. The TextToSpeech::synthesize method handles chunk iteration internally.

use supertonic::helper::{load_cfgs, load_voice_style, TextToSpeech};

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    let cfgs = load_cfgs("assets/onnx").await?;
    let style = load_voice_style(vec!["assets/voice_styles/M1.json"]).await?;
    let tts = TextToSpeech::new(cfgs, style);

    let long_text = "Your very long text content here...";
    let wav = tts.synthesize(long_text, "en", 8, 1.05, 0.3).await?;
    
    Ok(())
}

Swift

In the Swift binding, swift/Sources/Helper.swift lines 34-66 implements chunkText with the same hierarchical logic. The TextToSpeech class manages the synthesis loop and audio concatenation.

import Supertonic

let helper = Helper()
let cfgs = try await helper.loadCfgs(from: "onnx")
let style = try await helper.loadVoiceStyle(paths: ["voice_styles/M1.json"])
let tts = TextToSpeech(cfgs: cfgs, style: style)

let longText = "Your very long text content here..."
let wavData = try await tts.synthesize(
    text: longText,
    lang: "en",
    totalStep: 8,
    speed: 1.05,
    silenceDuration: 0.3
)

The Synthesis Loop and Audio Stitching

After chunking, Supertonic processes each text segment through the complete inference pipeline: duration prediction, encoding, latent estimation, and vocoding. The system inserts a configurable silence interval (default 0.3 seconds) between chunks to simulate natural pauses and prevent abrupt concatenation artifacts.

The TextToSpeech.call method in JavaScript demonstrates this workflow: it iterates over the chunk array, invokes _infer for each segment, prepends silence to all chunks after the first, and accumulates the waveforms into a single output buffer.

Configuration Flags and CLI Options

Automatic chunking is enabled by default in all Supertonic interfaces. However, users can disable this behavior when processing pre-segmented text using the --batch flag in the CLI, which treats each supplied text string as a single synthesis unit regardless of length.

The silence duration between chunks is configurable via the silenceDuration parameter in the call or synthesize methods across all language bindings.

Summary

  • Hierarchical chunking preserves semantic boundaries by splitting on paragraphs, then sentences, then commas, then spaces.
  • Configuration-driven limits derive from base_chunk_size and chunk_compress_factor in tts.json, yielding effective limits of 300 characters (or 120 for Korean/Japanese).
  • Cross-language consistency ensures identical chunk boundaries whether using JavaScript (web/helper.js), Rust (rust/src/helper.rs), or Swift (swift/Sources/Helper.swift).
  • Automatic stitching inserts 0.3 seconds of silence between chunks by default, with configurable intervals available.
  • CLI control via the --batch flag allows disabling chunking for batch processing of pre-segmented inputs.

Frequently Asked Questions

How does Supertonic prevent memory overruns when synthesizing long texts?

Supertonic prevents memory overruns by enforcing a maximum character limit per inference based on the acoustic encoder's capacity. The chunkText algorithm in web/helper.js lines 476-513 splits input hierarchically before synthesis, ensuring each chunk fits within the configured max_len (300 characters for most languages, 120 for Korean/Japanese). This segmentation occurs before the neural network processes the text, keeping memory usage bounded regardless of input length.

Can I customize the silence duration between chunks?

Yes, the silence duration between chunks is fully configurable through the silenceDuration parameter in the synthesis method. In JavaScript, pass the value to textToSpeech.call() as the sixth argument; in Rust, provide it to tts.synthesize(); in Swift, use the silenceDuration parameter of tts.synthesize(). The default value is 0.3 seconds, but you can adjust this to create longer pauses between paragraphs or shorter gaps for faster speech.

What happens if a single sentence exceeds the maximum chunk size?

When a single sentence exceeds the maximum chunk size, Supertonic implements recursive fallback splitting. First, it attempts to split the sentence on comma boundaries to preserve grammatical phrasing. If segments remain too long, it falls back to space-separated words as the final granularity. This ensures no text is rejected due to length while maintaining the largest possible semantic units, as implemented in the chunkText functions across web/helper.js, rust/src/helper.rs, and swift/Sources/Helper.swift.

How do I disable automatic chunking for batch processing?

To disable automatic chunking, use the --batch flag when running Supertonic from the command line. This flag instructs the engine to treat each input text as a single synthesis unit, bypassing the chunkText algorithm entirely. Use this option only when you have pre-segmented your text externally or when processing individual sentences that are guaranteed to fit within the acoustic encoder's limits.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →