How Supertonic's Automatic Text Chunking Works for Long-Form Speech Synthesis
Supertonic automatically splits long input texts into chunks of 300 characters (or 120 for Korean/Japanese) using a hierarchical paragraph → sentence → comma → word algorithm, then concatenates the synthesized audio with 0.3-second pauses to produce seamless long-form speech.
The supertone-inc/supertonic repository implements automatic text chunking to enable arbitrary-length text-to-speech synthesis without manual preprocessing. This system runs by default during single-text inference across all language bindings, ensuring that the underlying neural TTS model never receives inputs exceeding safe memory limits while preserving natural prosody.
When and How the Chunker Activates
Entry Point and Default Behavior
The chunking mechanism triggers automatically inside TextToSpeech::call in the Rust implementation (with equivalent functions in Python, Swift, and Java bindings). When you invoke inference on a long text, the system determines the maximum chunk length based on the language code and immediately passes your input to the chunk_text helper:
let max_len = if lang == "ko" || lang == "ja" { 120 } else { 300 };
let chunks = chunk_text(text, Some(max_len));
Source: TextToSpeech::call
The function returns a Vec<String> where each element respects the maximum length constraint. Notably, automatic chunking is disabled in --batch mode, allowing advanced users to manage chunking manually when processing multiple texts simultaneously.
Language-Specific Length Limits
Supertonic applies different character limits based on linguistic density:
- 300 characters for most languages (English, Spanish, French, etc.)
- 120 characters for Korean (
ko) and Japanese (ja)
These limits ensure that the TTS model receives manageable tensor sizes while maintaining semantic coherence within each chunk.
The Hierarchical Chunking Algorithm
The core logic resides in chunk_text (lines 330-495 in rust/src/helper.rs). The algorithm prioritizes linguistic boundaries to minimize awkward mid-sentence breaks.
Paragraph and Sentence Boundaries
The chunker first attempts to preserve structural integrity through a multi-tiered approach:
- Trim and validate – Removes leading/trailing whitespace and handles empty inputs defensively
- Paragraph split – Splits on double-newline (
\n\n) using regex to identify natural document breaks - Per-paragraph processing – If a paragraph fits within
max_len, it becomes a single chunk; otherwise, the algorithm invokessplit_sentencesto break it into sentences - Sentence accumulation – Appends sentences to a buffer until adding the next sentence would exceed
max_len, then flushes the buffer as a completed chunk
This accumulation strategy minimizes the total number of audio segments while strictly respecting the character limit.
Fallback Strategies for Oversized Segments
When individual sentences exceed the maximum length—common in legal or technical texts—the algorithm implements cascading fallbacks:
- Comma splitting – First attempts to break the long sentence on comma boundaries
- Word splitting – If comma segments remain too large, falls back to whitespace (word) boundaries
- Word-chunk construction – Builds chunks that respect
max_lenwithout breaking words
This ensures that even pathological inputs (e.g., thousand-character sentences) process correctly without crashing the inference engine.
Sentence Boundary Detection Logic
The split_sentences helper (lines 52-97 in rust/src/helper.rs) uses regex ([.!?])\s+ to identify potential sentence terminals, then validates against an abbreviation list (e.g., "Dr.", "e.g.", "vs.") to avoid false positives. This simple but effective approach maintains linguistic coherence without requiring heavy NLP dependencies.
Audio Concatenation and Synthesis Flow
After chunking, the system processes each segment through _infer and concatenates the resulting waveforms. The implementation inserts configurable silence between chunks to simulate natural breathing pauses:
if i == 0 {
wav_cat.extend_from_slice(wav_chunk);
dur_cat = dur;
} else {
let silence_len = (silence_duration * self.sample_rate as f32) as usize;
let silence = vec![0.0f32; silence_len];
wav_cat.extend_from_slice(&silence);
wav_cat.extend_from_slice(wav_chunk);
dur_cat += silence_duration + dur;
}
Source: chunk handling in call
The default silence_duration of 0.3 seconds provides perceptual continuity without noticeable gaps. Users can adjust this parameter to create longer pauses between paragraphs or maintain tighter pacing for continuous narration.
Cross-Platform Consistency
The identical chunking logic appears across all supported language bindings, ensuring deterministic behavior regardless of your SDK choice:
| Language | File | Key Functions |
|---|---|---|
| Rust | rust/src/helper.rs |
chunk_text (L330-L495), split_sentences (L52-L97) |
| Python | py/helper.py |
chunk_text, TextToSpeech.call |
| Swift | swift/Sources/Helper.swift |
chunkText(_:,maxLen:) |
| Java | java/Helper.java |
chunkText, call |
| Go | go/helper.go |
ChunkText, Call |
| Web | web/helper.js |
chunkText (L476-L512) |
All implementations follow the same paragraph → sentence → comma → word hierarchy, ensuring that automatic text chunking works uniformly across mobile, desktop, and server deployments.
Implementation Examples
The following examples demonstrate how automatic chunking operates transparently during inference. You never manually split text; the engine handles chunking internally.
Rust Implementation
use supertonic::rust::helper::{load_text_to_speech, load_voice_style};
fn main() -> anyhow::Result<()> {
let tts = load_text_to_speech("models/onnx", false)?;
let style = load_voice_style(&["models/style1.json"], true)?;
let long_text = "This is a very long text that will be automatically split \
into multiple chunks. The system will process each chunk \
separately and then concatenate them together with natural \
pauses between segments.";
// Automatic chunking occurs here
let (wav, total_sec) = tts.call(
long_text,
"en",
&style,
30, // total_step
1.0, // speed
0.3, // silence_duration (seconds)
)?;
supertonic::rust::helper::write_wav_file("output.wav", &wav, tts.sample_rate)?;
println!("Generated {} seconds of audio", total_sec);
Ok(())
}
Python Implementation
from supertonic import TextToSpeech, load_voice_style
tts = TextToSpeech.load("models/onnx")
style = load_voice_style(["models/style1.json"])
long_text = """This is a very long text that will be automatically split
into multiple chunks. The system will process each chunk separately
and then concatenate them together with natural pauses between segments."""
# Automatic chunking happens inside call()
wav, duration = tts.call(
text=long_text,
lang="en",
style=style,
total_step=30,
speed=1.0,
silence_duration=0.3,
)
tts.save_wav("output.wav", wav)
print(f"Generated {duration:.2f}s of speech")
Swift Implementation
let tts = try TextToSpeech.load(onnxDir: "Models/ONNX")
let style = try loadVoiceStyle(paths: ["Models/style1.json"])
let longText = """
This is a very long text that will be automatically split into multiple chunks...
"""
let (wav, totalSec) = try tts.call(
text: longText,
lang: "en",
style: style,
totalStep: 30,
speed: 1.0,
silenceDuration: 0.3
)
try tts.writeWavFile(to: "output.wav", audioData: wav)
Summary
- Automatic activation – The chunker runs by default in
TextToSpeech::callfor single-text inference, applying 300-character limits (120 for Korean/Japanese). - Hierarchical splitting – Prioritizes paragraph boundaries, then sentences, then commas, then words to maintain linguistic coherence.
- Memory safety – Prevents out-of-memory errors by ensuring the TTS model never processes inputs exceeding safe tensor sizes.
- Seamless output – Concatenates audio chunks with configurable silence (default 0.3s) to create natural-sounding long-form speech.
- Cross-platform parity – Identical logic across Rust, Python, Swift, Java, Go, and JavaScript implementations.
Frequently Asked Questions
What is the maximum chunk size in Supertonic?
Supertonic uses a default maximum of 300 characters for most languages, but reduces this to 120 characters for Korean and Japanese (ko and ja language codes). These limits are hardcoded in TextToSpeech::call within rust/src/helper.rs to optimize memory usage for character-dense languages.
How does Supertonic handle sentences longer than the chunk limit?
When a single sentence exceeds the maximum length, the algorithm first attempts to split on comma boundaries. If comma segments remain too long, it falls back to word-level splitting at whitespace characters. This ensures the TTS model never receives oversized inputs while avoiding mid-word breaks that would degrade speech quality.
Is automatic text chunking available in all language bindings?
Yes. The same chunking algorithm is implemented across all Supertonic SDKs including Rust, Python, Swift, Java, Go, and JavaScript. Each binding contains equivalent chunk_text (or chunkText) functions that follow the identical paragraph → sentence → comma → word hierarchy, ensuring consistent behavior across mobile, desktop, and web deployments.
Can I disable automatic chunking for batch processing?
Yes. Automatic text chunking is automatically disabled when using --batch mode or batch processing APIs. This allows you to manually control chunking strategies when processing multiple texts simultaneously, or to implement custom splitting logic optimized for your specific use case while leveraging the engine's inference capabilities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →