Supertonic Text Chunking Algorithm: How It Splits Text for TTS

Supertonic splits text by first dividing on paragraph boundaries, then breaking sentences using abbreviation-aware regex patterns, and finally assembling chunks that never exceed the 300-character model limit.

The supertone-inc/supertonic repository implements a robust text preprocessing pipeline that ensures long inputs fit within the ONNX TTS model's constraints. This article explains the Supertonic text chunking algorithm as implemented in both Python and Rust, detailing how it preserves semantic boundaries while enforcing strict length limits.

The Two-Stage Splitting Pipeline

Supertonic processes long text through a five-step pipeline designed to respect the model's maximum token length of approximately 300 characters. Both the Python (py/helper.py) and Rust (rust/src/helper.rs) implementations follow identical logical steps, differing only in language-specific regex capabilities.

The workflow proceeds as follows:

  1. Trim and validate – Leading and trailing whitespace is removed; empty strings immediately return a single empty chunk.
  2. Paragraph detection – The algorithm identifies paragraph boundaries using the pattern \n\s*\n+ (two or more line-feeds with optional whitespace).
  3. Sentence tokenization – Each paragraph is split into sentences using punctuation-aware logic that excludes common abbreviations.
  4. Chunk assembly – Sentences are appended to a running buffer until adding the next sentence would exceed max_len, at which point the buffer is stored as a complete chunk.
  5. Fallback fragmentation – If a single sentence exceeds the limit, the system falls back to splitting on commas, then spaces, ensuring no chunk ever exceeds the hard limit.

Sentence Splitting Implementation

The sentence boundary detection differs between languages due to regex engine capabilities, but both versions protect abbreviations like "Dr.", "Inc.", and "Ph.D." from being treated as sentence endings.

Python Implementation (py/helper.py)

In py/helper.py (lines 388-429), the Python implementation leverages the re module's negative lookbehind assertions to exclude abbreviations:

pattern = r"(?<!Mr\.)(?<!Mrs\.)(?<!Ms\.)(?<!Dr\.)(?<!Prof\.)(?<!Sr\.)(?<!Jr\.)(?<!Ph\.D\.)(?<!etc\.)(?<!e\.g\.)(?<!i\.e\.)(?<!vs\.)(?<!Inc\.)(?<!Ltd\.)(?<!Co\.)(?<!Corp\.)(?<!St\.)(?<!Ave\.)(?<!Blvd\.)(?<!\b[A-Z]\.)(?<=[.!?])\s+"
sentences = re.split(pattern, paragraph)

The pattern (?<=[.!?])\s+ identifies actual sentence boundaries (punctuation followed by whitespace), while the cascade of (?<!...) negative lookbehinds prevents splitting after honorifics, corporate suffixes, and single-letter abbreviations.

Rust Implementation (rust/src/helper.rs)

The Rust version in rust/src/helper.rs (lines 30-57) cannot use lookbehind assertions because the regex crate lacks support for this feature. Instead, it uses a two-step process:

let re = Regex::new(r"([.!?])\s+").unwrap();

After matching potential boundaries, the code checks the preceding text against a hard-coded ABBREVIATIONS list (defined in lines 24-28). If the split point matches an abbreviation, the algorithm ignores the match and continues searching; otherwise, it slices the text at that punctuation mark. This produces identical results to the Python implementation despite the regex limitation.

Chunk Assembly and Fallback Logic

Once split into sentences, the chunk_text function assembles final chunks by iterating through the sentence list. It maintains a running buffer string, appending sentences until the next addition would violate the max_len constraint (defaulting to 300 characters).

When a sentence itself exceeds max_len, Supertonic applies a hierarchical fallback strategy:

  • First attempt: Split the oversize sentence on comma characters to preserve clause boundaries.
  • Final resort: If comma splitting still produces fragments exceeding the limit, split on whitespace.
  • Guarantee: This ensures every output chunk respects the model's input constraints while minimizing semantic fragmentation.

Practical Examples

Python Usage

The chunk_text function in py/helper.py integrates directly into the inference pipeline:

from supertonic.py.helper import chunk_text

long_paragraph = """
Supertonic is a multi‑platform Text‑to‑Speech system. It supports many languages,
including English, Korean, and Japanese. Dr. Lee, the lead researcher, says:
“It's a breakthrough!”  The model works on CPU and GPU.
"""

chunks = chunk_text(long_paragraph, max_len=120)
for i, c in enumerate(chunks, 1):
    print(f"Chunk {i}: {c}")

This produces semantically coherent chunks that honor abbreviation boundaries:


Chunk 1: Supertonic is a multi‑platform Text‑to‑Speech system. It supports many languages, including English, Korean, and Japanese.
Chunk 2: Dr. Lee, the lead researcher, says: “It's a breakthrough!”
Chunk 3: The model works on CPU and GPU.

Rust Usage

The Rust implementation exposes an identical API through rust/src/helper.rs:

use supertonic::helper::chunk_text;

fn main() {
    let text = "Supertonic is a multi‑platform Text‑to‑Speech system. It supports many languages, including English, Korean, and Japanese. Dr. Lee, the lead researcher, says: \"It's a breakthrough!\" The model works on CPU and GPU.";
    let chunks = chunk_text(text, Some(120));
    for (i, c) in chunks.iter().enumerate() {
        println!("Chunk {}: {}", i + 1, c);
    }
}

The MAX_CHUNK_LENGTH constant ensures the Rust version defaults to the same 300-character limit as the Python implementation.

Summary

  • Supertonic implements a two-stage splitting pipeline (paragraph → sentence → chunk) to respect the ONNX TTS model's ~300 character input limit.
  • The abbreviation-aware regex prevents false splits after titles (Dr., Mr.) and corporate suffixes (Inc., Ltd.), with Python using negative lookbehind and Rust using a hard-coded exclusion list.
  • A hierarchical fallback splits oversized sentences first on commas, then on spaces, guaranteeing no chunk exceeds the maximum length.
  • Both py/helper.py and rust/src/helper.rs implement the chunk_text function with identical logic but language-appropriate regex techniques.

Frequently Asked Questions

What is the maximum chunk size in Supertonic?

Supertonic defaults to a maximum chunk length of approximately 300 characters, defined by the MAX_CHUNK_LENGTH constant in rust/src/helper.rs and applied consistently across both the Python and Rust implementations. This limit aligns with the ONNX TTS model's maximum token capacity for single inference steps.

How does Supertonic handle abbreviations like "Dr." or "Inc."?

The algorithm uses abbreviation-aware boundary detection to prevent splitting sentences after common abbreviations. In Python, this is implemented via negative lookbehind assertions in the regex pattern at lines 388-429 of py/helper.py. In Rust, the code checks matches against an ABBREVIATIONS list (lines 24-28 of rust/src/helper.rs) to filter out false positives before treating punctuation as a sentence boundary.

Why does the Rust implementation differ from Python?

The Rust implementation requires a manual abbreviation-checking step because the regex crate does not support lookbehind assertions, unlike Python's re module. According to the source code in rust/src/helper.rs (lines 30-57), Rust first matches ([.!?])\s+ patterns, then validates that the preceding text does not appear in the abbreviations list, achieving identical results to Python's single-regex approach.

What happens if a single sentence exceeds the maximum length?

Supertonic applies a cascading fallback strategy: if a sentence exceeds max_len, the system first attempts to split on comma characters to preserve clause structure. If segments remain too long, it falls back to splitting on whitespace. This ensures no chunk ever exceeds the model limit while maintaining the largest possible semantic units, as implemented in the buffer assembly logic of chunk_text.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →