How the chunk_text() Function Handles Long Text and Sentence Boundaries in Supertonic
The chunk_text() function splits arbitrary-length strings into speech-ready fragments by first separating paragraphs, then detecting sentences using abbreviation-aware regex patterns, and finally assembling chunks that respect a configurable maximum length while preserving natural boundaries.
The chunk_text() utility is implemented across the Supertonic repository (supertone-inc/supertonic) to preprocess text for text-to-speech (TTS) inference. Available in JavaScript for the web frontend, Swift for native iOS/macOS, and Rust for the backend, this function ensures that long inputs are divided into manageable segments without breaking mid-sentence or inside common abbreviations.
Algorithm Stages
The chunking process operates in three distinct stages to maintain natural speech flow.
Paragraph Separation
The input string is first trimmed and split on two or more newline characters (\n\s*\n). This preserves natural paragraph breaks as boundaries for the subsequent sentence detection phase.
In web/helper.js, this initial split occurs at line 81, creating an array of paragraph strings that still contain their internal sentence structure.
Sentence Detection
Each paragraph is divided at sentence-ending punctuation (., ?, !) followed by whitespace. The implementation uses a regular expression that specifically excludes common abbreviations such as Mr., Dr., and e.g., as well as single-letter capital abbreviations, preventing premature splits.
In the JavaScript implementation (web/helper.js line 92), this regex handles the filtering logic directly. The Swift version in swift/Sources/Helper.swift (line 68) uses a splitSentences helper that matches the pattern ([.!?])\s+ and then validates each token against an abbreviation list to avoid false boundaries.
Chunk Assembly
Sentences are concatenated into a temporary buffer (currentChunk). When adding the next sentence would exceed the configured maxLen parameter (default 300 characters), the buffer is emitted as a complete chunk and a new buffer is started.
This logic appears in web/helper.js at line 96. The Swift implementation follows the same principle, with an additional safeguard: if a single sentence exceeds maxLen, it is first split on commas, then on spaces, ensuring no chunk exceeds the limit.
Handling Long Sentences
When individual sentences exceed the maximum length limit, the Swift implementation (swift/Sources/Helper.swift line 86) applies a hierarchical fallback strategy. It first attempts to split on commas, and if the resulting segments are still too long, it falls back to splitting on spaces. This guarantees that every output chunk adheres to the maxLen constraint while minimizing semantic disruption.
The function also guarantees a non-empty return array, emitting [""] when the trimmed input is empty, ensuring consistent downstream handling.
Code Examples
JavaScript (Web Frontend)
Use the chunkText function from web/helper.js to prepare long text for TTS inference:
import { chunkText } from './helper.js';
const longText = `Lorem ipsum dolor sit amet. Consectetur adipiscing elit. Sed do eiusmod tempor incididunt.`;
const chunks = chunkText(longText, 300); // 300 is the default max length
console.log(chunks);
// → Array of strings, each ≤ 300 characters
Swift (iOS/macOS)
The native implementation provides identical behavior for Apple platforms:
let longText = "Lorem ipsum dolor sit amet. Consectetur adipiscing elit."
let chunks = chunkText(longText, maxLen: 300)
print(chunks)
// → [String] where each element ≤ 300 characters
Both implementations produce identical chunking behavior, which the TTS pipeline processes sequentially, inserting short silences between chunks for natural speech flow. The Rust backend in rust/src/helper.rs provides a chunk_text function that mirrors this logic for server-side processing.
Summary
- Three-stage pipeline: Paragraph separation → sentence detection → chunk assembly
- Abbreviation-aware: Regex patterns exclude Mr., Dr., e.g., and single-letter capitals to prevent false sentence boundaries
- Configurable limits: Default
maxLenof 300 characters, with hierarchical fallback splitting in Swift for oversized sentences - Cross-platform consistency: Identical logic implemented in JavaScript (
web/helper.js), Swift (swift/Sources/Helper.swift), and Rust (rust/src/helper.rs) - Non-empty guarantee: Returns
[""]for empty trimmed input to ensure predictable downstream processing
Frequently Asked Questions
What is the default maximum chunk length in chunk_text()?
The default maxLen parameter is 300 characters. This value balances TTS model sequence constraints with natural speech phrasing, though callers can configure this limit based on their specific model requirements.
How does chunk_text() prevent splitting inside abbreviations?
The function uses a regular expression that identifies sentence-ending punctuation (., ?, !) followed by whitespace, but excludes tokens matching common abbreviations like Mr., Dr., and e.g., as well as single-letter capital abbreviations. In Swift, the splitSentences helper validates each potential boundary against an abbreviation list before accepting the split.
What happens if a single sentence exceeds the maximum length?
In the Swift implementation (swift/Sources/Helper.swift), the function applies a hierarchical fallback: it first attempts to split the long sentence on commas, and if segments remain too large, it splits on spaces. The JavaScript implementation follows similar logic to ensure no chunk exceeds maxLen while preserving readability.
Does chunk_text() return an empty array for empty input?
No. According to the source code in both web/helper.js and swift/Sources/Helper.swift, the function guarantees a non-empty return by emitting [""] (an array containing a single empty string) when the trimmed input is empty, preventing null reference errors in downstream TTS processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →