# How the chunk_text() Function Handles Long Text and Sentence Boundaries in Supertonic

> Explore how Supertonic's chunk_text() function splits long text into speech-ready fragments, respecting sentence boundaries and configurable maximum lengths for natural output.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: deep-dive
- Published: 2026-06-14

---

**The `chunk_text()` function splits arbitrary-length strings into speech-ready fragments by first separating paragraphs, then detecting sentences using abbreviation-aware regex patterns, and finally assembling chunks that respect a configurable maximum length while preserving natural boundaries.**

The `chunk_text()` utility is implemented across the Supertonic repository (supertone-inc/supertonic) to preprocess text for text-to-speech (TTS) inference. Available in JavaScript for the web frontend, Swift for native iOS/macOS, and Rust for the backend, this function ensures that long inputs are divided into manageable segments without breaking mid-sentence or inside common abbreviations.

## Algorithm Stages

The chunking process operates in three distinct stages to maintain natural speech flow.

### Paragraph Separation

The input string is first trimmed and split on two or more newline characters (`\n\s*\n`). This preserves natural paragraph breaks as boundaries for the subsequent sentence detection phase.

In [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js), this initial split occurs at line 81, creating an array of paragraph strings that still contain their internal sentence structure.

### Sentence Detection

Each paragraph is divided at sentence-ending punctuation (`.`, `?`, `!`) followed by whitespace. The implementation uses a regular expression that specifically excludes common abbreviations such as *Mr.*, *Dr.*, and *e.g.*, as well as single-letter capital abbreviations, preventing premature splits.

In the JavaScript implementation ([`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) line 92), this regex handles the filtering logic directly. The Swift version in [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) (line 68) uses a `splitSentences` helper that matches the pattern `([.!?])\s+` and then validates each token against an abbreviation list to avoid false boundaries.

### Chunk Assembly

Sentences are concatenated into a temporary buffer (`currentChunk`). When adding the next sentence would exceed the configured `maxLen` parameter (default **300 characters**), the buffer is emitted as a complete chunk and a new buffer is started.

This logic appears in [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) at line 96. The Swift implementation follows the same principle, with an additional safeguard: if a single sentence exceeds `maxLen`, it is first split on commas, then on spaces, ensuring no chunk exceeds the limit.

## Handling Long Sentences

When individual sentences exceed the maximum length limit, the Swift implementation ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) line 86) applies a hierarchical fallback strategy. It first attempts to split on commas, and if the resulting segments are still too long, it falls back to splitting on spaces. This guarantees that every output chunk adheres to the `maxLen` constraint while minimizing semantic disruption.

The function also guarantees a non-empty return array, emitting `[""]` when the trimmed input is empty, ensuring consistent downstream handling.

## Code Examples

### JavaScript (Web Frontend)

Use the `chunkText` function from [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) to prepare long text for TTS inference:

```javascript
import { chunkText } from './helper.js';

const longText = `Lorem ipsum dolor sit amet. Consectetur adipiscing elit. Sed do eiusmod tempor incididunt.`;
const chunks = chunkText(longText, 300); // 300 is the default max length

console.log(chunks);
// → Array of strings, each ≤ 300 characters

```

### Swift (iOS/macOS)

The native implementation provides identical behavior for Apple platforms:

```swift
let longText = "Lorem ipsum dolor sit amet. Consectetur adipiscing elit."
let chunks = chunkText(longText, maxLen: 300)

print(chunks)
// → [String] where each element ≤ 300 characters

```

Both implementations produce identical chunking behavior, which the TTS pipeline processes sequentially, inserting short silences between chunks for natural speech flow. The Rust backend in [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) provides a `chunk_text` function that mirrors this logic for server-side processing.

## Summary

- **Three-stage pipeline**: Paragraph separation → sentence detection → chunk assembly
- **Abbreviation-aware**: Regex patterns exclude *Mr.*, *Dr.*, *e.g.*, and single-letter capitals to prevent false sentence boundaries
- **Configurable limits**: Default `maxLen` of 300 characters, with hierarchical fallback splitting in Swift for oversized sentences
- **Cross-platform consistency**: Identical logic implemented in JavaScript ([`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js)), Swift ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)), and Rust ([`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs))
- **Non-empty guarantee**: Returns `[""]` for empty trimmed input to ensure predictable downstream processing

## Frequently Asked Questions

### What is the default maximum chunk length in chunk_text()?

The default `maxLen` parameter is **300 characters**. This value balances TTS model sequence constraints with natural speech phrasing, though callers can configure this limit based on their specific model requirements.

### How does chunk_text() prevent splitting inside abbreviations?

The function uses a regular expression that identifies sentence-ending punctuation (`.`, `?`, `!`) followed by whitespace, but excludes tokens matching common abbreviations like *Mr.*, *Dr.*, and *e.g.*, as well as single-letter capital abbreviations. In Swift, the `splitSentences` helper validates each potential boundary against an abbreviation list before accepting the split.

### What happens if a single sentence exceeds the maximum length?

In the Swift implementation ([`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift)), the function applies a hierarchical fallback: it first attempts to split the long sentence on commas, and if segments remain too large, it splits on spaces. The JavaScript implementation follows similar logic to ensure no chunk exceeds `maxLen` while preserving readability.

### Does chunk_text() return an empty array for empty input?

No. According to the source code in both [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) and [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift), the function guarantees a non-empty return by emitting `[""]` (an array containing a single empty string) when the trimmed input is empty, preventing null reference errors in downstream TTS processing.