# Maximum Input Text Length and Text Chunking in Supertonic: Implementation Guide

> Discover Supertonic's maximum input text length limits and how it intelligently chunks long text into 300-character segments by default using sentence boundaries and a consistent algorithm.

- Repository: [Supertone Inc./supertonic](https://github.com/supertone-inc/supertonic)
- Tags: how-to-guide
- Published: 2026-06-14

---

**Supertonic limits input text chunks to 300 characters by default (120 for Korean/Japanese) and automatically segments long text on sentence boundaries using a consistent algorithm across all language bindings.**

The Supertonic library by Supertone Inc. processes arbitrarily long input text by enforcing strict maximum lengths and applying intelligent sentence-aware chunking. Understanding these limits is crucial for developers integrating the on-device text-to-speech models, as exceeding the thresholds triggers automatic segmentation that preserves linguistic coherence.

## Maximum Input Text Length Limits

### Default Character Limits

The default maximum input text length for a single chunk is **300 characters**. This limit applies to most languages and ensures that each processed segment remains within the bounds of the on-device model's context window. When text exceeds this threshold, the library automatically breaks it into smaller pieces before inference.

### Language-Specific Adjustments

For languages with denser scripts, the library reduces the limit to **120 characters**. This specifically applies to:
- **Korean** (language code: `ko`)
- **Japanese** (language code: `ja`)

This adjustment accounts for the higher information density per character in these scripts, as implemented in the `chunk_text` utility functions across all language bindings.

## How Text Chunking Works

### The Sentence-Aware Algorithm

All language implementations (Python, JavaScript, Swift, Rust, Go, and C++) share an identical chunking algorithm that respects sentence boundaries:

1. **Determine the per-language limit** (300 for most languages, 120 for Korean/Japanese)
2. **Iterate sentence-by-sentence**, identifying boundaries using punctuation markers (`.`, `!`, `?`)
3. **Accumulate sentences** into the current chunk, adding a trailing space after each
4. **Start a new chunk** when adding the next sentence would exceed `max_len`
5. **Return the list** of chunks, each guaranteed to be under the length limit

### Implementation Details Across Languages

The core logic resides in language-specific helper files that mirror the same algorithm:

- **Python**: [`py/helper.py`](https://github.com/supertone-inc/supertonic/blob/main/py/helper.py) lines 388-405 defines `chunk_text(text, max_len=300)`
- **JavaScript**: [`web/helper.js`](https://github.com/supertone-inc/supertonic/blob/main/web/helper.js) lines 522-546 implements `chunkText(text, maxLen = 300)`
- **Swift**: [`swift/Sources/Helper.swift`](https://github.com/supertone-inc/supertonic/blob/main/swift/Sources/Helper.swift) lines 334-351 provides `chunkText(_ text: String, maxLen: Int = 0)`
- **Rust**: [`rust/src/helper.rs`](https://github.com/supertone-inc/supertonic/blob/main/rust/src/helper.rs) lines 330-353 contains `pub fn chunk_text(text: &str, max_len: Option<usize>)`
- **Go**: [`go/helper.go`](https://github.com/supertone-inc/supertonic/blob/main/go/helper.go) lines 330-350 defines `func chunkText(text string, maxLen *int) []string`
- **C++**: [`cpp/helper.cpp`](https://github.com/supertone-inc/supertonic/blob/main/cpp/helper.cpp) implements the same sentence-aggregation approach in `chunkText`

## Practical Code Examples

### Python

```python
from supertonic.helper import chunk_text

long_text = "..."  # any length

chunks = chunk_text(long_text)          # default 300-char chunks

chunks_ko = chunk_text(long_text, max_len=120)  # explicit Korean/Japanese limit

```

### JavaScript

```javascript
import { chunkText } from "./helper.js";

const longText = "...";
const chunks = chunkText(longText);          // → array of ≤300-char strings
const chunksKo = chunkText(longText, 120);   // explicit 120-char limit

```

### Swift

```swift
import Supertonic

let longText = "..."
let chunks = Helper.chunkText(longText)          // default 300 chars
let chunksJA = Helper.chunkText(longText, maxLen: 120) // for Japanese

```

### Rust

```rust
use supertonic::helper::chunk_text;

let long_text = "...".to_string();
let chunks = chunk_text(&long_text, None);          // 300-char default
let chunks_ko = chunk_text(&long_text, Some(120));  // Korean/Japanese limit

```

## Summary

- Supertonic enforces a **300-character default limit** for text chunks, reduced to **120 characters** for Korean and Japanese
- The chunking algorithm operates **sentence-by-sentence**, splitting on punctuation to maintain linguistic coherence
- All six language bindings (Python, JavaScript, Swift, Rust, Go, C++) implement **identical chunking logic** in their respective helper modules
- The `chunk_text` (or `chunkText`) functions automatically handle segmentation, returning lists of strings that fit within model constraints

## Frequently Asked Questions

### What happens if my input text exceeds the maximum chunk length?

Supertonic automatically splits the text into multiple consecutive chunks at sentence boundaries. Each chunk is processed independently by the on-device models, and the results are concatenated to reconstruct the full audio output.

### Can I override the default 300-character limit?

Yes. While the library provides sensible defaults (300 characters for most languages, 120 for Korean/Japanese), you can pass an explicit `max_len` parameter to the `chunk_text` function in any language binding to adjust the limit for your specific use case.

### Why is the limit lower for Korean and Japanese text?

Korean and Japanese use denser scripts where a single character can represent entire syllables or complex concepts. The 120-character limit for these languages ensures that the semantic content per chunk remains comparable to the 300-character limit for alphabet-based languages, maintaining consistent model performance.

### Does chunking break words or only split at sentence boundaries?

The algorithm prioritizes sentence boundaries (ending with `.`, `!`, or `?`) to preserve natural language flow. It only breaks within a sentence if a single sentence exceeds the maximum length, though this edge case is rare for the 300-character threshold.