How Pocket‑TTS Handles Chunking and Sentence Splitting for Long Texts

Pocket‑TTS processes arbitrarily long inputs by hierarchically splitting text on primary sentence boundaries (periods, exclamation marks, question marks, and ellipses), falling back to secondary delimiters (commas, semicolons, and colons) for oversized segments, and greedily packing the resulting segments into token‑constrained chunks.

The kyutai-labs/pocket-tts repository implements a robust segmentation pipeline in pocket_tts/models/tts_model.py to ensure that text‑to‑speech generation never exceeds the model’s context window while preserving natural readability. The system automatically manages token budgets through a multi‑stage splitting algorithm that respects linguistic boundaries.

The Three‑Stage Splitting Pipeline

The core logic resides in the split_into_best_sentences function, which transforms raw text into a list of chunks ready for batched generation.

Text Pre‑processing

Before segmentation begins, the input string passes through prepare_text_prompt. This routine performs optional normalization, including space‑padding for very short inputs to improve acoustic quality and the optional removal of semicolons to prevent unintended pauses. This ensures the tokenizer receives clean, model‑friendly text.

Primary Sentence Boundary Detection

The pipeline first attempts to split the text into natural sentences using end‑of‑sentence punctuation. The process works as follows:

  • The full text is tokenized using the model’s SentencePiece‑based tokenizer.
  • The helper _find_boundary_indices (lines 945–962) scans the token stream for indices corresponding to ., !, ?, or ….
  • These indices feed into _segments_from_boundaries (lines 965–975), which slices the token stream into discrete sentence segments.

This stage preserves the integrity of complete thoughts, ensuring that generation boundaries align with natural speech pauses.

Fallback Sub‑Splitting for Oversized Sentences

When a single sentence exceeds the user‑supplied max_tokens limit, the system activates a fallback mechanism (lines 998–1008):

  • The function searches for secondary delimiters: commas (,), semicolons (;), and colons (:).
  • If these delimiters produce multiple sub‑segments, the sentence is split at those points.
  • If no suitable secondary boundaries exist, the original oversized sentence is retained and handled during the final chunk assembly.

This hierarchical approach prevents word skipping in run‑on sentences while attempting to maintain grammatical coherence.

Greedy Chunk Assembly Strategy

After establishing the segment hierarchy, the algorithm packs segments into final chunks using a greedy algorithm (lines 1023–1030):

  1. A new chunk initializes with the first available segment.
  2. Subsequent segments append to the current chunk while the cumulative token count remains ≤ max_tokens.
  3. When adding a segment would exceed the limit, the current chunk closes and a new chunk begins with the overflowing segment.

Finally, each assembled chunk undergoes verification (lines 1034–1042). If a chunk still exceeds max_tokens—typically because a single word or token sequence is longer than the limit—the system logs a warning rather than failing silently, alerting users to potential truncation risks.

Working with the Splitting API

You can interact with the chunking system either directly through the internal utility or via the high‑level TTSModel interface.

Direct Usage of the Splitter

For debugging or custom preprocessing, import the splitting function directly from the model module:

from pocket_tts.models.tts_model import split_into_best_sentences

tokenizer = model.tokenizer  # SentencePiece tokenizer instance

long_text = (
    "In the beginning the universe was created. "
    "This has made a lot of people very angry and been widely regarded as a bad move. "
    "Furthermore, the restaurant at the end of the universe serves excellent food."
)

chunks = split_into_best_sentences(
    tokenizer,
    text_to_generate=long_text,
    max_tokens=150,               # Typical model context size

    pad_with_spaces_for_short_inputs=False,
    remove_semicolons=False,
)

for i, chunk in enumerate(chunks, 1):
    token_count = len(tokenizer(chunk).tokens[0])
    print(f"Chunk {i} ({token_count} tokens): {chunk}")

Public API Integration

In standard usage, the splitter runs automatically inside TTSModel.generate_audio_stream:

from pocket_tts import TTSModel

model = TTSModel()
audio_stream = model.generate_audio_stream(
    text="A very long paragraph that exceeds the typical context window...",
    max_generated_tokens=150,  # Maps to internal max_tokens parameter

)

for audio_chunk in audio_stream:
    # Process or play audio_chunk

    pass

The audio stream yields chunks in the same order as the text segments produced by split_into_best_sentences, maintaining temporal alignment between the input text and generated speech.

Summary

  • Primary segmentation occurs at sentence boundaries (., !, ?, …) via _find_boundary_indices to preserve natural speech patterns.
  • Secondary fallback splits oversized sentences on ,, ;, and : to avoid exceeding token limits.
  • Greedy packing assembles segments into chunks that strictly respect the max_tokens ceiling.
  • Safety warnings log pathological cases where single tokens exceed the budget, as implemented in lines 1034–1042 of pocket_tts/models/tts_model.py.

Frequently Asked Questions

What delimiters does Pocket‑TTS use for sentence splitting?

The system uses a two‑tier delimiter hierarchy. Primary splits occur on end‑of‑sentence punctuation—periods, exclamation marks, question marks, and ellipses. If a resulting segment still exceeds max_tokens, the code falls back to secondary delimiters: commas, semicolons, and colons. This logic is implemented in split_into_best_sentences within pocket_tts/models/tts_model.py.

How does Pocket‑TTS handle sentences that exceed the token limit?

When a sentence surpasses max_tokens, the algorithm attempts to break it at secondary delimiters (commas, semicolons, or colons). If successful, it creates smaller sub‑segments; if no such delimiters exist, the sentence remains intact and the greedy chunk assembler attempts to place it in its own chunk. If the sentence itself is longer than max_tokens, a warning is logged during the final verification stage.

Can I customize the pre‑processing options like semicolon removal?

Yes. The split_into_best_sentences function exposes boolean parameters pad_with_spaces_for_short_inputs and remove_semicolons that control the prepare_text_prompt behavior. When using the high‑level TTSModel API, these options are typically managed internally, but you can access the low‑level splitter directly for granular control over text normalization.

Where are the unit tests for the splitting logic?

The test suite in tests/test_split_sentences.py (lines 1–136) documents expected behavior for normal sentence splitting, comma‑only fallback scenarios, and semicolon/colon handling. These tests verify that the chunking algorithm correctly respects token budgets across various linguistic edge cases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →