How Pocket‑TTS Handles Chunking and Sentence Splitting for Long Texts
Pocket‑TTS processes arbitrarily long inputs by hierarchically splitting text on primary sentence boundaries (periods, exclamation marks, question marks, and ellipses), falling back to secondary delimiters (commas, semicolons, and colons) for oversized segments, and greedily packing the resulting segments into token‑constrained chunks.
The kyutai-labs/pocket-tts repository implements a robust segmentation pipeline in pocket_tts/models/tts_model.py to ensure that text‑to‑speech generation never exceeds the model’s context window while preserving natural readability. The system automatically manages token budgets through a multi‑stage splitting algorithm that respects linguistic boundaries.
The Three‑Stage Splitting Pipeline
The core logic resides in the split_into_best_sentences function, which transforms raw text into a list of chunks ready for batched generation.
Text Pre‑processing
Before segmentation begins, the input string passes through prepare_text_prompt. This routine performs optional normalization, including space‑padding for very short inputs to improve acoustic quality and the optional removal of semicolons to prevent unintended pauses. This ensures the tokenizer receives clean, model‑friendly text.
Primary Sentence Boundary Detection
The pipeline first attempts to split the text into natural sentences using end‑of‑sentence punctuation. The process works as follows:
- The full text is tokenized using the model’s SentencePiece‑based tokenizer.
- The helper
_find_boundary_indices(lines 945–962) scans the token stream for indices corresponding to.,!,?, or…. - These indices feed into
_segments_from_boundaries(lines 965–975), which slices the token stream into discrete sentence segments.
This stage preserves the integrity of complete thoughts, ensuring that generation boundaries align with natural speech pauses.
Fallback Sub‑Splitting for Oversized Sentences
When a single sentence exceeds the user‑supplied max_tokens limit, the system activates a fallback mechanism (lines 998–1008):
- The function searches for secondary delimiters: commas (
,), semicolons (;), and colons (:). - If these delimiters produce multiple sub‑segments, the sentence is split at those points.
- If no suitable secondary boundaries exist, the original oversized sentence is retained and handled during the final chunk assembly.
This hierarchical approach prevents word skipping in run‑on sentences while attempting to maintain grammatical coherence.
Greedy Chunk Assembly Strategy
After establishing the segment hierarchy, the algorithm packs segments into final chunks using a greedy algorithm (lines 1023–1030):
- A new chunk initializes with the first available segment.
- Subsequent segments append to the current chunk while the cumulative token count remains ≤
max_tokens. - When adding a segment would exceed the limit, the current chunk closes and a new chunk begins with the overflowing segment.
Finally, each assembled chunk undergoes verification (lines 1034–1042). If a chunk still exceeds max_tokens—typically because a single word or token sequence is longer than the limit—the system logs a warning rather than failing silently, alerting users to potential truncation risks.
Working with the Splitting API
You can interact with the chunking system either directly through the internal utility or via the high‑level TTSModel interface.
Direct Usage of the Splitter
For debugging or custom preprocessing, import the splitting function directly from the model module:
from pocket_tts.models.tts_model import split_into_best_sentences
tokenizer = model.tokenizer # SentencePiece tokenizer instance
long_text = (
"In the beginning the universe was created. "
"This has made a lot of people very angry and been widely regarded as a bad move. "
"Furthermore, the restaurant at the end of the universe serves excellent food."
)
chunks = split_into_best_sentences(
tokenizer,
text_to_generate=long_text,
max_tokens=150, # Typical model context size
pad_with_spaces_for_short_inputs=False,
remove_semicolons=False,
)
for i, chunk in enumerate(chunks, 1):
token_count = len(tokenizer(chunk).tokens[0])
print(f"Chunk {i} ({token_count} tokens): {chunk}")
Public API Integration
In standard usage, the splitter runs automatically inside TTSModel.generate_audio_stream:
from pocket_tts import TTSModel
model = TTSModel()
audio_stream = model.generate_audio_stream(
text="A very long paragraph that exceeds the typical context window...",
max_generated_tokens=150, # Maps to internal max_tokens parameter
)
for audio_chunk in audio_stream:
# Process or play audio_chunk
pass
The audio stream yields chunks in the same order as the text segments produced by split_into_best_sentences, maintaining temporal alignment between the input text and generated speech.
Summary
- Primary segmentation occurs at sentence boundaries (
.,!,?,…) via_find_boundary_indicesto preserve natural speech patterns. - Secondary fallback splits oversized sentences on
,,;, and:to avoid exceeding token limits. - Greedy packing assembles segments into chunks that strictly respect the
max_tokensceiling. - Safety warnings log pathological cases where single tokens exceed the budget, as implemented in lines 1034–1042 of
pocket_tts/models/tts_model.py.
Frequently Asked Questions
What delimiters does Pocket‑TTS use for sentence splitting?
The system uses a two‑tier delimiter hierarchy. Primary splits occur on end‑of‑sentence punctuation—periods, exclamation marks, question marks, and ellipses. If a resulting segment still exceeds max_tokens, the code falls back to secondary delimiters: commas, semicolons, and colons. This logic is implemented in split_into_best_sentences within pocket_tts/models/tts_model.py.
How does Pocket‑TTS handle sentences that exceed the token limit?
When a sentence surpasses max_tokens, the algorithm attempts to break it at secondary delimiters (commas, semicolons, or colons). If successful, it creates smaller sub‑segments; if no such delimiters exist, the sentence remains intact and the greedy chunk assembler attempts to place it in its own chunk. If the sentence itself is longer than max_tokens, a warning is logged during the final verification stage.
Can I customize the pre‑processing options like semicolon removal?
Yes. The split_into_best_sentences function exposes boolean parameters pad_with_spaces_for_short_inputs and remove_semicolons that control the prepare_text_prompt behavior. When using the high‑level TTSModel API, these options are typically managed internally, but you can access the low‑level splitter directly for granular control over text normalization.
Where are the unit tests for the splitting logic?
The test suite in tests/test_split_sentences.py (lines 1–136) documents expected behavior for normal sentence splitting, comma‑only fallback scenarios, and semicolon/colon handling. These tests verify that the chunking algorithm correctly respects token budgets across various linguistic edge cases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →