How Pocket-TTS Handles Long Text Inputs with `split_into_best_sentences()`
Pocket-TTS automatically partitions long text into token-safe chunks by detecting sentence boundaries and applying intelligent fallback splitting, ensuring no input exceeds the model’s max_tokens limit while preserving natural phrasing.
The split_into_best_sentences() function in pocket_tts/models/tts_model.py serves as the gatekeeper for the generation pipeline, preventing context-window overflows when synthesizing arbitrary-length documents. This method implements a greedy, multi-stage segmentation strategy that prioritizes semantic coherence over arbitrary byte cuts.
The Splitting Pipeline Architecture
When you call model.generate() with a long prompt, the library delegates segmentation to split_into_best_sentences() L778-L845. The pipeline executes four distinct phases:
- Text normalization – Optionally pads short inputs with spaces and strips semicolons based on configuration flags.
- SentencePiece tokenization – Converts the cleaned string into a token ID sequence using the model’s SentencePiece tokenizer.
- Boundary detection – Identifies end-of-sentence markers (
.,!,?) to generate candidate split points. - Greedy assembly – Concatenates sentences until adding the next would exceed
max_tokens, then seals the chunk and starts a new one.
Tokenization and Primary Boundary Detection
After initial normalization, the function invokes _find_boundary_indices() L945-L962 to locate tokens that correspond to terminal punctuation. These indices represent hard semantic boundaries where speech naturally pauses. The algorithm treats each boundary as a potential split point, creating a preliminary list of sentence segments.
If every segment falls below the max_tokens threshold, the function proceeds directly to the assembly phase. This preserves complete sentences, which is critical for maintaining prosody and intonation in the generated audio.
Fallback Sub-Splitting for Oversized Sentences
When a single sentence exceeds max_tokens—common with complex compound statements or lists—the pipeline triggers fallback logic L998-L1011. Instead of splitting arbitrarily, the function re-tokenizes the oversized sentence using a secondary delimiter set: commas, semicolons, and colons (,;:).
The same boundary-finding algorithm runs on this narrower token stream, generating sub-sentence fragments at natural pauses. If no additional punctuation exists, the function retains the original oversized segment and emits a runtime warning L1415-L1432, allowing the generation to proceed with a potentially truncated context rather than failing silently.
Greedy Chunk Assembly and Token Limits
With the final list of sentence or sub-sentence tokens, split_into_best_sentences() performs a greedy accumulation:
- It initializes an empty buffer and iterates over the segment list.
- For each segment, it checks if
buffer_tokens + segment_tokens ≤ max_tokens. - If true, it appends the segment; otherwise, it closes the current chunk, adds it to the output list, and starts a fresh buffer with the current segment.
This approach minimizes the total number of chunks while strictly enforcing the token budget, ensuring that the downstream generate_audio_stream() method receives only sequences the transformer can process in a single forward pass.
Practical Implementation: Pre-chunking Text Inputs
While the public generate() API handles splitting transparently, you can access the internal logic for inspection or custom preprocessing:
from pocket_tts import TTSModel
from pocket_tts.utils.config import TTSConfig
# Initialize model (weights download on first use)
model = TTSModel(TTSConfig())
long_text = """
Natural language processing has revolutionized human-computer interaction.
However, transformer-based models enforce strict context limits that must be respected
during inference to avoid truncation artifacts or runtime errors.
"""
# Access the internal splitting helper (note: underscore prefix indicates internal API)
chunks = model._cached_split_into_best_sentences(
tokenizer=model.tokenizer,
text_to_generate=long_text,
max_tokens=200, # Typical safe limit for a single generation pass
pad_with_spaces_for_short_inputs=False,
remove_semicolons=False,
)
print(f"Generated {len(chunks)} chunks:")
for i, chunk in enumerate(chunks, 1):
token_count = len(model.tokenizer(chunk).tokens[0])
print(f" Chunk {i} ({token_count} tokens): {chunk[:50]}...")
The resulting chunks list contains strings guaranteed to tokenize to ≤ 200 tokens, which you can then feed individually into generate_audio_stream() or allow the high-level API to process automatically.
Key Source Files
pocket_tts/models/tts_model.py– Containssplit_into_best_sentences(),_find_boundary_indices(), and the greedy assembly logic.pocket_tts/conditioners/text.py– Implements the SentencePiece tokenizer used for all token counting operations.tests/test_split_sentences.py– Unit tests verifying boundary detection behavior for edge cases like consecutive punctuation and oversized sentences.
Summary
- Automatic segmentation occurs via
split_into_best_sentences()before any audio generation begins. - Sentence boundaries (
.,!,?) are preferred split points to maintain natural speech patterns. - Fallback splitting on
,;:handles oversized sentences that exceedmax_tokensindividually. - Greedy chunking minimizes API calls by packing multiple short sentences into single chunks up to the token limit.
- Safety warnings alert users when a single unbreakable sentence exceeds the allowed token budget.
Frequently Asked Questions
What happens if a single sentence exceeds max_tokens?
If a sentence is longer than the allowed token budget and contains no fallback punctuation (commas, semicolons, or colons), the function retains the full sentence in a single chunk and emits a warning. The generation pipeline will attempt to process it, though the model may truncate or exhibit degraded performance on that specific segment.
How does split_into_best_sentences choose split points?
The function first identifies end-of-sentence punctuation tokens (periods, exclamation marks, question marks) using _find_boundary_indices(). If a sentence remains too long, it re-scans for internal pauses marked by commas, semicolons, or colons. It never splits mid-word or at arbitrary character offsets, ensuring that each chunk remains semantically coherent.
Can I adjust the maximum token limit per chunk?
Yes. The max_tokens parameter is configurable when calling the internal splitting method or when initializing the generation configuration. Typical values range from 150–400 tokens depending on the specific Pocket-TTS checkpoint and available memory, though the default is tuned for the standard model release.
Does chunking affect audio continuity?
No. Because splits occur at natural sentence or clause boundaries, the generated audio streams concatenate seamlessly. The model does not insert pauses or artifacts between chunks; the output waveform is equivalent to synthesizing the full text in one pass, provided the text is reassembled in the correct order after generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →