# How Pocket-TTS Handles Infinitely Long Text Inputs: Chunking and Streaming Architecture

> Discover how Pocket-TTS manages long text inputs by chunking and streaming audio sequentially. Learn its architecture for efficient, continuous text-to-speech generation.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: internals
- Published: 2026-07-09

---

**Pocket-TTS cannot stream unbounded text as a single continuous prompt; instead, it automatically segments input into token-limited chunks at natural sentence boundaries and generates audio sequentially without carrying acoustic state between chunks.**

The open-source `kyutai-labs/pocket-tts` repository implements a pragmatic approach to handling arbitrarily long text inputs through intelligent chunking rather than true infinite-context streaming. When faced with text that exceeds the model's processing limits, the system breaks the input into manageable segments and processes them independently. This article examines the specific mechanisms in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) that enable pocket-tts to handle infinitely long text inputs while maintaining generation quality.

## Text Chunking Strategy

The core of long-text handling resides in the `split_into_best_sentences` method within [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (lines 778-845). This function tokenizes the entire input string using the model's SentencePiece tokenizer and identifies natural breaking points to create segments that respect both linguistic coherence and hard token limits.

### Token-Based Sentence Splitting

The chunking process begins by feeding the full text through the SentencePiece tokenizer defined in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py). The system enforces the `MAX_TOKEN_PER_CHUNK` constant—defaulting to approximately 200 tokens as specified in [`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py)—ensuring no single generation request exceeds the model's capacity. Each segment returned by `split_into_best_sentences` represents a discrete unit that can be processed independently by the autoregressive generation pipeline.

### Hierarchical Boundary Detection

To preserve semantic coherence, the algorithm implements a fallback hierarchy for boundary detection:

- **Primary boundaries**: Sentence terminators (`.`, `!`, `?`) are preferred to maintain natural prosody
- **Secondary boundaries**: For sentences exceeding the token limit, the system falls back to commas, semicolons, and colons
- **Forced splits**: If no punctuation exists within the token budget, the text is truncated at `MAX_TOKEN_PER_CHUNK` and a warning is logged

This approach ensures that pocket-tts can handle infinitely long text inputs in theory, though generation quality may degrade if the algorithm is forced to split mid-sentence.

## Streaming Generation Architecture

Once chunked, the `generate_audio_stream` method orchestrates sequential generation by iterating over the list of text segments and invoking the short-text pipeline for each chunk.

### Per-Chunk Processing Pipeline

For every chunk returned by `split_into_best_sentences`, the system executes `_generate_audio_stream_short_text`. This method:

1. Prepares the text tensor and estimates the required number of latent frames
2. Launches a background thread that decodes latent vectors into audio chunks
3. Yields audio tensors as they become available, creating the illusion of streaming

The architecture processes each chunk as an independent generation task, meaning the method can handle inputs of arbitrary length given sufficient computation time and memory.

### Autoregressive Generation Within Chunks

Within each chunk, the generation flow follows `generate_audio_stream` → `_generate` → `_autoregressive_generation`. This inner loop samples latents from the Flow-LM iteratively until either an EOS (end-of-sequence) token is encountered or the per-chunk token budget is exhausted. Because this autoregressive process operates on bounded inputs, it maintains stability and predictable memory usage regardless of the total input length.

## Limitations and Design Trade-offs

While the chunking approach enables processing of extremely long texts, the implementation makes specific architectural compromises that affect output continuity.

### Absence of Cross-Chunk Conditioning

The current implementation does not feed audio from previously generated chunks as conditioning for subsequent segments. According to comments in the source code, this is a deliberate simplification; true infinite-length streaming would require "teacher-forcing" or similar techniques to maintain acoustic continuity. A TODO comment in [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py) notes this as a future improvement, indicating that the current method may produce audible discontinuities at chunk boundaries when processing pocket-tts infinitely long text inputs.

### Practical Constraints

Because each chunk is processed with a fresh copy of the model state (or shared state without audio conditioning), the system cannot maintain prosodic context across sentence boundaries that fall in different chunks. Additionally, while the `MAX_TOKEN_PER_CHUNK` limit protects against out-of-memory errors, forcing splits within long sentences—particularly those without secondary punctuation—can result in unnatural pauses or intonation shifts.

## Practical Implementation Example

The following example demonstrates how pocket-tts automatically handles a very long input by chunking internally:

```python
from pocket_tts import TTSModel

# Load the pretrained model

model = TTSModel.load_model()

# Obtain a voice state (optional voice cloning)

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Very long prompt – the model will split it into chunks automatically

long_text = (
    "It was the best of times, it was the worst of times, "
    "it was the age of wisdom, it was the age of foolishness, "
    "… (many more sentences) …"
)

# Stream the audio – each yielded tensor is a small audio segment

for audio_chunk in model.generate_audio_stream(
    model_state=voice_state,
    text_to_generate=long_text,
    max_tokens=200,      # default; can be tuned

):
    # Do something with the chunk (e.g., play or write to a file)

    play(audio_chunk)   # pseudo-function

```

Key files involved in this process include [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) (containing the core logic), [`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py) (defining `MAX_TOKEN_PER_CHUNK`), and [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py) (providing the SentencePiece tokenizer). Test coverage for the chunking behavior is available in [`tests/test_split_sentences.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tests/test_split_sentences.py).

## Summary

- **Chunking mechanism**: Pocket-TTS handles long inputs via `split_into_best_sentences` in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), which segments text at natural boundaries while respecting the `MAX_TOKEN_PER_CHUNK` limit of approximately 200 tokens.
- **Sequential processing**: The `generate_audio_stream` method processes chunks sequentially using `_generate_audio_stream_short_text`, with each chunk undergoing independent autoregressive generation via `_autoregressive_generation` and the Flow-LM.
- **State isolation**: No acoustic state persists between chunks, meaning the system cannot condition later audio on previously generated speech, potentially affecting prosodic continuity at boundaries.
- **Theoretical limits**: Arbitrarily long inputs are supported in practice, but generation quality depends on sentence length relative to the token limit, with warnings issued for forced mid-sentence splits.

## Frequently Asked Questions

### Can pocket-tts stream infinite text in real-time without chunking?

No, pocket-tts cannot stream unbounded text as a single continuous prompt. According to the source code in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), the system must first tokenize and split the input using `split_into_best_sentences` before generation begins. This creates inherent latency between chunks, as each segment must be tokenized and processed independently.

### What happens if a single sentence exceeds the maximum token limit?

When a sentence exceeds `MAX_TOKEN_PER_CHUNK` (approximately 200 tokens by default), the system first attempts to find secondary punctuation such as commas and semicolons. If the sentence remains too long, it is forcibly split at the token boundary and a warning is logged to indicate potential quality degradation. This ensures the system can still process pocket-tts infinitely long text inputs without crashing, though with reduced prosodic naturalness.

### Does pocket-tts maintain voice consistency across chunk boundaries?

While the speaker embedding (voice state) remains consistent across all chunks, the acoustic generation does not condition on previous chunks' audio output. The code explicitly notes this as a current limitation, with a TODO comment suggesting future implementation of "teacher-forcing" to feed prior audio back as conditioning. Consequently, long inputs may exhibit slight discontinuities at chunk junctions.

### How can I adjust the chunk size for my specific use case?

You can modify the `max_tokens` parameter when calling `generate_audio_stream`, or adjust the `MAX_TOKEN_PER_CHUNK` constant in [`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py). However, increasing this value beyond the model's training context window risks degraded output quality or out-of-memory errors, while decreasing it may increase the frequency of prosodic discontinuities.