# Difference Between generate_audio() and generate_audio_stream() in Pocket-TTS

> Understand Pocket-TTS generate_audio() vs generate_audio_stream(). Get full audio tensors or progressive generator chunks for real-time playback with low latency.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-11

---

**The `generate_audio()` method returns the complete audio waveform as a single `torch.Tensor`, while `generate_audio_stream()` yields audio chunks progressively through a generator, enabling real-time playback with minimal latency.**

Both methods are implemented in the `TTSModel` class within the [kyutai-labs/pocket-tts](https://github.com/kyutai-labs/pocket-tts) repository in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py). While they produce identical audio output from the same text input, their internal architecture differs significantly depending on whether you need batch synthesis or streaming synthesis.

## Internal Architecture and Implementation

The relationship between these methods is hierarchical: `generate_audio()` acts as a convenience wrapper that consumes the streaming output of `generate_audio_stream()` and concatenates the results before returning.

### The Synchronous Wrapper: generate_audio()

Located at lines **[776‑785](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py#L776-L785)** in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), this method implements a simple loop:

```python
audio_chunks = []
for chunk in self.generate_audio_stream(...):
    audio_chunks.append(chunk)
return torch.cat(audio_chunks, dim=0)

```

This means you wait for the entire text to be processed before receiving any audio data. The method returns a tensor with shape `[channels, samples]`, making it suitable for immediate file writing or post-processing. Because it aggregates all chunks in memory before returning, it requires sufficient RAM to hold the complete audio waveform.

### The Streaming Pipeline: generate_audio_stream()

The actual synthesis logic resides in `generate_audio_stream()` at lines **[445‑456](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py#L445-L456)**. This method orchestrates a multi-threaded pipeline that includes:

- **Text Segmentation**: Uses `split_into_best_sentences` to divide input into optimal processing units
- **Worker Threading**: Spawns a **decoder worker thread** (`_decode_audio_worker`) that consumes latent vectors from a queue and decodes them into audio chunks
- **Concurrent Generation**: Runs `_generate()` in a separate thread to produce latents and push them into the same queue
- **Progressive Yield**: Returns each decoded chunk immediately via a generator as it becomes available

Each yielded chunk is a 1-D `torch.Tensor` with shape `[samples]` (single-channel, batch dimension removed), allowing downstream audio players to consume data while generation continues.

## Performance and Latency Characteristics

**`generate_audio()`** introduces higher perceived latency because it blocks until the entire text sequence is synthesized. This approach is ideal when you need the complete audio for non-real-time applications such as file export or batch processing.

**`generate_audio_stream()`** provides lower latency by yielding audio chunks as soon as they are decoded. This enables real-time playback scenarios where audio can begin playing while the model is still processing subsequent text segments. The streaming approach also maintains a lower memory footprint for long texts since it processes chunks incrementally rather than storing the entire waveform.

Both methods accept a `copy_state` parameter (default `True`). When enabled, the original `model_state` is deep-copied before generation, preserving it for reuse. In the streaming variant implemented via `_generate_audio_stream_short_text`, this copying occurs inside the pipeline for each short-text segment.

## Thread Safety Considerations

**Neither method is thread-safe.** Each `TTSModel` instance should be accessed by only one thread at a time. The streaming implementation spawns internal decoder threads, but these are managed within the method scope and joined before completion. Attempting to call either method concurrently from multiple threads on the same instance will result in race conditions.

## Practical Code Examples

### Batch Processing with generate_audio()

Use this approach when you need the complete waveform for file output or offline processing:

```python
from pocket_tts import TTSModel

# Initialize the model

model = TTSModel.load_model()

# Configure voice state from an audio sample

voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Generate complete audio tensor

audio = model.generate_audio(
    model_state=voice_state,
    text_to_generate="Hello world! This is a fully rendered audio example.",
    frames_after_eos=2,
)

print(f"Audio shape: {audio.shape}")  # → [channels, samples]

# Ready for torchaudio.save() or similar

```

### Real-Time Streaming with generate_audio_stream()

Use this for interactive applications where audio should play during generation:

```python
from pocket_tts import TTSModel

model = TTSModel.load_model()
voice_state = model.get_state_for_audio_prompt(
    "hf://kyutai/tts-voices/alba-mackenna/casual.wav"
)

# Stream audio chunk-by-chunk

for i, chunk in enumerate(
    model.generate_audio_stream(
        model_state=voice_state,
        text_to_generate="This is a long paragraph that will be streamed in real time.",
        frames_after_eos=None,
    )
):
    # chunk is a 1-D Tensor: [samples]

    print(f"Chunk {i}: {chunk.shape[0]} samples")
    # Send to audio player or write to circular buffer

```

## Summary

- **`generate_audio()`** is a synchronous wrapper located at lines **776‑785** in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) that returns a complete `[channels, samples]` tensor after processing all text
- **`generate_audio_stream()`** implements the core streaming logic at lines **445‑456**, yielding `[samples]` chunks immediately via a generator for low-latency applications
- Both methods support the `copy_state` parameter to preserve model state and are **not thread-safe**
- The streaming method uses internal worker threads (`_decode_audio_worker`) and text segmentation (`split_into_best_sentences`) to enable real-time audio production

## Frequently Asked Questions

### Can I use generate_audio_stream() for file output?

Yes, though it requires manual concatenation. You must iterate through the generator and accumulate chunks into a list before calling `torch.cat()`, essentially replicating what `generate_audio()` does internally. For direct file output, `generate_audio()` is more convenient as it returns the complete tensor ready for `torchaudio.save()`.

### Why does generate_audio_stream() return 1-D tensors instead of 2-D?

The streaming method yields single-channel mono audio with shape `[samples]` rather than `[1, samples]` to simplify integration with real-time audio pipelines. Most streaming audio consumers expect 1-D buffers. You can reshape the chunks to `[1, samples]` if your downstream processor requires explicit channel dimensions.

### Is generate_audio() slower than generate_audio_stream()?

The total processing time is nearly identical because `generate_audio()` simply wraps the streaming logic. However, `generate_audio()` has higher **perceived latency** because it blocks until completion. For long texts, `generate_audio_stream()` allows playback to begin immediately while generation continues in background threads.

### How does error handling differ between the two methods?

With `generate_audio()`, errors propagate after the full process completes. In contrast, `generate_audio_stream()` propagates errors immediately as they occur in the generation or decoding threads. The implementation ensures the decoder thread is cleanly joined via error handling logic at lines **[462‑470](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py#L462-L470)**, preventing resource leaks even when exceptions occur mid-stream.