# How to Debug Pocket-TTS Audio Generation Issues: A Complete Troubleshooting Guide

> Solve Pocket-TTS audio generation problems. This guide helps you debug by checking model weights, prompt encoding, and using debug tools. Get clear audio output.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: how-to-guide
- Published: 2026-07-09

---

**Debug Pocket-TTS audio generation issues by systematically verifying model weights, inspecting audio prompt encoding, enabling the `DEBUG_MIMI` environment variable, and tracing the FlowLM generation loop using utilities like `LoggingMode` and `display_execution_time`.**

Pocket-TTS generates speech by streaming latent vectors from a **FlowLM** transformer and decoding them with the **Mimi** neural audio codec. When output contains silence, noise bursts, or truncated audio, the root cause can originate in model loading, prompt encoding, KV-cache management, or the multi-threaded decoder. This guide provides exact source locations in the `kyutai-labs/pocket-tts` repository and runnable code snippets to isolate failures in the generation pipeline.

## Understanding the Generation Pipeline

The main generation path in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) follows four distinct stages:

1. **Prompt Handling** – Audio prompts are encoded via `_encode_audio` and merged into the FlowLM state.
2. **KV-Cache Management** – The transformer's key-value cache is expanded dynamically using `_expand_kv_cache`.
3. **Autoregressive Generation** – Latent vectors are generated step-by-step in `_autoregressive_generation` via `_run_flow_lm_and_increment_step`.
4. **Parallel Decoding** – Generated latents pass through a `queue.Queue` to the `_decode_audio_worker` thread for audio reconstruction.

## Systematic Debugging Steps

### 1. Verify Model Loading

Ensure the TTS model, FlowLM, and Mimi weights downloaded correctly and load onto the expected device. In [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) lines 33-41, `TTSModel.load_model()` initializes the pipeline.

```python
from pocket_tts import TTSModel

model = TTSModel.load_model()
print(f"Device: {model.device}, Sample rate: {model.sample_rate}")

```

If this fails, check your internet connection and cache directory permissions.

### 2. Inspect Audio Prompt Processing

Confirm that conditioning audio is correctly resampled and encoded. The `get_state_for_audio_prompt` method (lines 88-98) relies on `convert_audio` from [`pocket_tts/data/audio_utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/data/audio_utils.py).

```python
state = model.get_state_for_audio_prompt('my_voice.wav')
print(f"State size: {model.size_of_dict(state) // 1e6} MB")

```

A state size near zero indicates a failed audio read or encoding error.

### 3. Enable Low-Level Mimi Debugging

Set the `DEBUG_MIMI` environment variable to capture the intermediate latent-to-audio round-trip. This writes `debug_encoded_latent_decoded.wav` to your working directory, defined in [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py) lines 13-14.

```bash
export DEBUG_MIMI=1
python -c "
from pocket_tts import TTSModel
model = TTSModel.load_model()
state = model.get_state_for_audio_prompt('alba')
model.generate_audio(state, 'Test')
"

```

Listen to the debug file to verify the Mimi codec functions correctly.

### 4. Validate KV-Cache Expansion

Before autoregressive generation, the KV-cache must accommodate the full sequence length. The `_expand_kv_cache` method (lines 90-106) handles this allocation.

```python
required_len = 5000
model._expand_kv_cache(state, sequence_length=required_len)

# Inspect first layer cache shape

first_layer = next(iter(state['flow_lm'].values()))
print(f"Cache shape: {first_layer['cache'].shape}")

# Expected: [2, 1, 5000, heads, dim_per_head]

```

If the cache dimension is smaller than `required_len`, generation will crash or truncate.

### 5. Trace FlowLM Latent Generation

Use `LoggingMode` from [`pocket_tts/utils/debugging.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/debugging.py) (lines 16-27) to log every ATen call during the FlowLM forward pass. This reveals tensor shape mismatches or dtype errors.

```python
from pocket_tts.utils.debugging import LoggingMode

with LoggingMode():
    # Replace ... with actual text tokens from your conditioner

    model._run_flow_lm_and_increment_step(state, text_tokens=...)

```

This is particularly useful when debugging `NaN` propagation or unexpected sampling behavior.

### 6. Measure Generation Timing

Profile the generation bottleneck using `display_execution_time` from [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py) (lines 75-92). This reports per-step latency and real-time factor (RTF).

```python
from pocket_tts.utils.utils import display_execution_time

with display_execution_time('Full generation'):
    audio = model.generate_audio(state, 'Long text to synthesize')

```

Excessive step times often indicate decoder thread contention or GPU synchronization issues.

### 7. Debug the Decoder Worker Thread

The `_decode_audio_worker` (lines 41-69) runs in a separate thread consuming latents from a queue. Add instrumentation before the `result_queue.put(("chunk", audio_frame))` call to verify audio frames are being produced.

If the queue returns `("error", ...)` messages or the thread dies silently, check for exceptions in the Mimi decoder.

### 8. Check for Early EOS Termination

The `_autoregressive_generation` loop (lines 56-66) stops after `frames_after_eos` following the first EOS token. Unexpected truncation often occurs when the EOS token appears prematurely.

```python

# Add inside the generation loop or inspect the return value

print(f"EOS detected at step: {state.get('eos_step')}")

```

Increase `frames_after_eos` in `generate_audio()` to extend output beyond the end-of-sequence marker.

### 9. Detect NaN Propagation

Numerical instability in `_run_flow_lm_and_increment_step` (lines 43-45) can propagate silent failures. Check for NaNs in the latent output:

```python
import torch

next_latent = model._run_flow_lm_and_increment_step(state, text_tokens)
if torch.isnan(next_latent).any():
    print("NaN detected in latent generation")

```

NaNs in latents or KV-cache typically cause silent output or high-pitched noise bursts.

### 10. Test with Minimal Prompts

Isolate text-dependent issues by synthesizing a single word:

```python
audio = model.generate_audio(state, "Hi.")

```

If minimal prompts work but longer text fails, the issue likely relates to KV-cache sizing or EOS detection.

## Practical Debugging Examples

### Quick Run-and-Listen Script

This minimal example uses predefined voices from `_ORIGINS_OF_PREDEFINED_VOICES` in [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py) (lines 15-42):

```python
from pocket_tts import TTSModel
import soundfile as sf

model = TTSModel.load_model()
voice_state = model.get_state_for_audio_prompt('alba')  # Uses built-in voice

audio = model.generate_audio(voice_state, "Hello, this is a test.", frames_after_eos=3)

sf.write("out.wav", audio.cpu().numpy().T, model.sample_rate)

```

### Inspecting Cache Sizes

```python
from pocket_tts import TTSModel

model = TTSModel.load_model()
state = model.get_state_for_audio_prompt('alba')

# Prepare for 3000 tokens

model._expand_kv_cache(state, sequence_length=3000)
layer_cache = state['flow_lm']['layers.0']['cache']
print(f"Layer 0 cache shape: {layer_cache.shape}")

```

## Summary

- **Verify model loading** by checking `TTSModel.load_model()` output and device assignment in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py).
- **Debug audio prompts** using `get_state_for_audio_prompt` and confirm state size is non-zero.
- **Enable `DEBUG_MIMI`** to validate the Mimi codec round-trip independently of generation.
- **Expand KV-cache** explicitly with `_expand_kv_cache` to prevent truncation on long sequences.
- **Trace operations** with `LoggingMode` to catch tensor errors in FlowLM.
- **Monitor performance** using `display_execution_time` to identify decoder bottlenecks.
- **Check for NaNs** in latents and caches to catch numerical instability early.
- **Validate threading** by inspecting the `_decode_audio_worker` queue for error messages.

## Frequently Asked Questions

### Why is my Pocket-TTS output completely silent?

Silence typically indicates **NaN propagation** in the FlowLM latents or a **failed audio prompt encoding**. Check `torch.isnan(next_latent).any()` in `_run_flow_lm_and_increment_step` and verify the state size after `get_state_for_audio_prompt`. Also ensure `DEBUG_MIMI` produces a valid `debug_encoded_latent_decoded.wav` file.

### How do I fix truncated audio output in Pocket-TTS?

Truncation usually results from **insufficient KV-cache allocation** or **premature EOS detection**. Increase the `sequence_length` parameter in `_expand_kv_cache` and raise the `frames_after_eos` value in `generate_audio()` to extend generation beyond the end-of-sequence token.

### What causes noise bursts in generated audio?

Noise bursts often stem from **numerical instability** in the autoregressive generation or **corrupted model weights**. Verify no NaNs exist in the KV-cache after expansion, and confirm that `TTSModel.load_model()` completed without warnings about missing checkpoints.

### How can I profile Pocket-TTS generation speed?

Use the `display_execution_time` context manager from [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py) to measure real-time factor (RTF). If specific steps are slow, enable `LoggingMode` to identify which ATen operations are consuming the most time, typically revealing GPU synchronization or decoder thread contention.