How to Debug Pocket-TTS Audio Generation Issues: A Complete Troubleshooting Guide

Debug Pocket-TTS audio generation issues by systematically verifying model weights, inspecting audio prompt encoding, enabling the DEBUG_MIMI environment variable, and tracing the FlowLM generation loop using utilities like LoggingMode and display_execution_time.

Pocket-TTS generates speech by streaming latent vectors from a FlowLM transformer and decoding them with the Mimi neural audio codec. When output contains silence, noise bursts, or truncated audio, the root cause can originate in model loading, prompt encoding, KV-cache management, or the multi-threaded decoder. This guide provides exact source locations in the kyutai-labs/pocket-tts repository and runnable code snippets to isolate failures in the generation pipeline.

Understanding the Generation Pipeline

The main generation path in pocket_tts/models/tts_model.py follows four distinct stages:

  1. Prompt Handling – Audio prompts are encoded via _encode_audio and merged into the FlowLM state.
  2. KV-Cache Management – The transformer's key-value cache is expanded dynamically using _expand_kv_cache.
  3. Autoregressive Generation – Latent vectors are generated step-by-step in _autoregressive_generation via _run_flow_lm_and_increment_step.
  4. Parallel Decoding – Generated latents pass through a queue.Queue to the _decode_audio_worker thread for audio reconstruction.

Systematic Debugging Steps

1. Verify Model Loading

Ensure the TTS model, FlowLM, and Mimi weights downloaded correctly and load onto the expected device. In pocket_tts/models/tts_model.py lines 33-41, TTSModel.load_model() initializes the pipeline.

from pocket_tts import TTSModel

model = TTSModel.load_model()
print(f"Device: {model.device}, Sample rate: {model.sample_rate}")

If this fails, check your internet connection and cache directory permissions.

2. Inspect Audio Prompt Processing

Confirm that conditioning audio is correctly resampled and encoded. The get_state_for_audio_prompt method (lines 88-98) relies on convert_audio from pocket_tts/data/audio_utils.py.

state = model.get_state_for_audio_prompt('my_voice.wav')
print(f"State size: {model.size_of_dict(state) // 1e6} MB")

A state size near zero indicates a failed audio read or encoding error.

3. Enable Low-Level Mimi Debugging

Set the DEBUG_MIMI environment variable to capture the intermediate latent-to-audio round-trip. This writes debug_encoded_latent_decoded.wav to your working directory, defined in pocket_tts/utils/utils.py lines 13-14.

export DEBUG_MIMI=1
python -c "
from pocket_tts import TTSModel
model = TTSModel.load_model()
state = model.get_state_for_audio_prompt('alba')
model.generate_audio(state, 'Test')
"

Listen to the debug file to verify the Mimi codec functions correctly.

4. Validate KV-Cache Expansion

Before autoregressive generation, the KV-cache must accommodate the full sequence length. The _expand_kv_cache method (lines 90-106) handles this allocation.

required_len = 5000
model._expand_kv_cache(state, sequence_length=required_len)

# Inspect first layer cache shape

first_layer = next(iter(state['flow_lm'].values()))
print(f"Cache shape: {first_layer['cache'].shape}")

# Expected: [2, 1, 5000, heads, dim_per_head]

If the cache dimension is smaller than required_len, generation will crash or truncate.

5. Trace FlowLM Latent Generation

Use LoggingMode from pocket_tts/utils/debugging.py (lines 16-27) to log every ATen call during the FlowLM forward pass. This reveals tensor shape mismatches or dtype errors.

from pocket_tts.utils.debugging import LoggingMode

with LoggingMode():
    # Replace ... with actual text tokens from your conditioner

    model._run_flow_lm_and_increment_step(state, text_tokens=...)

This is particularly useful when debugging NaN propagation or unexpected sampling behavior.

6. Measure Generation Timing

Profile the generation bottleneck using display_execution_time from pocket_tts/utils/utils.py (lines 75-92). This reports per-step latency and real-time factor (RTF).

from pocket_tts.utils.utils import display_execution_time

with display_execution_time('Full generation'):
    audio = model.generate_audio(state, 'Long text to synthesize')

Excessive step times often indicate decoder thread contention or GPU synchronization issues.

7. Debug the Decoder Worker Thread

The _decode_audio_worker (lines 41-69) runs in a separate thread consuming latents from a queue. Add instrumentation before the result_queue.put(("chunk", audio_frame)) call to verify audio frames are being produced.

If the queue returns ("error", ...) messages or the thread dies silently, check for exceptions in the Mimi decoder.

8. Check for Early EOS Termination

The _autoregressive_generation loop (lines 56-66) stops after frames_after_eos following the first EOS token. Unexpected truncation often occurs when the EOS token appears prematurely.


# Add inside the generation loop or inspect the return value

print(f"EOS detected at step: {state.get('eos_step')}")

Increase frames_after_eos in generate_audio() to extend output beyond the end-of-sequence marker.

9. Detect NaN Propagation

Numerical instability in _run_flow_lm_and_increment_step (lines 43-45) can propagate silent failures. Check for NaNs in the latent output:

import torch

next_latent = model._run_flow_lm_and_increment_step(state, text_tokens)
if torch.isnan(next_latent).any():
    print("NaN detected in latent generation")

NaNs in latents or KV-cache typically cause silent output or high-pitched noise bursts.

10. Test with Minimal Prompts

Isolate text-dependent issues by synthesizing a single word:

audio = model.generate_audio(state, "Hi.")

If minimal prompts work but longer text fails, the issue likely relates to KV-cache sizing or EOS detection.

Practical Debugging Examples

Quick Run-and-Listen Script

This minimal example uses predefined voices from _ORIGINS_OF_PREDEFINED_VOICES in pocket_tts/utils/utils.py (lines 15-42):

from pocket_tts import TTSModel
import soundfile as sf

model = TTSModel.load_model()
voice_state = model.get_state_for_audio_prompt('alba')  # Uses built-in voice

audio = model.generate_audio(voice_state, "Hello, this is a test.", frames_after_eos=3)

sf.write("out.wav", audio.cpu().numpy().T, model.sample_rate)

Inspecting Cache Sizes

from pocket_tts import TTSModel

model = TTSModel.load_model()
state = model.get_state_for_audio_prompt('alba')

# Prepare for 3000 tokens

model._expand_kv_cache(state, sequence_length=3000)
layer_cache = state['flow_lm']['layers.0']['cache']
print(f"Layer 0 cache shape: {layer_cache.shape}")

Summary

  • Verify model loading by checking TTSModel.load_model() output and device assignment in pocket_tts/models/tts_model.py.
  • Debug audio prompts using get_state_for_audio_prompt and confirm state size is non-zero.
  • Enable DEBUG_MIMI to validate the Mimi codec round-trip independently of generation.
  • Expand KV-cache explicitly with _expand_kv_cache to prevent truncation on long sequences.
  • Trace operations with LoggingMode to catch tensor errors in FlowLM.
  • Monitor performance using display_execution_time to identify decoder bottlenecks.
  • Check for NaNs in latents and caches to catch numerical instability early.
  • Validate threading by inspecting the _decode_audio_worker queue for error messages.

Frequently Asked Questions

Why is my Pocket-TTS output completely silent?

Silence typically indicates NaN propagation in the FlowLM latents or a failed audio prompt encoding. Check torch.isnan(next_latent).any() in _run_flow_lm_and_increment_step and verify the state size after get_state_for_audio_prompt. Also ensure DEBUG_MIMI produces a valid debug_encoded_latent_decoded.wav file.

How do I fix truncated audio output in Pocket-TTS?

Truncation usually results from insufficient KV-cache allocation or premature EOS detection. Increase the sequence_length parameter in _expand_kv_cache and raise the frames_after_eos value in generate_audio() to extend generation beyond the end-of-sequence token.

What causes noise bursts in generated audio?

Noise bursts often stem from numerical instability in the autoregressive generation or corrupted model weights. Verify no NaNs exist in the KV-cache after expansion, and confirm that TTSModel.load_model() completed without warnings about missing checkpoints.

How can I profile Pocket-TTS generation speed?

Use the display_execution_time context manager from pocket_tts/utils/utils.py to measure real-time factor (RTF). If specific steps are slow, enable LoggingMode to identify which ATen operations are consuming the most time, typically revealing GPU synchronization or decoder thread contention.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →