# How to Debug KV Cache Issues with LINGBOT_DEBUG_KV in Ling Bot-Map

> Debug KV cache issues in Ling Bot-Map using LINGBOT_DEBUG_KV. Enable per-frame statistics for inference analysis, including hit rates and memory usage.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-26

---

**Set the `LINGBOT_DEBUG_KV` environment variable to enable per-frame KV cache statistics that print frame index, entry counts, hit rates, and memory usage during transformer inference.**

The `lingbot-map` repository by Robbyant implements streaming transformer models that rely on an attention key/value (KV) cache for efficient sequence processing. When this cache grows unbounded or suffers from poor hit rates, you can diagnose issues without modifying source code by enabling the built-in debug environment variable.

## What is LINGBOT_DEBUG_KV?

`LINGBOT_DEBUG_KV` is an environment variable that activates diagnostic logging for the attention KV cache inside `lingbot-map`'s streaming models. When set, the system prints statistics every *N* frames (defaulting to every frame if set to `1`) during the forward pass, revealing memory pressure, cache efficiency, and allocation patterns in real time.

According to the source code in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), the variable is read at import time and controls conditional debug blocks that format and print cache metrics.

## How the KV Cache Debugging Works Internally

Understanding the implementation helps you interpret the output correctly. The debug flow spans three stages across the streaming model files.

### Reading the Environment Variable

In [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) (lines 24-25), the code imports `os` and captures the variable at module load:

```python

# lingbot_map/models/gct_stream.py

_KV_DEBUG = os.environ.get("LINGBOT_DEBUG_KV", "")

```

If the variable is unset or empty, `_KV_DEBUG` evaluates to an empty string, and all debug code paths are bypassed to eliminate runtime overhead.

### Parsing the Interval Value

The variant file [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py) contains logic to parse the string into an integer interval. As documented in lines 35-36, the utility interprets the value as a print-every-*N* setting:

```python

# lingbot_map/models/gct_stream_window_v2.py

"""Parse LINGBOT_DEBUG_KV env var into a print-every-N interval."""

```

If you set `LINGBOT_DEBUG_KV=10`, the system logs statistics only on every 10th frame, reducing console noise during long sequences.

### Emitting Statistics During Forward Pass

Inside the forward pass of both streaming modules, conditional blocks check `_KV_DEBUG` and emit formatted strings containing:

- **Frame index**: Current position in the sequence
- **KV entries**: Number of cached key/value pairs allocated
- **Cache hit-rate**: Percentage of cache reuse efficiency
- **Memory used**: Cache footprint in MiB

The actual printing logic resides in the same model files, wrapped by `if _KV_DEBUG:` guards to ensure zero impact when disabled.

## Enabling KV Cache Debugging

Activate debugging by setting the variable before running any script that instantiates the streaming models.

### Basic Usage (Every Frame)

```bash
export LINGBOT_DEBUG_KV=1
python demo.py --input data/example/loop/

```

### Custom Interval (Every 20 Frames)

```bash
export LINGBOT_DEBUG_KV=20
python -m lingbot_map.demo --seq your_sequence.bin

```

### Capturing Output for Analysis

```bash
export LINGBOT_DEBUG_KV=1
python demo.py 2>&1 | tee kv_debug.log
grep "KV DEBUG" kv_debug.log > kv_stats.csv

```

## Understanding the Debug Output

When enabled, the system prints lines matching this format:

```

[KV DEBUG] frame=23  kv_entries=128  hit_rate=97.7%  mem=84.2MiB
[KV DEBUG] frame=24  kv_entries=128  hit_rate=98.0%  mem=84.2MiB

```

Analyzing these columns helps identify specific issues:

- **Rising `kv_entries`**: Indicates unbounded cache growth, often causing GPU OOM errors on long sequences
- **Low `hit_rate`**: Signals cache misses that force recomputation, explaining sudden inference slowdowns
- **Increasing `mem` values**: Helps profile memory scaling across different input resolutions or batch sizes

## When to Use KV Cache Debugging

Enable `LINGBOT_DEBUG_KV` in these specific scenarios:

- **Debugging GPU OOM errors** during long sequences to confirm if the KV cache grows beyond available VRAM
- **Investigating performance regressions** after model changes to verify cache reuse across sliding windows
- **Profiling memory footprints** when testing new input resolutions or batch sizes in [`flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/flashinfer_cache.py)
- **Validating windowed streaming logic** in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) to ensure entries are correctly evicted and reused

## Summary

- Set `LINGBOT_DEBUG_KV` to a positive integer to enable per-frame KV cache statistics in `lingbot-map` models
- The variable is read at import time in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) and parsed as a print interval in [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py)
- Output includes frame index, entry count, hit-rate percentage, and memory consumption in MiB
- Debugging incurs minimal overhead (string formatting and `print` calls) and does not alter cache behavior
- Use higher interval values (e.g., `10` or `20`) to reduce log verbosity during long inference runs

## Frequently Asked Questions

### What values can I set for LINGBOT_DEBUG_KV?

You can set `LINGBOT_DEBUG_KV` to any positive integer. The value specifies the frame interval for printing statistics. Setting it to `1` prints every frame, while `10` prints only on every 10th frame. If the variable is unset or empty, debugging remains disabled.

### Does enabling KV cache debugging slow down inference?

The performance impact is negligible. The code only performs string formatting and a `print` call when the debug flag is active, and it does not modify the underlying KV cache operations in [`flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/flashinfer_cache.py) or the transformer forward pass.

### Where exactly does the debug output come from?

The output originates from conditional blocks inside [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) and [`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py). These files check the `_KV_DEBUG` variable during the forward pass and emit formatted statistics lines when the current frame matches the specified interval.

### Why is my KV cache growing unbounded even with windowing?

If statistics show continuously increasing `kv_entries` across frames in windowed mode, verify that [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py) is correctly evicting old entries. The debug output helps distinguish between legitimate cache growth (new tokens) and memory leaks (failure to reuse or clear expired windows).