How to Debug KV Cache Issues with LINGBOT_DEBUG_KV in Ling Bot-Map

Set the LINGBOT_DEBUG_KV environment variable to enable per-frame KV cache statistics that print frame index, entry counts, hit rates, and memory usage during transformer inference.

The lingbot-map repository by Robbyant implements streaming transformer models that rely on an attention key/value (KV) cache for efficient sequence processing. When this cache grows unbounded or suffers from poor hit rates, you can diagnose issues without modifying source code by enabling the built-in debug environment variable.

What is LINGBOT_DEBUG_KV?

LINGBOT_DEBUG_KV is an environment variable that activates diagnostic logging for the attention KV cache inside lingbot-map's streaming models. When set, the system prints statistics every N frames (defaulting to every frame if set to 1) during the forward pass, revealing memory pressure, cache efficiency, and allocation patterns in real time.

According to the source code in lingbot_map/models/gct_stream.py, the variable is read at import time and controls conditional debug blocks that format and print cache metrics.

How the KV Cache Debugging Works Internally

Understanding the implementation helps you interpret the output correctly. The debug flow spans three stages across the streaming model files.

Reading the Environment Variable

In lingbot_map/models/gct_stream.py (lines 24-25), the code imports os and captures the variable at module load:


# lingbot_map/models/gct_stream.py

_KV_DEBUG = os.environ.get("LINGBOT_DEBUG_KV", "")

If the variable is unset or empty, _KV_DEBUG evaluates to an empty string, and all debug code paths are bypassed to eliminate runtime overhead.

Parsing the Interval Value

The variant file lingbot_map/models/gct_stream_window_v2.py contains logic to parse the string into an integer interval. As documented in lines 35-36, the utility interprets the value as a print-every-N setting:


# lingbot_map/models/gct_stream_window_v2.py

"""Parse LINGBOT_DEBUG_KV env var into a print-every-N interval."""

If you set LINGBOT_DEBUG_KV=10, the system logs statistics only on every 10th frame, reducing console noise during long sequences.

Emitting Statistics During Forward Pass

Inside the forward pass of both streaming modules, conditional blocks check _KV_DEBUG and emit formatted strings containing:

  • Frame index: Current position in the sequence
  • KV entries: Number of cached key/value pairs allocated
  • Cache hit-rate: Percentage of cache reuse efficiency
  • Memory used: Cache footprint in MiB

The actual printing logic resides in the same model files, wrapped by if _KV_DEBUG: guards to ensure zero impact when disabled.

Enabling KV Cache Debugging

Activate debugging by setting the variable before running any script that instantiates the streaming models.

Basic Usage (Every Frame)

export LINGBOT_DEBUG_KV=1
python demo.py --input data/example/loop/

Custom Interval (Every 20 Frames)

export LINGBOT_DEBUG_KV=20
python -m lingbot_map.demo --seq your_sequence.bin

Capturing Output for Analysis

export LINGBOT_DEBUG_KV=1
python demo.py 2>&1 | tee kv_debug.log
grep "KV DEBUG" kv_debug.log > kv_stats.csv

Understanding the Debug Output

When enabled, the system prints lines matching this format:


[KV DEBUG] frame=23  kv_entries=128  hit_rate=97.7%  mem=84.2MiB
[KV DEBUG] frame=24  kv_entries=128  hit_rate=98.0%  mem=84.2MiB

Analyzing these columns helps identify specific issues:

  • Rising kv_entries: Indicates unbounded cache growth, often causing GPU OOM errors on long sequences
  • Low hit_rate: Signals cache misses that force recomputation, explaining sudden inference slowdowns
  • Increasing mem values: Helps profile memory scaling across different input resolutions or batch sizes

When to Use KV Cache Debugging

Enable LINGBOT_DEBUG_KV in these specific scenarios:

  • Debugging GPU OOM errors during long sequences to confirm if the KV cache grows beyond available VRAM
  • Investigating performance regressions after model changes to verify cache reuse across sliding windows
  • Profiling memory footprints when testing new input resolutions or batch sizes in flashinfer_cache.py
  • Validating windowed streaming logic in gct_stream_window_v2.py to ensure entries are correctly evicted and reused

Summary

  • Set LINGBOT_DEBUG_KV to a positive integer to enable per-frame KV cache statistics in lingbot-map models
  • The variable is read at import time in lingbot_map/models/gct_stream.py and parsed as a print interval in gct_stream_window_v2.py
  • Output includes frame index, entry count, hit-rate percentage, and memory consumption in MiB
  • Debugging incurs minimal overhead (string formatting and print calls) and does not alter cache behavior
  • Use higher interval values (e.g., 10 or 20) to reduce log verbosity during long inference runs

Frequently Asked Questions

What values can I set for LINGBOT_DEBUG_KV?

You can set LINGBOT_DEBUG_KV to any positive integer. The value specifies the frame interval for printing statistics. Setting it to 1 prints every frame, while 10 prints only on every 10th frame. If the variable is unset or empty, debugging remains disabled.

Does enabling KV cache debugging slow down inference?

The performance impact is negligible. The code only performs string formatting and a print call when the debug flag is active, and it does not modify the underlying KV cache operations in flashinfer_cache.py or the transformer forward pass.

Where exactly does the debug output come from?

The output originates from conditional blocks inside lingbot_map/models/gct_stream.py and lingbot_map/models/gct_stream_window_v2.py. These files check the _KV_DEBUG variable during the forward pass and emit formatted statistics lines when the current frame matches the specified interval.

Why is my KV cache growing unbounded even with windowing?

If statistics show continuously increasing kv_entries across frames in windowed mode, verify that gct_stream_window_v2.py is correctly evicting old entries. The debug output helps distinguish between legitimate cache growth (new tokens) and memory leaks (failure to reuse or clear expired windows).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →