How to Debug KV Cache Issues with LINGBOT_DEBUG_KV in Ling Bot-Map
Set the LINGBOT_DEBUG_KV environment variable to enable per-frame KV cache statistics that print frame index, entry counts, hit rates, and memory usage during transformer inference.
The lingbot-map repository by Robbyant implements streaming transformer models that rely on an attention key/value (KV) cache for efficient sequence processing. When this cache grows unbounded or suffers from poor hit rates, you can diagnose issues without modifying source code by enabling the built-in debug environment variable.
What is LINGBOT_DEBUG_KV?
LINGBOT_DEBUG_KV is an environment variable that activates diagnostic logging for the attention KV cache inside lingbot-map's streaming models. When set, the system prints statistics every N frames (defaulting to every frame if set to 1) during the forward pass, revealing memory pressure, cache efficiency, and allocation patterns in real time.
According to the source code in lingbot_map/models/gct_stream.py, the variable is read at import time and controls conditional debug blocks that format and print cache metrics.
How the KV Cache Debugging Works Internally
Understanding the implementation helps you interpret the output correctly. The debug flow spans three stages across the streaming model files.
Reading the Environment Variable
In lingbot_map/models/gct_stream.py (lines 24-25), the code imports os and captures the variable at module load:
# lingbot_map/models/gct_stream.py
_KV_DEBUG = os.environ.get("LINGBOT_DEBUG_KV", "")
If the variable is unset or empty, _KV_DEBUG evaluates to an empty string, and all debug code paths are bypassed to eliminate runtime overhead.
Parsing the Interval Value
The variant file lingbot_map/models/gct_stream_window_v2.py contains logic to parse the string into an integer interval. As documented in lines 35-36, the utility interprets the value as a print-every-N setting:
# lingbot_map/models/gct_stream_window_v2.py
"""Parse LINGBOT_DEBUG_KV env var into a print-every-N interval."""
If you set LINGBOT_DEBUG_KV=10, the system logs statistics only on every 10th frame, reducing console noise during long sequences.
Emitting Statistics During Forward Pass
Inside the forward pass of both streaming modules, conditional blocks check _KV_DEBUG and emit formatted strings containing:
- Frame index: Current position in the sequence
- KV entries: Number of cached key/value pairs allocated
- Cache hit-rate: Percentage of cache reuse efficiency
- Memory used: Cache footprint in MiB
The actual printing logic resides in the same model files, wrapped by if _KV_DEBUG: guards to ensure zero impact when disabled.
Enabling KV Cache Debugging
Activate debugging by setting the variable before running any script that instantiates the streaming models.
Basic Usage (Every Frame)
export LINGBOT_DEBUG_KV=1
python demo.py --input data/example/loop/
Custom Interval (Every 20 Frames)
export LINGBOT_DEBUG_KV=20
python -m lingbot_map.demo --seq your_sequence.bin
Capturing Output for Analysis
export LINGBOT_DEBUG_KV=1
python demo.py 2>&1 | tee kv_debug.log
grep "KV DEBUG" kv_debug.log > kv_stats.csv
Understanding the Debug Output
When enabled, the system prints lines matching this format:
[KV DEBUG] frame=23 kv_entries=128 hit_rate=97.7% mem=84.2MiB
[KV DEBUG] frame=24 kv_entries=128 hit_rate=98.0% mem=84.2MiB
Analyzing these columns helps identify specific issues:
- Rising
kv_entries: Indicates unbounded cache growth, often causing GPU OOM errors on long sequences - Low
hit_rate: Signals cache misses that force recomputation, explaining sudden inference slowdowns - Increasing
memvalues: Helps profile memory scaling across different input resolutions or batch sizes
When to Use KV Cache Debugging
Enable LINGBOT_DEBUG_KV in these specific scenarios:
- Debugging GPU OOM errors during long sequences to confirm if the KV cache grows beyond available VRAM
- Investigating performance regressions after model changes to verify cache reuse across sliding windows
- Profiling memory footprints when testing new input resolutions or batch sizes in
flashinfer_cache.py - Validating windowed streaming logic in
gct_stream_window_v2.pyto ensure entries are correctly evicted and reused
Summary
- Set
LINGBOT_DEBUG_KVto a positive integer to enable per-frame KV cache statistics inlingbot-mapmodels - The variable is read at import time in
lingbot_map/models/gct_stream.pyand parsed as a print interval ingct_stream_window_v2.py - Output includes frame index, entry count, hit-rate percentage, and memory consumption in MiB
- Debugging incurs minimal overhead (string formatting and
printcalls) and does not alter cache behavior - Use higher interval values (e.g.,
10or20) to reduce log verbosity during long inference runs
Frequently Asked Questions
What values can I set for LINGBOT_DEBUG_KV?
You can set LINGBOT_DEBUG_KV to any positive integer. The value specifies the frame interval for printing statistics. Setting it to 1 prints every frame, while 10 prints only on every 10th frame. If the variable is unset or empty, debugging remains disabled.
Does enabling KV cache debugging slow down inference?
The performance impact is negligible. The code only performs string formatting and a print call when the debug flag is active, and it does not modify the underlying KV cache operations in flashinfer_cache.py or the transformer forward pass.
Where exactly does the debug output come from?
The output originates from conditional blocks inside lingbot_map/models/gct_stream.py and lingbot_map/models/gct_stream_window_v2.py. These files check the _KV_DEBUG variable during the forward pass and emit formatted statistics lines when the current frame matches the specified interval.
Why is my KV cache growing unbounded even with windowing?
If statistics show continuously increasing kv_entries across frames in windowed mode, verify that gct_stream_window_v2.py is correctly evicting old entries. The debug output helps distinguish between legitimate cache growth (new tokens) and memory leaks (failure to reuse or clear expired windows).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →