How to Reduce GPU Memory Usage with Limited VRAM in LingBot‑Map
LingBot‑Map reduces GPU memory usage with limited VRAM through patch‑aligned image sizing, sliding‑window KV caches with configurable keyframe intervals, flow‑based dynamic frame skipping, and aggressive CPU off‑loading of intermediate predictions.
LingBot‑Map is designed to run large‑scale video‑to‑3‑D reconstruction on a single GPU, even when processing thousands of frames. To achieve this, the repository implements several memory‑efficient mechanisms that keep the peak GPU memory footprint low without sacrificing reconstruction quality. This guide explains how to leverage these techniques to reduce GPU memory usage with limited VRAM hardware.
Patch‑Aligned Image Sizing
The first line of defense against wasted GPU memory is patch‑aligned image sizing. In scripts/benchmark_gct_memory.py, the patch_align() function automatically reduces input resolution to a multiple of the patch size (default = 14). This guarantees that the internal token grid fits exactly, avoiding wasted memory on partially filled patches.
# From scripts/benchmark_gct_memory.py
def patch_align(x, patch_size=14):
# Aligns height and width to patch size multiples
return x[..., :patch_size*(x.shape[-2]//patch_size),
:patch_size*(x.shape[-1]//patch_size)]
When configuring your pipeline, choose modest dimensions that align to your patch size. For example, use --height 384 --width 518 rather than arbitrary resolutions that leave incomplete patch rows in memory.
Sliding‑Window KV Cache Management
The core memory optimization lives in lingbot_map/models/gct_stream_window.py. Instead of caching key‑value tensors for every frame in a sequence, LingBot‑Map implements a sliding‑window KV cache that limits growth to O(window size) rather than O(sequence length).
Configuring the Window Size
Set the kv_cache_sliding_window parameter to constrain how many past frames remain in GPU memory. The eviction logic is handled by _set_defer_eviction() and _execute_deferred_eviction():
# From gct_stream_window.py
def _set_defer_eviction(self, kv_cache):
# Marks old blocks for eviction when window slides
pass
def _execute_deferred_eviction(self):
# Actually frees the GPU memory
pass
For limited VRAM budgets, reduce the window aggressively:
--kv-cache-sliding-window 32 # Default is 64; 32 uses ~50% less cache memory
Keyframe Interval Optimization
The inference_streaming() method supports a keyframe_interval parameter that stores KV entries only for every N‑th frame. Non‑keyframes trigger _set_skip_append(True), preventing cache writes entirely:
# From gct_stream_window.py lines 100-106
if frame_idx % keyframe_interval != 0:
self._set_skip_append(True) # Skip KV cache update
Increase the interval to reduce cached frames by a factor of keyframe_interval:
--keyframe-interval 4 # Stores only 1/4 of frames in cache
Flow‑Based Dynamic Keyframes
For scenes with variable motion, enable flow‑based dynamic keyframe detection. The model computes cheap optical‑flow magnitude on‑the‑fly; frames with motion below --flow-threshold are treated as non‑keyframes automatically:
# From gct_stream_window.py lines 63-94
flow_magnitude = compute_flow(prev_frame, curr_frame)
if flow_magnitude < flow_threshold:
# Treat as non-keyframe, skip cache write
self._set_skip_append(True)
This adapts cache size to scene dynamics, saving significant memory when the camera is static.
CPU Off‑Loading Strategies
When you only need final reconstruction results—not intermediate tensors—configure CPU off‑loading via the output_device parameter. In inference_streaming(), the _to_out() helper moves predictions to CPU before concatenation:
# From gct_stream_window.py lines 18-38
def _to_out(self, tensor):
if self.output_device is not None:
return tensor.to(self.output_device)
return tensor
Set --output-device cpu to keep the GPU reserved strictly for the KV cache and current frame processing, while accumulated predictions reside in system RAM.
Input Materialization and Allocator Configuration
Prevent PyTorch from pre‑allocating large static memory pools with two additional techniques found in scripts/benchmark_gct_memory.py.
Synthetic Input Materialization: The --materialize-inputs flag (CLI) forces each frame tensor to live on CPU until just before the forward pass. In make_source_images(), this prevents a huge pre‑allocation on GPU during long sequences.
CUDA Allocator Tuning: The benchmark script sets PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to let the allocator grow only as needed:
# From benchmark_gct_memory.py lines 38-40
os.environ.setdefault(
"PYTORCH_CUDA_ALLOC_CONF",
"expandable_segments:True"
)
Explicit Cleanup: After each run, call torch.cuda.empty_cache() and torch.cuda.reset_peak_memory_stats() (wrapped in the cleanup() function) to release stray allocations and guarantee a fresh memory state.
Complete Configuration Example
Combine these techniques to run large sequences on under 4 GiB of VRAM:
python scripts/benchmark_gct_memory.py \
--height 384 --width 518 \
--patch-size 14 \
--frame-counts 64 128 256 512 1024 \
--kv-cache-sliding-window 32 \
--keyframe-interval 4 \
--flow-threshold 0.3 \
--output-device cpu \
--output gct_memory_limited.csv
For programmatic inference:
import torch
from lingbot_map.models.gct_stream import GCTStream
model = GCTStream(
img_size=518,
patch_size=14,
kv_cache_sliding_window=32,
kv_cache_scale_frames=8,
kv_cache_cross_frame_special=True,
kv_cache_include_scale_frames=True,
use_sdpa=False,
)
model.eval().to('cuda')
# images is a (B, S, 3, H, W) tensor on CPU
pred = model.inference_streaming(
images,
keyframe_interval=4,
flow_threshold=0.3,
output_device=torch.device('cpu'),
)
Summary
- Patch‑align inputs to avoid partially filled patches wasting memory in
scripts/benchmark_gct_memory.py. - Limit KV cache growth with
kv_cache_sliding_windowand automatic eviction inlingbot_map/models/gct_stream_window.py. - Reduce cache entries by increasing
keyframe_intervalor enabling flow‑based dynamic keyframes to skip static frames. - Off‑load predictions to CPU via
output_deviceto free GPU memory immediately after processing. - Control allocator behavior with
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:Trueand explicittorch.cuda.empty_cache()calls.
Frequently Asked Questions
What is the minimum VRAM required to run LingBot‑Map?
With aggressive configuration—patch‑aligned resolution of 384×518, a sliding window of 32, keyframe interval of 4, and CPU off‑loading—you can process thousands of frames using under 4 GiB of VRAM. The exact requirement scales linearly with window size and resolution.
How does the sliding‑window KV cache differ from standard transformer caches?
Standard transformers cache all past key‑value pairs, causing memory usage to grow with sequence length. LingBot‑Map’s sliding window, implemented in gct_stream_window.py, evicts older blocks via _execute_deferred_eviction() once the window fills, bounding memory to O(window size) regardless of video length.
Can I use flow‑based keyframes with a fixed keyframe interval?
Yes. The inference_streaming() method checks both conditions: it skips cache writes for frames below the flow_threshold motion value, and also respects the modulo logic of keyframe_interval. These mechanisms work together to minimize cache writes for static or redundant frames.
Where should I place the CPU off‑loading in my inference pipeline?
Set output_device=torch.device('cpu') when calling inference_streaming() in gct_stream_window.py. The internal _to_out() helper handles the transfer immediately after each forward pass, ensuring the GPU never accumulates the full output tensor history.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →