How to Run LingBot-Map on GPUs with Limited VRAM: Complete 8GB Configuration Guide
Run LingBot-Map on 8GB GPUs by reducing the KV cache size via --keyframe_interval, offloading tensors to CPU with --offload_to_cpu, and using windowed inference to bound memory complexity from linear to constant.
LingBot-Map is a feed-forward 3D foundation model that processes thousands of video frames using a paged KV cache of transformer activations. While this architecture enables streaming inference over long sequences, the cache becomes the primary VRAM consumer on consumer hardware. According to the Robbyant/lingbot-map source code, several command-line flags in demo.py allow you to trade minor computational overhead for significant memory savings, making 8GB GPUs viable for production use.
Architecture-Level Memory Controls
The LingBot-Map inference pipeline exposes explicit knobs in demo.py (lines 44–89) to control resident GPU memory. These adjustments target the two largest allocation sources: the transformer KV cache and the bidirectional scale-phase activations.
Reduce Keyframe Retention with --keyframe_interval
The KV cache stores key/value tensors for every keyframe—frames retained for attention lookups during streaming inference. By default, most frames become keyframes after the initial scale phase.
Increase --keyframe_interval N to retain only every N-th frame. This reduces cache size linearly from #keyframes × hidden_dim × head_dim to approximately #keyframes/N. For 8GB GPUs, use --keyframe_interval 4 to quarter the cache footprint.
Offload Per-Frame Predictions to CPU with --offload_to_cpu
The camera and depth heads generate per-frame tensors (depth maps, world points, pose estimates) that are not required for subsequent forward passes. The --offload_to_cpu flag triggers torch.to("cpu") operations (implemented in demo.py line 543) immediately after each frame’s prediction, freeing GPU memory before the next iteration.
Minimize Scale-Frame Overhead with --num_scale_frames
The initial scale phase processes bidirectional frames for global alignment, creating a temporary activation stack whose size scales with --num_scale_frames (default 8). Reducing this to 2 via --num_scale_frames 2 cuts the activation peak by approximately 2–3GB, as noted in the comments around lines 47–51 of demo.py.
Select Efficient Attention Backends
LingBot-Map supports two attention implementations:
- FlashInfer (recommended): Provides a paged KV cache that releases unused pages dynamically. Install via
pip install flashinfer-pythonto enable this backend. - SDPA (fallback): Allocates a dense cache and requires larger sliding windows. Activate with
--use_sdpaonly if FlashInfer installation fails.
FlashInfer’s paging mechanism is essential for bounded memory usage during long sequences.
Enable Windowed Inference for Long Sequences
For videos exceeding 3000 frames, use --mode windowed to reset the KV cache after each window. This changes memory complexity from O(total_frames) to O(window_size).
Configure with:
--window_size 64(default): Maximum keyframes per window--overlap_keyframes 8: Context frames preserved across windows to maintain pose continuity (seelingbot_map/models/gct_stream_window.py)
Step-by-Step Workflow for 8GB GPUs
Follow this sequence to configure a constrained GPU environment according to the repository's Performance & Memory documentation (README.md lines 93–108).
-
Install dependencies with CUDA 12.8 support and FlashInfer:
conda create -n lingbot python=3.10 conda activate lingbot pip install torch==2.8.0 --index-url https://download.pytorch.org/whl/cu128 pip install flashinfer-python -
Run streaming inference for standard videos (under 3000 frames):
python demo.py \ --model_path ./lingbot-map.pt \ --video_path ./sequence.mp4 \ --fps 10 \ --mode streaming \ --keyframe_interval 4 \ --offload_to_cpu \ --num_scale_frames 2 -
Switch to windowed mode for very long videos to bound cache size:
python demo.py \ --model_path ./lingbot-map.pt \ --video_path ./long_walkthrough.mp4 \ --fps 15 \ --mode windowed \ --window_size 64 \ --overlap_keyframes 8 \ --keyframe_interval 4 \ --offload_to_cpu \ --num_scale_frames 2 -
Optional accuracy trade-off: Reduce camera refinement iterations to lower intermediate tensor allocations:
--camera_num_iterations 1
Technical Deep Dive: Why These Optimizations Work
Understanding the implementation details in lingbot_map/models/gct_stream.py and gct_stream_window.py clarifies why these specific flags resolve VRAM bottlenecks.
KV-Cache Paging Mechanism: FlashInfer stores key/value pairs in paginated buffers rather than contiguous tensors. When combined with --keyframe_interval 4, the system allocates only 25% of the original pages, directly shrinking the working set size.
CPU Offloading Strategy: The per-frame depth and pose tensors consume significant memory but have no dependencies in the autograd graph for future frames. By moving these to host RAM immediately after generation (line 543 in demo.py), the GPU retains only the essential KV cache and model weights.
Windowed Reset Logic: gct_stream_window.py explicitly clears the cache between windows, ensuring that memory usage plateaus at the level required for --window_size keyframes regardless of input video length. The overlap logic maintains geometric consistency without requiring full-sequence attention.
Mixed-Precision Optimization: Lines 48–61 of demo.py automatically cast the DINOv2-style backbone to bfloat16 (or fp16 on older GPUs) while keeping camera and depth heads in fp32. This selective precision saves approximately 2GB without degrading output quality.
Summary
- Reduce keyframe density using
--keyframe_interval 4to shrink the KV cache linearly. - Move predictions to CPU immediately via
--offload_to_cputo free GPU memory after each frame. - Lower scale frames from 8 to 2 using
--num_scale_frames 2to cut initial activation peaks by 2–3GB. - Install FlashInfer for paged KV-cache management; avoid
--use_sdpaunless necessary. - Use windowed inference (
--mode windowed) with--window_size 64to achieve constant memory complexity for videos of any length. - Apply mixed precision automatically handled in
demo.pyfor an additional ~2GB savings.
Frequently Asked Questions
Can LingBot-Map run on a GPU with only 6GB of VRAM?
Yes, though you must use aggressive settings. Set --keyframe_interval 3 or higher, ensure --offload_to_cpu is enabled, and use --mode windowed with --window_size 48 to strictly bound the cache. You may also need to process at lower resolution or use --use_sdpa with a reduced sliding window if FlashInfer page sizes are too large for your specific card.
What is the performance impact of using --offload_to_cpu?
The impact is typically minimal (under 10% slowdown) because only prediction tensors are moved, not model parameters or the KV cache. The operation uses non-blocking torch.to("cpu") calls that overlap with the next frame’s preprocessing. However, if your CPU RAM is constrained, this may trigger paging to disk, which would significantly degrade performance.
How does windowed mode affect mapping accuracy compared to streaming mode?
Windowed mode maintains high accuracy through --overlap_keyframes, which preserves geometric context across window boundaries. According to the implementation in gct_stream_window.py, overlapping 8 keyframes provides sufficient continuity for smooth pose stitching. While extreme window sizes (below 32) may introduce slight drift, the default 64-frame window with 8 overlaps matches streaming quality for most practical sequences.
Is FlashInfer required, or can I use standard PyTorch attention?
FlashInfer is strongly recommended for limited VRAM because its paged KV cache releases memory dynamically. The fallback --use_sdpa allocates dense tensors that remain resident until explicitly cleared. If you must use SDPA, increase --keyframe_interval to 2 or 3 and reduce --window_size to 32 to compensate for the lack of paging.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →