How to Optimize LingBot-MAP for Limited VRAM Using `--offload_to_cpu` and `--num_scale_frames`

Use --num_scale_frames to reduce the transformer KV cache size and --offload_to_cpu to move per-frame predictions to CPU memory after each forward pass, significantly lowering GPU memory consumption while maintaining inference capability.

LingBot-MAP is a 3D reconstruction framework that processes video streams through a transformer-based architecture, requiring substantial GPU memory to maintain key-value caches and intermediate outputs. When running inference on hardware with constrained VRAM—such as 4GB-class consumer GPUs—memory optimization becomes critical to prevent out-of-memory errors. The repository provides two specific flags, --offload_to_cpu and --num_scale_frames, that directly control memory allocation strategies within the GCTStream model architecture.

Understanding VRAM Bottlenecks in LingBot-MAP

The LingBot-MAP model consumes GPU memory through two primary mechanisms: the key-value (KV) cache that stores attention states for previous frames, and the per-frame prediction tensors containing poses, depth maps, and point clouds. According to the source code in lingbot_map/models/gct_base.py, the KV cache grows linearly with the number of frames retained for scale processing, while the output tensors from each forward pass can dominate peak memory usage during long sequences.

Reducing Memory with --num_scale_frames

The --num_scale_frames parameter controls how many initial "scale frames" are processed together before streaming begins, directly impacting the size of the transformer KV cache.

How Scale Frames Affect KV Cache Allocation

In demo.py (lines 31-48), this flag is passed to the GCTStream constructor as kv_cache_scale_frames. These frames are duplicated in the KV cache for each subsequent frame during inference. Lower values reduce the cache size linearly—setting --num_scale_frames 4 creates a significantly smaller cache than the default 8 or 16 frames. The cache is allocated once per frame and persists until the sequence ends, making this parameter the most direct lever for controlling base memory consumption.

Integration with Inference Methods

The parameter flows through to model.inference_streaming and model.inference_windowed (lines 45-53 in demo.py), where it influences cache allocation within the transformer layers defined in lingbot_map/models/gct_stream.py. Reducing this value limits the context available for initial reconstruction but frees substantial VRAM for the actual model weights and temporary activations.

Offloading Predictions with --offload_to_cpu

While reducing the KV cache saves memory, the per-frame outputs—dense depth maps and point cloud tensors—can still exhaust available VRAM during processing.

CPU Memory Management Strategy

The --offload_to_cpu flag, parsed in demo.py (lines 88-93), sets output_device = torch.device("cpu") (lines 43-45). This device parameter is passed to the inference functions and respected in lingbot_map/models/gct_stream.py, where the model explicitly calls tensor.to(output_device) before returning predictions. This ensures that large output tensors immediately move to system RAM rather than accumulating in GPU memory between frames.

Performance Characteristics

The offload operation introduces PCI-e transfer overhead, but the impact is typically minimal compared to the memory savings achieved. The model weights remain on GPU for accelerated computation, while only the final predictions migrate to CPU. This strategy is particularly effective when processing high-resolution outputs or long video sequences where intermediate results would otherwise accumulate.

Practical Configuration Examples

Configure these flags based on your available VRAM. For systems with approximately 4GB of GPU memory, use conservative settings that minimize both cache and resident outputs.

Minimal VRAM Configuration (4GB GPUs):

python demo.py \
    --model_path checkpoint.pt \
    --image_folder path/to/images \
    --mode streaming \
    --num_scale_frames 4 \
    --offload_to_cpu \
    --keyframe_interval 2

Balanced Quality and Memory (6-8GB GPUs):

python demo.py \
    --model_path models/gct.pt \
    --video_path my_video.mp4 \
    --mode streaming \
    --num_scale_frames 6 \
    --offload_to_cpu \
    --keyframe_interval 3

Monitor actual memory usage by checking torch.cuda.memory_allocated() output printed by the demo script to find the optimal balance for your specific hardware.

Summary

  • --num_scale_frames directly limits the transformer KV cache size by controlling how many initial anchor frames populate the attention mechanism, with lower values linearly reducing GPU memory requirements.
  • --offload_to_cpu moves prediction tensors (poses, depth, point clouds) to CPU memory immediately after generation, preventing accumulation of large intermediate results in VRAM.
  • These flags can be combined safely: reduce scale frames to minimize persistent cache, then offload outputs to handle temporary peak usage.
  • The implementation in demo.py and lingbot_map/models/gct_stream.py ensures device placement happens correctly without modifying the underlying model architecture.

Frequently Asked Questions

What is the KV cache in LingBot-MAP and why does it consume so much VRAM?

The KV cache stores key and value tensors from the transformer's self-attention layers for previously processed frames, enabling the model to maintain temporal coherence in 3D reconstruction. Because LingBot-MAP duplicates these cached states for each scale frame and retains them throughout the sequence, the memory grows proportionally with the number of frames and the --num_scale_frames parameter, often consuming gigabytes of VRAM in default configurations.

How low can I set --num_scale_frames without breaking the reconstruction?

You can set --num_scale_frames as low as 2, though this reduces the initial context available for establishing world scale and may degrade reconstruction quality for the first several frames. The model requires at least 2 frames to compute initial relative poses, but values below 4 may produce less stable camera trajectories in inference_streaming mode according to the implementation in lingbot_map/models/gct_stream.py.

Does --offload_to_cpu affect the quality of depth predictions or pose estimation?

No, --offload_to_cpu does not affect computational precision or reconstruction quality because it only changes where tensors are stored after the forward pass completes. The model performs all calculations in full precision on the GPU; the tensor.to(output_device) operation in lingbot_map/models/gct_stream.py merely copies results to CPU memory without quantization or modification.

Can I use these flags with the windowed inference mode instead of streaming?

Yes, both flags work with --mode windowed. The kv_cache_scale_frames parameter controls cache initialization in model.inference_windowed (called in demo.py lines 45-53), while output_device ensures predictions move to CPU regardless of whether you use streaming or windowed processing. Windowed mode may require different tuning for --num_scale_frames depending on your window size and overlap settings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →