# How to Optimize LingBot-MAP for Limited VRAM Using `--offload_to_cpu` and `--num_scale_frames`

> Optimize LingBot-MAP for limited VRAM by using --offload_to_cpu and --num_scale_frames. Reduce memory usage and maintain inference speed on your GPU.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-31

---

**Use `--num_scale_frames` to reduce the transformer KV cache size and `--offload_to_cpu` to move per-frame predictions to CPU memory after each forward pass, significantly lowering GPU memory consumption while maintaining inference capability.**

LingBot-MAP is a 3D reconstruction framework that processes video streams through a transformer-based architecture, requiring substantial GPU memory to maintain key-value caches and intermediate outputs. When running inference on hardware with constrained VRAM—such as 4GB-class consumer GPUs—memory optimization becomes critical to prevent out-of-memory errors. The repository provides two specific flags, `--offload_to_cpu` and `--num_scale_frames`, that directly control memory allocation strategies within the `GCTStream` model architecture.

## Understanding VRAM Bottlenecks in LingBot-MAP

The LingBot-MAP model consumes GPU memory through two primary mechanisms: the **key-value (KV) cache** that stores attention states for previous frames, and the **per-frame prediction tensors** containing poses, depth maps, and point clouds. According to the source code in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py), the KV cache grows linearly with the number of frames retained for scale processing, while the output tensors from each forward pass can dominate peak memory usage during long sequences.

## Reducing Memory with `--num_scale_frames`

The `--num_scale_frames` parameter controls how many initial "scale frames" are processed together before streaming begins, directly impacting the size of the transformer KV cache.

### How Scale Frames Affect KV Cache Allocation

In [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (lines 31-48), this flag is passed to the `GCTStream` constructor as `kv_cache_scale_frames`. These frames are duplicated in the KV cache for each subsequent frame during inference. Lower values reduce the cache size linearly—setting `--num_scale_frames 4` creates a significantly smaller cache than the default 8 or 16 frames. The cache is allocated once per frame and persists until the sequence ends, making this parameter the most direct lever for controlling base memory consumption.

### Integration with Inference Methods

The parameter flows through to `model.inference_streaming` and `model.inference_windowed` (lines 45-53 in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py)), where it influences cache allocation within the transformer layers defined in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py). Reducing this value limits the context available for initial reconstruction but frees substantial VRAM for the actual model weights and temporary activations.

## Offloading Predictions with `--offload_to_cpu`

While reducing the KV cache saves memory, the per-frame outputs—dense depth maps and point cloud tensors—can still exhaust available VRAM during processing.

### CPU Memory Management Strategy

The `--offload_to_cpu` flag, parsed in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (lines 88-93), sets `output_device = torch.device("cpu")` (lines 43-45). This device parameter is passed to the inference functions and respected in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), where the model explicitly calls `tensor.to(output_device)` before returning predictions. This ensures that large output tensors immediately move to system RAM rather than accumulating in GPU memory between frames.

### Performance Characteristics

The offload operation introduces PCI-e transfer overhead, but the impact is typically minimal compared to the memory savings achieved. The model weights remain on GPU for accelerated computation, while only the final predictions migrate to CPU. This strategy is particularly effective when processing high-resolution outputs or long video sequences where intermediate results would otherwise accumulate.

## Practical Configuration Examples

Configure these flags based on your available VRAM. For systems with approximately 4GB of GPU memory, use conservative settings that minimize both cache and resident outputs.

**Minimal VRAM Configuration (4GB GPUs):**

```bash
python demo.py \
    --model_path checkpoint.pt \
    --image_folder path/to/images \
    --mode streaming \
    --num_scale_frames 4 \
    --offload_to_cpu \
    --keyframe_interval 2

```

**Balanced Quality and Memory (6-8GB GPUs):**

```bash
python demo.py \
    --model_path models/gct.pt \
    --video_path my_video.mp4 \
    --mode streaming \
    --num_scale_frames 6 \
    --offload_to_cpu \
    --keyframe_interval 3

```

Monitor actual memory usage by checking `torch.cuda.memory_allocated()` output printed by the demo script to find the optimal balance for your specific hardware.

## Summary

- **`--num_scale_frames`** directly limits the transformer KV cache size by controlling how many initial anchor frames populate the attention mechanism, with lower values linearly reducing GPU memory requirements.
- **`--offload_to_cpu`** moves prediction tensors (poses, depth, point clouds) to CPU memory immediately after generation, preventing accumulation of large intermediate results in VRAM.
- These flags can be combined safely: reduce scale frames to minimize persistent cache, then offload outputs to handle temporary peak usage.
- The implementation in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) ensures device placement happens correctly without modifying the underlying model architecture.

## Frequently Asked Questions

### What is the KV cache in LingBot-MAP and why does it consume so much VRAM?

The KV cache stores key and value tensors from the transformer's self-attention layers for previously processed frames, enabling the model to maintain temporal coherence in 3D reconstruction. Because LingBot-MAP duplicates these cached states for each scale frame and retains them throughout the sequence, the memory grows proportionally with the number of frames and the `--num_scale_frames` parameter, often consuming gigabytes of VRAM in default configurations.

### How low can I set `--num_scale_frames` without breaking the reconstruction?

You can set `--num_scale_frames` as low as 2, though this reduces the initial context available for establishing world scale and may degrade reconstruction quality for the first several frames. The model requires at least 2 frames to compute initial relative poses, but values below 4 may produce less stable camera trajectories in `inference_streaming` mode according to the implementation in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py).

### Does `--offload_to_cpu` affect the quality of depth predictions or pose estimation?

No, `--offload_to_cpu` does not affect computational precision or reconstruction quality because it only changes where tensors are stored after the forward pass completes. The model performs all calculations in full precision on the GPU; the `tensor.to(output_device)` operation in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) merely copies results to CPU memory without quantization or modification.

### Can I use these flags with the windowed inference mode instead of streaming?

Yes, both flags work with `--mode windowed`. The `kv_cache_scale_frames` parameter controls cache initialization in `model.inference_windowed` (called in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) lines 45-53), while `output_device` ensures predictions move to CPU regardless of whether you use streaming or windowed processing. Windowed mode may require different tuning for `--num_scale_frames` depending on your window size and overlap settings.