How to Troubleshoot Out-of-Memory Errors with offload_to_cpu and num_scale_frames in LingBot‑Map

Enable --offload_to_cpu (default) and reduce --num_scale_frames from 8 to 2–4 to keep GPU memory usage constant during long sequences, allowing LingBot‑Map to run on GPUs with as little as 8 GB of VRAM.

LingBot‑Map’s streaming 3‑D reconstruction maintains a paged KV‑cache that grows with every stored keyframe, often triggering out‑of-memory (OOM) errors on consumer GPUs. The demo.py script provides two critical command‑line flags—--offload_to_cpu and --num_scale_frames—that directly control peak VRAM consumption by managing where intermediate tensors reside and how many frames participate in initial scale estimation.

Understanding the KV‑Cache Bottleneck

The LingBot‑Map pipeline implements a paged KV‑cache that stores activations for every keyframe (or “scale frame”) processed during streaming reconstruction. As this cache expands to accommodate long sequences, it can exceed available VRAM, causing the process to abort with a CUDA OOM error.

According to the Robbyant/lingbot-map source code, memory pressure manifests primarily in two phases:

  • Scale-phase spike: The initial global scale estimation processes multiple frames simultaneously (default = 8), creating a temporary memory peak before the main streaming loop begins.
  • Per-frame accumulation: Without explicit memory management, dense depth and color map predictions accumulate on the GPU across thousands of frames.

The Two Key Memory Management Flags

--offload_to_cpu

The --offload_to_cpu flag (enabled by default via store_true in the argument parser) moves per-frame predictions from GPU to host CPU memory immediately after each forward pass. According to the implementation in demo.py, this invokes tensor.to('cpu') on the dense depth/color maps and their gradients, freeing GPU memory that would otherwise be occupied by full-resolution activations.

  • Default: Enabled (use --no-offload_to_cpu to disable)
  • Impact: Keeps GPU footprint roughly constant regardless of image resolution or sequence length
  • Trade-off: Minimal latency increase from host-device transfers

--num_scale_frames

The --num_scale_frames N parameter controls how many frames the “scale” stage samples to estimate global scene scale before streaming begins. Reducing this value lowers the temporary activation memory during initialization, though it may slightly reduce initial scale accuracy.

  • Default: 8
  • Recommended for OOM: 2–4
  • Implementation: Parsed in demo.py and forwarded to the LingBotMap inference class

Step-by-Step Troubleshooting Configuration

When encountering OOM errors, apply these configurations in demo.py starting from the most conservative:

1. Reduce scale-phase memory (minimal quality impact):

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --num_scale_frames 2

2. Ensure CPU offloading is active (default behavior):

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --offload_to_cpu \
    --num_scale_frames 2

3. Combine with keyframe interval tuning:

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --keyframe_interval 4 \
    --offload_to_cpu \
    --num_scale_frames 2

4. Verify offloading is disabled (only for high-VRAM GPUs):

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --no-offload_to_cpu

Implementation Details

The actual flag parsing resides in demo.py, where the argument parser registers these switches and forwards values to the LingBotMap inference class. The same parameters are respected in batch processing workflows via benchmark/run.py.

As documented in the Performance & Memory section of README.md, these flags work hierarchically:

  1. --keyframe_interval limits how many frames persist in the KV‑cache
  2. --offload_to_cpu ensures non-keyframe predictions do not accumulate on GPU
  3. --num_scale_frames reduces the initial memory spike before streaming commences

This three-tier approach enables processing of thousands of frames on GPUs with limited VRAM.

Summary

  • Out-of-memory errors in LingBot‑Map stem from an unbounded KV‑cache that grows with stored keyframes and scale frames.
  • --offload_to_cpu (default: enabled) moves predictions to CPU after each forward pass, maintaining constant GPU memory usage.
  • --num_scale_frames 2 reduces the initial scale-estimation memory peak from 8 frames to 2, with negligible quality loss.
  • Both flags are parsed in demo.py and passed to the LingBotMap class, with similar implementations available in benchmark/run.py.
  • Together, these settings allow the pipeline to run on GPUs with as little as 8 GB of VRAM while processing long sequences.

Frequently Asked Questions

What is the default value of --num_scale_frames?

The default value is 8, as defined in the argument parser in demo.py. This provides robust initial scale estimation but creates higher memory usage during the first phase of reconstruction.

Does --offload_to_cpu reduce reconstruction quality?

No. This flag only affects memory placement by calling tensor.to('cpu') after predictions are computed. The full-precision tensors are preserved on the host and can be moved back to GPU if needed for subsequent operations.

How does --keyframe_interval interact with these memory flags?

While --num_scale_frames and --offload_to_cpu manage temporary activation memory, --keyframe_interval controls the growth rate of the persistent KV‑cache. A larger interval (e.g., 4 instead of 1) reduces the total number of stored keyframes, compounding the memory savings from the other two flags.

Can I run LingBot‑Map without CPU offloading?

Yes, by passing --no-offload_to_cpu. This is only recommended if your GPU has abundant VRAM (typically 16 GB+), as it allows faster tensor access without host-device transfer overhead. For 8 GB GPUs, keep the default offloading enabled.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →