# How to Troubleshoot Out-of-Memory Errors with offload_to_cpu and num_scale_frames in LingBot‑Map

> Troubleshoot out-of-memory errors in LingBot-Map by adjusting offload_to_cpu and num_scale_frames. Run on GPUs with 8GB VRAM by optimizing memory usage for long sequences.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-29

---

**Enable `--offload_to_cpu` (default) and reduce `--num_scale_frames` from 8 to 2–4 to keep GPU memory usage constant during long sequences, allowing LingBot‑Map to run on GPUs with as little as 8 GB of VRAM.**

LingBot‑Map’s streaming 3‑D reconstruction maintains a paged KV‑cache that grows with every stored keyframe, often triggering out‑of-memory (OOM) errors on consumer GPUs. The [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) script provides two critical command‑line flags—`--offload_to_cpu` and `--num_scale_frames`—that directly control peak VRAM consumption by managing where intermediate tensors reside and how many frames participate in initial scale estimation.

## Understanding the KV‑Cache Bottleneck

The LingBot‑Map pipeline implements a **paged KV‑cache** that stores activations for every keyframe (or “scale frame”) processed during streaming reconstruction. As this cache expands to accommodate long sequences, it can exceed available VRAM, causing the process to abort with a CUDA OOM error.

According to the Robbyant/lingbot-map source code, memory pressure manifests primarily in two phases:

- **Scale-phase spike**: The initial global scale estimation processes multiple frames simultaneously (default = 8), creating a temporary memory peak before the main streaming loop begins.
- **Per-frame accumulation**: Without explicit memory management, dense depth and color map predictions accumulate on the GPU across thousands of frames.

## The Two Key Memory Management Flags

### --offload_to_cpu

The `--offload_to_cpu` flag (enabled by default via `store_true` in the argument parser) moves per-frame predictions from GPU to host CPU memory immediately after each forward pass. According to the implementation in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py), this invokes `tensor.to('cpu')` on the dense depth/color maps and their gradients, freeing GPU memory that would otherwise be occupied by full-resolution activations.

- **Default**: Enabled (use `--no-offload_to_cpu` to disable)
- **Impact**: Keeps GPU footprint roughly constant regardless of image resolution or sequence length
- **Trade-off**: Minimal latency increase from host-device transfers

### --num_scale_frames

The `--num_scale_frames N` parameter controls how many frames the “scale” stage samples to estimate global scene scale before streaming begins. Reducing this value lowers the temporary activation memory during initialization, though it may slightly reduce initial scale accuracy.

- **Default**: 8
- **Recommended for OOM**: 2–4
- **Implementation**: Parsed in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and forwarded to the `LingBotMap` inference class

## Step-by-Step Troubleshooting Configuration

When encountering OOM errors, apply these configurations in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) starting from the most conservative:

**1. Reduce scale-phase memory (minimal quality impact):**

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --num_scale_frames 2

```

**2. Ensure CPU offloading is active (default behavior):**

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --offload_to_cpu \
    --num_scale_frames 2

```

**3. Combine with keyframe interval tuning:**

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --keyframe_interval 4 \
    --offload_to_cpu \
    --num_scale_frames 2

```

**4. Verify offloading is disabled (only for high-VRAM GPUs):**

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/university \
    --mask_sky \
    --no-offload_to_cpu

```

## Implementation Details

The actual flag parsing resides in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py), where the argument parser registers these switches and forwards values to the `LingBotMap` inference class. The same parameters are respected in batch processing workflows via [`benchmark/run.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/run.py).

As documented in the **Performance & Memory** section of [`README.md`](https://github.com/Robbyant/lingbot-map/blob/main/README.md), these flags work hierarchically:

1. `--keyframe_interval` limits how many frames persist in the KV‑cache
2. `--offload_to_cpu` ensures non-keyframe predictions do not accumulate on GPU
3. `--num_scale_frames` reduces the initial memory spike before streaming commences

This three-tier approach enables processing of thousands of frames on GPUs with limited VRAM.

## Summary

- **Out-of-memory errors** in LingBot‑Map stem from an unbounded KV‑cache that grows with stored keyframes and scale frames.
- **`--offload_to_cpu`** (default: enabled) moves predictions to CPU after each forward pass, maintaining constant GPU memory usage.
- **`--num_scale_frames 2`** reduces the initial scale-estimation memory peak from 8 frames to 2, with negligible quality loss.
- Both flags are parsed in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and passed to the `LingBotMap` class, with similar implementations available in [`benchmark/run.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/run.py).
- Together, these settings allow the pipeline to run on GPUs with as little as 8 GB of VRAM while processing long sequences.

## Frequently Asked Questions

### What is the default value of `--num_scale_frames`?

The default value is **8**, as defined in the argument parser in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py). This provides robust initial scale estimation but creates higher memory usage during the first phase of reconstruction.

### Does `--offload_to_cpu` reduce reconstruction quality?

No. This flag only affects memory placement by calling `tensor.to('cpu')` after predictions are computed. The full-precision tensors are preserved on the host and can be moved back to GPU if needed for subsequent operations.

### How does `--keyframe_interval` interact with these memory flags?

While `--num_scale_frames` and `--offload_to_cpu` manage temporary activation memory, `--keyframe_interval` controls the growth rate of the persistent KV‑cache. A larger interval (e.g., 4 instead of 1) reduces the total number of stored keyframes, compounding the memory savings from the other two flags.

### Can I run LingBot‑Map without CPU offloading?

Yes, by passing `--no-offload_to_cpu`. This is only recommended if your GPU has abundant VRAM (typically 16 GB+), as it allows faster tensor access without host-device transfer overhead. For 8 GB GPUs, keep the default offloading enabled.