# How to Reduce GPU Memory Usage with Limited VRAM in LingBot‑Map

> Learn how to reduce GPU memory usage with limited VRAM in LingBot-Map. Discover techniques like patch-aligned sizing, sliding-window KV caches, dynamic frame skipping, and CPU offloading.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-26

---

**LingBot‑Map reduces GPU memory usage with limited VRAM through patch‑aligned image sizing, sliding‑window KV caches with configurable keyframe intervals, flow‑based dynamic frame skipping, and aggressive CPU off‑loading of intermediate predictions.**

LingBot‑Map is designed to run large‑scale video‑to‑3‑D reconstruction on a single GPU, even when processing thousands of frames. To achieve this, the repository implements several memory‑efficient mechanisms that keep the peak GPU memory footprint low without sacrificing reconstruction quality. This guide explains how to leverage these techniques to reduce GPU memory usage with limited VRAM hardware.

## Patch‑Aligned Image Sizing

The first line of defense against wasted GPU memory is **patch‑aligned image sizing**. In [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py), the `patch_align()` function automatically reduces input resolution to a multiple of the patch size (default = 14). This guarantees that the internal token grid fits exactly, avoiding wasted memory on partially filled patches.

```python

# From scripts/benchmark_gct_memory.py

def patch_align(x, patch_size=14):
    # Aligns height and width to patch size multiples

    return x[..., :patch_size*(x.shape[-2]//patch_size), 
               :patch_size*(x.shape[-1]//patch_size)]

```

When configuring your pipeline, choose modest dimensions that align to your patch size. For example, use `--height 384 --width 518` rather than arbitrary resolutions that leave incomplete patch rows in memory.

## Sliding‑Window KV Cache Management

The core memory optimization lives in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py). Instead of caching key‑value tensors for every frame in a sequence, LingBot‑Map implements a **sliding‑window KV cache** that limits growth to *O(window size)* rather than *O(sequence length)*.

### Configuring the Window Size

Set the `kv_cache_sliding_window` parameter to constrain how many past frames remain in GPU memory. The eviction logic is handled by `_set_defer_eviction()` and `_execute_deferred_eviction()`:

```python

# From gct_stream_window.py

def _set_defer_eviction(self, kv_cache):
    # Marks old blocks for eviction when window slides

    pass

def _execute_deferred_eviction(self):
    # Actually frees the GPU memory

    pass

```

For limited VRAM budgets, reduce the window aggressively:

```bash
--kv-cache-sliding-window 32  # Default is 64; 32 uses ~50% less cache memory

```

### Keyframe Interval Optimization

The `inference_streaming()` method supports a `keyframe_interval` parameter that stores KV entries only for every *N*‑th frame. Non‑keyframes trigger `_set_skip_append(True)`, preventing cache writes entirely:

```python

# From gct_stream_window.py lines 100-106

if frame_idx % keyframe_interval != 0:
    self._set_skip_append(True)  # Skip KV cache update

```

Increase the interval to reduce cached frames by a factor of `keyframe_interval`:

```bash
--keyframe-interval 4  # Stores only 1/4 of frames in cache

```

### Flow‑Based Dynamic Keyframes

For scenes with variable motion, enable **flow‑based dynamic keyframe detection**. The model computes cheap optical‑flow magnitude on‑the‑fly; frames with motion below `--flow-threshold` are treated as non‑keyframes automatically:

```python

# From gct_stream_window.py lines 63-94

flow_magnitude = compute_flow(prev_frame, curr_frame)
if flow_magnitude < flow_threshold:
    # Treat as non-keyframe, skip cache write

    self._set_skip_append(True)

```

This adapts cache size to scene dynamics, saving significant memory when the camera is static.

## CPU Off‑Loading Strategies

When you only need final reconstruction results—not intermediate tensors—configure **CPU off‑loading** via the `output_device` parameter. In `inference_streaming()`, the `_to_out()` helper moves predictions to CPU before concatenation:

```python

# From gct_stream_window.py lines 18-38

def _to_out(self, tensor):
    if self.output_device is not None:
        return tensor.to(self.output_device)
    return tensor

```

Set `--output-device cpu` to keep the GPU reserved strictly for the KV cache and current frame processing, while accumulated predictions reside in system RAM.

## Input Materialization and Allocator Configuration

Prevent PyTorch from pre‑allocating large static memory pools with two additional techniques found in [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py).

**Synthetic Input Materialization:** The `--materialize-inputs` flag (CLI) forces each frame tensor to live on CPU until just before the forward pass. In `make_source_images()`, this prevents a huge pre‑allocation on GPU during long sequences.

**CUDA Allocator Tuning:** The benchmark script sets `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` to let the allocator grow only as needed:

```python

# From benchmark_gct_memory.py lines 38-40

os.environ.setdefault(
    "PYTORCH_CUDA_ALLOC_CONF", 
    "expandable_segments:True"
)

```

**Explicit Cleanup:** After each run, call `torch.cuda.empty_cache()` and `torch.cuda.reset_peak_memory_stats()` (wrapped in the `cleanup()` function) to release stray allocations and guarantee a fresh memory state.

## Complete Configuration Example

Combine these techniques to run large sequences on under 4 GiB of VRAM:

```bash
python scripts/benchmark_gct_memory.py \
    --height 384 --width 518 \
    --patch-size 14 \
    --frame-counts 64 128 256 512 1024 \
    --kv-cache-sliding-window 32 \
    --keyframe-interval 4 \
    --flow-threshold 0.3 \
    --output-device cpu \
    --output gct_memory_limited.csv

```

For programmatic inference:

```python
import torch
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    kv_cache_sliding_window=32,
    kv_cache_scale_frames=8,
    kv_cache_cross_frame_special=True,
    kv_cache_include_scale_frames=True,
    use_sdpa=False,
)
model.eval().to('cuda')

# images is a (B, S, 3, H, W) tensor on CPU

pred = model.inference_streaming(
    images,
    keyframe_interval=4,
    flow_threshold=0.3,
    output_device=torch.device('cpu'),
)

```

## Summary

- **Patch‑align** inputs to avoid partially filled patches wasting memory in [`scripts/benchmark_gct_memory.py`](https://github.com/Robbyant/lingbot-map/blob/main/scripts/benchmark_gct_memory.py).
- **Limit KV cache growth** with `kv_cache_sliding_window` and automatic eviction in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py).
- **Reduce cache entries** by increasing `keyframe_interval` or enabling flow‑based dynamic keyframes to skip static frames.
- **Off‑load predictions** to CPU via `output_device` to free GPU memory immediately after processing.
- **Control allocator behavior** with `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` and explicit `torch.cuda.empty_cache()` calls.

## Frequently Asked Questions

### What is the minimum VRAM required to run LingBot‑Map?

With aggressive configuration—patch‑aligned resolution of 384×518, a sliding window of 32, keyframe interval of 4, and CPU off‑loading—you can process thousands of frames using under 4 GiB of VRAM. The exact requirement scales linearly with window size and resolution.

### How does the sliding‑window KV cache differ from standard transformer caches?

Standard transformers cache all past key‑value pairs, causing memory usage to grow with sequence length. LingBot‑Map’s sliding window, implemented in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py), evicts older blocks via `_execute_deferred_eviction()` once the window fills, bounding memory to *O(window size)* regardless of video length.

### Can I use flow‑based keyframes with a fixed keyframe interval?

Yes. The `inference_streaming()` method checks both conditions: it skips cache writes for frames below the `flow_threshold` motion value, and also respects the modulo logic of `keyframe_interval`. These mechanisms work together to minimize cache writes for static or redundant frames.

### Where should I place the CPU off‑loading in my inference pipeline?

Set `output_device=torch.device('cpu')` when calling `inference_streaming()` in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py). The internal `_to_out()` helper handles the transfer immediately after each forward pass, ensuring the GPU never accumulates the full output tensor history.