# Can LingBot-Map Achieve 20 FPS Streaming 3D Reconstruction?

> Discover if LingBot-Map achieves 20 FPS streaming 3D reconstruction. Learn about its impressive performance, even on long sequences, maintaining speed and quality.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: performance
- Published: 2026-07-27

---

**Yes, LingBot-Map achieves approximately 20 FPS streaming 3D reconstruction at 518 × 378 resolution, maintaining this performance even for sequences exceeding 10,000 frames.**

LingBot-Map is a real-time neural mapping system from the Robbyant/lingbot-map repository designed for long-sequence 3D reconstruction. The system sustains **20 FPS streaming 3D reconstruction** through a feed-forward transformer architecture and optimized CUDA kernels that eliminate per-frame iterative optimization bottlenecks.

## Architectural Innovations Enabling Real-Time Performance

### Geometric Context Transformer

The core of LingBot-Map's speed comes from its **Geometric Context Transformer**, implemented in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) at line 36. Unlike traditional SLAM systems that rely on iterative bundle adjustment, this feed-forward transformer unifies coordinate grounding, dense geometry cues, and long-range drift correction in a single forward pass. By avoiding iterative optimization loops, the architecture guarantees constant-time inference regardless of scene complexity.

### Paged KV-Cache Attention with FlashInfer

The `FlashInferAttention` class (line 52 of [`attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/attention.py)) implements a custom paged KV-cache using FlashInfer's highly optimized kernels. This backend provides up to **3× speed-up** over the native PyTorch SDPA fallback by fusing attention operations into custom CUDA kernels. The system automatically selects FlashInfer when available, falling back to standard attention only when necessary.

### Sliding-Window Eviction Strategy

To handle arbitrary sequence lengths without memory explosion, LingBot-Map implements `_apply_kv_cache_eviction_causal` at line 99 of [`attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/attention.py). This mechanism maintains two critical parameters:

- **`kv_cache_sliding_window`**: Default value of 64 frames
- **`kv_cache_scale_frames`**: Default value of 8 frames

Together, these ensure a **bounded memory footprint** and constant-time attention computation regardless of total sequence length, enabling the system to process 10,000+ frame sequences without degradation.

## Memory Optimization Techniques

### Keyframe Interval Configuration

The `--keyframe_interval` parameter reduces cache size by storing only every *N*-th frame in the KV cache. For very long videos, increasing this interval (e.g., to 4 or 10) significantly reduces memory traffic while preserving reconstruction quality. This is particularly effective when combined with the sliding-window mechanism.

### Page-Aligned Tensor Layout

In [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py) at line 74, the `_write_patch_page` function manages a page-aligned tensor layout that avoids padding and enables direct GPU memory writes. This design minimizes memory copies and improves cache locality, contributing to the sustained **20 FPS throughput** during streaming inference.

## Running LingBot-Map at 20 FPS

### Interactive Real-Time Demo

Run the browser-based viewer with default settings to achieve approximately 20 FPS:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --image_folder example/courthouse \
    --mask_sky \
    --keyframe_interval 2

```

The demo automatically selects the FlashInfer backend when available.

### Force FlashInfer for Maximum Performance

To explicitly enable the high-performance path and configure the sliding-window cache:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --image_folder example/loop \
    --backend flashinfer \
    --keyframe_interval 4 \
    --window_size 128

```

This configuration sets `kv_cache_sliding_window=64` (2× per-frame tokens) and leverages the optimized kernels in `FlashInferAttention`.

### Offline Rendering for Long Sequences

Process videos with 25,000+ frames using windowed inference:

```bash
python demo_render/batch_demo.py \
    --video_path /data/indoor_travel.MP4 \
    --output_folder /data/out/indoor_travel \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 10 \
    --overlap_keyframes 8

```

The `--mode windowed` flag triggers the sliding-window inference that keeps the KV cache bounded while preserving ~20 FPS throughput.

## Summary

- **LingBot-Map achieves 20 FPS** at 518 × 378 resolution through its feed-forward Geometric Context Transformer, eliminating iterative optimization bottlenecks.
- **FlashInfer integration** in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py) delivers 3× speed-up over PyTorch SDPA via kernel-fused attention.
- **Paged KV-cache** with sliding-window eviction (default 64 frames) ensures constant memory usage regardless of sequence length.
- **Keyframe intervals** and page-aligned tensor layouts in [`flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/flashinfer_cache.py) minimize memory traffic for long videos.
- The system requires **CUDA 12.8**, **PyTorch 2.8**, and optional **FlashInfer** for optimal performance.

## Frequently Asked Questions

### What hardware is required to achieve 20 FPS streaming 3D reconstruction?

LingBot-Map requires a CUDA-capable GPU with CUDA 12.8 and PyTorch 2.8 installed. For the full 20 FPS performance, install the optional `flashinfer-python` package to enable the optimized FlashInfer backend, which provides up to 3× speed-up over the CPU/GPU fallback.

### How does LingBot-Map handle sequences longer than 10,000 frames without slowing down?

The system uses a **sliding-window KV cache** implemented in `_apply_kv_cache_eviction_causal` that evicts older frames while retaining "scale" frames for geometric consistency. This bounds memory usage to the configured window size (default 64 frames) regardless of total video length, ensuring constant-time inference per frame.

### Can I adjust the trade-off between reconstruction quality and FPS?

Yes. Increase `--keyframe_interval` to store fewer frames in the cache, boosting throughput at the cost of some temporal consistency. Alternatively, adjust `--window_size` to change the effective `kv_cache_sliding_window` (typically half the window size in tokens). For maximum quality at reduced speed, decrease the keyframe interval or use the PyTorch SDPA backend instead of FlashInfer.

### Where is the core attention mechanism implemented?

The transformer attention modules—including `Attention`, `CausalAttention`, `FlashInferAttention`, and `SDPAAttention`—are located in [`lingbot_map/layers/attention.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/attention.py). The low-level paged cache manager used by FlashInfer is implemented in [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py), specifically in functions like `_write_patch_page`.