# How to Configure the Offline Rendering Pipeline for Long Video Sequences in LingBot-Map

> Learn to configure the offline rendering pipeline for long video sequences in LingBot-Map. Process vast video data efficiently with headless batch, caching, and GPU acceleration.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-27

---

**The offline rendering pipeline provides a headless batch driver that processes arbitrarily long video sequences through sliding-window inference, configurable keyframe caching, and GPU-accelerated point-cloud rasterization, keeping memory usage bounded even for tens of thousands of frames.**

LingBot-Map is an open-source geometric context transformer for real-time 3D scene reconstruction. While the interactive [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) viewer handles short clips, processing very long video sequences—such as 25,000-frame indoor walkthroughs—requires configuring the dedicated offline rendering pipeline. This system reuses the same model, FlashInfer cache, and checkpoint stack as the interactive demo, but adds batch-style drivers that perform sliding-window inference, optional sky-masking, and final MP4 compositing without real-time viewer constraints.

## Prerequisites and Installation

Before configuring the pipeline, you must install the core package along with rendering-specific dependencies and compile the required CUDA extensions.

### Core Dependencies

Install the base `lingbot-map` package and PyTorch with CUDA support:

```bash
conda create -n lingbot-map python=3.10 -y
conda activate lingbot-map
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128
pip install -e .
pip install --index-url https://pypi.org/simple flashinfer-python

```

### Rendering Extras and CUDA Extensions

The pipeline requires `open3d`, `pyyaml`, `onnxruntime-gpu`, and NVIDIA Kaolin for voxelization and frustum culling. Install these along with the custom CUDA kernels:

```bash
pip install -e ".[vis,render]"
pip install onnxruntime-gpu
pip install --index-url https://pypi.org/simple \
    kaolin -f https://nvidia-kaolin.s3.us-east-2.amazonaws.com/torch-2.8.0_cu128.html

# Build the CUDA extensions (required once)

cd demo_render/render_cuda_ext && python setup.py build_ext --inplace && cd ../..

```

The `voxel_morton_ext` and `frustum_cull_ext` extensions in `demo_render/render_cuda_ext/` are imported by the rasterizer (`rgbd_render`) to accelerate point-cloud processing.

## Understanding the Pipeline Architecture

The offline rendering pipeline solves three critical constraints that prevent the interactive viewer from handling very long sequences.

### Memory Constraints via Paged KV-Cache

The **Geometric Context Transformer** maintains a paged KV cache whose size scales with the number of keyframes. In [`lingbot_map/layers/flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/flashinfer_cache.py), the cache implementation stores attention keys and values for geometric context windows. Without mitigation, this cache would exceed GPU memory after a few hundred frames.

### RoPE Training Limitations

The model was trained on RoPE (Rotary Position Embedding) windows of approximately 320 frames. Beyond this limit, positional encoding accuracy degrades. The batch driver resets RoPE state between windows to preserve pose quality.

### Batch-Mode Compositing

Rendering a point-cloud fly-through requires all predictions before compositing. The interactive viewer cannot accumulate 10,000+ frames of depth predictions in memory. The offline pipeline stores per-frame predictions as NPZ files during inference, then runs a single GPU-accelerated rasterization pass to generate the final MP4.

## Configuring Sliding-Window Inference

For sequences longer than the training window, you must enable sliding-window mode to keep memory bounded.

### Window Size and Overlap

In [`demo_render/batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo_render/batch_demo.py), the `--mode windowed` flag activates the sliding-window driver. The system splits the video stream into overlapping windows, resetting KV-cache slots per window:

```bash
python demo_render/batch_demo.py \
    --video_path /data/long_video.mp4 \
    --output_folder /data/outputs/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8

```

**Critical parameters for long sequences:**
- **`--window_size`**: Sets the number of KV-cache slots per window (typically 128, comprising 8 scale frames + 120 keyframe slots).
- **`--overlap_keyframes`**: Shares N keyframes between consecutive windows (e.g., 8) to maintain pose alignment across window boundaries.

### Keyframe Interval Optimization

The **`--keyframe_interval`** flag controls cache growth by storing only every k-th frame in the KV cache. Non-keyframes still produce depth predictions but do not expand the cache:

```bash
python demo_render/batch_demo.py \
    --video_path long_seq.mp4 \
    --output_folder ./out/ \
    --model_path ./lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 13 \
    --save_predictions

```

Setting `--keyframe_interval 13` for a 25,000-frame sequence reduces cache entries from 25,000 to approximately 1,923, keeping GPU memory usage constant regardless of sequence length.

## Step-by-Step Configuration Guide

Follow this workflow to configure the pipeline for a 25,000-frame indoor walkthrough.

### Step 1: Prepare Input Data

The pipeline accepts either video files (MP4) or folders of images. Ensure your data is accessible at a high-throughput location:

```bash

# Video input

python demo_render/batch_demo.py \
    --video_path /data/demo_videos/indoor_walk.mp4 \
    --output_folder /data/outputs/indoor_walk/ \
    [other flags...]

# Image folder input

python demo_render/batch_demo.py \
    --image_folder ./example/university/ \
    --output_folder ./out/university/ \
    [other flags...]

```

### Step 2: Select a YAML Configuration

YAML files in `demo_render/config/` decouple camera trajectory design from CLI arguments. The [`indoor.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/indoor.yaml) preset provides sensible defaults for long indoor walks:

```yaml

# demo_render/config/indoor.yaml (excerpt)

camera:
  fov: 60.0
  transition: 30
  segments:
    - mode: follow
      frames: [0, 3000]
      back_offset: 0.25
      up_offset: 0.1
      look_offset: 0.5
    - mode: birdeye
      frames: [3000, 3500]
      reveal_height_mult: 3.0
    - mode: follow
      frames: [3500, -1]
      back_offset: 0.3
      up_offset: 0.08
      look_offset: 0.4

```

Reference this configuration with `--config demo_render/config/indoor.yaml`.

### Step 3: Execute the Batch Driver

Combine all flags for a complete long-sequence render:

```bash
python demo_render/batch_demo.py \
    --video_path /data/demo_videos/indoor_walk.mp4 \
    --output_folder /data/outputs/indoor_walk/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 13 \
    --overlap_keyframes 8 \
    --mask_sky \
    --sky_mask_dir /data/outputs/sky_masks \
    --sky_mask_visualization_dir /data/outputs/sky_mask_viz \
    --camera_vis default \
    --keyframes_only_points \
    --frame_tag \
    --frame_tag_position top_right \
    --save_predictions

```

## Advanced Configuration Options

### Sky Masking for Outdoor Scenes

Enable the ONNX sky-segmentation model to filter sky points and improve visual quality for outdoor or large-scale indoor scenes:

```bash
python demo_render/batch_demo.py \
    --video_path outdoor_seq.mp4 \
    --output_folder ./out/ \
    --model_path ./lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 10 \
    --mask_sky

```

The [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) script loads `skyseg.onnx` via `onnxruntime-gpu` when `--mask_sky` is supplied, running GPU-accelerated segmentation before point-cloud generation.

### Memory-Efficient Point Clouds

The **`--keyframes_only_points`** flag restricts depth unprojection to keyframe indices only, producing sparse point clouds that consume significantly less disk space and GPU memory during rendering.

### Persistent Predictions

Always include **`--save_predictions`** when processing long sequences. This persists per-frame NPZ files to disk, allowing you to re-render with different camera trajectories or visualization settings without re-running the expensive model inference:

```bash

# First pass: save predictions

python demo_render/batch_demo.py ... --save_predictions

# Second pass: re-render with different config (predictions already exist)

python demo_render/batch_demo.py \
    --video_path long_seq.mp4 \
    --output_folder ./out_v2/ \
    --config demo_render/config/alternate.yaml \
    [other flags...]

```

### Output Artifacts

The pipeline generates several files in the output folder:
- `*_pointcloud.mp4`: The final rendered point-cloud fly-through
- `*_pointcloud_rgb.mp4`: Copy of the original RGB video
- `*_pointcloud_config.yaml`: Snapshot of the complete configuration used
- [`batch_results.json`](https://github.com/Robbyant/lingbot-map/blob/main/batch_results.json): JSON summary of processing statistics

## Summary

- **Install rendering extras**: Use `pip install -e ".[vis,render]"` and build CUDA extensions in `demo_render/render_cuda_ext/` before processing long sequences.
- **Enable sliding windows**: Set `--mode windowed` with `--window_size 128` and `--overlap_keyframes 8` to handle sequences beyond the 320-frame RoPE limit.
- **Control cache growth**: Use `--keyframe_interval 13` to limit KV-cache entries to every 13th frame, keeping GPU memory bounded for 25,000+ frame videos.
- **Persist predictions**: Include `--save_predictions` to cache NPZ files, enabling re-rendering without re-inference.
- **Optimize for outdoors**: Add `--mask_sky` to filter sky points using the ONNX segmentation model.

## Frequently Asked Questions

### What is the maximum video length the offline pipeline can handle?

The offline rendering pipeline can process arbitrarily long video sequences—tested on 25,000-frame (13-minute) walkthroughs—because the sliding-window mode (`--mode windowed`) resets the KV cache and RoPE positional encodings at each window boundary. Memory usage remains constant regardless of input length.

### How do I prevent GPU memory errors when processing thousands of frames?

Configure `--keyframe_interval` to store only every Nth frame in the KV cache (e.g., 13), set `--window_size` to 128 slots, and enable `--mode windowed`. These settings ensure the [`flashinfer_cache.py`](https://github.com/Robbyant/lingbot-map/blob/main/flashinfer_cache.py) implementation never allocates more than 128 slots per window, regardless of total video duration.

### Can I re-render the point cloud with a different camera path without re-running inference?

Yes. Include `--save_predictions` during the initial run to store per-frame NPZ files. You can then run [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) again with a different `--config` YAML file specifying alternative camera trajectories; the pipeline will detect existing predictions and skip model inference, proceeding directly to the rendering phase.

### Why does the interactive demo crash on long videos while the batch pipeline succeeds?

The interactive [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) viewer accumulates all frames in memory for real-time display, exceeding GPU memory after a few hundred frames. The offline pipeline in [`batch_demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/batch_demo.py) uses a batch-style driver that processes frames in sliding windows, writes intermediate results to disk, and performs final MP4 compositing only after all predictions are complete.