How to Configure the Offline Rendering Pipeline for Long Video Sequences in LingBot-Map

The offline rendering pipeline provides a headless batch driver that processes arbitrarily long video sequences through sliding-window inference, configurable keyframe caching, and GPU-accelerated point-cloud rasterization, keeping memory usage bounded even for tens of thousands of frames.

LingBot-Map is an open-source geometric context transformer for real-time 3D scene reconstruction. While the interactive demo.py viewer handles short clips, processing very long video sequences—such as 25,000-frame indoor walkthroughs—requires configuring the dedicated offline rendering pipeline. This system reuses the same model, FlashInfer cache, and checkpoint stack as the interactive demo, but adds batch-style drivers that perform sliding-window inference, optional sky-masking, and final MP4 compositing without real-time viewer constraints.

Prerequisites and Installation

Before configuring the pipeline, you must install the core package along with rendering-specific dependencies and compile the required CUDA extensions.

Core Dependencies

Install the base lingbot-map package and PyTorch with CUDA support:

conda create -n lingbot-map python=3.10 -y
conda activate lingbot-map
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128
pip install -e .
pip install --index-url https://pypi.org/simple flashinfer-python

Rendering Extras and CUDA Extensions

The pipeline requires open3d, pyyaml, onnxruntime-gpu, and NVIDIA Kaolin for voxelization and frustum culling. Install these along with the custom CUDA kernels:

pip install -e ".[vis,render]"
pip install onnxruntime-gpu
pip install --index-url https://pypi.org/simple \
    kaolin -f https://nvidia-kaolin.s3.us-east-2.amazonaws.com/torch-2.8.0_cu128.html

# Build the CUDA extensions (required once)

cd demo_render/render_cuda_ext && python setup.py build_ext --inplace && cd ../..

The voxel_morton_ext and frustum_cull_ext extensions in demo_render/render_cuda_ext/ are imported by the rasterizer (rgbd_render) to accelerate point-cloud processing.

Understanding the Pipeline Architecture

The offline rendering pipeline solves three critical constraints that prevent the interactive viewer from handling very long sequences.

Memory Constraints via Paged KV-Cache

The Geometric Context Transformer maintains a paged KV cache whose size scales with the number of keyframes. In lingbot_map/layers/flashinfer_cache.py, the cache implementation stores attention keys and values for geometric context windows. Without mitigation, this cache would exceed GPU memory after a few hundred frames.

RoPE Training Limitations

The model was trained on RoPE (Rotary Position Embedding) windows of approximately 320 frames. Beyond this limit, positional encoding accuracy degrades. The batch driver resets RoPE state between windows to preserve pose quality.

Batch-Mode Compositing

Rendering a point-cloud fly-through requires all predictions before compositing. The interactive viewer cannot accumulate 10,000+ frames of depth predictions in memory. The offline pipeline stores per-frame predictions as NPZ files during inference, then runs a single GPU-accelerated rasterization pass to generate the final MP4.

Configuring Sliding-Window Inference

For sequences longer than the training window, you must enable sliding-window mode to keep memory bounded.

Window Size and Overlap

In demo_render/batch_demo.py, the --mode windowed flag activates the sliding-window driver. The system splits the video stream into overlapping windows, resetting KV-cache slots per window:

python demo_render/batch_demo.py \
    --video_path /data/long_video.mp4 \
    --output_folder /data/outputs/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --overlap_keyframes 8

Critical parameters for long sequences:

  • --window_size: Sets the number of KV-cache slots per window (typically 128, comprising 8 scale frames + 120 keyframe slots).
  • --overlap_keyframes: Shares N keyframes between consecutive windows (e.g., 8) to maintain pose alignment across window boundaries.

Keyframe Interval Optimization

The --keyframe_interval flag controls cache growth by storing only every k-th frame in the KV cache. Non-keyframes still produce depth predictions but do not expand the cache:

python demo_render/batch_demo.py \
    --video_path long_seq.mp4 \
    --output_folder ./out/ \
    --model_path ./lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 13 \
    --save_predictions

Setting --keyframe_interval 13 for a 25,000-frame sequence reduces cache entries from 25,000 to approximately 1,923, keeping GPU memory usage constant regardless of sequence length.

Step-by-Step Configuration Guide

Follow this workflow to configure the pipeline for a 25,000-frame indoor walkthrough.

Step 1: Prepare Input Data

The pipeline accepts either video files (MP4) or folders of images. Ensure your data is accessible at a high-throughput location:


# Video input

python demo_render/batch_demo.py \
    --video_path /data/demo_videos/indoor_walk.mp4 \
    --output_folder /data/outputs/indoor_walk/ \
    [other flags...]

# Image folder input

python demo_render/batch_demo.py \
    --image_folder ./example/university/ \
    --output_folder ./out/university/ \
    [other flags...]

Step 2: Select a YAML Configuration

YAML files in demo_render/config/ decouple camera trajectory design from CLI arguments. The indoor.yaml preset provides sensible defaults for long indoor walks:


# demo_render/config/indoor.yaml (excerpt)

camera:
  fov: 60.0
  transition: 30
  segments:
    - mode: follow
      frames: [0, 3000]
      back_offset: 0.25
      up_offset: 0.1
      look_offset: 0.5
    - mode: birdeye
      frames: [3000, 3500]
      reveal_height_mult: 3.0
    - mode: follow
      frames: [3500, -1]
      back_offset: 0.3
      up_offset: 0.08
      look_offset: 0.4

Reference this configuration with --config demo_render/config/indoor.yaml.

Step 3: Execute the Batch Driver

Combine all flags for a complete long-sequence render:

python demo_render/batch_demo.py \
    --video_path /data/demo_videos/indoor_walk.mp4 \
    --output_folder /data/outputs/indoor_walk/ \
    --model_path /path/to/lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 13 \
    --overlap_keyframes 8 \
    --mask_sky \
    --sky_mask_dir /data/outputs/sky_masks \
    --sky_mask_visualization_dir /data/outputs/sky_mask_viz \
    --camera_vis default \
    --keyframes_only_points \
    --frame_tag \
    --frame_tag_position top_right \
    --save_predictions

Advanced Configuration Options

Sky Masking for Outdoor Scenes

Enable the ONNX sky-segmentation model to filter sky points and improve visual quality for outdoor or large-scale indoor scenes:

python demo_render/batch_demo.py \
    --video_path outdoor_seq.mp4 \
    --output_folder ./out/ \
    --model_path ./lingbot-map.pt \
    --config demo_render/config/indoor.yaml \
    --mode windowed \
    --window_size 128 \
    --keyframe_interval 10 \
    --mask_sky

The batch_demo.py script loads skyseg.onnx via onnxruntime-gpu when --mask_sky is supplied, running GPU-accelerated segmentation before point-cloud generation.

Memory-Efficient Point Clouds

The --keyframes_only_points flag restricts depth unprojection to keyframe indices only, producing sparse point clouds that consume significantly less disk space and GPU memory during rendering.

Persistent Predictions

Always include --save_predictions when processing long sequences. This persists per-frame NPZ files to disk, allowing you to re-render with different camera trajectories or visualization settings without re-running the expensive model inference:


# First pass: save predictions

python demo_render/batch_demo.py ... --save_predictions

# Second pass: re-render with different config (predictions already exist)

python demo_render/batch_demo.py \
    --video_path long_seq.mp4 \
    --output_folder ./out_v2/ \
    --config demo_render/config/alternate.yaml \
    [other flags...]

Output Artifacts

The pipeline generates several files in the output folder:

  • *_pointcloud.mp4: The final rendered point-cloud fly-through
  • *_pointcloud_rgb.mp4: Copy of the original RGB video
  • *_pointcloud_config.yaml: Snapshot of the complete configuration used
  • batch_results.json: JSON summary of processing statistics

Summary

  • Install rendering extras: Use pip install -e ".[vis,render]" and build CUDA extensions in demo_render/render_cuda_ext/ before processing long sequences.
  • Enable sliding windows: Set --mode windowed with --window_size 128 and --overlap_keyframes 8 to handle sequences beyond the 320-frame RoPE limit.
  • Control cache growth: Use --keyframe_interval 13 to limit KV-cache entries to every 13th frame, keeping GPU memory bounded for 25,000+ frame videos.
  • Persist predictions: Include --save_predictions to cache NPZ files, enabling re-rendering without re-inference.
  • Optimize for outdoors: Add --mask_sky to filter sky points using the ONNX segmentation model.

Frequently Asked Questions

What is the maximum video length the offline pipeline can handle?

The offline rendering pipeline can process arbitrarily long video sequences—tested on 25,000-frame (13-minute) walkthroughs—because the sliding-window mode (--mode windowed) resets the KV cache and RoPE positional encodings at each window boundary. Memory usage remains constant regardless of input length.

How do I prevent GPU memory errors when processing thousands of frames?

Configure --keyframe_interval to store only every Nth frame in the KV cache (e.g., 13), set --window_size to 128 slots, and enable --mode windowed. These settings ensure the flashinfer_cache.py implementation never allocates more than 128 slots per window, regardless of total video duration.

Can I re-render the point cloud with a different camera path without re-running inference?

Yes. Include --save_predictions during the initial run to store per-frame NPZ files. You can then run batch_demo.py again with a different --config YAML file specifying alternative camera trajectories; the pipeline will detect existing predictions and skip model inference, proceeding directly to the rendering phase.

Why does the interactive demo crash on long videos while the batch pipeline succeeds?

The interactive demo.py viewer accumulates all frames in memory for real-time display, exceeding GPU memory after a few hundred frames. The offline pipeline in batch_demo.py uses a batch-style driver that processes frames in sliding windows, writes intermediate results to disk, and performs final MP4 compositing only after all predictions are complete.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →