LingBot-Map Pretrained Models: `lingbot-map` vs `lingbot-map-long` Comparison

The Robbyant/lingbot-map repository provides two official pretrained checkpoints—lingbot-map for balanced short and long video performance and lingbot-map-long optimized specifically for extended sequences and large-scale scenes—both available as standalone .pt files from Hugging Face and ModelScope.

The LingBot-Map open-source codebase ships with distinct pretrained model variants designed to handle different video reconstruction scenarios. Understanding the differences between the balanced lingbot-map checkpoint and the long-sequence lingbot-map-long variant ensures optimal 3D reconstruction results across both short clips and massive outdoor drive sequences.

Available Pretrained LingBot-Map Models

According to the model-download table in the repository's README (lines 135-138), two standalone checkpoints are officially distributed:

lingbot-map: The Balanced Checkpoint

This is the default recommended model used in the original paper and standard benchmarks. The lingbot-map.pt file provides balanced performance across both short videos and moderately long sequences, making it the versatile choice for general-purpose 3D reconstruction tasks.

  • Hugging Face: robbyant/lingbot-map
  • ModelScope: Robbyant/lingbot-map
  • File: lingbot-map.pt

lingbot-map-long: Long-Sequence Optimization

The lingbot-map-long.pt checkpoint is specifically optimized for long sequences and large-scale scenes. Use this variant when processing very long indoor walkthroughs or extended outdoor drives where memory management and temporal consistency across thousands of frames become critical.

  • Hugging Face: robbyant/lingbot-map (same repository, different file)
  • ModelScope: Robbyant/lingbot-map
  • File: lingbot-map-long.pt

Downloading the Checkpoints from Hugging Face

Both models are distributed as stand-alone .pt files that do not require additional configuration files. Download them directly using standard HTTP tools:


# Download the balanced checkpoint (recommended default)

wget https://huggingface.co/robbyant/lingbot-map/resolve/main/lingbot-map.pt \
     -O /path/to/lingbot-map.pt

# Download the long-sequence-optimized checkpoint

wget https://huggingface.co/robbyant/lingbot-map/resolve/main/lingbot-map-long.pt \
     -O /path/to/lingbot-map-long.pt

Both files can also be retrieved from ModelScope using the equivalent URLs in the Robbyant/lingbot-map repository.

Running Inference with Different Model Variants

The demo scripts accept the --model_path argument to specify which checkpoint to load. Pass the absolute path to your downloaded .pt file to override any default behavior.

Standard Short-Video Reconstruction

Use the balanced lingbot-map.pt for typical reconstruction tasks:

python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --image_folder example/courthouse \
    --mask_sky

Extended Sequence Processing (25K+ Frames)

For very long indoor walkthroughs or massive datasets, switch to the long-optimized checkpoint and enable windowed processing mode:

python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --video_path /data/indoor_walkthrough.mp4 \
    --mode windowed --window_size 128 --keyframe_interval 2 \
    --mask_sky

Offline Batch Rendering

The demo_render/batch_demo.py pipeline also supports model selection via --model_path, as implemented in the repository's offline rendering utilities:

python demo_render/batch_demo.py \
    --video_path /data/drive.mp4 \
    --output_folder /data/outputs/drive/ \
    --model_path /path/to/lingbot-map-long.pt \
    --config demo_render/config/outdoor_drive.yaml \
    --mode windowed --window_size 128 --keyframe_interval 10 \
    --mask_sky

Model Loading Architecture

All inference scripts rely on the same loading routine defined in lingbot_map/models/gct_stream.py. This file contains the core streaming model class (GCTStream) that restores the architecture and weights from the saved checkpoint.

When you specify --model_path, the demo scripts instantiate GCTStream and load the state dictionary from the .pt file, rebuilding the full streaming architecture regardless of which variant (balanced or long) you select. This design ensures API consistency between both lingbot-map and lingbot-map-long checkpoints.

Summary

  • Two official variants: lingbot-map.pt (balanced) and lingbot-map-long.pt (long sequences)
  • Distribution: Both hosted on Hugging Face and ModelScope under robbyant/lingbot-map
  • File format: Stand-alone PyTorch .pt files requiring no additional configuration
  • Loading mechanism: Unified through lingbot_map/models/gct_stream.py via the --model_path argument
  • Selection criteria: Use lingbot-map for general tasks; use lingbot-map-long for videos exceeding several thousand frames or large-scale outdoor scenes

Frequently Asked Questions

What is the difference between lingbot-map and lingbot-map-long?

The lingbot-map checkpoint provides balanced performance suitable for both short and long videos, serving as the default model used in paper benchmarks. The lingbot-map-long variant is specifically optimized for extended sequences and large-scale scenes, featuring architectural adjustments that improve stability when processing 25,000+ frames.

Where are the LingBot-Map pretrained models hosted?

Both checkpoints are available from Hugging Face at huggingface.co/robbyant/lingbot-map and from ModelScope at modelscope.cn/models/Robbyant/lingbot-map. They are distributed as direct-download .pt files rather than Hugging Face Transformers-style repositories.

How do I load a custom checkpoint in the demo scripts?

Pass the absolute path to your downloaded .pt file using the --model_path argument in either demo.py or demo_render/batch_demo.py. The scripts automatically detect the file and load it through the GCTStream class defined in lingbot_map/models/gct_stream.py.

Which model should I use for processing very long video sequences?

Use lingbot-map-long.pt for sequences exceeding several thousand frames, particularly for large-scale outdoor drives or extended indoor walkthroughs. Combine this checkpoint with windowed processing mode (--mode windowed --window_size 128) to manage memory efficiently across long temporal windows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →