# LingBot-Map Checkpoints Explained: lingbot-map vs lingbot-map-long vs lingbot-map-stage1

> Understand lingbot map checkpoints. Explore lingbot-map for short videos, lingbot-map-long for city-scale sequences, and lingbot-map-stage1 for pre-training from Robbyant/lingbot-map.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-26

---

**The `lingbot-map-long` checkpoint handles city-scale sequences over 10,000 frames with a large KV-cache and drift correction, `lingbot-map` optimizes for short-to-medium videos with balanced memory usage (~320 frames), and `lingbot-map-stage1` provides stage-one pre-training weights for research and fine-tuning.**

The Robbyant/lingbot-map repository provides three specialized model weights designed for distinct mapping workloads. Understanding the differences between these **lingbot-map checkpoints** is essential for matching your video length, GPU constraints, and inference requirements to the correct trained weights.

## Checkpoint Overview

The three variants share the same codebase but differ in their trained weights, KV-cache configurations, and training stages. As documented in the repository's [`README.md`](https://github.com/Robbyant/lingbot-map/blob/main/README.md) Model Download section, each targets a specific deployment scenario, from real-time short videos to research experimentation.

### lingbot-map-long: City-Scale Long Sequences

The `lingbot-map-long` checkpoint is optimized for **very long video sequences** exceeding 10,000 frames and large-scale outdoor scenes such as city-scale reconstructions. It implements the full **Geometric Context Transformer** with anchor context, pose-reference windows, and trajectory memory without sacrificing speed. This variant uses a significantly larger KV-cache and stronger drift-correction mechanisms to maintain accuracy over extended trajectories. Despite its focus on long sequences, it is the **recommended default** for most users, balancing approximately **20 FPS** inference speed with the highest reconstruction quality on long runs.

### lingbot-map: Balanced General-Purpose Inference

The standard `lingbot-map` checkpoint serves as the **general-purpose** option for short to medium-length video streams ranging from a few hundred to a few thousand frames. It shares the same architecture as the long variant but employs a more modest KV-cache size—limited to approximately **320 frames**—making it lighter on GPU memory. This checkpoint trades a small amount of long-sequence robustness for slightly better per-frame pose accuracy on shorter clips, ideal when your target scene fits comfortably within the cache limits.

### lingbot-map-stage1: Research and Pre-training

The `lingbot-map-stage1` checkpoint represents the **stage-one training weights** after the initial pre-training of the Vision-Guided Geometric Transformer (VGGT). Intended strictly for research and development, it enables **bidirectional inference** (camera-to-world, `c2w`) and serves as a foundation for fine-tuning or experimenting with alternative loss functions. This checkpoint is **not meant for end-user inference**; it requires additional training steps to reach the final performance levels of the production checkpoints.

## Loading Checkpoints in demo.py

The [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) entry point accepts any checkpoint via the `--model_path` argument. The underlying model class—either `GCTStream` or `GCTStreamWindow`—is selected automatically based on the `--mode` flag, but you must explicitly specify which checkpoint weights to load.

```bash

# lingbot-map-long: Use for city-scale or >10,000 frame sequences

python demo.py --model_path /path/to/lingbot-map-long.pt \
               --image_folder example/courthouse --mask_sky

```

```bash

# lingbot-map: Use for short-to-medium sequences (~320 frame cache)

python demo.py --model_path /path/to/lingbot-map.pt \
               --image_folder example/university --mask_sky

```

```bash

# lingbot-map-stage1: Use for research and bidirectional pose extraction

python demo.py --model_path /path/to/lingbot-map-stage1.pt \
               --image_folder example/loop --mask_sky

```

All commands assume the repository is installed in editable mode (`pip install -e .`).

## Core Implementation Files

The checkpoint behavior is determined by the underlying model implementations in the repository:

- **[`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py)**: Contains the `GCTStream` class used by both `lingbot-map` and `lingbot-map-long` for streaming inference.
- **[`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)**: Implements windowed inference mechanisms required by `lingbot-map-long` to manage extended sequences without memory overflow.
- **[`benchmark/configs/oxford_long.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/configs/oxford_long.yaml)**: Example configuration file demonstrating how to evaluate the long checkpoint on large-scale outdoor benchmarks.

## Summary

- **`lingbot-map-long`**: Deploy for sequences exceeding 10,000 frames; features full Geometric Context Transformer, large KV-cache, and drift correction at ~20 FPS.
- **`lingbot-map`**: Select for short-to-medium sequences (up to ~320 frames) requiring optimal per-frame accuracy with moderate GPU memory.
- **`lingbot-map-stage1`**: Use exclusively for research, fine-tuning, or bidirectional pose encoding; requires additional training before deployment.

## Frequently Asked Questions

### What is the KV-cache limit for the standard lingbot-map checkpoint?

The `lingbot-map` checkpoint operates with a KV-cache limit of approximately **320 frames**, making it suitable for short-to-medium video sequences. For footage exceeding this length, use `lingbot-map-long`, which implements a larger KV-cache and windowed attention mechanisms via `GCTStreamWindow` to handle 10,000+ frames without drift.

### Can I use lingbot-map-stage1 for end-user inference?

No. The `lingbot-map-stage1` checkpoint contains weights after the first training phase of the Vision-Guided Geometric Transformer (VGGT) and lacks the full refinement of the complete model. It is intended for researchers who need to continue training, experiment with loss functions, or extract intermediate bidirectional pose encodings.

### Which checkpoint offers the best speed and quality trade-off?

According to the repository documentation, `lingbot-map-long` is the recommended default for most users because it balances speed—processing at approximately **20 FPS**—with the highest reconstruction quality on long runs. While optimized for extended sequences, its architecture maintains efficiency across various lengths.

### Where are the checkpoint weights stored and loaded?

All weights are downloaded from HuggingFace or ModelScope and loaded into the model classes defined in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py). The [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) script handles loading via the `--model_path` parameter, automatically instantiating the correct architecture based on your configuration flags.