# Lingbot-Map-Long vs Lingbot-Map Checkpoints: Architecture and Use Case Differences

> Explore lingbot-map-long vs lingbot-map checkpoints. Understand architecture and use case differences for handling long video sequences versus achieving per-frame accuracy on shorter runs.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-22

---

**The `lingbot-map-long` checkpoint is optimized for very long video sequences exceeding 10,000 frames with enhanced drift correction and a larger KV-cache, while `lingbot-map` provides balanced performance for short-to-medium sequences with lower GPU memory requirements and superior per-frame accuracy on shorter runs.**

The Robbyant/lingbot-map repository ships three distinct model checkpoints tailored to different inference scenarios and research needs. Understanding the specific differences between **lingbot-map-long and lingbot-map checkpoints** ensures you select the optimal configuration for your sequence length and accuracy requirements.

## Checkpoint Overview and Intended Scenarios

### Lingbot-Map-Long for Extended Sequences

The `lingbot-map-long` checkpoint targets **very long video sequences** (greater than 10,000 frames) and large-scale outdoor reconstructions such as city-scale mapping. According to the repository's README, this checkpoint implements the full **Geometric Context Transformer** architecture with anchor context, pose-reference windows, and trajectory memory.

Key characteristics include:

- Optimized KV-cache configuration for long-term dependency tracking without truncation
- Advanced drift-correction mechanisms essential for extended trajectories
- Processing speed of approximately 20 FPS on standard hardware
- Recommended as the default for production deployments involving lengthy video streams

### Lingbot-Map for General-Purpose Inference

The standard `lingbot-map` checkpoint offers a **balanced architecture** designed for short to medium-length video streams ranging from a few hundred to a few thousand frames. While sharing the same core architecture as the long variant, this checkpoint uses a more modest KV-cache size.

This configuration provides:

- Reduced GPU memory footprint compared to the long variant
- Superior per-frame pose accuracy for sequences within the cache limit (~320 frames)
- Ideal performance for indoor scenes or shorter outdoor captures where extreme length handling is unnecessary

### Lingbot-Map-Stage1 for Research and Fine-Tuning

The `lingbot-map-stage1` checkpoint represents the **stage-1 training weights** after pre-training of the Vision-Guided Geometric Transformer (VGGT). This variant supports **bidirectional inference** (camera-to-world, `c2w`) and serves as a foundation for researchers continuing training or experimenting with custom loss functions. It is not intended for end-user inference without additional training steps.

## Technical Architecture Differences

The primary distinction between `lingbot-map-long` and `lingbot-map` lies in their **KV-cache management** and **trajectory memory** implementation. As implemented in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) and [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py), the long checkpoint maintains an expanded cache to preserve geometric context across thousands of frames.

The standard checkpoint employs windowed attention mechanisms optimized for approximately 320 frames, trading long-sequence robustness for computational efficiency. Both checkpoints utilize the same `GCTStream` model class, but the weight files differ in their learned attention biases for handling sequential drift.

## Loading Checkpoints in Practice

You can load any checkpoint using the [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) entry point by specifying the `--model_path` argument. The script automatically selects the appropriate model class (`GCTStream` or `GCTStreamWindow`) based on the `--mode` flag.

```bash

# Long-sequence checkpoint for city-scale reconstruction

python demo.py --model_path /path/to/lingbot-map-long.pt \
               --image_folder example/courthouse --mask_sky

# Balanced checkpoint for standard sequences

python demo.py --model_path /path/to/lingbot-map.pt \
               --image_folder example/university --mask_sky

# Stage-1 checkpoint for research/fine-tuning

python demo.py --model_path /path/to/lingbot-map-stage1.pt \
               --image_folder example/loop --mask_sky

```

All commands assume the repository is installed via `pip install -e .`.

## Key Implementation Files

- **[`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py)**: Entry point for interactive inference demonstrating checkpoint loading via `--model_path`.
- **[`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py)**: Core model implementation used by both `lingbot-map` and `lingbot-map-long` checkpoints.
- **[`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)**: Windowed inference variant specifically leveraged for very long sequences.
- **[`benchmark/configs/oxford_long.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/configs/oxford_long.yaml)**: Configuration example for evaluating the long checkpoint on large-scale outdoor datasets.

## Summary

- **`lingbot-map-long`** is optimized for sequences exceeding 10,000 frames with enhanced drift correction and full trajectory memory, processing at ~20 FPS.
- **`lingbot-map`** provides balanced performance for sequences up to ~320 frames with lower GPU memory usage and better per-frame accuracy on shorter inputs.
- **`lingbot-map-stage1`** contains stage-1 pre-training weights for research use and bidirectional pose encoding, requiring additional training for deployment.
- All checkpoints share the same codebase in `lingbot_map/models/` and are loaded via [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) using the `--model_path` argument.

## Frequently Asked Questions

### Can I use the lingbot-map-long checkpoint for short video sequences?

Yes, but it is not optimal. While `lingbot-map-long` will process short sequences correctly, its larger KV-cache and drift-correction mechanisms introduce unnecessary computational overhead. For sequences under 320 frames, `lingbot-map` delivers better per-frame accuracy with lower GPU memory consumption.

### What is the frame limit for the standard lingbot-map checkpoint?

The standard `lingbot-map` checkpoint operates effectively within a **KV-cache limit of approximately 320 frames**. Beyond this threshold, the model may discard earlier geometric context, potentially leading to accumulated drift in longer trajectories that the `lingbot-map-long` checkpoint handles more robustly.

### When should I use the stage1 checkpoint instead of the full models?

Use `lingbot-map-stage1` when you need to **continue training** or extract intermediate bidirectional camera-to-world (`c2w`) pose encodings from the Vision-Guided Geometric Transformer. This checkpoint represents the weights after the first training phase and is intended for research experimentation rather than end-user inference pipelines.

### Do all checkpoints support the same inference script?

Yes. All three checkpoints are compatible with [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) and the underlying `GCTStream` architecture. The selection is determined solely by the `--model_path` argument, allowing seamless switching between checkpoints without modifying the inference code in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py).