# LingBot-Map Training Checkpoints: Comparing lingbot-map, lingbot-map-long, and lingbot-map-stage1

> Explore lingbot-map training checkpoints: lingbot-map, lingbot-map-long, and lingbot-map-stage1. Understand their unique strengths for pose estimation, long sequences, and VGGT backbone applications.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-30

---

**The Robbyant/lingbot-map repository provides three distinct training checkpoints—`lingbot-map` (balanced default), `lingbot-map-long` (long-sequence optimized), and `lingbot-map-stage1` (VGGT backbone stage-1)—each engineered for specific inference scenarios ranging from general pose estimation to extended video streams and bidirectional camera-to-world regression.**

Selecting the correct **LingBot-Map training checkpoint** is critical for optimal performance in your specific deployment environment. Each variant targets different sequence lengths, memory constraints, and architectural backends as defined in the repository's model download specifications.

## Overview of the Three LingBot-Map Checkpoints

According to the **Model Download** table in [`README.md`](https://github.com/Robbyant/lingbot-map/blob/main/README.md) (lines 137-140), the repository maintains three distinct weight files optimized for different computational requirements.

### lingbot-map (Balanced Default)

The **`lingbot-map`** checkpoint serves as the general-purpose, balanced option suitable for most research and production use cases. It delivers optimal performance across both short and long video sequences and represents the exact checkpoint utilized in the paper, benchmark suite, and offline demo. As configured in [`benchmark/configs/methods/lingbot_map.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/configs/methods/lingbot_map.yaml), this is the default weight file expected by evaluation scripts.

### lingbot-map-long (Long-Sequence Optimized)

The **`lingbot-map-long`** checkpoint specifically targets **very long sequences** and large-scale scenes. It incorporates expanded capacity for the **KV-cache** and employs a longer training window, providing greater stability when streaming thousands of consecutive frames. This variant is essential for extended video processing where standard checkpoints might encounter memory pressure or temporal drift.

### lingbot-map-stage1 (Stage-1 VGGT Backbone)

The **`lingbot-map-stage1`** checkpoint captures the first stage of the two-stage training pipeline. Unlike the final checkpoints, this weight is specifically intended for **bidirectional inference** (camera-to-world or `c2w`) when loaded into the **VGGT backbone**. It enables coarse-to-fine pose regression within the VGGT architecture rather than the standard LingBot-Map inference path.

## Technical Differences and Architecture

All three checkpoints instantiate the same core architecture defined in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py), ensuring weight compatibility across variants. However, their training regimes and intended deployment contexts differ significantly:

- **`lingbot-map`** offers the best overall performance for mixed-length sequences, balancing accuracy and computational efficiency.
- **`lingbot-map-long`** sacrifices some generalization for extended context windows, making it robust for continuous streaming applications.
- **`lingbot-map-stage1`** represents an intermediate training state compatible with the VGGT model, enabling specialized `c2w` bidirectional pose regression that the standard checkpoints do not support.

## How to Load and Use Each Checkpoint

The CLI interface remains consistent across all variants. You select the desired behavior by specifying the path to the appropriate `.pt` file via the `--model_path` argument.

To use the balanced default checkpoint for standard inference:

```bash
python demo.py \
    --model_path /path/to/lingbot-map.pt \
    --image_folder example/courthouse --mask_sky

```

For extended video sequences requiring additional KV-cache capacity:

```bash
python demo.py \
    --model_path /path/to/lingbot-map-long.pt \
    --image_folder example/long_video \
    --keyframe_interval 2

```

To load the stage-1 checkpoint into the VGGT model for bidirectional `c2w` inference:

```bash
python demo.py \
    --model_path /path/to/lingbot-map-stage1.pt \
    --mode c2w --image_folder example/scene

```

## Key Source Files and Configuration

Several critical source files govern how these checkpoints are referenced and loaded:

- **[`README.md`](https://github.com/Robbyant/lingbot-map/blob/main/README.md)** (lines 137-140): Contains the authoritative description of each checkpoint's intended use case and training characteristics.
- **[`benchmark/configs/methods/lingbot_map.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/configs/methods/lingbot_map.yaml)**: Defines the `_checkpoint` field that specifies which weight file the benchmark suite loads during evaluation.
- **[`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py)**: Implements the base model architecture that all three checkpoints instantiate.
- **[`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py)**: Wraps the upstream package and consumes the checkpoint path defined in the YAML configuration, handling the actual model initialization for benchmark runs.

## Summary

- **`lingbot-map`**: The balanced, default checkpoint optimized for general-purpose use across short and long sequences; used in the paper and standard benchmarks.
- **`lingbot-map-long`**: A high-capacity variant with expanded KV-cache and longer training windows for stable inference on very long video streams and large scenes.
- **`lingbot-map-stage1`**: An early-training checkpoint designed specifically for bidirectional `c2w` inference within the VGGT backbone architecture.
- All checkpoints utilize the same underlying architecture from [`gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_base.py) and are interchangeable via the `--model_path` argument in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py).

## Frequently Asked Questions

### What is the default checkpoint for most LingBot-Map experiments?

The **`lingbot-map`** checkpoint serves as the default for most experiments. It provides balanced performance across sequence lengths and is the specific weight used in the paper, benchmark suite, and offline demo applications according to the repository documentation.

### When should I use lingbot-map-long instead of the standard checkpoint?

Use **`lingbot-map-long`** when processing very long video sequences or large-scale scenes that span thousands of frames. Its expanded KV-cache capacity and optimized training window prevent instability and memory issues during extended streaming inference.

### Can I use lingbot-map-stage1 with the standard inference pipeline?

No, **`lingbot-map-stage1`** is specifically intended for the VGGT backbone and bidirectional camera-to-world (`c2w`) inference. It represents an intermediate training stage and requires the VGGT architecture to function correctly, rather than the standard LingBot-Map inference path defined in [`gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_base.py).

### Where are checkpoint paths configured in the benchmark suite?

Checkpoint paths are configured in **[`benchmark/configs/methods/lingbot_map.yaml`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/configs/methods/lingbot_map.yaml)** via the `_checkpoint` field. The **[`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py)** wrapper reads this configuration to load the specified weights during evaluation runs.