Lingbot-Map-Long vs Lingbot-Map Checkpoints: Architecture and Use Case Differences
The lingbot-map-long checkpoint is optimized for very long video sequences exceeding 10,000 frames with enhanced drift correction and a larger KV-cache, while lingbot-map provides balanced performance for short-to-medium sequences with lower GPU memory requirements and superior per-frame accuracy on shorter runs.
The Robbyant/lingbot-map repository ships three distinct model checkpoints tailored to different inference scenarios and research needs. Understanding the specific differences between lingbot-map-long and lingbot-map checkpoints ensures you select the optimal configuration for your sequence length and accuracy requirements.
Checkpoint Overview and Intended Scenarios
Lingbot-Map-Long for Extended Sequences
The lingbot-map-long checkpoint targets very long video sequences (greater than 10,000 frames) and large-scale outdoor reconstructions such as city-scale mapping. According to the repository's README, this checkpoint implements the full Geometric Context Transformer architecture with anchor context, pose-reference windows, and trajectory memory.
Key characteristics include:
- Optimized KV-cache configuration for long-term dependency tracking without truncation
- Advanced drift-correction mechanisms essential for extended trajectories
- Processing speed of approximately 20 FPS on standard hardware
- Recommended as the default for production deployments involving lengthy video streams
Lingbot-Map for General-Purpose Inference
The standard lingbot-map checkpoint offers a balanced architecture designed for short to medium-length video streams ranging from a few hundred to a few thousand frames. While sharing the same core architecture as the long variant, this checkpoint uses a more modest KV-cache size.
This configuration provides:
- Reduced GPU memory footprint compared to the long variant
- Superior per-frame pose accuracy for sequences within the cache limit (~320 frames)
- Ideal performance for indoor scenes or shorter outdoor captures where extreme length handling is unnecessary
Lingbot-Map-Stage1 for Research and Fine-Tuning
The lingbot-map-stage1 checkpoint represents the stage-1 training weights after pre-training of the Vision-Guided Geometric Transformer (VGGT). This variant supports bidirectional inference (camera-to-world, c2w) and serves as a foundation for researchers continuing training or experimenting with custom loss functions. It is not intended for end-user inference without additional training steps.
Technical Architecture Differences
The primary distinction between lingbot-map-long and lingbot-map lies in their KV-cache management and trajectory memory implementation. As implemented in lingbot_map/models/gct_stream.py and lingbot_map/models/gct_stream_window.py, the long checkpoint maintains an expanded cache to preserve geometric context across thousands of frames.
The standard checkpoint employs windowed attention mechanisms optimized for approximately 320 frames, trading long-sequence robustness for computational efficiency. Both checkpoints utilize the same GCTStream model class, but the weight files differ in their learned attention biases for handling sequential drift.
Loading Checkpoints in Practice
You can load any checkpoint using the demo.py entry point by specifying the --model_path argument. The script automatically selects the appropriate model class (GCTStream or GCTStreamWindow) based on the --mode flag.
# Long-sequence checkpoint for city-scale reconstruction
python demo.py --model_path /path/to/lingbot-map-long.pt \
--image_folder example/courthouse --mask_sky
# Balanced checkpoint for standard sequences
python demo.py --model_path /path/to/lingbot-map.pt \
--image_folder example/university --mask_sky
# Stage-1 checkpoint for research/fine-tuning
python demo.py --model_path /path/to/lingbot-map-stage1.pt \
--image_folder example/loop --mask_sky
All commands assume the repository is installed via pip install -e ..
Key Implementation Files
demo.py: Entry point for interactive inference demonstrating checkpoint loading via--model_path.lingbot_map/models/gct_stream.py: Core model implementation used by bothlingbot-mapandlingbot-map-longcheckpoints.lingbot_map/models/gct_stream_window.py: Windowed inference variant specifically leveraged for very long sequences.benchmark/configs/oxford_long.yaml: Configuration example for evaluating the long checkpoint on large-scale outdoor datasets.
Summary
lingbot-map-longis optimized for sequences exceeding 10,000 frames with enhanced drift correction and full trajectory memory, processing at ~20 FPS.lingbot-mapprovides balanced performance for sequences up to ~320 frames with lower GPU memory usage and better per-frame accuracy on shorter inputs.lingbot-map-stage1contains stage-1 pre-training weights for research use and bidirectional pose encoding, requiring additional training for deployment.- All checkpoints share the same codebase in
lingbot_map/models/and are loaded viademo.pyusing the--model_pathargument.
Frequently Asked Questions
Can I use the lingbot-map-long checkpoint for short video sequences?
Yes, but it is not optimal. While lingbot-map-long will process short sequences correctly, its larger KV-cache and drift-correction mechanisms introduce unnecessary computational overhead. For sequences under 320 frames, lingbot-map delivers better per-frame accuracy with lower GPU memory consumption.
What is the frame limit for the standard lingbot-map checkpoint?
The standard lingbot-map checkpoint operates effectively within a KV-cache limit of approximately 320 frames. Beyond this threshold, the model may discard earlier geometric context, potentially leading to accumulated drift in longer trajectories that the lingbot-map-long checkpoint handles more robustly.
When should I use the stage1 checkpoint instead of the full models?
Use lingbot-map-stage1 when you need to continue training or extract intermediate bidirectional camera-to-world (c2w) pose encodings from the Vision-Guided Geometric Transformer. This checkpoint represents the weights after the first training phase and is intended for research experimentation rather than end-user inference pipelines.
Do all checkpoints support the same inference script?
Yes. All three checkpoints are compatible with demo.py and the underlying GCTStream architecture. The selection is determined solely by the --model_path argument, allowing seamless switching between checkpoints without modifying the inference code in lingbot_map/models/gct_stream.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →