# Lingbot-Map vs Lingbot-Map-Long vs Lingbot-Map-Stage1: Key Architecture Differences

> Understand key architecture differences between lingbot-map, lingbot-map-long, and lingbot-map-stage1 models. Explore streaming KV cache, windowed inference, and lightweight stage-I training.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-27

---

**The `lingbot-map` repository provides three distinct inference modes: the base `lingbot-map` for streaming frame-by-frame processing up to 200 frames using a persistent KV cache, `lingbot-map-long` for windowed inference on extended sequences via cache trimming in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py), and `lingbot-map-stage1` for lightweight stage-I training with reduced transformer blocks.**

The Robbyant/lingbot-map repository implements a General-Purpose Convolutional Transformer (GCT) architecture designed for video mapping and streaming inference. Each variant—**lingbot-map**, **lingbot-map-long**, and **lingbot-map-stage1**—optimizes the same underlying backbone for different sequence lengths and computational constraints. Understanding these distinctions ensures you select the appropriate model configuration for your specific video processing requirements.

## Base Model: Lingbot-Map for Standard Streaming

### Frame-by-Frame Processing Architecture

The base **lingbot-map** model processes video streams incrementally rather than requiring complete sequences upfront. Implemented in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py), this variant utilizes a **key-value (KV) cache** across transformer layers to store and reuse past attention computations. By maintaining temporal state between frames, the model achieves efficient streaming inference for sequences of approximately 200 frames or fewer.

### Full KV Cache Persistence

During inference, the base model retains the complete KV cache for the entire sequence duration. This approach eliminates redundant attention calculations but requires memory proportional to sequence length. For standard video segments under 200 frames, this trade-off delivers optimal throughput without memory pressure.

## Extended Sequences: Lingbot-Map-Long Windowed Inference

### Overlapping Window Strategy

**Lingbot-map-long** addresses the memory limitations of the base model for sequences significantly exceeding 200 frames. The implementation in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py) (and the updated [`gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window_v2.py)) divides input video into overlapping windows—typically 500 frames each with configurable overlap parameters. Each window processes independently using the identical GCT backbone as the base model.

### Cache Trimming and Memory Management

To prevent unbounded memory growth during long-sequence processing, the windowed implementation applies aggressive **KV cache trimming** after processing each window. The source code specifically implements an "SDPA aggregator cache — trim last frame" mechanism that discards stale cache entries while preserving necessary temporal context. Predictions from successive windows are concatenated to produce seamless outputs, enabling processing of videos thousands of frames long without out-of-memory errors.

## Lightweight Training: Lingbot-Map-Stage1

### Reduced Transformer Depth

**Lingbot-map-stage1** provides a lightweight variant for rapid prototyping and early-stage training workflows. Defined within the same [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py) file as the base model, this variant instantiates the `GCTStream` class with restricted depth. By setting the `max_stage=1` parameter, the model executes only the first transformer stage, significantly reducing parameter count and computational overhead.

### Resource Optimization

This configuration is ideal for debugging pipeline integrity, validating data loaders, or performing initial training iterations where full model capacity is unnecessary. The reduced depth maintains the identical input-output interface as the full model, ensuring seamless integration with existing training scripts while accelerating iteration cycles.

## Code Implementation Examples

Initialize the base streaming model for standard sequences:

```python
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(num_random_frames=0)
outputs = model(video_frames)  # video_frames shape: (T, H, W, 3)

```

Configure the windowed long-sequence model:

```python
from lingbot_map.models.gct_stream_window import GCTStreamWindow

model_long = GCTStreamWindow(
    window_size=500,
    overlap=50,
    num_random_frames=0
)
outputs_long = model_long(video_frames)

```

Instantiate the lightweight stage-I variant:

```python
from lingbot_map.models.gct_stream import GCTStream

model_stage1 = GCTStream(num_random_frames=0, max_stage=1)
outputs_stage1 = model_stage1(video_frames)

```

## Summary

- **Lingbot-map** (base): Processes sequences up to 200 frames with full KV cache persistence in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py), ideal for standard streaming applications.
- **Lingbot-map-long** (windowed): Handles extended sequences via overlapping windows and cache trimming in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py), essential for videos exceeding 500 frames.
- **Lingbot-map-stage1** (lightweight): Limits execution to the first transformer stage using `max_stage=1` in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py), optimized for rapid prototyping and resource-constrained environments.

## Frequently Asked Questions

### What is the maximum sequence length for the base lingbot-map model?

The base `lingbot-map` model efficiently processes sequences of approximately 200 frames or fewer. Beyond this threshold, the persistent KV cache in [`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py) consumes excessive GPU memory, necessitating the windowed approach of `lingbot-map-long`.

### When should I use lingbot-map-long instead of the base model?

Switch to `lingbot-map-long` when processing videos longer than 200-500 frames. The windowed implementation in [`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py) prevents out-of-memory errors by trimming the KV cache after each window while maintaining temporal coherence through overlapping regions.

### How does lingbot-map-stage1 differ from the full model?

`Lingbot-map-stage1` executes only the first transformer stage by setting `max_stage=1` in the `GCTStream` constructor, whereas the full model processes all stages. This reduces parameter count and accelerates inference, making it suitable for debugging and early training phases.

### Can I switch between models without modifying my data pipeline?

Yes, all three variants maintain identical input-output interfaces expecting tensors of shape `(T, H, W, 3)`. You can substitute `GCTStream` with `GCTStreamWindow` or toggle the `max_stage` parameter without changing preprocessing or postprocessing logic.