Lingbot-Map vs Lingbot-Map-Long vs Lingbot-Map-Stage1: Key Architecture Differences

The lingbot-map repository provides three distinct inference modes: the base lingbot-map for streaming frame-by-frame processing up to 200 frames using a persistent KV cache, lingbot-map-long for windowed inference on extended sequences via cache trimming in gct_stream_window.py, and lingbot-map-stage1 for lightweight stage-I training with reduced transformer blocks.

The Robbyant/lingbot-map repository implements a General-Purpose Convolutional Transformer (GCT) architecture designed for video mapping and streaming inference. Each variant—lingbot-map, lingbot-map-long, and lingbot-map-stage1—optimizes the same underlying backbone for different sequence lengths and computational constraints. Understanding these distinctions ensures you select the appropriate model configuration for your specific video processing requirements.

Base Model: Lingbot-Map for Standard Streaming

Frame-by-Frame Processing Architecture

The base lingbot-map model processes video streams incrementally rather than requiring complete sequences upfront. Implemented in lingbot_map/models/gct_stream.py, this variant utilizes a key-value (KV) cache across transformer layers to store and reuse past attention computations. By maintaining temporal state between frames, the model achieves efficient streaming inference for sequences of approximately 200 frames or fewer.

Full KV Cache Persistence

During inference, the base model retains the complete KV cache for the entire sequence duration. This approach eliminates redundant attention calculations but requires memory proportional to sequence length. For standard video segments under 200 frames, this trade-off delivers optimal throughput without memory pressure.

Extended Sequences: Lingbot-Map-Long Windowed Inference

Overlapping Window Strategy

Lingbot-map-long addresses the memory limitations of the base model for sequences significantly exceeding 200 frames. The implementation in lingbot_map/models/gct_stream_window.py (and the updated gct_stream_window_v2.py) divides input video into overlapping windows—typically 500 frames each with configurable overlap parameters. Each window processes independently using the identical GCT backbone as the base model.

Cache Trimming and Memory Management

To prevent unbounded memory growth during long-sequence processing, the windowed implementation applies aggressive KV cache trimming after processing each window. The source code specifically implements an "SDPA aggregator cache — trim last frame" mechanism that discards stale cache entries while preserving necessary temporal context. Predictions from successive windows are concatenated to produce seamless outputs, enabling processing of videos thousands of frames long without out-of-memory errors.

Lightweight Training: Lingbot-Map-Stage1

Reduced Transformer Depth

Lingbot-map-stage1 provides a lightweight variant for rapid prototyping and early-stage training workflows. Defined within the same lingbot_map/models/gct_stream.py file as the base model, this variant instantiates the GCTStream class with restricted depth. By setting the max_stage=1 parameter, the model executes only the first transformer stage, significantly reducing parameter count and computational overhead.

Resource Optimization

This configuration is ideal for debugging pipeline integrity, validating data loaders, or performing initial training iterations where full model capacity is unnecessary. The reduced depth maintains the identical input-output interface as the full model, ensuring seamless integration with existing training scripts while accelerating iteration cycles.

Code Implementation Examples

Initialize the base streaming model for standard sequences:

from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(num_random_frames=0)
outputs = model(video_frames)  # video_frames shape: (T, H, W, 3)

Configure the windowed long-sequence model:

from lingbot_map.models.gct_stream_window import GCTStreamWindow

model_long = GCTStreamWindow(
    window_size=500,
    overlap=50,
    num_random_frames=0
)
outputs_long = model_long(video_frames)

Instantiate the lightweight stage-I variant:

from lingbot_map.models.gct_stream import GCTStream

model_stage1 = GCTStream(num_random_frames=0, max_stage=1)
outputs_stage1 = model_stage1(video_frames)

Summary

  • Lingbot-map (base): Processes sequences up to 200 frames with full KV cache persistence in gct_stream.py, ideal for standard streaming applications.
  • Lingbot-map-long (windowed): Handles extended sequences via overlapping windows and cache trimming in gct_stream_window.py, essential for videos exceeding 500 frames.
  • Lingbot-map-stage1 (lightweight): Limits execution to the first transformer stage using max_stage=1 in gct_stream.py, optimized for rapid prototyping and resource-constrained environments.

Frequently Asked Questions

What is the maximum sequence length for the base lingbot-map model?

The base lingbot-map model efficiently processes sequences of approximately 200 frames or fewer. Beyond this threshold, the persistent KV cache in gct_stream.py consumes excessive GPU memory, necessitating the windowed approach of lingbot-map-long.

When should I use lingbot-map-long instead of the base model?

Switch to lingbot-map-long when processing videos longer than 200-500 frames. The windowed implementation in gct_stream_window.py prevents out-of-memory errors by trimming the KV cache after each window while maintaining temporal coherence through overlapping regions.

How does lingbot-map-stage1 differ from the full model?

Lingbot-map-stage1 executes only the first transformer stage by setting max_stage=1 in the GCTStream constructor, whereas the full model processes all stages. This reduces parameter count and accelerates inference, making it suitable for debugging and early training phases.

Can I switch between models without modifying my data pipeline?

Yes, all three variants maintain identical input-output interfaces expecting tensors of shape (T, H, W, 3). You can substitute GCTStream with GCTStreamWindow or toggle the max_stage parameter without changing preprocessing or postprocessing logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →