How the Geometric Context Transformer Works in LingBot-Map: Architecture Deep Dive

The Geometric Context Transformer (GCT) is a Vision-Transformer-style architecture that processes RGB video streams to estimate camera poses, depth maps, and 3-D point clouds using DINOv2 patch embeddings, multi-scale transformers with optional 3-D RoPE, and task-specific prediction heads.

LingBot-Map implements the Geometric Context Transformer as its core 3-D scene understanding backbone. This architecture combines a DINOv2-derived patch embedder with streaming transformer blocks and dense prediction transformers to deliver real-time geometric context for robotics and AR/VR pipelines.

Core Architecture Components

The GCT architecture consists of three primary stages orchestrated by the abstract GCTBase class. Each stage handles a specific aspect of converting raw video frames into structured 3-D geometry.

Patch Embedding and Tokenization

The first stage converts input images into token sequences suitable for transformer processing. In lingbot_map/layers/patch_embed.py, the system uses a DINOv2-derived patch embedder (patch_embed='dinov2_vitl14_reg' by default) to split each RGB frame into non-overlapping patches.

The default configuration uses an img_size of 518 pixels, a patch_size of 14 pixels, and produces an embed_dim of 1024. This yields a sequence of visual tokens that preserves spatial relationships while reducing the computational footprint compared to raw pixel processing.

Transformer Aggregator with KV-Cache

The middle stage, implemented in lingbot_map/aggregator/stream.py, processes token sequences through multi-scale transformer blocks. This AggregatorStream module supports two critical features for temporal consistency:

  • 3-D Rotary Positional Encoding (RoPE): When enable_3d_rope=True, the system injects temporal sinusoids into attention scores to maintain geometric consistency across frames.
  • KV-Cache: The kv_cache_sliding_window parameter enables causal attention during streaming inference, allowing each new frame to attend only to cached tokens from previous frames rather than recomputing the full sequence.

Prediction Heads

The final stage decodes aggregated tokens into geometric outputs through specialized heads in lingbot_map/heads/:

  • Depth Head (dpt_head.py): A Dense Prediction Transformer (DPT) that outputs 2-channel maps containing depth values and confidence scores.
  • Point Head (dpt_head.py): Another DPT head with 4 output channels (XYZ coordinates plus confidence) for world-space point clouds.
  • Camera Head (camera_head.py): A causal transformer (CameraCausalHead) that refines 9-dimensional pose encodings (camera center plus quaternion) per frame.

The GCTBase Abstract Class

All GCT variants inherit from GCTBase, defined in lingbot_map/models/gct_base.py. This base class serves as the architectural scaffold:

class GCTBase(nn.Module, PyTorchModelHubMixin, ABC):
    def __init__(self,
                 img_size: int = 518,
                 patch_size: int = 14,
                 embed_dim: int = 1024,
                 patch_embed: str = 'dinov2_vitl14_reg',
                 enable_camera: bool = True,
                 enable_point: bool = True,
                 enable_depth: bool = True,
                 enable_3d_rope: bool = False,
                 use_gradient_checkpoint: bool = True):
        self.aggregator = self._build_aggregator()
        self.camera_head = self._build_camera_head() if enable_camera else None
        self.point_head = self._build_point_head() if enable_point else None
        self.depth_head = self._build_depth_head() if enable_depth else None

The class stores hyper-parameters and builds the aggregator and heads via abstract methods (_build_aggregator, _build_camera_head, etc.). During the forward pass (lines 87-120 in gct_base.py), it calls _aggregate_features followed by _predict_* helpers to generate the final dictionary containing pose_enc, depth, and world_points.

Streaming Inference Variants

LingBot-Map provides two concrete implementations for different inference scenarios: online streaming and windowed batch processing.

GCTStream for Online Processing

GCTStream (in lingbot_map/models/gct_stream.py) specializes the base class for real-time applications. It constructs an AggregatorStream with KV-cache support and optional flash attention (use_flashinfer):

class GCTStream(GCTBase):
    def _build_aggregator(self) -> nn.Module:
        return AggregatorStream(
            img_size=self.img_size,
            patch_size=self.patch_size,
            embed_dim=self.embed_dim,
            use_flashinfer=not self.use_sdpa,
            kv_cache_sliding_window=self.kv_cache_sliding_window,
            ...
        )

The inference_streaming method (lines 459-511) handles the causal processing loop: it normalizes inputs, initializes the KV-cache, processes initial scale frames with bidirectional attention, then streams subsequent frames while optionally skipping KV-cache writes for non-keyframes based on the keyframe_interval parameter.

GCTStreamWindow for Long Videos

For arbitrarily long sequences, GCTStreamWindow (in lingbot_map/models/gct_stream_window.py) extends the streaming core with windowed processing. It splits videos into overlapping windows, processes each with a fresh KV-cache, then aligns results via _pairwise_alignment and stitches them using _stitch_windows (lines 1120-1190).

This approach preserves a global coordinate frame while maintaining the memory efficiency of the streaming architecture.

End-to-End Processing Pipeline

The complete data flow through the Geometric Context Transformer follows these steps:

  1. Tokenization: The patch embedder in patch_embed.py converts input frames (B, S, 3, H, W) into token sequences.
  2. Aggregation: The AggregatorStream processes tokens through multi-scale blocks, applying 3-D RoPE for temporal consistency and maintaining the KV-cache for causal attention.
  3. Decoding: Task-specific heads convert aggregated features into:
    • pose_enc: 9-dimensional camera pose encodings
    • depth: Metric depth maps with confidence
    • world_points: 3-D point clouds in world coordinates
  4. Post-processing: Utilities in lingbot_map/utils/pose_enc.py convert encodings to extrinsic matrices, while lingbot_map/utils/geometry.py handles SE(3) transformations.

Implementation Examples

Instantiating a streaming model with temporal consistency:

from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    patch_embed='dinov2_vitl14_reg',
    enable_camera=True,
    enable_depth=True,
    enable_3d_rope=True,          # Enable 3-D RoPE

    sliding_window_size=64,       # KV-cache window

    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
).cuda().eval()

Running streaming inference on video tensors:

import torch

# frames shape: (B, S, 3, H, W), values in [0, 1]

frames = torch.randn(1, 120, 3, 518, 518).cuda()

preds = model.inference_streaming(
    images=frames,
    num_scale_frames=1,
    keyframe_interval=4,
    output_device=torch.device('cpu')
)

pose_enc = preds['pose_enc']         # (B, S, 9)

depth = preds['depth']               # (B, S, H, W, 1)

world_points = preds['world_points'] # (B, S, H, W, 3)

Processing long videos with windowed inference:

preds = model.inference_windowed(
    images=frames,
    window_size=16,
    overlap_size=8,
    keyframe_interval=2,
    flow_threshold=1.0,
    output_device=torch.device('cpu')
)

Summary

  • The Geometric Context Transformer combines DINOv2 patch embeddings with streaming transformers to estimate camera poses, depth, and point clouds from video.
  • GCTBase provides the shared scaffold for all variants, handling hyper-parameters and prediction head construction in lingbot_map/models/gct_base.py.
  • KV-cache and 3-D RoPE enable efficient causal inference with temporal consistency, implemented in the AggregatorStream class.
  • GCTStream supports real-time online processing, while GCTStreamWindow handles arbitrarily long videos via alignment and stitching routines.
  • DPT heads decode tokens into dense geometric predictions, with separate heads for depth, points, and camera poses.

Frequently Asked Questions

What is the difference between GCTStream and GCTStreamWindow?

GCTStream processes video sequentially using a KV-cache for constant memory usage during online inference, making it suitable for real-time applications. GCTStreamWindow extends this by splitting long videos into overlapping segments, processing each independently, then aligning and stitching the results to maintain global geometric consistency across the entire sequence.

How does the KV-cache improve inference performance?

The KV-cache stores key and value tensors from previous frames, enabling causal attention where each new frame only attends to cached history rather than the full sequence. This reduces GPU memory consumption from quadratic to linear with sequence length, as implemented in lingbot_map/aggregator/stream.py via the kv_cache_sliding_window parameter.

What is 3-D RoPE and why is it used in the Geometric Context Transformer?

3-D Rotary Positional Encoding (RoPE) injects temporal position information into the transformer's attention mechanism using sinusoidal embeddings in 3-D space. When enable_3d_rope=True, the system encodes temporal continuity directly into the attention scores, improving geometric consistency across video frames without requiring explicit temporal convolution layers.

Which geometric outputs can the GCT architecture produce?

The architecture supports three primary outputs controlled by boolean flags in the constructor: camera pose (9-dimensional encoding via CameraCausalHead), metric depth (2-channel depth + confidence via DPT head), and world point clouds (4-channel XYZ + confidence via DPT head). Each head operates on the shared token stream produced by the aggregator, allowing flexible configuration based on application requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →