# How the Geometric Context Transformer Works in LingBot-Map: Architecture Deep Dive

> Explore the Geometric Context Transformer in LingBot-Map. Understand its architecture for estimating camera poses, depth maps, and 3D point clouds using DINOv2 and multi-scale transformers.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: deep-dive
- Published: 2026-07-27

---

**The Geometric Context Transformer (GCT) is a Vision-Transformer-style architecture that processes RGB video streams to estimate camera poses, depth maps, and 3-D point clouds using DINOv2 patch embeddings, multi-scale transformers with optional 3-D RoPE, and task-specific prediction heads.**

LingBot-Map implements the Geometric Context Transformer as its core 3-D scene understanding backbone. This architecture combines a DINOv2-derived patch embedder with streaming transformer blocks and dense prediction transformers to deliver real-time geometric context for robotics and AR/VR pipelines.

## Core Architecture Components

The GCT architecture consists of three primary stages orchestrated by the abstract `GCTBase` class. Each stage handles a specific aspect of converting raw video frames into structured 3-D geometry.

### Patch Embedding and Tokenization

The first stage converts input images into token sequences suitable for transformer processing. In [`lingbot_map/layers/patch_embed.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/patch_embed.py), the system uses a DINOv2-derived patch embedder (`patch_embed='dinov2_vitl14_reg'` by default) to split each RGB frame into non-overlapping patches.

The default configuration uses an `img_size` of 518 pixels, a `patch_size` of 14 pixels, and produces an `embed_dim` of 1024. This yields a sequence of visual tokens that preserves spatial relationships while reducing the computational footprint compared to raw pixel processing.

### Transformer Aggregator with KV-Cache

The middle stage, implemented in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py), processes token sequences through multi-scale transformer blocks. This `AggregatorStream` module supports two critical features for temporal consistency:

- **3-D Rotary Positional Encoding (RoPE)**: When `enable_3d_rope=True`, the system injects temporal sinusoids into attention scores to maintain geometric consistency across frames.
- **KV-Cache**: The `kv_cache_sliding_window` parameter enables causal attention during streaming inference, allowing each new frame to attend only to cached tokens from previous frames rather than recomputing the full sequence.

### Prediction Heads

The final stage decodes aggregated tokens into geometric outputs through specialized heads in `lingbot_map/heads/`:

- **Depth Head** ([`dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/dpt_head.py)): A Dense Prediction Transformer (DPT) that outputs 2-channel maps containing depth values and confidence scores.
- **Point Head** ([`dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/dpt_head.py)): Another DPT head with 4 output channels (XYZ coordinates plus confidence) for world-space point clouds.
- **Camera Head** ([`camera_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/camera_head.py)): A causal transformer (`CameraCausalHead`) that refines 9-dimensional pose encodings (camera center plus quaternion) per frame.

## The GCTBase Abstract Class

All GCT variants inherit from `GCTBase`, defined in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py). This base class serves as the architectural scaffold:

```python
class GCTBase(nn.Module, PyTorchModelHubMixin, ABC):
    def __init__(self,
                 img_size: int = 518,
                 patch_size: int = 14,
                 embed_dim: int = 1024,
                 patch_embed: str = 'dinov2_vitl14_reg',
                 enable_camera: bool = True,
                 enable_point: bool = True,
                 enable_depth: bool = True,
                 enable_3d_rope: bool = False,
                 use_gradient_checkpoint: bool = True):
        self.aggregator = self._build_aggregator()
        self.camera_head = self._build_camera_head() if enable_camera else None
        self.point_head = self._build_point_head() if enable_point else None
        self.depth_head = self._build_depth_head() if enable_depth else None

```

The class stores hyper-parameters and builds the aggregator and heads via abstract methods (`_build_aggregator`, `_build_camera_head`, etc.). During the forward pass (lines 87-120 in [`gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_base.py)), it calls `_aggregate_features` followed by `_predict_*` helpers to generate the final dictionary containing `pose_enc`, `depth`, and `world_points`.

## Streaming Inference Variants

LingBot-Map provides two concrete implementations for different inference scenarios: online streaming and windowed batch processing.

### GCTStream for Online Processing

`GCTStream` (in [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py)) specializes the base class for real-time applications. It constructs an `AggregatorStream` with KV-cache support and optional flash attention (`use_flashinfer`):

```python
class GCTStream(GCTBase):
    def _build_aggregator(self) -> nn.Module:
        return AggregatorStream(
            img_size=self.img_size,
            patch_size=self.patch_size,
            embed_dim=self.embed_dim,
            use_flashinfer=not self.use_sdpa,
            kv_cache_sliding_window=self.kv_cache_sliding_window,
            ...
        )

```

The `inference_streaming` method (lines 459-511) handles the causal processing loop: it normalizes inputs, initializes the KV-cache, processes initial scale frames with bidirectional attention, then streams subsequent frames while optionally skipping KV-cache writes for non-keyframes based on the `keyframe_interval` parameter.

### GCTStreamWindow for Long Videos

For arbitrarily long sequences, `GCTStreamWindow` (in [`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)) extends the streaming core with windowed processing. It splits videos into overlapping windows, processes each with a fresh KV-cache, then aligns results via `_pairwise_alignment` and stitches them using `_stitch_windows` (lines 1120-1190).

This approach preserves a global coordinate frame while maintaining the memory efficiency of the streaming architecture.

## End-to-End Processing Pipeline

The complete data flow through the Geometric Context Transformer follows these steps:

1. **Tokenization**: The patch embedder in [`patch_embed.py`](https://github.com/Robbyant/lingbot-map/blob/main/patch_embed.py) converts input frames `(B, S, 3, H, W)` into token sequences.
2. **Aggregation**: The `AggregatorStream` processes tokens through multi-scale blocks, applying 3-D RoPE for temporal consistency and maintaining the KV-cache for causal attention.
3. **Decoding**: Task-specific heads convert aggregated features into:
   - `pose_enc`: 9-dimensional camera pose encodings
   - `depth`: Metric depth maps with confidence
   - `world_points`: 3-D point clouds in world coordinates
4. **Post-processing**: Utilities in [`lingbot_map/utils/pose_enc.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/pose_enc.py) convert encodings to extrinsic matrices, while [`lingbot_map/utils/geometry.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/utils/geometry.py) handles SE(3) transformations.

## Implementation Examples

**Instantiating a streaming model with temporal consistency:**

```python
from lingbot_map.models.gct_stream import GCTStream

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    patch_embed='dinov2_vitl14_reg',
    enable_camera=True,
    enable_depth=True,
    enable_3d_rope=True,          # Enable 3-D RoPE

    sliding_window_size=64,       # KV-cache window

    kv_cache_sliding_window=64,
    kv_cache_scale_frames=8,
).cuda().eval()

```

**Running streaming inference on video tensors:**

```python
import torch

# frames shape: (B, S, 3, H, W), values in [0, 1]

frames = torch.randn(1, 120, 3, 518, 518).cuda()

preds = model.inference_streaming(
    images=frames,
    num_scale_frames=1,
    keyframe_interval=4,
    output_device=torch.device('cpu')
)

pose_enc = preds['pose_enc']         # (B, S, 9)

depth = preds['depth']               # (B, S, H, W, 1)

world_points = preds['world_points'] # (B, S, H, W, 3)

```

**Processing long videos with windowed inference:**

```python
preds = model.inference_windowed(
    images=frames,
    window_size=16,
    overlap_size=8,
    keyframe_interval=2,
    flow_threshold=1.0,
    output_device=torch.device('cpu')
)

```

## Summary

- The **Geometric Context Transformer** combines DINOv2 patch embeddings with streaming transformers to estimate camera poses, depth, and point clouds from video.
- **`GCTBase`** provides the shared scaffold for all variants, handling hyper-parameters and prediction head construction in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py).
- **KV-cache** and **3-D RoPE** enable efficient causal inference with temporal consistency, implemented in the `AggregatorStream` class.
- **`GCTStream`** supports real-time online processing, while **`GCTStreamWindow`** handles arbitrarily long videos via alignment and stitching routines.
- **DPT heads** decode tokens into dense geometric predictions, with separate heads for depth, points, and camera poses.

## Frequently Asked Questions

### What is the difference between GCTStream and GCTStreamWindow?

**GCTStream** processes video sequentially using a KV-cache for constant memory usage during online inference, making it suitable for real-time applications. **GCTStreamWindow** extends this by splitting long videos into overlapping segments, processing each independently, then aligning and stitching the results to maintain global geometric consistency across the entire sequence.

### How does the KV-cache improve inference performance?

The KV-cache stores key and value tensors from previous frames, enabling **causal attention** where each new frame only attends to cached history rather than the full sequence. This reduces GPU memory consumption from quadratic to linear with sequence length, as implemented in [`lingbot_map/aggregator/stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/aggregator/stream.py) via the `kv_cache_sliding_window` parameter.

### What is 3-D RoPE and why is it used in the Geometric Context Transformer?

**3-D Rotary Positional Encoding (RoPE)** injects temporal position information into the transformer's attention mechanism using sinusoidal embeddings in 3-D space. When `enable_3d_rope=True`, the system encodes temporal continuity directly into the attention scores, improving geometric consistency across video frames without requiring explicit temporal convolution layers.

### Which geometric outputs can the GCT architecture produce?

The architecture supports three primary outputs controlled by boolean flags in the constructor: **camera pose** (9-dimensional encoding via `CameraCausalHead`), **metric depth** (2-channel depth + confidence via DPT head), and **world point clouds** (4-channel XYZ + confidence via DPT head). Each head operates on the shared token stream produced by the aggregator, allowing flexible configuration based on application requirements.