# Where to Find the Base Model Class for LingBot-Map: Complete Implementation Guide

> Locate the GCTBase model class for LingBot-Map in lingbot_map/models/gct_base.py. This implementation guide shows you where to find the abstract foundation for all GCT models.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: implementation-guide
- Published: 2026-07-28

---

**The base model class for LingBot-Map is `GCTBase`, defined in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py), which serves as the abstract foundation for all Generalized Cross-Transformer (GCT) models in the repository.**

The LingBot-Map repository implements a family of vision transformers for robotic mapping and localization tasks. Understanding the base model architecture is essential for extending the framework or debugging prediction pipelines, as every concrete model variant inherits from this single abstract interface.

## Location of the Base Model Class in LingBot-Map

The foundational class for all GCT models resides in a single file within the models package.

### The GCTBase Abstract Class

According to the LingBot-Map source code, the base model class **`GCTBase`** is located at:

- **[`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py)**

This file defines the abstract base class that inherits from both `torch.nn.Module` and `PyTorchModelHubMixin`, enabling standard PyTorch training workflows alongside Hugging Face Hub integration. The class implements the shared architecture for LingBot-Map models while requiring subclasses to provide specific implementations for feature aggregation and camera pose estimation.

## Architecture and Key Components of GCTBase

The `GCTBase` class orchestrates the multi-task prediction pipeline through several specialized components.

### Inheritance and Mixins

As implemented in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py), `GCTBase` extends:

- **`torch.nn.Module`** – Provides standard PyTorch model functionality
- **`PyTorchModelHubMixin`** – Enables model serialization and Hugging Face Hub compatibility

This dual inheritance allows LingBot-Map models to support both custom training loops and standardized model distribution.

### Abstract Methods for Subclasses

Concrete implementations must override two critical methods defined in the base class:

- **`_build_aggregator()`** – Constructs the feature aggregation mechanism (e.g., streaming vs. windowed attention)
- **`_build_camera_head()`** – Creates the camera pose prediction head specific to the model variant

These abstract methods enforce a consistent interface while allowing architectural specialization across different input processing strategies.

### Prediction Heads and Forward Pass

The base class automatically constructs common prediction heads using `DPTHead` from [`lingbot_map/heads/dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/dpt_head.py):

- **Camera head** – Pose estimation and encoding
- **Depth head** – Metric or relative depth prediction
- **Point head** – 3D point cloud generation
- **Local-point head** – Local coordinate frame predictions

The **`forward()`** method in `GCTBase` handles input normalization, feature aggregation, and sequential execution of enabled prediction heads, returning a dictionary containing `pose_enc`, `depth`, and `world_points` tensors.

## Concrete Implementations and Subclasses

Because `GCTBase` is abstract, you cannot instantiate it directly. Instead, use one of the concrete subclasses that implement the required aggregation and camera head logic.

### GCTStream Implementation

The streaming variant processes frames sequentially without windowing constraints:

- **File:** [`lingbot_map/models/gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream.py)
- **Class:** `GCTStream`
- **Use case:** Real-time inference with continuous video streams

This implementation provides a concrete `_build_aggregator()` suitable for online processing where frame history accumulates indefinitely.

### Windowed Variants

For memory-constrained scenarios, LingBot-Map provides windowed implementations:

- **[`lingbot_map/models/gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window.py)** – Basic windowed GCT with fixed history buffers
- **[`lingbot_map/models/gct_stream_window_v2.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_stream_window_v2.py)** – Updated windowed implementation with additional optimization features

Both variants inherit from `GCTBase` and override the abstract methods to implement sliding window attention mechanisms that limit computational complexity during long sequences.

## Working with GCTBase in Practice

Instantiate concrete models directly or subclass `GCTBase` for custom architectures.

### Instantiating a Pre-built Model

Use the streaming implementation for standard inference tasks:

```python
from lingbot_map.models.gct_stream import GCTStream
import torch

# Create the model with default configuration

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    enable_camera=True,
    enable_depth=True,
    enable_point=True,
)

# Dummy input: a batch of 2 frames, each 3×H×W (values in [0, 1])

images = torch.rand(2, 3, 224, 224)

# Forward pass – returns a dict with pose, depth, and point predictions

outputs = model(images)

# Access specific predictions

camera_pose = outputs["pose_enc"]
depth_map = outputs["depth"]
world_pts = outputs["world_points"]

```

### Creating Custom Subclasses

Extend `GCTBase` to implement novel aggregation strategies:

```python
from lingbot_map.models.gct_base import GCTBase

class MyCustomModel(GCTBase):
    def _build_aggregator(self):
        # Build a custom transformer encoder for feature aggregation

        pass

    def _build_camera_head(self):
        # Provide a camera head compatible with the base class interface

        pass

# Instantiate your custom model

my_model = MyCustomModel()

```

## Summary

- **`GCTBase`** in [`lingbot_map/models/gct_base.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/models/gct_base.py) is the abstract base model class for LingBot-Map.
- It inherits from `torch.nn.Module` and `PyTorchModelHubMixin` for PyTorch and Hugging Face compatibility.
- Subclasses must implement `_build_aggregator()` and `_build_camera_head()` to define specific architectures.
- Concrete implementations include `GCTStream` ([`gct_stream.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream.py)) and windowed variants ([`gct_stream_window.py`](https://github.com/Robbyant/lingbot-map/blob/main/gct_stream_window.py)).
- The base class manages common prediction heads for camera, depth, and point estimation using `DPTHead`.

## Frequently Asked Questions

### Can I instantiate GCTBase directly?

No, `GCTBase` is an abstract class that defines the interface and common functionality for all LingBot-Map models. You must use concrete subclasses like `GCTStream` or `GCTStreamWindow`, or create your own subclass that implements the abstract methods `_build_aggregator()` and `_build_camera_head()`.

### What is the difference between GCTStream and GCTStreamWindow?

**`GCTStream`** processes frames with unbounded history accumulation, suitable for short sequences or systems with ample memory. **`GCTStreamWindow`** and **`GCTStreamWindowV2`** implement sliding window attention that limits the temporal context to a fixed buffer size, reducing memory consumption during long sequences or real-time deployment.

### How does GCTBase handle multiple prediction tasks?

The base class initializes separate prediction heads for each enabled task (camera, depth, points) during construction. The `forward()` method orchestrates these heads sequentially, normalizing inputs, aggregating features through the subclass-defined aggregator, and returning a dictionary with keys like `pose_enc`, `depth`, and `world_points` containing the respective predictions.