Where to Find the Base Model Class for LingBot-Map: Complete Implementation Guide

The base model class for LingBot-Map is GCTBase, defined in lingbot_map/models/gct_base.py, which serves as the abstract foundation for all Generalized Cross-Transformer (GCT) models in the repository.

The LingBot-Map repository implements a family of vision transformers for robotic mapping and localization tasks. Understanding the base model architecture is essential for extending the framework or debugging prediction pipelines, as every concrete model variant inherits from this single abstract interface.

Location of the Base Model Class in LingBot-Map

The foundational class for all GCT models resides in a single file within the models package.

The GCTBase Abstract Class

According to the LingBot-Map source code, the base model class GCTBase is located at:

This file defines the abstract base class that inherits from both torch.nn.Module and PyTorchModelHubMixin, enabling standard PyTorch training workflows alongside Hugging Face Hub integration. The class implements the shared architecture for LingBot-Map models while requiring subclasses to provide specific implementations for feature aggregation and camera pose estimation.

Architecture and Key Components of GCTBase

The GCTBase class orchestrates the multi-task prediction pipeline through several specialized components.

Inheritance and Mixins

As implemented in lingbot_map/models/gct_base.py, GCTBase extends:

  • torch.nn.Module – Provides standard PyTorch model functionality
  • PyTorchModelHubMixin – Enables model serialization and Hugging Face Hub compatibility

This dual inheritance allows LingBot-Map models to support both custom training loops and standardized model distribution.

Abstract Methods for Subclasses

Concrete implementations must override two critical methods defined in the base class:

  • _build_aggregator() – Constructs the feature aggregation mechanism (e.g., streaming vs. windowed attention)
  • _build_camera_head() – Creates the camera pose prediction head specific to the model variant

These abstract methods enforce a consistent interface while allowing architectural specialization across different input processing strategies.

Prediction Heads and Forward Pass

The base class automatically constructs common prediction heads using DPTHead from lingbot_map/heads/dpt_head.py:

  • Camera head – Pose estimation and encoding
  • Depth head – Metric or relative depth prediction
  • Point head – 3D point cloud generation
  • Local-point head – Local coordinate frame predictions

The forward() method in GCTBase handles input normalization, feature aggregation, and sequential execution of enabled prediction heads, returning a dictionary containing pose_enc, depth, and world_points tensors.

Concrete Implementations and Subclasses

Because GCTBase is abstract, you cannot instantiate it directly. Instead, use one of the concrete subclasses that implement the required aggregation and camera head logic.

GCTStream Implementation

The streaming variant processes frames sequentially without windowing constraints:

This implementation provides a concrete _build_aggregator() suitable for online processing where frame history accumulates indefinitely.

Windowed Variants

For memory-constrained scenarios, LingBot-Map provides windowed implementations:

Both variants inherit from GCTBase and override the abstract methods to implement sliding window attention mechanisms that limit computational complexity during long sequences.

Working with GCTBase in Practice

Instantiate concrete models directly or subclass GCTBase for custom architectures.

Instantiating a Pre-built Model

Use the streaming implementation for standard inference tasks:

from lingbot_map.models.gct_stream import GCTStream
import torch

# Create the model with default configuration

model = GCTStream(
    img_size=518,
    patch_size=14,
    embed_dim=1024,
    enable_camera=True,
    enable_depth=True,
    enable_point=True,
)

# Dummy input: a batch of 2 frames, each 3×H×W (values in [0, 1])

images = torch.rand(2, 3, 224, 224)

# Forward pass – returns a dict with pose, depth, and point predictions

outputs = model(images)

# Access specific predictions

camera_pose = outputs["pose_enc"]
depth_map = outputs["depth"]
world_pts = outputs["world_points"]

Creating Custom Subclasses

Extend GCTBase to implement novel aggregation strategies:

from lingbot_map.models.gct_base import GCTBase

class MyCustomModel(GCTBase):
    def _build_aggregator(self):
        # Build a custom transformer encoder for feature aggregation

        pass

    def _build_camera_head(self):
        # Provide a camera head compatible with the base class interface

        pass

# Instantiate your custom model

my_model = MyCustomModel()

Summary

  • GCTBase in lingbot_map/models/gct_base.py is the abstract base model class for LingBot-Map.
  • It inherits from torch.nn.Module and PyTorchModelHubMixin for PyTorch and Hugging Face compatibility.
  • Subclasses must implement _build_aggregator() and _build_camera_head() to define specific architectures.
  • Concrete implementations include GCTStream (gct_stream.py) and windowed variants (gct_stream_window.py).
  • The base class manages common prediction heads for camera, depth, and point estimation using DPTHead.

Frequently Asked Questions

Can I instantiate GCTBase directly?

No, GCTBase is an abstract class that defines the interface and common functionality for all LingBot-Map models. You must use concrete subclasses like GCTStream or GCTStreamWindow, or create your own subclass that implements the abstract methods _build_aggregator() and _build_camera_head().

What is the difference between GCTStream and GCTStreamWindow?

GCTStream processes frames with unbounded history accumulation, suitable for short sequences or systems with ample memory. GCTStreamWindow and GCTStreamWindowV2 implement sliding window attention that limits the temporal context to a fixed buffer size, reducing memory consumption during long sequences or real-time deployment.

How does GCTBase handle multiple prediction tasks?

The base class initializes separate prediction heads for each enabled task (camera, depth, points) during construction. The forward() method orchestrates these heads sequentially, normalizing inputs, aggregating features through the subclass-defined aggregator, and returning a dictionary with keys like pose_enc, depth, and world_points containing the respective predictions.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →