Where to Find the Base Model Class for LingBot-Map: Complete Implementation Guide
The base model class for LingBot-Map is GCTBase, defined in lingbot_map/models/gct_base.py, which serves as the abstract foundation for all Generalized Cross-Transformer (GCT) models in the repository.
The LingBot-Map repository implements a family of vision transformers for robotic mapping and localization tasks. Understanding the base model architecture is essential for extending the framework or debugging prediction pipelines, as every concrete model variant inherits from this single abstract interface.
Location of the Base Model Class in LingBot-Map
The foundational class for all GCT models resides in a single file within the models package.
The GCTBase Abstract Class
According to the LingBot-Map source code, the base model class GCTBase is located at:
This file defines the abstract base class that inherits from both torch.nn.Module and PyTorchModelHubMixin, enabling standard PyTorch training workflows alongside Hugging Face Hub integration. The class implements the shared architecture for LingBot-Map models while requiring subclasses to provide specific implementations for feature aggregation and camera pose estimation.
Architecture and Key Components of GCTBase
The GCTBase class orchestrates the multi-task prediction pipeline through several specialized components.
Inheritance and Mixins
As implemented in lingbot_map/models/gct_base.py, GCTBase extends:
torch.nn.Module– Provides standard PyTorch model functionalityPyTorchModelHubMixin– Enables model serialization and Hugging Face Hub compatibility
This dual inheritance allows LingBot-Map models to support both custom training loops and standardized model distribution.
Abstract Methods for Subclasses
Concrete implementations must override two critical methods defined in the base class:
_build_aggregator()– Constructs the feature aggregation mechanism (e.g., streaming vs. windowed attention)_build_camera_head()– Creates the camera pose prediction head specific to the model variant
These abstract methods enforce a consistent interface while allowing architectural specialization across different input processing strategies.
Prediction Heads and Forward Pass
The base class automatically constructs common prediction heads using DPTHead from lingbot_map/heads/dpt_head.py:
- Camera head – Pose estimation and encoding
- Depth head – Metric or relative depth prediction
- Point head – 3D point cloud generation
- Local-point head – Local coordinate frame predictions
The forward() method in GCTBase handles input normalization, feature aggregation, and sequential execution of enabled prediction heads, returning a dictionary containing pose_enc, depth, and world_points tensors.
Concrete Implementations and Subclasses
Because GCTBase is abstract, you cannot instantiate it directly. Instead, use one of the concrete subclasses that implement the required aggregation and camera head logic.
GCTStream Implementation
The streaming variant processes frames sequentially without windowing constraints:
- File:
lingbot_map/models/gct_stream.py - Class:
GCTStream - Use case: Real-time inference with continuous video streams
This implementation provides a concrete _build_aggregator() suitable for online processing where frame history accumulates indefinitely.
Windowed Variants
For memory-constrained scenarios, LingBot-Map provides windowed implementations:
lingbot_map/models/gct_stream_window.py– Basic windowed GCT with fixed history bufferslingbot_map/models/gct_stream_window_v2.py– Updated windowed implementation with additional optimization features
Both variants inherit from GCTBase and override the abstract methods to implement sliding window attention mechanisms that limit computational complexity during long sequences.
Working with GCTBase in Practice
Instantiate concrete models directly or subclass GCTBase for custom architectures.
Instantiating a Pre-built Model
Use the streaming implementation for standard inference tasks:
from lingbot_map.models.gct_stream import GCTStream
import torch
# Create the model with default configuration
model = GCTStream(
img_size=518,
patch_size=14,
embed_dim=1024,
enable_camera=True,
enable_depth=True,
enable_point=True,
)
# Dummy input: a batch of 2 frames, each 3×H×W (values in [0, 1])
images = torch.rand(2, 3, 224, 224)
# Forward pass – returns a dict with pose, depth, and point predictions
outputs = model(images)
# Access specific predictions
camera_pose = outputs["pose_enc"]
depth_map = outputs["depth"]
world_pts = outputs["world_points"]
Creating Custom Subclasses
Extend GCTBase to implement novel aggregation strategies:
from lingbot_map.models.gct_base import GCTBase
class MyCustomModel(GCTBase):
def _build_aggregator(self):
# Build a custom transformer encoder for feature aggregation
pass
def _build_camera_head(self):
# Provide a camera head compatible with the base class interface
pass
# Instantiate your custom model
my_model = MyCustomModel()
Summary
GCTBaseinlingbot_map/models/gct_base.pyis the abstract base model class for LingBot-Map.- It inherits from
torch.nn.ModuleandPyTorchModelHubMixinfor PyTorch and Hugging Face compatibility. - Subclasses must implement
_build_aggregator()and_build_camera_head()to define specific architectures. - Concrete implementations include
GCTStream(gct_stream.py) and windowed variants (gct_stream_window.py). - The base class manages common prediction heads for camera, depth, and point estimation using
DPTHead.
Frequently Asked Questions
Can I instantiate GCTBase directly?
No, GCTBase is an abstract class that defines the interface and common functionality for all LingBot-Map models. You must use concrete subclasses like GCTStream or GCTStreamWindow, or create your own subclass that implements the abstract methods _build_aggregator() and _build_camera_head().
What is the difference between GCTStream and GCTStreamWindow?
GCTStream processes frames with unbounded history accumulation, suitable for short sequences or systems with ample memory. GCTStreamWindow and GCTStreamWindowV2 implement sliding window attention that limits the temporal context to a fixed buffer size, reducing memory consumption during long sequences or real-time deployment.
How does GCTBase handle multiple prediction tasks?
The base class initializes separate prediction heads for each enabled task (camera, depth, points) during construction. The forward() method orchestrates these heads sequentially, normalizing inputs, aggregating features through the subclass-defined aggregator, and returning a dictionary with keys like pose_enc, depth, and world_points containing the respective predictions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →