DPTHead Architecture for Depth Prediction: Implementation Guide in Ling-Bot-Map

The DPTHead class implements a dense-prediction transformer head that converts Vision Transformer patch tokens into depth maps through multi-scale feature fusion, supporting both depth estimation and confidence prediction.

The DPTHead architecture for depth prediction in the Ling-Bot-Map repository follows the design from "Vision Transformers for Dense Prediction" (arXiv:2103.13413) and the Depth-Anything-V2 project. Located in lingbot_map/heads/dpt_head.py, this module transforms intermediate transformer tokens into high-resolution depth predictions using a series of projection layers, upsampling operations, and residual fusion blocks.

Core Components of the DPTHead Architecture

The DPTHead class aggregates features from multiple transformer layers and progressively fuses them to produce dense depth predictions. The implementation centers on several key architectural components defined in the __init__ method of lingbot_map/heads/dpt_head.py.

Input Normalization and Token Projection

The head begins by normalizing incoming patch tokens using layer normalization. The self.norm = nn.LayerNorm(dim_in) layer processes flattened token vectors before they enter the projection pipeline. Following normalization, self.projects—a nn.ModuleList of 1×1 convolutions—maps the transformer dimension (dim_in) to multiple output channel sizes specified in out_channels (typically [256, 512, 1024, 1024]).

Upsampling and Resize Layers

Each projected feature map undergoes spatial restoration through self.resize_layers, another nn.ModuleList containing transposed convolutions or strided convolutions. These layers upsample the patch-level representations (spatial size H/patch_size × W/patch_size) back toward the original image resolution, or a down-scaled version when down_ratio > 1.

Positional Embedding Injection

When pos_embed=True, the head injects UV-grid based positional encodings into feature maps via the _apply_pos_embed method. This utility function, supported by create_uv_grid from lingbot_map/heads/utils.py, generates coordinate grids that help the network maintain spatial awareness during the token-to-image conversion process.

Feature Fusion via the Scratch Network

The "scratch" fusion pipeline constitutes the core of the DPTHead architecture for depth prediction. The _make_scratch helper constructs a lightweight network containing:

  • Per-scale refinement layers (layer1_rn through layer4_rn) implemented as 1×1 convolutions
  • FeatureFusionBlock modules created by _make_fusion_block that perform residual convolutions and bilinear upsampling
  • Progressive merging of multi-scale features through refinenet1 to refinenet4

Each FeatureFusionBlock utilizes ResidualConvUnit modules—simple residual blocks with two 3×3 convolutions—to refine features before upsampling.

Output Convolution and Activation

The final prediction stage uses self.scratch.output_conv2 to reduce the fused feature channels to output_dim (default 4). The activate_head function from lingbot_map/heads/head_act.py then splits these logits into depth predictions and confidence maps. By default, the implementation uses inv_log activation for depth values and expp1 for confidence scores.

Chunked Inference Support

For long video sequences, the forward method implements chunked processing via the frames_chunk_size parameter. This optional logic in _forward_impl splits the batch along the temporal dimension to reduce GPU memory consumption during inference.

Forward Pass Execution Flow

The DPTHead.forward method executes an 8-stage pipeline to transform transformer tokens into depth predictions:

  1. Token Selection: For each layer index in intermediate_layer_idx, the head extracts patch tokens from aggregated_tokens_list[layer_idx][:, :, patch_start_idx:], skipping extra tokens like class or camera embeddings.

  2. Reshaping: Tokens are rearranged from [B, S, N, C] to [B·S, C, Hₚ, Wₚ] where Hₚ = H/patch_size and Wₚ = W/patch_size.

  3. Normalization: self.norm applies layer normalization to the flattened token vectors.

  4. Projection and Upsampling: The self.projects 1×1 convolutions map channels to the target dimensions, followed by self.resize_layers upsampling to the target resolution.

  5. Positional Encoding: If enabled, _apply_pos_embed fuses UV-grid positional encodings into the feature maps.

  6. Multi-Scale Fusion: The scratch_forward method passes all four upsampled maps through the scratch network's FeatureFusionBlock modules, progressively merging scales from low to high resolution.

  7. Safe Interpolation: custom_interpolate resizes the fused features to the exact target resolution, chunking the operation if the tensor size would exceed INT_MAX limits.

  8. Output Generation: Either returns fused features directly (when feature_only=True) or processes through output_conv2 and activate_head to produce final depth and confidence tensors.

Helper Modules and Utilities

The dpt_head.py file contains several supporting classes that enable the DPTHead architecture for depth prediction:

  • _make_fusion_block: Wraps FeatureFusionBlock for building the refinement network
  • _make_scratch: Constructs the scratch network architecture with random initialization ("from scratch")
  • ResidualConvUnit: Provides basic residual connectivity with optional batch normalization
  • custom_interpolate: Memory-safe wrapper around nn.functional.interpolate for high-resolution outputs

Implementation Examples

Basic Instantiation

import torch
from lingbot_map.heads.dpt_head import DPTHead

# Configuration matching the repository defaults

head = DPTHead(
    dim_in=768,           # Input token dimension from transformer backbone

    patch_size=14,
    output_dim=4,        # 3-channel depth + 1-channel confidence

    activation="inv_log",
    conf_activation="expp1",
    features=256,
    out_channels=[256, 512, 1024, 1024],
    intermediate_layer_idx=[0, 1, 2, 3],
    pos_embed=True,
    feature_only=False,
    down_ratio=1,
)

Forward Pass on Dummy Batch


# Simulated aggregated token list from 4 transformer layers

# Shape: [B, S, N_total, C] where N_total = patch_tokens + extra tokens

B, S, C = 2, 5, 768
patch_tokens = 196               # 14×14 patches per image

extra_tokens = 2                 # class / camera tokens

N_total = patch_tokens + extra_tokens

agg_tokens = [
    torch.randn(B, S, N_total, C) for _ in range(4)
]

# Dummy images [B, S, 3, H, W]

images = torch.rand(B, S, 3, 224, 224)
patch_start_idx = extra_tokens  # Skip non-patch tokens

# Returns depth predictions and confidence maps

depth, confidence = head(agg_tokens, images, patch_start_idx)
print(depth.shape, confidence.shape)   

# → torch.Size([2, 5, 1, 224, 224]) torch.Size([2, 5, 1, 224, 224])

Feature-Only Mode for Downstream Tasks

head_feat = DPTHead(
    dim_in=768,
    feature_only=True,  # Skip output_conv2 and activation head

)

features = head_feat(agg_tokens, images, patch_start_idx)
print(features.shape)   # → torch.Size([2, 5, C_fused, 224, 224])

Chunked Inference for Video Sequences


# Process 30 frames in chunks of 8 to manage GPU memory

depth, conf = head(
    agg_tokens,                 # List of 4 tensors, each [B, 30, N, C]

    images,                     # [B, 30, 3, 224, 224]

    patch_start_idx,
    frames_chunk_size=8,
)

Key Source Files

The DPTHead architecture for depth prediction spans several files in the lingbot-map repository:

Summary

  • The DPTHead class in lingbot_map/heads/dpt_head.py implements a dense-prediction head following the DPT (Dense Prediction Transformer) architecture.
  • It processes intermediate transformer tokens through layer normalization, 1×1 projections, and transposed convolutions to restore spatial resolution.
  • The scratch network fuses multi-scale features using FeatureFusionBlock modules with residual connections.
  • Positional embeddings based on UV grids can be injected to enhance spatial coherence.
  • The head supports chunked inference via frames_chunk_size for processing long video sequences without memory overflow.
  • Output activations (inv_log for depth, expp1 for confidence) are handled by activate_head in head_act.py.

Frequently Asked Questions

What input does the DPTHead expect from the transformer backbone?

The DPTHead expects a list of aggregated token tensors (aggregated_tokens_list) from specific transformer layers indexed by intermediate_layer_idx. Each tensor has shape [B, S, N_total, C] where B is batch size, S is sequence length (frames), N_total includes both patch tokens and extra tokens (like class or camera embeddings), and C is the token dimension (dim_in). The patch_start_idx parameter tells the head which index to start slicing from to exclude non-patch tokens.

How does the DPTHead handle memory constraints with high-resolution images?

The implementation includes two memory-saving mechanisms. First, custom_interpolate safely handles upsampling by chunking tensors if the output size would exceed INT_MAX limits. Second, the frames_chunk_size parameter in the forward method enables chunked inference, processing temporal sequences in smaller chunks to reduce peak GPU memory usage during video depth estimation.

What is the difference between feature_only mode and standard depth prediction?

When feature_only=True, the DPTHead returns the fused multi-scale features directly after the scratch network's interpolation step, skipping the final output_conv2 convolution and activation functions. This mode produces high-level features suitable for downstream tasks rather than explicit depth maps. When feature_only=False (default), the head applies the output convolution and activate_head to generate depth predictions and confidence values.

Which activation functions does the DPTHead use for depth and confidence outputs?

According to the implementation in lingbot_map/heads/head_act.py, the default configuration uses inv_log (inverse logarithmic) activation for depth predictions and expp1 (exponential plus one) for confidence maps. These activations are applied by the activate_head function after the final output_conv2 layer, converting raw logits into physically meaningful depth and uncertainty estimates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →