# DPTHead Architecture for Depth Prediction: Implementation Guide in Ling-Bot-Map

> Learn about the DPTHead architecture for depth prediction. This implementation fuses multi-scale features from Vision Transformer patch tokens to generate depth maps and confidence predictions. Explore the Robbyant/lingbot-map ...

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: architecture
- Published: 2026-07-25

---

**The `DPTHead` class implements a dense-prediction transformer head that converts Vision Transformer patch tokens into depth maps through multi-scale feature fusion, supporting both depth estimation and confidence prediction.**

The `DPTHead` architecture for depth prediction in the [Ling-Bot-Map](https://github.com/Robbyant/lingbot-map) repository follows the design from "Vision Transformers for Dense Prediction" (arXiv:2103.13413) and the Depth-Anything-V2 project. Located in [`lingbot_map/heads/dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/dpt_head.py), this module transforms intermediate transformer tokens into high-resolution depth predictions using a series of projection layers, upsampling operations, and residual fusion blocks.

## Core Components of the DPTHead Architecture

The `DPTHead` class aggregates features from multiple transformer layers and progressively fuses them to produce dense depth predictions. The implementation centers on several key architectural components defined in the `__init__` method of [`lingbot_map/heads/dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/dpt_head.py).

### Input Normalization and Token Projection

The head begins by normalizing incoming patch tokens using **layer normalization**. The `self.norm = nn.LayerNorm(dim_in)` layer processes flattened token vectors before they enter the projection pipeline. Following normalization, `self.projects`—a `nn.ModuleList` of 1×1 convolutions—maps the transformer dimension (`dim_in`) to multiple output channel sizes specified in `out_channels` (typically `[256, 512, 1024, 1024]`).

### Upsampling and Resize Layers

Each projected feature map undergoes spatial restoration through `self.resize_layers`, another `nn.ModuleList` containing transposed convolutions or strided convolutions. These layers upsample the patch-level representations (spatial size `H/patch_size × W/patch_size`) back toward the original image resolution, or a down-scaled version when `down_ratio > 1`.

### Positional Embedding Injection

When `pos_embed=True`, the head injects UV-grid based positional encodings into feature maps via the `_apply_pos_embed` method. This utility function, supported by `create_uv_grid` from [`lingbot_map/heads/utils.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/utils.py), generates coordinate grids that help the network maintain spatial awareness during the token-to-image conversion process.

### Feature Fusion via the Scratch Network

The **"scratch" fusion pipeline** constitutes the core of the DPTHead architecture for depth prediction. The `_make_scratch` helper constructs a lightweight network containing:
- Per-scale refinement layers (`layer1_rn` through `layer4_rn`) implemented as 1×1 convolutions
- `FeatureFusionBlock` modules created by `_make_fusion_block` that perform residual convolutions and bilinear upsampling
- Progressive merging of multi-scale features through `refinenet1` to `refinenet4`

Each `FeatureFusionBlock` utilizes `ResidualConvUnit` modules—simple residual blocks with two 3×3 convolutions—to refine features before upsampling.

### Output Convolution and Activation

The final prediction stage uses `self.scratch.output_conv2` to reduce the fused feature channels to `output_dim` (default 4). The `activate_head` function from [`lingbot_map/heads/head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/head_act.py) then splits these logits into depth predictions and confidence maps. By default, the implementation uses `inv_log` activation for depth values and `expp1` for confidence scores.

### Chunked Inference Support

For long video sequences, the `forward` method implements chunked processing via the `frames_chunk_size` parameter. This optional logic in `_forward_impl` splits the batch along the temporal dimension to reduce GPU memory consumption during inference.

## Forward Pass Execution Flow

The `DPTHead.forward` method executes an 8-stage pipeline to transform transformer tokens into depth predictions:

1. **Token Selection**: For each layer index in `intermediate_layer_idx`, the head extracts patch tokens from `aggregated_tokens_list[layer_idx][:, :, patch_start_idx:]`, skipping extra tokens like class or camera embeddings.

2. **Reshaping**: Tokens are rearranged from `[B, S, N, C]` to `[B·S, C, Hₚ, Wₚ]` where `Hₚ = H/patch_size` and `Wₚ = W/patch_size`.

3. **Normalization**: `self.norm` applies layer normalization to the flattened token vectors.

4. **Projection and Upsampling**: The `self.projects` 1×1 convolutions map channels to the target dimensions, followed by `self.resize_layers` upsampling to the target resolution.

5. **Positional Encoding**: If enabled, `_apply_pos_embed` fuses UV-grid positional encodings into the feature maps.

6. **Multi-Scale Fusion**: The `scratch_forward` method passes all four upsampled maps through the scratch network's `FeatureFusionBlock` modules, progressively merging scales from low to high resolution.

7. **Safe Interpolation**: `custom_interpolate` resizes the fused features to the exact target resolution, chunking the operation if the tensor size would exceed INT_MAX limits.

8. **Output Generation**: Either returns fused features directly (when `feature_only=True`) or processes through `output_conv2` and `activate_head` to produce final depth and confidence tensors.

## Helper Modules and Utilities

The [`dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/dpt_head.py) file contains several supporting classes that enable the DPTHead architecture for depth prediction:

- **`_make_fusion_block`**: Wraps `FeatureFusionBlock` for building the refinement network
- **`_make_scratch`**: Constructs the scratch network architecture with random initialization ("from scratch")
- **`ResidualConvUnit`**: Provides basic residual connectivity with optional batch normalization
- **`custom_interpolate`**: Memory-safe wrapper around `nn.functional.interpolate` for high-resolution outputs

## Implementation Examples

### Basic Instantiation

```python
import torch
from lingbot_map.heads.dpt_head import DPTHead

# Configuration matching the repository defaults

head = DPTHead(
    dim_in=768,           # Input token dimension from transformer backbone

    patch_size=14,
    output_dim=4,        # 3-channel depth + 1-channel confidence

    activation="inv_log",
    conf_activation="expp1",
    features=256,
    out_channels=[256, 512, 1024, 1024],
    intermediate_layer_idx=[0, 1, 2, 3],
    pos_embed=True,
    feature_only=False,
    down_ratio=1,
)

```

### Forward Pass on Dummy Batch

```python

# Simulated aggregated token list from 4 transformer layers

# Shape: [B, S, N_total, C] where N_total = patch_tokens + extra tokens

B, S, C = 2, 5, 768
patch_tokens = 196               # 14×14 patches per image

extra_tokens = 2                 # class / camera tokens

N_total = patch_tokens + extra_tokens

agg_tokens = [
    torch.randn(B, S, N_total, C) for _ in range(4)
]

# Dummy images [B, S, 3, H, W]

images = torch.rand(B, S, 3, 224, 224)
patch_start_idx = extra_tokens  # Skip non-patch tokens

# Returns depth predictions and confidence maps

depth, confidence = head(agg_tokens, images, patch_start_idx)
print(depth.shape, confidence.shape)   

# → torch.Size([2, 5, 1, 224, 224]) torch.Size([2, 5, 1, 224, 224])

```

### Feature-Only Mode for Downstream Tasks

```python
head_feat = DPTHead(
    dim_in=768,
    feature_only=True,  # Skip output_conv2 and activation head

)

features = head_feat(agg_tokens, images, patch_start_idx)
print(features.shape)   # → torch.Size([2, 5, C_fused, 224, 224])

```

### Chunked Inference for Video Sequences

```python

# Process 30 frames in chunks of 8 to manage GPU memory

depth, conf = head(
    agg_tokens,                 # List of 4 tensors, each [B, 30, N, C]

    images,                     # [B, 30, 3, 224, 224]

    patch_start_idx,
    frames_chunk_size=8,
)

```

## Key Source Files

The DPTHead architecture for depth prediction spans several files in the lingbot-map repository:

- **[`lingbot_map/heads/dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/dpt_head.py)**: Core implementation containing the `DPTHead` class, `FeatureFusionBlock`, and `ResidualConvUnit`
- **[`lingbot_map/heads/head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/head_act.py)**: Defines `activate_head` for converting logits to depth/confidence values
- **[`lingbot_map/heads/utils.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/utils.py)**: Provides `create_uv_grid` and `position_grid_to_embed` for positional encodings
- **[`lingbot_map/layers/vision_transformer.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/layers/vision_transformer.py)**: Supplies the transformer backbone that feeds tokens to the head
- **[`benchmark/methods/lingbot_map.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/methods/lingbot_map.py)**: Demonstrates end-to-end integration within the benchmark pipeline

## Summary

- The **DPTHead** class in [`lingbot_map/heads/dpt_head.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/dpt_head.py) implements a dense-prediction head following the DPT (Dense Prediction Transformer) architecture.
- It processes intermediate transformer tokens through **layer normalization**, **1×1 projections**, and **transposed convolutions** to restore spatial resolution.
- The **scratch network** fuses multi-scale features using `FeatureFusionBlock` modules with residual connections.
- **Positional embeddings** based on UV grids can be injected to enhance spatial coherence.
- The head supports **chunked inference** via `frames_chunk_size` for processing long video sequences without memory overflow.
- Output activations (`inv_log` for depth, `expp1` for confidence) are handled by `activate_head` in [`head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/head_act.py).

## Frequently Asked Questions

### What input does the DPTHead expect from the transformer backbone?

The `DPTHead` expects a list of aggregated token tensors (`aggregated_tokens_list`) from specific transformer layers indexed by `intermediate_layer_idx`. Each tensor has shape `[B, S, N_total, C]` where `B` is batch size, `S` is sequence length (frames), `N_total` includes both patch tokens and extra tokens (like class or camera embeddings), and `C` is the token dimension (`dim_in`). The `patch_start_idx` parameter tells the head which index to start slicing from to exclude non-patch tokens.

### How does the DPTHead handle memory constraints with high-resolution images?

The implementation includes two memory-saving mechanisms. First, `custom_interpolate` safely handles upsampling by chunking tensors if the output size would exceed INT_MAX limits. Second, the `frames_chunk_size` parameter in the `forward` method enables chunked inference, processing temporal sequences in smaller chunks to reduce peak GPU memory usage during video depth estimation.

### What is the difference between `feature_only` mode and standard depth prediction?

When `feature_only=True`, the `DPTHead` returns the fused multi-scale features directly after the scratch network's interpolation step, skipping the final `output_conv2` convolution and activation functions. This mode produces high-level features suitable for downstream tasks rather than explicit depth maps. When `feature_only=False` (default), the head applies the output convolution and `activate_head` to generate depth predictions and confidence values.

### Which activation functions does the DPTHead use for depth and confidence outputs?

According to the implementation in [`lingbot_map/heads/head_act.py`](https://github.com/Robbyant/lingbot-map/blob/main/lingbot_map/heads/head_act.py), the default configuration uses `inv_log` (inverse logarithmic) activation for depth predictions and `expp1` (exponential plus one) for confidence maps. These activations are applied by the `activate_head` function after the final `output_conv2` layer, converting raw logits into physically meaningful depth and uncertainty estimates.