What Is the Role of the Detect Head in YOLOv5? Architecture and Implementation Explained

The Detect head in YOLOv5 serves as the final neural network layer that transforms backbone feature maps into concrete object detection predictions by applying anchor-based decoding, grid generation, and bounding box regression.

The Detect head is the critical bridge between raw convolutional features and usable detection results in the ultralytics/yolov5 repository. Understanding the role of the Detect head in YOLOv5 is essential for anyone customizing the architecture, exporting to ONNX/TensorRT, or debugging prediction quality.

Core Responsibilities of the Detect Head

Located in models/yolo.py, the Detect class (lines 72–89) performs several specialized operations that convert multi-scale feature maps into the final detection tensor format [x, y, w, h, confidence, class_scores].

Anchor Box Management

The head stores predefined anchor boxes for each detection scale (P3, P4, P5) as registered buffers to ensure they are saved with the model state but not updated during backpropagation:

self.register_buffer("anchors", torch.tensor(anchors).float().view(self.nl, -1, 2))

Here, self.nl represents the number of detection layers (typically 3), and the anchors are reshaped to [nl, na, 2] where na is the number of anchors per layer.

Convolutional Projection

Each feature map channel set undergoes a 1×1 convolution that projects the backbone features into the detection output space. In models/yolo.py line 89, the module list is constructed as:

self.m = nn.ModuleList(nn.Conv2d(x, self.no * self.na, 1) for x in ch)

Where:

  • ch is the list of input channel depths (e.g., [256, 512, 1024] for YOLOv5s)
  • self.no = nc + 5 (number of classes plus 5 box parameters)
  • self.na = number of anchors per layer (typically 3)

Tensor Reshaping and Grid Generation

During the forward pass (lines 96–103), the raw convolutional output is reshaped from [bs, self.no*na, ny, nx] to [bs, na, ny, nx, self.no] and permuted for efficient memory layout:

x[i] = x[i].view(bs, self.na, self.no, ny, nx).permute(0, 1, 3, 4, 2).contiguous()

Simultaneously, the head creates coordinate grids and anchor grids using _make_grid:

self.grid[i], self.anchor_grid[i] = self._make_grid(nx, ny, i)

These grids map relative box coordinates to absolute image coordinates and are recomputed dynamically when input strides change, which is critical for ONNX and TensorRT export compatibility.

Prediction Decoding

The final decoding step (lines 110–113) splits the prediction tensor into components and applies the sigmoid activation and geometric transformations:

xy, wh, conf = x[i].sigmoid().split((2, 2, self.nc + 1), 4)
xy = (xy * 2 + self.grid[i]) * self.stride[i]  # xy coordinates

wh = (wh * 2) ** 2 * self.anchor_grid[i]        # width/height

This converts the network's raw outputs into normalized bounding box coordinates, objectness scores, and class probabilities.

Training vs. Inference Modes

The Detect head exhibits different behavior depending on the model mode, controlled by the return statement in lines 115–117:

return x if self.training else (torch.cat(z, 1),) if self.export else (torch.cat(z, 1), x)
  • Training mode: Returns the raw prediction list x so that the loss function (ComputeLoss) can operate on unprocessed tensors
  • Inference mode: Concatenates all scale predictions into a single tensor [bs, N, no] where N is the total number of predictions across all anchors and scales
  • Export mode: Returns only the concatenated predictions without the auxiliary training tensors, optimizing for ONNX/TensorRT deployment

Extension for Instance Segmentation

YOLOv5 extends the Detect head for segmentation tasks through the Segment class (lines 130–141 in models/yolo.py), which inherits from Detect:

class Segment(Detect):
    def __init__(self, nc=80, anchors=(), nm=32, npr=256, ch=(), inplace=True):
        super().__init__(nc, anchors, ch, inplace)
        self.nm = nm  # number of masks

        self.npr = npr  # number of protos

        self.m = nn.ModuleList(nn.Conv2d(x, self.no * self.na, 1) for x in ch)  # output conv

        self.proto = Proto(ch[0], self.npr, self.nm)  # protos

This subclass adds mask prototype coefficients to the detection output, enabling simultaneous object detection and instance segmentation while reusing the core anchor-based decoding logic.

Practical Code Examples

Loading and Running Inference

import torch
from models.common import DetectMultiBackend

# Load pretrained YOLOv5 model (Detect head included internally)

model = DetectMultiBackend('yolov5s.pt', device='cpu')

# Dummy input tensor (batch=1, 3 channels, 640×640)

img = torch.randn(1, 3, 640, 640)

# Run inference – Detect head processes feature maps internally

with torch.no_grad():
    preds = model(img)[0]  # Output: [1, 25200, 85]

print(preds.shape)  # 25200 predictions = 3 scales × 80×80 + 40×40 + 20×20 anchors

Manual Detect Head Initialization

from models.yolo import Detect

# Configuration matching YOLOv5s

num_classes = 80
anchors = [
    [(10,13), (16,30), (33,23)],      # P3/8

    [(30,61), (62,45), (59,119)],     # P4/16

    [(116,90), (156,198), (373,326)]  # P5/32

]
channels = [256, 512, 1024]  # From C3 blocks in backbone

# Instantiate head

detect = Detect(nc=num_classes, anchors=anchors, ch=channels)

# Simulate backbone output

feature_maps = [
    torch.randn(1, 256, 80, 80),   # P3

    torch.randn(1, 512, 40, 40),   # P4

    torch.randn(1, 1024, 20, 20)   # P5

]

# Forward pass (training mode returns raw tensors)

outputs = detect(feature_maps)
print([o.shape for o in outputs])  # [1, 3, 80, 80, 85], [1, 3, 40, 40, 85], etc.

Summary

  • The Detect head in models/yolo.py serves as the final processing layer that converts backbone feature maps into bounding box predictions, objectness scores, and class probabilities.

  • It employs 1×1 convolutions to project features into the output space, followed by tensor reshaping and grid generation to map relative coordinates to absolute image positions.

  • Anchor-based decoding scales width and height predictions using predefined anchor boxes, while sigmoid activations normalize center coordinates and confidences.

  • The head operates in dual modes: returning raw tensors for loss computation during training, and concatenated, decoded predictions during inference or export.

  • The Segment subclass extends Detect for instance segmentation by adding mask prototype coefficients, demonstrating the head's extensibility for multi-task learning.

Frequently Asked Questions

How does the Detect head differ from the backbone and neck in YOLOv5?

The backbone (CSPDarknet) extracts hierarchical visual features from input images, while the neck (PANet) fuses multi-scale features to create rich detection representations. The Detect head is the final stage that operates on these fused features to produce concrete predictions—bounding box coordinates, objectness scores, and class probabilities—using 1×1 convolutions and anchor-based decoding.

Why does the Detect head use three different scales (P3, P4, P5)?

YOLOv5 employs multi-scale detection to handle objects of varying sizes effectively. P3 (stride 8) captures fine-grained details for small objects, P4 (stride 16) handles medium-sized objects, and P5 (stride 32) detects large objects with global context. The Detect head processes each scale independently using scale-specific anchor boxes, then concatenates the results during inference to produce a comprehensive detection output.

Can the Detect head be modified for custom object detection tasks?

Yes, the Detect head is highly configurable through the model YAML files (e.g., yolov5s.yaml). You can adjust the number of anchors per scale, modify the anchor box dimensions to match your dataset's object aspect ratios, or change the number of classes (nc). For advanced modifications, you can subclass Detect in models/yolo.py to implement custom decoding logic, such as anchor-free detection or additional prediction heads for keypoint estimation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →