# What Is the Role of the Detect Head in YOLOv5? Architecture and Implementation Explained

> Discover the Detect head in YOLOv5. Understand its role in object detection, anchor decoding, grid generation, and bounding box regression for precise predictions. Learn how it transforms feature maps into outputs.

- Repository: [Ultralytics/yolov5](https://github.com/ultralytics/yolov5)
- Tags: internals
- Published: 2026-03-06

---

**The Detect head in YOLOv5 serves as the final neural network layer that transforms backbone feature maps into concrete object detection predictions by applying anchor-based decoding, grid generation, and bounding box regression.**

The Detect head is the critical bridge between raw convolutional features and usable detection results in the `ultralytics/yolov5` repository. Understanding the role of the Detect head in YOLOv5 is essential for anyone customizing the architecture, exporting to ONNX/TensorRT, or debugging prediction quality.

## Core Responsibilities of the Detect Head

Located in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py), the `Detect` class (lines 72–89) performs several specialized operations that convert multi-scale feature maps into the final detection tensor format `[x, y, w, h, confidence, class_scores]`.

### Anchor Box Management

The head stores predefined anchor boxes for each detection scale (P3, P4, P5) as registered buffers to ensure they are saved with the model state but not updated during backpropagation:

```python
self.register_buffer("anchors", torch.tensor(anchors).float().view(self.nl, -1, 2))

```

Here, `self.nl` represents the number of detection layers (typically 3), and the anchors are reshaped to `[nl, na, 2]` where `na` is the number of anchors per layer.

### Convolutional Projection

Each feature map channel set undergoes a 1×1 convolution that projects the backbone features into the detection output space. In [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py) line 89, the module list is constructed as:

```python
self.m = nn.ModuleList(nn.Conv2d(x, self.no * self.na, 1) for x in ch)

```

Where:
- `ch` is the list of input channel depths (e.g., `[256, 512, 1024]` for YOLOv5s)
- `self.no = nc + 5` (number of classes plus 5 box parameters)
- `self.na` = number of anchors per layer (typically 3)

### Tensor Reshaping and Grid Generation

During the forward pass (lines 96–103), the raw convolutional output is reshaped from `[bs, self.no*na, ny, nx]` to `[bs, na, ny, nx, self.no]` and permuted for efficient memory layout:

```python
x[i] = x[i].view(bs, self.na, self.no, ny, nx).permute(0, 1, 3, 4, 2).contiguous()

```

Simultaneously, the head creates coordinate grids and anchor grids using `_make_grid`:

```python
self.grid[i], self.anchor_grid[i] = self._make_grid(nx, ny, i)

```

These grids map relative box coordinates to absolute image coordinates and are recomputed dynamically when input strides change, which is critical for ONNX and TensorRT export compatibility.

### Prediction Decoding

The final decoding step (lines 110–113) splits the prediction tensor into components and applies the sigmoid activation and geometric transformations:

```python
xy, wh, conf = x[i].sigmoid().split((2, 2, self.nc + 1), 4)
xy = (xy * 2 + self.grid[i]) * self.stride[i]  # xy coordinates

wh = (wh * 2) ** 2 * self.anchor_grid[i]        # width/height

```

This converts the network's raw outputs into normalized bounding box coordinates, objectness scores, and class probabilities.

## Training vs. Inference Modes

The Detect head exhibits different behavior depending on the model mode, controlled by the return statement in lines 115–117:

```python
return x if self.training else (torch.cat(z, 1),) if self.export else (torch.cat(z, 1), x)

```

- **Training mode**: Returns the raw prediction list `x` so that the loss function (`ComputeLoss`) can operate on unprocessed tensors
- **Inference mode**: Concatenates all scale predictions into a single tensor `[bs, N, no]` where N is the total number of predictions across all anchors and scales
- **Export mode**: Returns only the concatenated predictions without the auxiliary training tensors, optimizing for ONNX/TensorRT deployment

## Extension for Instance Segmentation

YOLOv5 extends the Detect head for segmentation tasks through the `Segment` class (lines 130–141 in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py)), which inherits from `Detect`:

```python
class Segment(Detect):
    def __init__(self, nc=80, anchors=(), nm=32, npr=256, ch=(), inplace=True):
        super().__init__(nc, anchors, ch, inplace)
        self.nm = nm  # number of masks

        self.npr = npr  # number of protos

        self.m = nn.ModuleList(nn.Conv2d(x, self.no * self.na, 1) for x in ch)  # output conv

        self.proto = Proto(ch[0], self.npr, self.nm)  # protos

```

This subclass adds mask prototype coefficients to the detection output, enabling simultaneous object detection and instance segmentation while reusing the core anchor-based decoding logic.

## Practical Code Examples

### Loading and Running Inference

```python
import torch
from models.common import DetectMultiBackend

# Load pretrained YOLOv5 model (Detect head included internally)

model = DetectMultiBackend('yolov5s.pt', device='cpu')

# Dummy input tensor (batch=1, 3 channels, 640×640)

img = torch.randn(1, 3, 640, 640)

# Run inference – Detect head processes feature maps internally

with torch.no_grad():
    preds = model(img)[0]  # Output: [1, 25200, 85]

print(preds.shape)  # 25200 predictions = 3 scales × 80×80 + 40×40 + 20×20 anchors

```

### Manual Detect Head Initialization

```python
from models.yolo import Detect

# Configuration matching YOLOv5s

num_classes = 80
anchors = [
    [(10,13), (16,30), (33,23)],      # P3/8

    [(30,61), (62,45), (59,119)],     # P4/16

    [(116,90), (156,198), (373,326)]  # P5/32

]
channels = [256, 512, 1024]  # From C3 blocks in backbone

# Instantiate head

detect = Detect(nc=num_classes, anchors=anchors, ch=channels)

# Simulate backbone output

feature_maps = [
    torch.randn(1, 256, 80, 80),   # P3

    torch.randn(1, 512, 40, 40),   # P4

    torch.randn(1, 1024, 20, 20)   # P5

]

# Forward pass (training mode returns raw tensors)

outputs = detect(feature_maps)
print([o.shape for o in outputs])  # [1, 3, 80, 80, 85], [1, 3, 40, 40, 85], etc.

```

## Summary

- The **Detect head** in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py) serves as the final processing layer that converts backbone feature maps into bounding box predictions, objectness scores, and class probabilities.

- It employs **1×1 convolutions** to project features into the output space, followed by tensor reshaping and grid generation to map relative coordinates to absolute image positions.

- **Anchor-based decoding** scales width and height predictions using predefined anchor boxes, while sigmoid activations normalize center coordinates and confidences.

- The head operates in **dual modes**: returning raw tensors for loss computation during training, and concatenated, decoded predictions during inference or export.

- The **Segment** subclass extends Detect for instance segmentation by adding mask prototype coefficients, demonstrating the head's extensibility for multi-task learning.

## Frequently Asked Questions

### How does the Detect head differ from the backbone and neck in YOLOv5?

The **backbone** (CSPDarknet) extracts hierarchical visual features from input images, while the **neck** (PANet) fuses multi-scale features to create rich detection representations. The **Detect head** is the final stage that operates on these fused features to produce concrete predictions—bounding box coordinates, objectness scores, and class probabilities—using 1×1 convolutions and anchor-based decoding.

### Why does the Detect head use three different scales (P3, P4, P5)?

YOLOv5 employs **multi-scale detection** to handle objects of varying sizes effectively. P3 (stride 8) captures fine-grained details for small objects, P4 (stride 16) handles medium-sized objects, and P5 (stride 32) detects large objects with global context. The Detect head processes each scale independently using scale-specific anchor boxes, then concatenates the results during inference to produce a comprehensive detection output.

### Can the Detect head be modified for custom object detection tasks?

Yes, the Detect head is highly configurable through the model YAML files (e.g., [`yolov5s.yaml`](https://github.com/ultralytics/yolov5/blob/main/yolov5s.yaml)). You can adjust the **number of anchors** per scale, modify the **anchor box dimensions** to match your dataset's object aspect ratios, or change the **number of classes** (`nc`). For advanced modifications, you can subclass `Detect` in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py) to implement custom decoding logic, such as anchor-free detection or additional prediction heads for keypoint estimation.