# How to Perform Transfer Learning with YOLOv5 Using Pre-trained Weights

> Learn transfer learning with YOLOv5 using pre-trained weights. Load checkpoints and fine-tune detection heads on custom datasets for faster, more accurate object detection.

- Repository: [Ultralytics/yolov5](https://github.com/ultralytics/yolov5)
- Tags: tutorial
- Published: 2026-03-06

---

**You can perform transfer learning with YOLOv5 by loading a pre-trained checkpoint via the `--weights` argument and optionally freezing backbone layers with `--freeze N` to fine-tune the detection head on your custom dataset.**

The **ultralytics/yolov5** repository provides built-in support for transfer learning out of the box. By leveraging pre-trained weights from the COCO dataset, you can significantly reduce training time and improve detection accuracy on custom datasets with limited labeled data.

## Understanding YOLOv5 Transfer Learning Architecture

Transfer learning in YOLOv5 relies on three core mechanisms: checkpoint loading, layer freezing, and the CSP backbone architecture.

### Pre-trained Checkpoint Loading

The training script ([`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py)) handles weight initialization through the `--weights` argument. When you specify a `.pt` file such as `yolov5s.pt`, the script invokes `attempt_load()` in [`models/common.py`](https://github.com/ultralytics/yolov5/blob/main/models/common.py) to load both the model architecture and learned parameters. This occurs around line 120 in [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py), where the model is instantiated before training begins.

### Freezing Backbone Layers

The **freeze** mechanism allows you to lock specific layers during training. In [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) around line 224, the `model.freeze()` method iterates through the model's modules and sets `requires_grad=False` for the first *N* layers. This excludes frozen parameters from optimizer updates, preserving generic visual features while allowing the detection head to adapt to your specific classes.

### CSP Bottleneck Blocks

The backbone consists of **CSP (Cross-Stage-Partial)** modules defined in [`models/common.py`](https://github.com/ultralytics/yolov5/blob/main/models/common.py) (specifically the `BottleneckCSP` class around line 182). These blocks balance accuracy and speed. When freezing layers, you typically target these CSP stages in the backbone while leaving the detection head in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py) trainable.

## Step-by-Step Transfer Learning Workflow

### Dataset Preparation

Create a YAML configuration file that defines your custom dataset structure. This file specifies paths to training and validation images, number of classes, and class names.

```yaml

# data/custom.yaml

train: /path/to/train/images
val: /path/to/val/images
nc: 3  # number of classes

names: ['cat', 'dog', 'rabbit']

```

### Selecting Pre-trained Weights

Choose an appropriate checkpoint based on your computational constraints:
- **yolov5n.pt**: Nano model for edge devices
- **yolov5s.pt**: Small model for faster training
- **yolov5m/l/x.pt**: Medium, large, or extra-large for higher accuracy

The weights file contains both the backbone features learned from COCO and the detection head architecture.

### Configuring Layer Freezing

Use the `--freeze` argument to specify how many layers to freeze. For example, `--freeze 10` locks the first ten layers of the CSP backbone, which typically represents approximately 70% of the backbone parameters. This preserves low-level features (edges, textures) while allowing the model to learn high-level patterns specific to your dataset.

## CLI Command Examples

Execute transfer learning directly from the command line using [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py). The following example freezes ten backbone layers and fine-tunes for 50 epochs:

```bash
python train.py \
  --img 640 \
  --batch 16 \
  --epochs 50 \
  --data data/custom.yaml \
  --weights yolov5s.pt \
  --freeze 10 \
  --project runs/train \
  --name custom_transfer

```

Key parameters explained:
- `--img 640`: Input resolution must match the model stride requirements
- `--batch 16`: Adjust based on available GPU memory
- `--freeze 10`: Freezes the first ten CSP backbone layers (set to `0` to unfreeze all)
- `--weights yolov5s.pt`: Loads pre-trained COCO weights

During initialization, the console displays: `freeze: 10 layers (70.3%)`, confirming that gradient computation is disabled for those parameters.

## Programmatic Implementation

For custom training pipelines, you can load checkpoints and freeze layers programmatically using the `DetectMultiBackend` class:

```python
import torch
from models.common import DetectMultiBackend

# Load pretrained model

model = DetectMultiBackend('yolov5s.pt', device='cpu')

# Freeze the first N backbone layers

freeze_n = 10
for i, (name, param) in enumerate(model.model.named_parameters()):
    if i < freeze_n:
        param.requires_grad = False

# Verify frozen status

frozen = sum(p.requires_grad == False for p in model.model.parameters())
print(f'Frozen layers: {frozen} / {len(list(model.model.parameters()))}')

```

This approach mirrors the internal logic of [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) but allows integration into existing Python workflows. The `DetectMultiBackend` class handles loading for PyTorch, ONNX, and TensorRT backends.

## Hyper-parameter Considerations

YOLOv5 provides optimized hyper-parameter configurations in the `data/hyps/` directory. For transfer learning with smaller models like `yolov5s`, use [`hyp.scratch-low.yaml`](https://github.com/ultralytics/yolov5/blob/main/hyp.scratch-low.yaml), which contains conservative learning rates suitable for fine-tuning. Larger models may benefit from [`hyp.scratch-high.yaml`](https://github.com/ultralytics/yolov5/blob/main/hyp.scratch-high.yaml).

## Summary

- **Transfer learning with YOLOv5** requires only the `--weights` argument to load pre-trained checkpoints from the COCO dataset
- **Freeze backbone layers** using `--freeze N` to preserve learned visual features while training the detection head on your custom classes
- The **CSP backbone** in [`models/common.py`](https://github.com/ultralytics/yolov5/blob/main/models/common.py) provides the architecture for the layers you typically freeze
- **Training scripts** in [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) automatically handle layer freezing via `requires_grad=False` and adjust the optimizer accordingly
- **Fine-tuning epochs** can be significantly fewer (30-50) compared to training from scratch (300+)

## Frequently Asked Questions

### How many layers should I freeze when doing transfer learning with YOLOv5?

Freeze 10 layers for the small model (`yolov5s`) to preserve the CSP backbone while training the detection head. For larger models like `yolov5x`, you may freeze up to 20-30 layers. If your dataset is very small (< 1000 images), freezing more layers prevents overfitting. If your dataset is large and diverse, set `--freeze 0` to fine-tune the entire network.

### Can I use custom pre-trained weights instead of the official YOLOv5 checkpoints?

Yes. Simply point the `--weights` argument to your custom `.pt` file path. The [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) script in the ultralytics/yolov5 repository uses `attempt_load()` to handle any valid PyTorch state dictionary, allowing you to start from previously fine-tuned models or weights trained on different datasets.

### Why does YOLOv5 freeze layers by default during transfer learning?

Freezing the CSP backbone preserves generic features learned from the COCO dataset (edges, textures, shapes) that transfer well to most visual tasks. This reduces training time, lowers GPU memory requirements, and prevents overfitting on small custom datasets. The detection head in [`models/yolo.py`](https://github.com/ultralytics/yolov5/blob/main/models/yolo.py) remains trainable to map these frozen features to your specific class labels.

### What hyper-parameters should I adjust for transfer learning?

Reduce the learning rate compared to training from scratch. The default optimizer (Adam) in [`train.py`](https://github.com/ultralytics/yolov5/blob/main/train.py) automatically handles frozen parameters by skipping updates where `requires_grad=False`. Use [`data/hyps/hyp.scratch-low.yaml`](https://github.com/ultralytics/yolov5/blob/main/data/hyps/hyp.scratch-low.yaml) for conservative fine-tuning, and reduce the number of epochs to 30-50 instead of 300, as the model converges faster when starting from pre-trained weights.