How to Perform Object Detection with YOLOv5 on a Custom Dataset

To perform object detection with YOLOv5 on a custom dataset, organize your images and normalized bounding-box annotations in YOLO format, define a dataset YAML configuration file, execute train.py to fit a DetectionModel, and run detect.py with the saved weights to detect objects in new images.

YOLOv5 is a modular PyTorch-based object detector maintained by Ultralytics. Training the model on custom objects—whether medical imagery, retail products, or wildlife—requires organizing data in a specific structure and invoking the repository’s optimized training scripts. This guide explains the complete pipeline using the actual source code paths and functions implemented in the ultralytics/yolov5 repository.

Prepare Your Custom Dataset in YOLO Format

Directory Structure and Annotation Files

YOLOv5 expects the YOLO annotation format, where each image has a corresponding text file containing one line per object instance:


<class_id> <x_center> <y_center> <width> <height>

All coordinate values must be normalized between 0 and 1 relative to the image dimensions. Organize your files so that images and labels are paired by filename:

dataset/
├── images/
│   ├── train/
│   │   └── image001.jpg
│   └── val/
│       └── image002.jpg
└── labels/
    ├── train/
    │   └── image001.txt   # Contains: 0 0.5 0.5 0.3 0.4

    └── val/
        └── image002.txt

The validation function check_dataset in utils/general.py (lines 60–70) verifies that every image has a corresponding label file and that the YAML configuration is valid.

Dataset Configuration YAML

Create a YAML file (e.g., my_data.yaml) that describes the dataset paths, number of classes, and class names. The paths can be absolute or relative to the repository root:

train: /path/to/dataset/images/train
val: /path/to/dataset/images/val
nc: 3                     # number of classes

names: ['cat', 'dog', 'bird']

According to the source code in utils/general.py → check_dataset, this file is parsed at the start of training to resolve paths and validate that the directory structure matches the expected format.

Train YOLOv5 on Your Custom Dataset

Command-Line Training Execution

After installing dependencies (pip install -r requirements.txt), launch training by specifying your dataset YAML, a model architecture configuration, and hyperparameters:

python train.py \
  --data my_data.yaml \
  --cfg yolov5s.yaml \
  --weights '' \
  --epochs 100 \
  --batch-size 16 \
  --imgsz 640 \
  --project runs/train \
  --name my_experiment

Key parameters:

  • --data: Path to your dataset YAML file.
  • --cfg: Model architecture (e.g., yolov5s.yaml, yolov5m.yaml from the models/ directory).
  • --weights: Use '' to train from scratch, or provide a .pt file to fine-tune.

Internal Training Pipeline

The train() function in train.py (around lines 1000–1060) orchestrates the following steps:

  • Data Loading: create_dataloader in utils/dataloaders.py builds a PyTorch DataLoader that yields augmented images and target tensors.
  • Model Construction: DetectionModel in models/yolo.py parses the YAML configuration and initializes the backbone, neck, and Detect head layers.
  • Loss Computation: ComputeLoss in utils/loss.py calculates the classification, box, and objectness losses during the forward pass.
  • EMA and Checkpointing: An exponential moving average model (ModelEMA) is maintained, and checkpoints (last.pt, best.pt) are saved to the project directory.

Run Inference with the Trained Model

Executing detect.py

Once training completes, the best weights are saved to runs/train/my_experiment/weights/best.pt. Run inference on images, videos, or streams using detect.py:

python detect.py \
  --weights runs/train/my_experiment/weights/best.pt \
  --source /path/to/test/images \
  --conf-thres 0.25 \
  --iou-thres 0.45 \
  --save-txt \
  --save-csv \
  --project runs/detect \
  --name my_test

The script accepts multiple source types including single images, folders, webcams, or YouTube URLs via the LoadImages and LoadStreams classes in utils/dataloaders.py.

Inference Backend Architecture

Inside detect.py → run() (lines 84–110), the pipeline executes the following:

  1. Model Loading: DetectMultiBackend in models/common.py abstracts the model weights, supporting PyTorch, ONNX, TensorRT, and other formats.
  2. Forward Pass: The model processes batched images and returns raw predictions.
  3. NMS Filtering: non_max_suppression in utils/general.py filters overlapping bounding boxes using confidence and IoU thresholds.
  4. Output Generation: The Annotator class draws boxes on images, while optional flags write results to .txt files or a predictions.csv summary.

Complete End-to-End Workflow

The following bash commands demonstrate the full pipeline from data preparation to inference:


# 1. Prepare directory structure and YAML

mkdir -p data/custom/images/train data/custom/images/val

# Copy images and create corresponding *.txt label files

cat > data/custom.yaml <<EOF
train: data/custom/images/train
val: data/custom/images/val
nc: 2
names: ['apple', 'orange']
EOF

# 2. Train the model

python train.py --data data/custom.yaml --cfg yolov5s.yaml --epochs 150 --batch-size 16

# 3. Detect objects in new images

python detect.py \
  --weights runs/train/exp/weights/best.pt \
  --source data/custom/images/test \
  --save-txt \
  --save-csv

Results are written to runs/detect/exp/, including annotated images, individual label files in labels/, and a CSV summary if requested.

Summary

Frequently Asked Questions

What is the YOLO annotation format for custom datasets?

The YOLO format requires one text file per image with each line representing a single object as <class_id> <x_center> <y_center> <width> <height>, where all spatial values are normalized between 0 and 1. Class IDs are zero-indexed integers corresponding to the names list in your dataset YAML file.

How do I choose the right YOLOv5 model size (s, m, l, x)?

Select yolov5s.yaml for speed on edge devices, yolov5m.yaml for balanced accuracy and inference time, or yolov5l.yaml/yolov5x.yaml for maximum accuracy when GPU memory permits. The DetectionModel class in models/yolo.py scales the depth and width multipliers based on the configuration file you pass to train.py.

Can I fine-tune a pre-trained YOLOv5 model instead of training from scratch?

Yes. Instead of --weights '', provide a path to a pre-trained checkpoint such as --weights yolov5s.pt. The training script will load the weights into the DetectionModel, freeze backbone layers if configured, and resume optimization for your specific classes, significantly reducing convergence time compared to training from random initialization.

How does YOLOv5 handle overlapping detections during inference?

The non_max_suppression function in utils/general.py filters raw model outputs by first removing boxes below the confidence threshold, then iteratively selecting the highest-confidence box and removing overlapping candidates that exceed the IoU threshold (default 0.45). This process is executed within the inference loop in detect.py before visualization or file export.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →