How to Perform Object Detection with YOLOv5 on a Custom Dataset
To perform object detection with YOLOv5 on a custom dataset, organize your images and normalized bounding-box annotations in YOLO format, define a dataset YAML configuration file, execute train.py to fit a DetectionModel, and run detect.py with the saved weights to detect objects in new images.
YOLOv5 is a modular PyTorch-based object detector maintained by Ultralytics. Training the model on custom objects—whether medical imagery, retail products, or wildlife—requires organizing data in a specific structure and invoking the repository’s optimized training scripts. This guide explains the complete pipeline using the actual source code paths and functions implemented in the ultralytics/yolov5 repository.
Prepare Your Custom Dataset in YOLO Format
Directory Structure and Annotation Files
YOLOv5 expects the YOLO annotation format, where each image has a corresponding text file containing one line per object instance:
<class_id> <x_center> <y_center> <width> <height>
All coordinate values must be normalized between 0 and 1 relative to the image dimensions. Organize your files so that images and labels are paired by filename:
dataset/
├── images/
│ ├── train/
│ │ └── image001.jpg
│ └── val/
│ └── image002.jpg
└── labels/
├── train/
│ └── image001.txt # Contains: 0 0.5 0.5 0.3 0.4
└── val/
└── image002.txt
The validation function check_dataset in utils/general.py (lines 60–70) verifies that every image has a corresponding label file and that the YAML configuration is valid.
Dataset Configuration YAML
Create a YAML file (e.g., my_data.yaml) that describes the dataset paths, number of classes, and class names. The paths can be absolute or relative to the repository root:
train: /path/to/dataset/images/train
val: /path/to/dataset/images/val
nc: 3 # number of classes
names: ['cat', 'dog', 'bird']
According to the source code in utils/general.py → check_dataset, this file is parsed at the start of training to resolve paths and validate that the directory structure matches the expected format.
Train YOLOv5 on Your Custom Dataset
Command-Line Training Execution
After installing dependencies (pip install -r requirements.txt), launch training by specifying your dataset YAML, a model architecture configuration, and hyperparameters:
python train.py \
--data my_data.yaml \
--cfg yolov5s.yaml \
--weights '' \
--epochs 100 \
--batch-size 16 \
--imgsz 640 \
--project runs/train \
--name my_experiment
Key parameters:
--data: Path to your dataset YAML file.--cfg: Model architecture (e.g.,yolov5s.yaml,yolov5m.yamlfrom themodels/directory).--weights: Use''to train from scratch, or provide a.ptfile to fine-tune.
Internal Training Pipeline
The train() function in train.py (around lines 1000–1060) orchestrates the following steps:
- Data Loading:
create_dataloaderinutils/dataloaders.pybuilds a PyTorchDataLoaderthat yields augmented images and target tensors. - Model Construction:
DetectionModelinmodels/yolo.pyparses the YAML configuration and initializes the backbone, neck, andDetecthead layers. - Loss Computation:
ComputeLossinutils/loss.pycalculates the classification, box, and objectness losses during the forward pass. - EMA and Checkpointing: An exponential moving average model (
ModelEMA) is maintained, and checkpoints (last.pt,best.pt) are saved to the project directory.
Run Inference with the Trained Model
Executing detect.py
Once training completes, the best weights are saved to runs/train/my_experiment/weights/best.pt. Run inference on images, videos, or streams using detect.py:
python detect.py \
--weights runs/train/my_experiment/weights/best.pt \
--source /path/to/test/images \
--conf-thres 0.25 \
--iou-thres 0.45 \
--save-txt \
--save-csv \
--project runs/detect \
--name my_test
The script accepts multiple source types including single images, folders, webcams, or YouTube URLs via the LoadImages and LoadStreams classes in utils/dataloaders.py.
Inference Backend Architecture
Inside detect.py → run() (lines 84–110), the pipeline executes the following:
- Model Loading:
DetectMultiBackendinmodels/common.pyabstracts the model weights, supporting PyTorch, ONNX, TensorRT, and other formats. - Forward Pass: The model processes batched images and returns raw predictions.
- NMS Filtering:
non_max_suppressioninutils/general.pyfilters overlapping bounding boxes using confidence and IoU thresholds. - Output Generation: The
Annotatorclass draws boxes on images, while optional flags write results to.txtfiles or apredictions.csvsummary.
Complete End-to-End Workflow
The following bash commands demonstrate the full pipeline from data preparation to inference:
# 1. Prepare directory structure and YAML
mkdir -p data/custom/images/train data/custom/images/val
# Copy images and create corresponding *.txt label files
cat > data/custom.yaml <<EOF
train: data/custom/images/train
val: data/custom/images/val
nc: 2
names: ['apple', 'orange']
EOF
# 2. Train the model
python train.py --data data/custom.yaml --cfg yolov5s.yaml --epochs 150 --batch-size 16
# 3. Detect objects in new images
python detect.py \
--weights runs/train/exp/weights/best.pt \
--source data/custom/images/test \
--save-txt \
--save-csv
Results are written to runs/detect/exp/, including annotated images, individual label files in labels/, and a CSV summary if requested.
Summary
- Data Format: YOLOv5 requires paired image and
.txtfiles with normalized center coordinates, plus a dataset YAML defining paths and class names validated byutils/general.py→check_dataset. - Training: Execute
train.pywith your data YAML and model config; the script usesDetectionModel(models/yolo.py),ComputeLoss(utils/loss.py), and saves EMA-smoothed checkpoints. - Inference: Use
detect.pywith trained weights;DetectMultiBackend(models/common.py) handles model loading whilenon_max_suppression(utils/general.py) refines raw predictions. - Key Files:
train.pyfor training loops,detect.pyfor inference,models/yolo.pyfor model architecture, andutils/dataloaders.pyfor data ingestion.
Frequently Asked Questions
What is the YOLO annotation format for custom datasets?
The YOLO format requires one text file per image with each line representing a single object as <class_id> <x_center> <y_center> <width> <height>, where all spatial values are normalized between 0 and 1. Class IDs are zero-indexed integers corresponding to the names list in your dataset YAML file.
How do I choose the right YOLOv5 model size (s, m, l, x)?
Select yolov5s.yaml for speed on edge devices, yolov5m.yaml for balanced accuracy and inference time, or yolov5l.yaml/yolov5x.yaml for maximum accuracy when GPU memory permits. The DetectionModel class in models/yolo.py scales the depth and width multipliers based on the configuration file you pass to train.py.
Can I fine-tune a pre-trained YOLOv5 model instead of training from scratch?
Yes. Instead of --weights '', provide a path to a pre-trained checkpoint such as --weights yolov5s.pt. The training script will load the weights into the DetectionModel, freeze backbone layers if configured, and resume optimization for your specific classes, significantly reducing convergence time compared to training from random initialization.
How does YOLOv5 handle overlapping detections during inference?
The non_max_suppression function in utils/general.py filters raw model outputs by first removing boxes below the confidence threshold, then iteratively selecting the highest-confidence box and removing overlapping candidates that exceed the IoU threshold (default 0.45). This process is executed within the inference loop in detect.py before visualization or file export.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →