How to Analyze YOLOv5 Training Results and Performance Metrics
YOLOv5 automatically logs comprehensive training metrics to results.csv and generates visual artifacts including confusion matrices, PR curves, and loss plots to evaluate object detection performance.
To effectively analyze YOLOv5 training results, you need to understand the three core components in the ultralytics/yolov5 repository: the training loop in train.py, the metric computation engine in utils/metrics.py, and the visualization suite in utils/plots.py. This guide explains how to extract actionable insights from the numerical data and visual outputs produced during and after training.
Where Training Metrics Are Stored
The train() function in train.py orchestrates the entire training process. At the end of every epoch, it invokes the validation routine via val.run() (see train.py lines 53–55). This call returns a tuple containing precision, recall, mean Average Precision (mAP), loss components, and timing breakdowns.
The training script aggregates these values and writes them to runs/train/<exp>/results.csv, where <exp> represents your experiment folder. This CSV serves as the primary source for quantitative analysis, while the save_dir directory stores visual artifacts like PNG plots and sample predictions.
Decoding the results.csv File Structure
The results.csv file contains epoch-wise records of both training and validation metrics. According to the implementation in utils/plots.py (lines 30–38), the CSV includes the following columns:
epoch– The training epoch indextrain/box_loss,train/obj_loss,train/cls_loss– Bounding box regression, objectness, and classification training lossesval/precision,val/recall– Global precision and recall across all classesval/mAP_0.5– mAP calculated at IoU threshold 0.5val/mAP_0.5:0.95– mAP averaged across IoU thresholds from 0.5 to 0.95 (COCO standard)val/box_loss,val/obj_loss,val/cls_loss– Corresponding validation loss componentsval/fitness– Weighted fitness score used for checkpoint selectionspeed/preprocess,speed/inference,speed/NMS– Milliseconds per image for each processing stage (seetrain.pylines 18–20)
Per-Class Statistics and Validation Logic
Deep performance analysis requires looking beyond global aggregates. The validation routine in val.py constructs a ConfusionMatrix object (defined in utils/metrics.py, lines 32–34) and calls ap_per_class() (lines 32–38) to compute class-specific metrics.
The ap_per_class() function returns:
tpandfp– True positives and false positives per classpandr– Precision and recall curves across confidence thresholdsf1– F1-score curvesap– Average precision per class, with the first column representing AP@0.5
When verbose=True or when the dataset contains 50 or fewer classes, val.py (lines 13–15) prints these per-class statistics to the console, enabling detailed bottleneck identification.
Visual Artifacts Generated by plots.py
The repository automatically generates several diagnostic images in your run directory to facilitate qualitative analysis:
confusion_matrix.png– Normalized confusion matrix showing class-wise prediction accuracy (seeConfusionMatrix.plot()inutils/metrics.py, lines 99–108)PR_curve.png– Precision-Recall curves for each class (seeplot_pr_curve()inutils/plots.py, lines 45–66)P_curve.png,R_curve.png,F1_curve.png– Metric-versus-confidence curves generated byplot_mc_curve()val_batch*_labels.jpgandval_batch*_pred.jpg– Side-by-side visualizations of ground truth labels and model predictions on validation batches (seeplot_images()inutils/plots.py, lines 51–84)results.png– Consolidated plot showing loss convergence and metric trends over epochs (seeplot_results()inutils/plots.py, lines 30–57)
Understanding the Fitness Score and Early Stopping
The fitness score determines which checkpoint is saved as best.pt. Implemented in utils/metrics.fitness (lines 18–22), it calculates a weighted sum of [precision, recall, mAP@0.5, mAP@0.5:0.95] using weights [0.0, 0.0, 0.1, 0.9]. This means the score prioritizes mAP@0.5:0.95 (strict localization) over raw precision and recall.
The training loop tracks best_fitness and triggers early stopping via the EarlyStopping callback if no improvement occurs for 100 epochs (see train.py lines 58–59).
How to Interpret Key Performance Indicators
| Metric | Interpretation | Target Range |
|---|---|---|
| Precision | Fraction of predicted boxes that are correct; high values indicate few false positives | ≥ 0.80 |
| Recall | Fraction of ground-truth boxes detected; high values indicate few missed objects | ≥ 0.80 |
| mAP@0.5 | Average precision at IoU = 0.5; baseline metric for detection tasks | 0.50 – 0.70 |
| mAP@0.5:0.95 | Strict metric averaging AP across multiple IoU thresholds; reflects localization quality | 0.40 – 0.60 |
| Fitness | Weighted combination (0.1 × mAP@0.5 + 0.9 × mAP@0.5:0.95) used for model selection | Higher is better; look for plateau |
| Loss Components | Declining box_loss, obj_loss, and cls_loss indicate stable training |
Near zero at convergence |
| Speed Metrics | Preprocessing, inference, and NMS latency per image in milliseconds | ≤ 5 ms/img for GPU deployment |
Practical Code Examples for Analysis
Loading and Plotting Custom Metric Curves
Use pandas and matplotlib to analyze trends from the CSV output:
import pandas as pd
import matplotlib.pyplot as plt
from pathlib import Path
# Locate your training run directory
run_dir = Path("runs/train/exp")
csv_path = run_dir / "results.csv"
df = pd.read_csv(csv_path)
# Visualize box loss convergence
plt.figure(figsize=(10, 6))
plt.plot(df["epoch"], df["train/box_loss"], label="Training")
plt.plot(df["epoch"], df["val/box_loss"], label="Validation")
plt.xlabel("Epoch")
plt.ylabel("Box Loss")
plt.legend()
plt.title("Bounding Box Loss Convergence")
plt.grid(True)
plt.show()
Extracting Per-Class AP Scores
Access detailed class performance programmatically using the validation API:
import torch
from utils.metrics import ap_per_class
from val import run as validate
# Run validation on your trained weights
results, maps, _ = validate(
data="path/to/data.yaml",
weights=Path("runs/train/exp/weights/best.pt"),
batch_size=32,
imgsz=640,
device="0"
)
# Extract per-class AP@0.5
tp, fp, p, r, f1, ap, classes = ap_per_class(
*results[:6],
plot=False,
save_dir=None,
names={i: name for i, name in enumerate(model.names)}
)
for i, class_idx in enumerate(classes):
ap50 = ap[i, 0] # First column is AP@0.5
print(f"Class {model.names[class_idx]}: AP@0.5 = {ap50:.3f}")
Programmatically Accessing the Confusion Matrix
Load the generated confusion matrix image for custom reporting:
import matplotlib.pyplot as plt
from pathlib import Path
cm_path = Path("runs/train/exp/confusion_matrix.png")
plt.figure(figsize=(8, 8))
img = plt.imread(cm_path)
plt.imshow(img)
plt.axis("off")
plt.title("Validation Confusion Matrix")
plt.tight_layout()
plt.show()
Running Hyperparameter Evolution Analysis
Generate and visualize the fitness landscape across hyperparameter combinations:
python train.py --data coco.yaml --weights yolov5s.pt --evolve 50
Then analyze the results:
from utils.plots import plot_evolve
from pathlib import Path
evolve_csv = Path("runs/train/exp/evolve.csv")
plot_evolve(evolve_csv) # Saves evolve.png with per-parameter fitness clouds
Summary
- Data Location: YOLOv5 writes epoch-wise metrics to
runs/train/<exp>/results.csvand saves visual diagnostics to the same directory. - Core Functions:
utils/metrics.pycontainsap_per_class()andConfusionMatrixfor statistical computation, whileutils/plots.pyhandles visualization viaplot_results()andplot_pr_curve(). - Fitness Metric: The
fitness()function weights mAP@0.5:0.95 at 90% and mAP@0.5 at 10%, driving theEarlyStoppinglogic intrain.py. - File References: Key analysis targets include
confusion_matrix.pngfor class confusion,PR_curve.pngfor precision-recall tradeoffs, andresults.csvfor loss convergence tracking. - Actionable Analysis: Load CSV data with
pandas, extract per-class AP via the validation API, and useplot_evolve()to optimize hyperparameters based on fitness scores.
Frequently Asked Questions
How do I know if my YOLOv5 model is overfitting?
Compare the train/*_loss and val/*_loss columns in results.csv. If training loss continues to decrease while validation loss increases after epoch 50, your model is likely overfitting. Check val/mAP_0.5:0.95 for performance degradation; if it drops while training loss improves, apply augmentation or reduce epochs.
What is the difference between mAP@0.5 and mAP@0.5:0.95?
mAP@0.5 calculates average precision using a single IoU threshold of 0.5, rewarding rough localization. mAP@0.5:0.95 averages AP across ten IoU thresholds from 0.5 to 0.95 in 0.05 steps, strictly penalizing poor bounding box localization. According to utils/metrics.py, the fitness function weights mAP@0.5:0.95 at 90% because it better reflects real-world detection quality.
How can I extract per-class precision and recall values?
Set verbose=True when calling val.run() or use the ap_per_class() function directly. The function returns precision (p) and recall (r) arrays shaped [n_classes, n_iou_thresholds]. For class-specific analysis at IoU=0.5, index the first column: p[:, 0] for precision and r[:, 0] for recall values.
Why is my fitness score increasing but mAP@0.5 is decreasing?
The fitness function in utils/metrics.py uses weights [0.0, 0.0, 0.1, 0.9], meaning it ignores precision and recall entirely and weights mAP@0.5:0.95 nine times higher than mAP@0.5. Your mAP@0.5:0.95 is likely improving due to better localization (higher IoU), even if mAP@0.5 drops slightly due to classification threshold changes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →