How to Analyze YOLOv5 Training Results and Performance Metrics

YOLOv5 automatically logs comprehensive training metrics to results.csv and generates visual artifacts including confusion matrices, PR curves, and loss plots to evaluate object detection performance.

To effectively analyze YOLOv5 training results, you need to understand the three core components in the ultralytics/yolov5 repository: the training loop in train.py, the metric computation engine in utils/metrics.py, and the visualization suite in utils/plots.py. This guide explains how to extract actionable insights from the numerical data and visual outputs produced during and after training.

Where Training Metrics Are Stored

The train() function in train.py orchestrates the entire training process. At the end of every epoch, it invokes the validation routine via val.run() (see train.py lines 53–55). This call returns a tuple containing precision, recall, mean Average Precision (mAP), loss components, and timing breakdowns.

The training script aggregates these values and writes them to runs/train/<exp>/results.csv, where <exp> represents your experiment folder. This CSV serves as the primary source for quantitative analysis, while the save_dir directory stores visual artifacts like PNG plots and sample predictions.

Decoding the results.csv File Structure

The results.csv file contains epoch-wise records of both training and validation metrics. According to the implementation in utils/plots.py (lines 30–38), the CSV includes the following columns:

  • epoch – The training epoch index
  • train/box_loss, train/obj_loss, train/cls_loss – Bounding box regression, objectness, and classification training losses
  • val/precision, val/recall – Global precision and recall across all classes
  • val/mAP_0.5 – mAP calculated at IoU threshold 0.5
  • val/mAP_0.5:0.95 – mAP averaged across IoU thresholds from 0.5 to 0.95 (COCO standard)
  • val/box_loss, val/obj_loss, val/cls_loss – Corresponding validation loss components
  • val/fitness – Weighted fitness score used for checkpoint selection
  • speed/preprocess, speed/inference, speed/NMS – Milliseconds per image for each processing stage (see train.py lines 18–20)

Per-Class Statistics and Validation Logic

Deep performance analysis requires looking beyond global aggregates. The validation routine in val.py constructs a ConfusionMatrix object (defined in utils/metrics.py, lines 32–34) and calls ap_per_class() (lines 32–38) to compute class-specific metrics.

The ap_per_class() function returns:

  • tp and fp – True positives and false positives per class
  • p and r – Precision and recall curves across confidence thresholds
  • f1 – F1-score curves
  • ap – Average precision per class, with the first column representing AP@0.5

When verbose=True or when the dataset contains 50 or fewer classes, val.py (lines 13–15) prints these per-class statistics to the console, enabling detailed bottleneck identification.

Visual Artifacts Generated by plots.py

The repository automatically generates several diagnostic images in your run directory to facilitate qualitative analysis:

  • confusion_matrix.png – Normalized confusion matrix showing class-wise prediction accuracy (see ConfusionMatrix.plot() in utils/metrics.py, lines 99–108)
  • PR_curve.png – Precision-Recall curves for each class (see plot_pr_curve() in utils/plots.py, lines 45–66)
  • P_curve.png, R_curve.png, F1_curve.png – Metric-versus-confidence curves generated by plot_mc_curve()
  • val_batch*_labels.jpg and val_batch*_pred.jpg – Side-by-side visualizations of ground truth labels and model predictions on validation batches (see plot_images() in utils/plots.py, lines 51–84)
  • results.png – Consolidated plot showing loss convergence and metric trends over epochs (see plot_results() in utils/plots.py, lines 30–57)

Understanding the Fitness Score and Early Stopping

The fitness score determines which checkpoint is saved as best.pt. Implemented in utils/metrics.fitness (lines 18–22), it calculates a weighted sum of [precision, recall, mAP@0.5, mAP@0.5:0.95] using weights [0.0, 0.0, 0.1, 0.9]. This means the score prioritizes mAP@0.5:0.95 (strict localization) over raw precision and recall.

The training loop tracks best_fitness and triggers early stopping via the EarlyStopping callback if no improvement occurs for 100 epochs (see train.py lines 58–59).

How to Interpret Key Performance Indicators

Metric Interpretation Target Range
Precision Fraction of predicted boxes that are correct; high values indicate few false positives ≥ 0.80
Recall Fraction of ground-truth boxes detected; high values indicate few missed objects ≥ 0.80
mAP@0.5 Average precision at IoU = 0.5; baseline metric for detection tasks 0.50 – 0.70
mAP@0.5:0.95 Strict metric averaging AP across multiple IoU thresholds; reflects localization quality 0.40 – 0.60
Fitness Weighted combination (0.1 × mAP@0.5 + 0.9 × mAP@0.5:0.95) used for model selection Higher is better; look for plateau
Loss Components Declining box_loss, obj_loss, and cls_loss indicate stable training Near zero at convergence
Speed Metrics Preprocessing, inference, and NMS latency per image in milliseconds ≤ 5 ms/img for GPU deployment

Practical Code Examples for Analysis

Loading and Plotting Custom Metric Curves

Use pandas and matplotlib to analyze trends from the CSV output:

import pandas as pd
import matplotlib.pyplot as plt
from pathlib import Path

# Locate your training run directory

run_dir = Path("runs/train/exp")
csv_path = run_dir / "results.csv"

df = pd.read_csv(csv_path)

# Visualize box loss convergence

plt.figure(figsize=(10, 6))
plt.plot(df["epoch"], df["train/box_loss"], label="Training")
plt.plot(df["epoch"], df["val/box_loss"], label="Validation")
plt.xlabel("Epoch")
plt.ylabel("Box Loss")
plt.legend()
plt.title("Bounding Box Loss Convergence")
plt.grid(True)
plt.show()

Extracting Per-Class AP Scores

Access detailed class performance programmatically using the validation API:

import torch
from utils.metrics import ap_per_class
from val import run as validate

# Run validation on your trained weights

results, maps, _ = validate(
    data="path/to/data.yaml",
    weights=Path("runs/train/exp/weights/best.pt"),
    batch_size=32,
    imgsz=640,
    device="0"
)

# Extract per-class AP@0.5

tp, fp, p, r, f1, ap, classes = ap_per_class(
    *results[:6], 
    plot=False, 
    save_dir=None, 
    names={i: name for i, name in enumerate(model.names)}
)

for i, class_idx in enumerate(classes):
    ap50 = ap[i, 0]  # First column is AP@0.5

    print(f"Class {model.names[class_idx]}: AP@0.5 = {ap50:.3f}")

Programmatically Accessing the Confusion Matrix

Load the generated confusion matrix image for custom reporting:

import matplotlib.pyplot as plt
from pathlib import Path

cm_path = Path("runs/train/exp/confusion_matrix.png")

plt.figure(figsize=(8, 8))
img = plt.imread(cm_path)
plt.imshow(img)
plt.axis("off")
plt.title("Validation Confusion Matrix")
plt.tight_layout()
plt.show()

Running Hyperparameter Evolution Analysis

Generate and visualize the fitness landscape across hyperparameter combinations:

python train.py --data coco.yaml --weights yolov5s.pt --evolve 50

Then analyze the results:

from utils.plots import plot_evolve
from pathlib import Path

evolve_csv = Path("runs/train/exp/evolve.csv")
plot_evolve(evolve_csv)  # Saves evolve.png with per-parameter fitness clouds

Summary

  • Data Location: YOLOv5 writes epoch-wise metrics to runs/train/<exp>/results.csv and saves visual diagnostics to the same directory.
  • Core Functions: utils/metrics.py contains ap_per_class() and ConfusionMatrix for statistical computation, while utils/plots.py handles visualization via plot_results() and plot_pr_curve().
  • Fitness Metric: The fitness() function weights mAP@0.5:0.95 at 90% and mAP@0.5 at 10%, driving the EarlyStopping logic in train.py.
  • File References: Key analysis targets include confusion_matrix.png for class confusion, PR_curve.png for precision-recall tradeoffs, and results.csv for loss convergence tracking.
  • Actionable Analysis: Load CSV data with pandas, extract per-class AP via the validation API, and use plot_evolve() to optimize hyperparameters based on fitness scores.

Frequently Asked Questions

How do I know if my YOLOv5 model is overfitting?

Compare the train/*_loss and val/*_loss columns in results.csv. If training loss continues to decrease while validation loss increases after epoch 50, your model is likely overfitting. Check val/mAP_0.5:0.95 for performance degradation; if it drops while training loss improves, apply augmentation or reduce epochs.

What is the difference between mAP@0.5 and mAP@0.5:0.95?

mAP@0.5 calculates average precision using a single IoU threshold of 0.5, rewarding rough localization. mAP@0.5:0.95 averages AP across ten IoU thresholds from 0.5 to 0.95 in 0.05 steps, strictly penalizing poor bounding box localization. According to utils/metrics.py, the fitness function weights mAP@0.5:0.95 at 90% because it better reflects real-world detection quality.

How can I extract per-class precision and recall values?

Set verbose=True when calling val.run() or use the ap_per_class() function directly. The function returns precision (p) and recall (r) arrays shaped [n_classes, n_iou_thresholds]. For class-specific analysis at IoU=0.5, index the first column: p[:, 0] for precision and r[:, 0] for recall values.

Why is my fitness score increasing but mAP@0.5 is decreasing?

The fitness function in utils/metrics.py uses weights [0.0, 0.0, 0.1, 0.9], meaning it ignores precision and recall entirely and weights mAP@0.5:0.95 nine times higher than mAP@0.5. Your mAP@0.5:0.95 is likely improving due to better localization (higher IoU), even if mAP@0.5 drops slightly due to classification threshold changes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →