How to Resume YOLOv5 Training from a Checkpoint

Add --resume to your train.py command to automatically restore the model, optimizer, and EMA state from last.pt and continue training from the next epoch.

YOLOv5 includes a robust checkpointing system that makes it easy to resume interrupted training sessions without losing progress. According to the ultralytics/yolov5 source code, the framework automatically saves training state to last.pt and best.pt files in runs/train/exp*/weights/ after every epoch. You can restart training from exactly where you left off using the built-in --resume flag, which handles both local and remote checkpoints seamlessly.

How the Resume Flag Works

In train.py (lines 576-585), the --resume argument is defined with parser.add_argument("--resume", ...) and accepts either a boolean or a file path. When supplied without a value, it defaults to True, triggering an automatic search for the most recent checkpoint. When given a specific path, the script loads that checkpoint directly.

Locating Checkpoints Automatically

If opt.resume is set to True, the script calls get_latest_run() from utils/general.py (lines 12-16) to scan the runs/ directory for the newest last*.pt file. If you provide a string path instead, the script verifies the file exists using check_file() before proceeding.

Loading and Restoring State

Once the checkpoint is identified, the script loads it using torch_load(weights, map_location="cpu") in train.py. This restores the model weights, optimizer state, Exponential Moving Average (EMA), and the saved epoch number. The smart_resume() function in utils/torch_utils.py (lines 94-113) then copies the saved optimizer and EMA states into the current objects and recalculates the start epoch. Training continues from the next epoch (ckpt["epoch"] + 1), with logging output indicating that the session has been resumed.

Common Scenarios for Resuming Training

Resume the Most Recent Run

To automatically find and resume from the latest last.pt in your runs/train/ directory:

python train.py --data data/coco.yaml --cfg yolov5s.yaml --weights yolov5s.pt \
    --batch-size 16 --epochs 100 --resume

The script locates the most recent runs/train/exp*/weights/last.pt and continues training without requiring a specific file path.

Resume From a Specific Checkpoint

To resume from a specific checkpoint file rather than the latest one:

python train.py --data data/coco.yaml --cfg yolov5s.yaml \
    --resume runs/train/exp15/weights/last.pt \
    --epochs 150

When you specify an explicit path, the script loads that checkpoint directly. Note that the --epochs argument specifies the total number of epochs, and the script will add any previously completed epochs to this count.

Fine-Tune for Additional Epochs

If you want to extend training for extra epochs beyond your original run, use the --resume flag with a new epoch count:

python train.py --weights yolov5s.pt --epochs 300 --resume

The smart_resume() function detects the previous epoch count from the checkpoint and adds it to your new --epochs value, effectively training for 300 + previous_epochs total iterations.

Resume With Remote Artifacts (W&B and Comet)

YOLOv5 supports resuming from remote experiment trackers like Weights & Biases (W&B) and Comet. The logger code in utils/loggers/wandb/wandb_utils.py and utils/loggers/comet/comet_utils.py downloads the latest checkpoint artifact before smart_resume() processes it locally:

python train.py --data data/coco.yaml --resume --wandb --project my-yolov5

This command fetches the most recent last.pt from your W&B project and continues training seamlessly.

Key Source Files and Functions

  • train.py (lines 576-585): Defines the --resume argument and orchestrates the loading process.
  • utils/general.py → get_latest_run() (lines 12-16): Scans the runs/ directory to locate the newest checkpoint when resuming automatically.
  • utils/torch_utils.py → smart_resume() (lines 94-113): Restores optimizer state, EMA, and adjusts epoch counts for fine-tuning scenarios.
  • utils/loggers/wandb/wandb_utils.py: Handles downloading remote checkpoints from W&B artifacts when resuming with --wandb.

Summary

  • YOLOv5 saves checkpoints to runs/train/exp*/weights/last.pt and best.pt after every epoch.
  • Use --resume alone to automatically find and resume the most recent training run.
  • Provide a specific path with --resume path/to/last.pt to resume from a particular checkpoint.
  • The smart_resume() function in utils/torch_utils.py restores the optimizer, EMA, and calculates the correct starting epoch.
  • When fine-tuning with --resume, the script adds previous epochs to your new --epochs count.
  • Remote checkpoint resuming is supported via W&B and Comet integrations.

Frequently Asked Questions

What file does YOLOv5 use to resume training?

YOLOv5 uses last.pt (the most recent checkpoint) by default, located in runs/train/exp*/weights/. You can also resume from best.pt or any other saved checkpoint by providing the explicit file path to --resume.

Will resuming training overwrite my previous best model?

No, resuming training preserves your existing best.pt file. The training loop continues from the next epoch, and the best model is only updated if the resumed training achieves a higher fitness score than the previously saved best.

Can I resume training on a different machine or GPU?

Yes, you can resume training on different hardware. The checkpoint stores the model state dict and optimizer state, which are loaded with map_location="cpu" before being moved to the appropriate device. However, ensure your environment has the same YOLOv5 version and dependencies to avoid compatibility issues.

How do I resume training from a Weights & Biases (W&B) artifact?

Use the --resume flag combined with --wandb and your project name. The W&B logger in utils/loggers/wandb/wandb_utils.py automatically downloads the latest last.pt artifact from your W&B project before the training script restores the state via smart_resume().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →