YOLOv5 Model Sizes (n, s, m, l, x): Complete Guide to Speed and Accuracy Trade-offs
YOLOv5 provides five standardized model variants—nano (n), small (s), medium (m), large (l), and extra-large (x)—that scale network depth and width through YAML configuration files to optimize the balance between inference speed, parameter count, and mean Average Precision (mAP).
The Ultralytics YOLOv5 repository implements a unified architecture across all model sizes, differing only in scaling multipliers defined within individual YAML files. This design allows developers to deploy the same detection pipeline on hardware ranging from Raspberry Pi devices to multi-GPU servers by simply switching the model configuration.
YOLOv5 Model Variants and Performance Specifications
Each YOLOv5 size applies specific depth and width multipliers to the base v6.0 architecture, directly impacting the parameter count and computational complexity:
| Model | depth_multiple | width_multiple | Parameters (M) | FLOPs @ 640 | mAPval 50-95 | Inference Speed (V100 b1) |
|---|---|---|---|---|---|---|
| YOLOv5n | 0.33 | 0.25 | 1.9 | 4.5 | 28.0 | 6.3 ms |
| YOLOv5s | 0.33 | 0.50 | 7.2 | 16.5 | 37.4 | 6.4 ms |
| YOLOv5m | 0.67 | 0.75 | 21.2 | 49.0 | 45.4 | 8.2 ms |
| YOLOv5l | 1.00 | 1.00 | 46.5 | 109.1 | 49.0 | 10.1 ms |
| YOLOv5x | 1.33 | 1.25 | 86.7 | 205.7 | 50.7 | 12.1 ms |
Data sourced from the repository's performance table in the README and model definition files.
Architecture Scaling Mechanics
The scaling mechanism operates through two key parameters in each YAML configuration:
- depth_multiple: Scales the number of layers in the backbone and head (e.g., C3 block repetitions)
- width_multiple: Scales the number of channels in each layer (e.g., convolution filter counts)
For example, in models/yolov5n.yaml, the values depth_multiple: 0.33 and width_multiple: 0.25 create the smallest variant by using one-third the layers and one-quarter the channels of the baseline architecture defined in models/yolov5l.yaml.
Selecting the Right YOLOv5 Model Size for Your Use Case
Edge Devices and Embedded Systems (YOLOv5n)
YOLOv5n (nano) is optimized for Raspberry Pi, Jetson Nano, and mobile devices where memory and thermal constraints are critical. With only 1.9 million parameters and 4.5 GFLOPs, this variant achieves real-time inference under 10 ms on low-power GPUs while maintaining sufficient accuracy for basic detection tasks.
CPU Inference and Low-End GPUs (YOLOv5s)
YOLOv5s (small) provides the best balance for consumer-grade hardware and CPU-based inference pipelines. Doubling the width multiplier to 0.50 (while keeping the same depth as nano) increases parameters to 7.2M, delivering significantly higher mAP (37.4%) with minimal latency increase (6.4 ms on V100).
Standard Workstation Training (YOLOv5m)
YOLOv5m (medium) is the recommended default for single-GPU training on moderate datasets. The configuration in models/yolov5m.yaml scales both depth (0.67) and width (0.75), resulting in 21.2M parameters that fit comfortably within 8 GB GPU memory while achieving 45.4% mAP.
High-Accuracy Production Systems (YOLOv5l)
YOLOv5l (large) serves as the baseline architecture with depth_multiple: 1.0 and width_multiple: 1.0. This 46.5M parameter model reaches 49.0% mAP and is ideal for server-side batch processing or cloud APIs where GPU resources are ample but latency must remain under 11 ms per image.
Maximum Precision Applications (YOLOv5x)
YOLOv5x (extra-large) maximizes detection accuracy for critical applications like medical imaging or satellite analysis. Defined in models/yolov5x.yaml with depth_multiple: 1.33 and width_multiple: 1.25, this 86.7M parameter variant achieves 50.7% mAP at the cost of 205.7 GFLOPs, making it suitable for offline processing or high-throughput GPU clusters.
Loading and Running Different YOLOv5 Sizes
PyTorch Hub Implementation
Load any variant dynamically using the Ultralytics repository through PyTorch Hub:
import torch
# Select model size: 'n', 's', 'm', 'l', or 'x'
model_size = 'm'
model_name = f'yolov5{model_size}'
# Load pretrained weights from ultralytics/yolov5
model = torch.hub.load('ultralytics/yolov5', model_name, pretrained=True)
# Optional: Enable half-precision for faster GPU inference
if torch.cuda.is_available():
model = model.half().cuda()
# Run inference on image URL, local path, or numpy array
results = model('https://ultralytics.com/images/zidane.jpg')
# Output results
results.print() # Print detection table
results.save('runs/detect/') # Save annotated images
Command-Line Detection
Use the detect.py script to process images, videos, or streams with any model size:
# Detect using YOLOv5s (small)
python detect.py --weights yolov5s.pt --source data/images/bus.jpg
# Detect using YOLOv5x (extra-large) on a webcam
python detect.py --weights yolov5x.pt --source 0
The detect.py utility automatically handles model loading, input preprocessing, and NMS (Non-Maximum Suppression) for all five YOLOv5 configurations.
Summary
- YOLOv5n through YOLOv5x share the same underlying v6.0 architecture but scale via
depth_multipleandwidth_multipleparameters in their respective YAML files (models/yolov5n.yamltomodels/yolov5x.yaml). - Parameter counts range from 1.9M (nano) to 86.7M (extra-large), directly correlating with inference speed and detection accuracy on COCO benchmarks.
- Nano and Small variants target edge devices and CPU inference, while Medium through Extra-Large serve workstation and server environments requiring higher precision.
- Switching between sizes requires only changing the model name in PyTorch Hub or the
--weightsargument indetect.py, with no code changes to the inference pipeline.
Frequently Asked Questions
What do the letters n, s, m, l, and x stand for in YOLOv5?
These letters represent nano, small, medium, large, and extra-large, corresponding to the model's capacity and physical size on disk. Each letter maps to specific scaling multipliers in the YAML configuration files that control layer depth and channel width.
How do I switch between YOLOv5 model sizes in my Python code?
Pass the desired size letter as part of the model name when loading via torch.hub.load(). For example, use 'yolov5n' for nano or 'yolov5x' for extra-large. The function automatically downloads the corresponding pretrained weights (yolov5n.pt through yolov5x.pt) from the Ultralytics release assets.
Can I create a custom YOLOv5 size between the standard variants?
Yes, you can create intermediate sizes by manually editing the depth_multiple and width_multiple values in any of the YAML files (e.g., setting width_multiple: 0.6 in a copy of yolov5s.yaml). However, the five official variants are pre-optimized to balance memory alignment and computational efficiency on standard hardware.
Which YOLOv5 model size is best for real-time detection on a CPU?
YOLOv5s (small) is generally recommended for CPU-based real-time applications, as it provides the best accuracy-to-speed ratio for processors without dedicated GPU acceleration. YOLOv5n offers higher frame rates for extremely limited hardware but sacrifices significant accuracy (28.0% mAP vs. 37.4% mAP).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →