YOLOv5 Image Segmentation: A Complete Training and Implementation Guide
YOLOv5 supports instance segmentation through the SegmentationModel class, which extends standard detection architectures with a mask-head prototype branch, enabling pixel-wise object masks while maintaining the same fast inference pipeline used for object detection.
The ultralytics/yolov5 repository includes a dedicated instance segmentation pipeline that extends the framework's detection capabilities. YOLOv5 image segmentation leverages a prototype-based mask generation system to produce high-quality pixel-wise object boundaries without sacrificing the real-time performance that makes YOLO models popular for production deployments.
How YOLOv5 Image Segmentation Works
Architecture Overview
The segmentation implementation centers on the SegmentationModel class defined in models/yolo.py (lines 36-44). This class inherits from DetectionModel and reuses the same backbone, neck, and detection head architecture, then appends a prototype mask branch that generates per-object masks through a coefficient-based approach.
The architecture produces masks by combining a low-resolution prototype tensor (protos) with per-object coefficient vectors (masks_in). This design, implemented in utils/segment/general.py (lines 24-41), multiplies the up-sampled prototype tensor by object-specific coefficients to generate full-resolution masks, which are then cropped to each detected bounding box.
Mask Post-Processing Pipeline
After inference, the crop_mask, process_mask_*, and scale_image functions in utils/segment/general.py (lines 9-23) handle the conversion from model outputs to final mask predictions. These utilities manage up-sampling, resizing to original image dimensions, and cropping masks to fit precise object boundaries.
Loss Function Implementation
The training loss combines standard detection objectives (box, objectness, classification) with a dedicated mask loss. The ComputeLoss class in utils/segment/loss.py implements this using binary cross-entropy to optimize the prototype-coefficient formulation, ensuring accurate mask boundaries alongside detection metrics.
Training a YOLOv5 Segmentation Model
Dataset Preparation
Training requires images paired with binary segmentation masks. Organize your data with separate images/ and masks/ directories, where mask filenames match their corresponding images. Each mask should be a binary image where object pixels are marked for segmentation targets.
Create a data configuration YAML file specifying paths, class count (nc), and class names. The utils/segment/dataloaders.py module loads these image-mask pairs, applies augmentations, and returns tensors in the format (imgs, targets, paths, _, masks).
Training Script Configuration
Use segment/train.py as the dedicated training entry point. This script mirrors the standard detection trainer but adds segmentation-specific arguments including --mask-ratio (controls ground-truth mask down-sampling, default 4) and --no-overlap (disables overlapping masks for faster training).
python segment/train.py \
--data my_seg.yaml \
--weights yolov5s-seg.pt \
--cfg yolov5s-seg.yaml \
--batch-size 16 \
--imgsz 640 \
--epochs 100 \
--device 0
Model configurations reside in models/segment/*.yaml (such as yolov5s-seg.yaml), defining the architecture backbone and head specifically for segmentation tasks.
Running Inference with Segmentation
For prediction, load pretrained segmentation weights and enable mask output:
import torch
# Load model
model = torch.hub.load('ultralytics/yolov5', 'custom', path='yolov5s-seg.pt')
# Run inference with masks enabled
results = model('image.jpg', size=640, masks=True)
results.show() # Display with overlaid masks
To convert output masks to polygon format for COCO-style evaluation, use the masks2segments utility:
from utils.segment.general import masks2segments
# masks shape: (N, H, W)
segments = masks2segments(masks, strategy='largest')
# Returns list of (M_i, 2) numpy arrays
Key Files and Implementation Details
segment/train.py: Training entry point with segmentation-specific data loading and loss computationsegment/predict.py: CLI inference tool for mask generationmodels/yolo.py: ContainsSegmentationModelclass definition extending detection architectureutils/segment/dataloaders.py: Handles image-mask pair loading and augmentationutils/segment/loss.py: Implements combined detection and mask loss functionsutils/segment/general.py: Post-processing utilities includingcrop_maskand prototype generationmodels/segment/*.yaml: Architecture configurations (e.g.,yolov5s-seg.yaml)
Summary
- YOLOv5 image segmentation extends the detection architecture via
SegmentationModel, adding a prototype-based mask head tomodels/yolo.py - Training uses
segment/train.pywith binary mask annotations stored alongside images, loaded byutils/segment/dataloaders.py - The loss function in
utils/segment/loss.pycombines detection metrics with binary cross-entropy mask optimization - Post-processing utilities in
utils/segment/general.pyhandle mask up-sampling, cropping, and scaling to original image dimensions - Pretrained weights (e.g.,
yolov5s-seg.pt) and configuration files inmodels/segment/enable immediate training or fine-tuning
Frequently Asked Questions
What is the difference between YOLOv5 detection and segmentation models?
Detection models predict bounding boxes and class labels, while segmentation models add pixel-wise mask predictions. The SegmentationModel class in models/yolo.py inherits from DetectionModel and appends a prototype mask branch, enabling instance segmentation that identifies exact object boundaries rather than just rectangular regions.
How do I prepare mask annotations for YOLOv5 training?
Store binary masks as PNG files in a masks/ directory parallel to your images/ folder, using identical filenames to link masks with images. Each mask should contain binary pixel values indicating object presence. The dataloader in utils/segment/dataloaders.py automatically pairs these files during training batch generation.
Can I use pretrained detection weights for segmentation training?
Yes, you can initialize segmentation training with detection backbone weights, though using dedicated segmentation pretrained weights (e.g., yolov5s-seg.pt) typically yields better results. The backbone architecture is shared between detection and segmentation models, allowing knowledge transfer from detection datasets to segmentation tasks.
What hardware requirements are needed for training segmentation models?
Segmentation training requires more GPU memory than detection due to the additional mask processing and prototype generation overhead. A modern NVIDIA GPU with at least 8GB VRAM is recommended for training yolov5s-seg at 640px resolution with batch size 16. Reduce --batch-size or increase --mask-ratio (default 4) if encountering memory constraints.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →