YOLOv5 Image Segmentation: A Complete Training and Implementation Guide

YOLOv5 supports instance segmentation through the SegmentationModel class, which extends standard detection architectures with a mask-head prototype branch, enabling pixel-wise object masks while maintaining the same fast inference pipeline used for object detection.

The ultralytics/yolov5 repository includes a dedicated instance segmentation pipeline that extends the framework's detection capabilities. YOLOv5 image segmentation leverages a prototype-based mask generation system to produce high-quality pixel-wise object boundaries without sacrificing the real-time performance that makes YOLO models popular for production deployments.

How YOLOv5 Image Segmentation Works

Architecture Overview

The segmentation implementation centers on the SegmentationModel class defined in models/yolo.py (lines 36-44). This class inherits from DetectionModel and reuses the same backbone, neck, and detection head architecture, then appends a prototype mask branch that generates per-object masks through a coefficient-based approach.

The architecture produces masks by combining a low-resolution prototype tensor (protos) with per-object coefficient vectors (masks_in). This design, implemented in utils/segment/general.py (lines 24-41), multiplies the up-sampled prototype tensor by object-specific coefficients to generate full-resolution masks, which are then cropped to each detected bounding box.

Mask Post-Processing Pipeline

After inference, the crop_mask, process_mask_*, and scale_image functions in utils/segment/general.py (lines 9-23) handle the conversion from model outputs to final mask predictions. These utilities manage up-sampling, resizing to original image dimensions, and cropping masks to fit precise object boundaries.

Loss Function Implementation

The training loss combines standard detection objectives (box, objectness, classification) with a dedicated mask loss. The ComputeLoss class in utils/segment/loss.py implements this using binary cross-entropy to optimize the prototype-coefficient formulation, ensuring accurate mask boundaries alongside detection metrics.

Training a YOLOv5 Segmentation Model

Dataset Preparation

Training requires images paired with binary segmentation masks. Organize your data with separate images/ and masks/ directories, where mask filenames match their corresponding images. Each mask should be a binary image where object pixels are marked for segmentation targets.

Create a data configuration YAML file specifying paths, class count (nc), and class names. The utils/segment/dataloaders.py module loads these image-mask pairs, applies augmentations, and returns tensors in the format (imgs, targets, paths, _, masks).

Training Script Configuration

Use segment/train.py as the dedicated training entry point. This script mirrors the standard detection trainer but adds segmentation-specific arguments including --mask-ratio (controls ground-truth mask down-sampling, default 4) and --no-overlap (disables overlapping masks for faster training).

python segment/train.py \
  --data my_seg.yaml \
  --weights yolov5s-seg.pt \
  --cfg yolov5s-seg.yaml \
  --batch-size 16 \
  --imgsz 640 \
  --epochs 100 \
  --device 0

Model configurations reside in models/segment/*.yaml (such as yolov5s-seg.yaml), defining the architecture backbone and head specifically for segmentation tasks.

Running Inference with Segmentation

For prediction, load pretrained segmentation weights and enable mask output:

import torch

# Load model

model = torch.hub.load('ultralytics/yolov5', 'custom', path='yolov5s-seg.pt')

# Run inference with masks enabled

results = model('image.jpg', size=640, masks=True)
results.show()  # Display with overlaid masks

To convert output masks to polygon format for COCO-style evaluation, use the masks2segments utility:

from utils.segment.general import masks2segments

# masks shape: (N, H, W)

segments = masks2segments(masks, strategy='largest')

# Returns list of (M_i, 2) numpy arrays

Key Files and Implementation Details

Summary

  • YOLOv5 image segmentation extends the detection architecture via SegmentationModel, adding a prototype-based mask head to models/yolo.py
  • Training uses segment/train.py with binary mask annotations stored alongside images, loaded by utils/segment/dataloaders.py
  • The loss function in utils/segment/loss.py combines detection metrics with binary cross-entropy mask optimization
  • Post-processing utilities in utils/segment/general.py handle mask up-sampling, cropping, and scaling to original image dimensions
  • Pretrained weights (e.g., yolov5s-seg.pt) and configuration files in models/segment/ enable immediate training or fine-tuning

Frequently Asked Questions

What is the difference between YOLOv5 detection and segmentation models?

Detection models predict bounding boxes and class labels, while segmentation models add pixel-wise mask predictions. The SegmentationModel class in models/yolo.py inherits from DetectionModel and appends a prototype mask branch, enabling instance segmentation that identifies exact object boundaries rather than just rectangular regions.

How do I prepare mask annotations for YOLOv5 training?

Store binary masks as PNG files in a masks/ directory parallel to your images/ folder, using identical filenames to link masks with images. Each mask should contain binary pixel values indicating object presence. The dataloader in utils/segment/dataloaders.py automatically pairs these files during training batch generation.

Can I use pretrained detection weights for segmentation training?

Yes, you can initialize segmentation training with detection backbone weights, though using dedicated segmentation pretrained weights (e.g., yolov5s-seg.pt) typically yields better results. The backbone architecture is shared between detection and segmentation models, allowing knowledge transfer from detection datasets to segmentation tasks.

What hardware requirements are needed for training segmentation models?

Segmentation training requires more GPU memory than detection due to the additional mask processing and prototype generation overhead. A modern NVIDIA GPU with at least 8GB VRAM is recommended for training yolov5s-seg at 640px resolution with batch size 16. Reduce --batch-size or increase --mask-ratio (default 4) if encountering memory constraints.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →