YOLOv5 AutoBatch: Automatic Batch Size Optimization Explained

AutoBatch in YOLOv5 automatically calculates the largest training batch size that fits within available GPU memory, maximizing hardware throughput while preventing out-of-memory crashes.

YOLOv5 includes an automatic batch size optimization feature that eliminates manual tuning by dynamically adapting to your GPU's available memory. Implemented in the ultralytics/yolov5 repository, this utility profiles memory consumption in real-time to select the optimal batch size for training. Users simply pass --batch-size -1 to the training script, and the algorithm handles the rest.

How AutoBatch Works in YOLOv5

The core algorithm resides in utils/autobatch.py and follows a systematic approach to determine safe memory limits without user intervention.

Device Detection and Memory Inspection

AutoBatch begins by identifying the target device using device = next(model.parameters()).device. It then queries the GPU's memory state through PyTorch CUDA functions according to the YOLOv5 source code.

  • torch.cuda.get_device_properties retrieves total memory capacity
  • memory_reserved and memory_allocated calculate currently used memory
  • Free memory f is computed as the difference between total and allocated/reserved space

This runtime inspection allows the same code to work across 8 GB, 16 GB, and 24 GB devices regardless of current system load.

Profiling and Linear Modeling

To estimate memory requirements, AutoBatch profiles progressively larger batch sizes. The profile function from utils/torch_utils.py creates dummy tensors for batch sizes [1, 2, 4, 8, 16] and measures actual GPU memory consumption.

Using np.polyfit, the system builds a linear model y = p[0] * batch + p[1] where:

  • y represents memory usage
  • p[0] is the slope (memory per batch item)
  • p[1] is the intercept (base model overhead)

Batch Size Calculation and Safety Limits

The optimal batch size b is calculated to utilize a user-defined memory fraction (default 80%) without exceeding available free memory f:

b = int((f * fraction - p[1]) / p[0])

Safety mechanisms in utils/autobatch.py enforce strict boundaries:

  • If profiling fails at any step, AutoBatch falls back to the last successful batch size
  • Results are clamped between 1 and 1024 to prevent extreme values
  • Anomalies trigger a fallback to the default batch size of 16

Using AutoBatch in YOLOv5

Command Line Interface

The most common usage occurs through the training CLI. When you specify --batch-size -1 in train.py or segment/train.py, the script invokes check_train_batch_size, which automatically runs the AutoBatch algorithm.

python train.py \
  --data coco.yaml \
  --weights yolov5s.pt \
  --batch-size -1

This command triggers utils/autobatch.py to analyze your GPU and select the optimal batch size before training begins.

Python API Integration

For custom training loops, use the check_train_batch_size wrapper to determine the optimal size programmatically:

import torch
from utils.autobatch import check_train_batch_size

# Load model without autoshape

model = torch.hub.load('ultralytics/yolov5', 'yolov5s', autoshape=False)

# Calculate optimal batch size for 640px images

optimal_bs = check_train_batch_size(model, imgsz=640, amp=True)
print(f'Optimal batch size: {optimal_bs}')

The function returns an integer representing the maximum safe batch size for your current hardware configuration.

Direct autobatch Function Access

For debugging or specialized pipelines, invoke the core autobatch function directly:

from utils.autobatch import autobatch
import torch

model = torch.hub.load('ultralytics/yolov5', 'yolov5s', autoshape=False)

# Custom fraction and starting batch size

batch = autobatch(model, imgsz=640, fraction=0.85, batch_size=16)
print(f'Chosen batch size: {batch}')

Key Source Files and Functions

Understanding the file structure helps when customizing or debugging automatic batch size optimization:

  • utils/autobatch.py – Contains the core autobatch() function and check_train_batch_size() wrapper that implement the linear modeling and safety checks described above.

  • train.py – Parses the --batch-size argument and routes to check_train_batch_size when the value is -1, as implemented in the ultralytics/yolov5 repository.

  • segment/train.py – Mirrors the same AutoBatch integration for segmentation training tasks.

  • utils/torch_utils.py – Provides the profile utility function that measures memory consumption during the batch size profiling phase.

Summary

  • AutoBatch eliminates manual GPU memory tuning by automatically selecting the largest safe batch size for training.
  • The algorithm profiles batch sizes 1 through 16 to build a linear memory model, then solves for the batch size that uses approximately 80% of available GPU memory.
  • Implementation resides primarily in utils/autobatch.py, with entry points in train.py and segment/train.py when using --batch-size -1.
  • Safety mechanisms include linear regression validation, clamping to 1-1024, and fallback to batch size 16 if anomalies occur.
  • The system works across different GPU architectures (8 GB, 16 GB, 24 GB) by reading actual runtime memory availability rather than using hardcoded values.

Frequently Asked Questions

What triggers AutoBatch in YOLOv5 training?

AutoBatch activates when you pass --batch-size -1 to the training script. This flag signals train.py to call check_train_batch_size(), which invokes the autobatch() function from utils/autobatch.py to calculate the optimal size based on current GPU memory availability.

How does AutoBatch prevent out-of-memory errors?

The algorithm builds a linear model of memory consumption using test batches of sizes 1, 2, 4, 8, and 16. It then calculates a batch size that uses only a fraction (default 80%) of available GPU memory, leaving headroom for CUDA overhead and system fluctuations. If calculations fail, it falls back to the last successful batch size or defaults to 16.

Can I use AutoBatch with multiple GPUs?

Yes. AutoBatch detects the device using next(model.parameters()).device and reads memory properties for that specific GPU. When using DataParallel or DistributedDataParallel, you should ensure the model is moved to the appropriate device before calling check_train_batch_size(), as the function profiles memory on the device where the model parameters reside.

Why does AutoBatch use only 80% of available GPU memory?

The default fraction parameter of 0.8 leaves approximately 20% of GPU memory free to accommodate CUDA kernel overhead, temporary gradients, and memory fragmentation that occurs during backpropagation. You can adjust this safety margin by calling autobatch() directly with a custom fraction argument between 0.0 and 1.0.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →