How Soup CLI Handles Automatic Batch Sizing: A Technical Deep Dive
Soup CLI automatically determines the maximum feasible batch size by selecting between a fast static formula or a physical GPU memory probe, choosing the appropriate method based on your hardware configuration.
The MakazhanAlpamys/Soup framework eliminates manual batch size tuning through intelligent automatic batch sizing. When you leave training.batch_size set to "auto" in your configuration, the CLI computes the largest batch that fits within available memory using one of three selectable strategies. This process ensures optimal GPU utilization without the risk of out-of-memory crashes during training runs.
Configuration Schema for Automatic Batch Detection
The automatic batch sizing behavior is controlled through the TrainingConfig schema defined in src/soup_cli/config/schema.py. The system recognizes two critical configuration fields that govern memory allocation.
The batch_size parameter accepts Union[int, Literal["auto"]] with a default value of "auto" (lines 26-30). When set to "auto", Soup CLI delegates batch size calculation to the autotuning system rather than using a fixed integer.
The auto_batch_size_strategy field (lines 43-50) determines which algorithm executes during initialization. This field defaults to "auto", enabling the framework to select the most appropriate method based on the detected accelerator type.
The Three Auto-Batch Sizing Strategies
Soup CLI implements three distinct strategies for memory-aware batch calculation, each suited to different hardware environments:
static — Uses a deterministic formula based on the model's VRAM budget and per-sample memory requirements without attempting actual device allocation. This approach provides fast, predictable results ideal for CPU-only training runs or environments where physical GPU probing is undesirable.
probe — Performs real-world memory allocation trials by attempting to create dummy tensors on the GPU. When an out-of-memory (OOM) error occurs, the algorithm halves the batch size and retries until successful allocation confirms the maximum feasible size. This method guarantees the largest possible batch on CUDA devices accounting for quantization overhead and data type variations.
auto (default) — Automatically selects the optimal approach: executing the probe loop on CUDA devices while falling back to the static formula on CPU-only systems. This dual-path logic ensures both accuracy on GPUs and efficiency on CPU architectures.
Step-by-Step Batch Size Resolution Process
The resolution process follows a five-phase pipeline that transforms the "auto" placeholder into a concrete integer value:
-
Configuration parsing — The system validates
TrainingConfig.batch_sizeas either an explicit integer or the literal"auto"string during YAML loading. -
Strategy selection — Based on
auto_batch_size_strategy, Soup CLI determines whether to use the static formula or physical probing. The default"auto"setting checks for CUDA availability to choose between methods. -
Static calculation — For CPU targets or explicit
staticselection, the system estimates per-GPU consumption from model dimensions and available VRAM, computing the largest batch staying under the target memory budget. -
Probe loop — For GPU targets using the
probestrategy, the framework iteratively tests batch allocation. Starting from an estimated maximum, it catches OOM exceptions, reduces the batch size by half, and repeats until successful allocation confirms the hardware limit. -
Result application — The resolved integer replaces
"auto"in the configuration object before passing to the underlying trainer (HF Trainer or TRL Trainer) for the actual training execution.
Practical Configuration Examples
Override automatic detection in your soup.yaml to match your infrastructure constraints:
Default GPU-aware automatic sizing:
training:
batch_size: auto # Automatically pick the biggest fitting batch
auto_batch_size_strategy: auto # Probe on GPU, static on CPU
Force static calculation for headless servers:
training:
batch_size: auto
auto_batch_size_strategy: static # Fast formula without allocation attempts
Explicit fixed batch size (disable auto-sizing):
training:
batch_size: 4 # Fixed size, bypass auto-selection entirely
Access resolved values programmatically:
from soup_cli.config import load_config
cfg = load_config("soup.yaml")
print(f"Resolved batch size: {cfg.training.batch_size}")
Implementation Details in Source Code
The batch sizing logic spans multiple modules within the repository:
src/soup_cli/config/schema.py— DefinesTrainingConfig.batch_sizeasUnion[int, Literal["auto"]]andauto_batch_size_strategywith Pydantic validation.src/soup_cli/utils/batch_autosize.py— Contains the static memory formula and probe allocation loop implementations.src/soup_cli/trainer.py— Reads the resolved integer from the configuration and injects it into the HF Trainer or TRL Trainer initialization.
Summary
- Soup CLI defaults to
batch_size: "auto"insrc/soup_cli/config/schema.py, removing manual tuning requirements. - Three strategies exist: static (formula-based), probe (physical allocation testing), and auto (hardware-aware selection).
- The probe strategy guarantees maximum GPU utilization by testing actual memory allocation, while static provides fast CPU-compatible sizing.
- Configuration occurs through
training.batch_sizeandtraining.auto_batch_size_strategyfields insoup.yaml. - Resolved batch sizes replace the
"auto"placeholder before reaching the underlying Hugging Face or TRL trainer.
Frequently Asked Questions
How does Soup CLI prevent out-of-memory errors during automatic batch sizing?
The probe strategy prevents OOM crashes during actual training by catching allocation failures during the initialization phase. When the system attempts to create a dummy batch on GPU memory and receives a CUDA out-of-memory error, it immediately halves the batch size and retries until finding a size that fits safely within available VRAM.
What is the difference between the static and probe auto-batch strategies?
The static strategy estimates memory usage using a mathematical formula based on model parameters and data types without touching the GPU, making it fast but potentially less accurate with quantization. The probe strategy physically allocates memory tensors to verify available capacity, providing exact measurements but requiring CUDA device access during configuration initialization.
Can I use automatic batch sizing on CPU-only machines?
Yes. When auto_batch_size_strategy is set to "auto" (the default) and no CUDA device is detected, Soup CLI automatically selects the static strategy. This CPU-safe approach calculates batch sizes using memory estimation formulas rather than physical allocation attempts, which would be unnecessary overhead on CPU architectures.
Where does Soup CLI store the resolved batch size after auto-detection?
The framework modifies the TrainingConfig object in-place during initialization, replacing the string "auto" with the calculated integer value. This resolved value persists in the configuration instance passed to src/soup_cli/trainer.py and subsequent training loops, ensuring downstream components receive a concrete batch size rather than the placeholder.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →