How to Troubleshoot VRAM Issues with Soup Doctor: A Complete Guide

Soup doctor prevents out-of-memory crashes by running a VRAM pre-flight that predicts GPU memory requirements and refuses to start training when the predicted peak exceeds available VRAM.

soup doctor serves as the first-line health check for any Soup training run. Its built-in VRAM prediction system analyzes your model configuration, batch size, and sequence length to estimate peak memory consumption before execution begins. Understanding how this system works—and how to interpret its warnings—lets you resolve memory bottlenecks without trial-and-error crashes.

How Soup Doctor's VRAM Pre-flight Works

The VRAM check operates across three core components in the Soup codebase:

Core Components

Component Purpose Source File
CLI entry point Calls GPU detection and prints results src/soup_cli/commands/doctor.py
Static predictor Analytically estimates peak VRAM src/soup_cli/utils/hardware_fit.py
Streaming runtime Enforces VRAM limits during training src/soup_cli/utils/layer_stream.py

Step-by-Step VRAM Verification

The pre-flight follows a four-stage pipeline:

  1. GPU detection – _check_gpu() reads torch.cuda.get_device_properties(...).total_memory to determine total GPU memory per device.

  2. Peak prediction – estimate_peak_vram_gb() in hardware_fit.py constructs a HardwareFitInput from your model config, calculating memory for weights, optimizer states, and single-layer activations (the key optimization in Soup's streaming architecture). It applies a 10% safety margin (VRAM_SAFETY_MARGIN = 0.10) and returns a detailed VRAMBreakdown.

  3. Free memory comparison – Available VRAM is measured via torch.cuda.memory_allocated (with fallback methods), stored as available_bytes.

  4. Streaming decision – decide_stream_fit in layer_stream.py returns True only if predicted_bytes ≤ available_bytes. On failure, it logs a warning and aborts training.

Common VRAM Issues and Solutions

Doctor Reports "Ran Out of VRAM"

This occurs when the predicted peak exceeds free memory. The model config, batch size, or precision setting exceeds your hardware capacity.

Immediate fixes:

  • Reduce batch_size or max_seq_len via --batch or --max-length arguments
  • Enable quantization with --quantize 4bit to load compressed weights
  • Activate gradient checkpointing with --gradient-checkpointing (the estimator automatically adjusts for lower activation memory)

Doctor Shows "VRAM … GB (free)" Without Actual Measurement

The free-VRAM probe failed to execute—common in CI environments or when CUDA runtime is unavailable.

Resolution: Use the measured VRAM path only on physical GPUs. For headless environments, set TRAINING_VRAM_OVERRIDE=<GB> to manually specify available memory.

Doctor Warns: "GPU Hardware Present but Torch Is the CPU Build"

A GPU is detected, but PyTorch was installed without CUDA support.

Fix: Reinstall PyTorch with the CUDA-specific index:

pip install torch --index-url https://download.pytorch.org/whl/cu121

This check originates in _detect_gpu_hw_without_torch_cuda() within doctor.py.

Disk-Type Refusal with --disk Flag

Layer streaming's overflow tier requires NVMe SSD throughput. Spinning HDDs or SATA SSDs are rejected to prevent training stalls.

Options: Use an NVMe drive or disable overflow tier with --disk false.

NCCL Bandwidth Warning

Low NCCL bandwidth (often from poor interconnects or single-GPU setups) triggers a red status.

Resolution: Disable NCCL checks with --nccl false, or upgrade to PCIe 4.0/NVLink for multi-GPU training.

Diagnostic Commands

Use these commands to isolate and verify VRAM issues:


# Full health check with VRAM prediction

soup doctor

# VRAM-only check (skip disk and NCCL probes)

soup doctor --disk false --nccl false

# Quick fit test with reduced precision

soup train --model mymodel --quantize 4bit --batch 4

# Verify available VRAM without training

python -c "import torch; print(torch.cuda.get_device_properties(0).total_memory/1e9)"

Advanced VRAM Tuning

Override the Analytical Predictor

Force a custom free-VRAM value for heterogeneous clusters:

export TRAINING_VRAM_OVERRIDE=40  # Treat system as having 40GB free

Adjust Safety Margins

The VRAM_SAFETY_MARGIN constant in hardware_fit.py defaults to 0.10 (10%). Reducing this increases usable memory but raises OOM risk. Modify with caution.

Analyze VRAM Breakdown Output

After a failed check, Soup prints a VRAMBreakdown table showing memory distribution across:

  • Weights
  • Gradients
  • Optimizer states
  • Activation buffers

Use this to target optimizations. For example, if optimizer memory dominates, enable LoRA to reduce trainable parameters.

Key Source Files

File Role
src/soup_cli/commands/doctor.py CLI entry, GPU detection, dependency verification
src/soup_cli/utils/hardware_fit.py Analytical predictor (estimate_peak_vram_gb)
src/soup_cli/utils/layer_stream.py Runtime enforcement (decide_stream_fit)
src/soup_cli/utils/grad_accum.py VRAM pressure tracking and recommendations
tests/test_v07203.py Validation suite for pre-flight logic

Summary

  • Soup doctor prevents training crashes by predicting VRAM requirements before execution.
  • The three-component architecture separates detection (doctor.py), prediction (hardware_fit.py), and enforcement (layer_stream.py).
  • Common fixes include reducing batch size, enabling 4-bit quantization, activating gradient checkpointing, and installing CUDA-enabled PyTorch.
  • Environment overrides like TRAINING_VRAM_OVERRIDE support non-standard deployments.
  • The modular design makes it straightforward to extend checks for custom hardware or memory policies.

Frequently Asked Questions

How accurate is Soup doctor's VRAM prediction?

The analytical predictor in hardware_fit.py calculates peak usage for a single decoder layer plus gradients, then applies a 10% safety margin. It accounts for weights, optimizer states, and activations at your specified precision. While conservative, this approach reliably prevents OOM crashes in Soup's layer-streaming architecture.

Can I disable the VRAM check entirely?

There is no explicit disable flag. However, you can bypass enforcement by setting TRAINING_VRAM_OVERRIDE to a value exceeding your actual VRAM—though this risks runtime crashes. The check exists specifically to prevent wasted compute on doomed training runs.

Why does Soup doctor report less free VRAM than nvidia-smi?

The estimator uses torch.cuda.memory_allocated and accounts for existing PyTorch cache fragmentation and the 10% safety margin. Other tools may show raw unallocated memory without these training-specific adjustments.

Does gradient checkpointing always reduce VRAM usage in Soup?

Yes—when enabled via --gradient-checkpointing, the estimate_peak_vram_gb() function recalculates activation memory requirements. The VRAMBreakdown reflects this reduction before training begins, allowing larger effective batch sizes within the same hardware budget.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →