How to Configure Multi‑GPU Training with DeepSpeed or FSDP in Soup

Soup provides first‑class DeepSpeed and FSDP utilities that let you scale training across multiple GPUs using simple preset flags, with automatic config resolution and GPU‑count adaptation.

The Soup CLI framework offers two complementary strategies for distributed training: Microsoft DeepSpeed (with ZeRO‑2/3 sharding and off‑loading) and PyTorch FSDP (Fully‑Sharded Data Parallel). Both are implemented as helper modules that translate human‑readable presets into complete configuration objects, handling temporary files, GPU scaling, and validation automatically.

DeepSpeed: ZeRO Sharding with Preset Configs

DeepSpeed integration lives in [src/soup_cli/utils/deepspeed.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/deepspeed.py). This module exposes three core functions:

  • get_deepspeed_config(preset) — Returns a mutable dictionary containing the full DeepSpeed JSON schema for presets like zero2, zero3, or zero3_offload.
  • resolve_deepspeed_config(cfg, gpu_count=n) — Adapts batch sizes and gradient accumulation steps to match your target GPU count.
  • write_deepspeed_config(cfg) — Persists the config to a temporary .json file that the trainer consumes.

Enabling DeepSpeed via CLI

Pass --deepspeed <preset> to the soup train command:

soup train \
  --model meta-llama/Llama-2-7b-chat-hf \
  --dataset my_dataset \
  --deepspeed zero3 \
  --gpus 4 \
  --output_dir ./checkpoints

Available presets include:

Preset Description
zero2 ZeRO‑2 (optimizer state sharding)
zero3 ZeRO‑3 (parameter + gradient + optimizer sharding)
zero3_offload ZeRO‑3 with optimizer state off‑loaded to CPU/NVMe

Programmatic DeepSpeed Configuration

Customize configs in Python before writing to disk:

from soup_cli.utils.deepspeed import (
    get_deepspeed_config,
    resolve_deepspeed_config,
    write_deepspeed_config,
)

# Load preset configuration

cfg = get_deepspeed_config("zero3_offload")

# Adapt for 2 GPUs (adjusts micro_batch_size and grad_accum)

cfg, notes = resolve_deepspeed_config(cfg, gpu_count=2)

# Write temporary JSON for trainer consumption

config_path = write_deepspeed_config(cfg)
print(f"DeepSpeed config: {config_path}")

FSDP: PyTorch Native Sharding

For cases where you prefer PyTorch's built‑in distributed strategy, Soup provides [src/soup_cli/utils/fsdp.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/fsdp.py). The primary entry point is:

  • get_fsdp_config(preset) — Returns a configuration dict with the "fsdp" key containing sharding_strategy, auto_wrap_policy, and offload_params settings.

Enabling FSDP via CLI

Use the --fsdp flag with a preset name:

soup train \
  --model meta-llama/Llama-2-7b-chat-hf \
  --dataset my_dataset \
  --fsdp full_offload \
  --gpus 4 \
  --output_dir ./checkpoints

FSDP presets map directly to PyTorch sharding strategies:

Preset PyTorch Equivalent Use Case
full_shard FULL_SHARD Maximum memory efficiency
shard_grad SHARD_GRAD_OP Lower communication overhead
full_offload FULL_SHARD + CPU off‑load Extreme memory constraints

Programmatic FSDP Configuration

from soup_cli.utils.fsdp import get_fsdp_config

fsdp_cfg = get_fsdp_config("shard_grad")
print(fsdp_cfg["fsdp"]["sharding_strategy"])

# Output: "SHARD_GRAD_OP"

Choosing Between DeepSpeed and FSDP

Both strategies are mutually exclusive. Soup validates CLI arguments and raises a clear error if you attempt to enable both simultaneously. This validation is enforced in [tests/test_multi_gpu.py](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_multi_gpu.py) and [tests/test_performance.py](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_performance.py).

DeepSpeed excels when you need:

  • CPU/NVMe off‑loading for massive models
  • ZeRO‑Infinity for training models larger than aggregate GPU memory
  • Built‑in DeepSpeed‑specific optimizers (AdamW, OneBitAdam)

FSDP is preferable when you want:

  • Native PyTorch integration without additional dependencies
  • Simpler debugging with standard torch.distributed tooling
  • Faster setup for medium‑scale models that fit with standard sharding

Source Files and Architecture

File Purpose
[src/soup_cli/utils/deepspeed.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/deepspeed.py) DeepSpeed config loading, GPU adaptation, and JSON serialization
[src/soup_cli/utils/fsdp.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/fsdp.py) FSDP preset definitions and config generation
[src/soup_cli/commands/train.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/commands/train.py) CLI argument parsing and trainer wiring
[tests/test_multi_gpu.py](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_multi_gpu.py) DeepSpeed config resolution and conflict detection tests
[tests/test_performance.py](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_performance.py) FSDP preset validation and structure verification
[tests/test_issue336_deepspeed_lora.py](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_issue336_deepspeed_lora.py) GPU‑count adaptation logic for DeepSpeed + LoRA combinations

Summary

Frequently Asked Questions

What happens if I specify both --deepspeed and --fsdp?

Soup raises a validation error before training starts. The CLI explicitly checks for conflicting flags in the argument parser, as verified by [tests/test_multi_gpu.py](https://github.com/MakazhanAlpamys/Soup/blob/main/tests/test_multi_gpu.py). You must choose one distributed strategy per run.

Does Soup support DeepSpeed ZeRO‑Infinity or NVMe off‑loading?

Yes. The zero3_offload preset in [deepspeed.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/deepspeed.py) includes configuration keys for CPU and NVMe off‑loading. Extend the preset or use get_deepspeed_config() as a base for custom ZeRO‑Infinity configurations.

Can I use a custom DeepSpeed JSON file instead of presets?

Absolutely. Pass --deepspeed-config /path/to/config.json directly to soup train. The CLI skips preset resolution and uses your file verbatim. The [deepspeed.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/deepspeed.py) helpers remain available for inspection or modification if needed.

How does Soup determine batch sizes for different GPU counts?

The resolve_deepspeed_config(cfg, gpu_count=n) function in [deepspeed.py](https://github.com/MakazhanAlpamys/Soup/blob/main/src/soup_cli/utils/deepspeed.py) recalculates train_micro_batch_size_per_gpu and gradient_accumulation_steps—`

` bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash `` bash bash ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` ` bash ``

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →