How to Set Up and Run Multi-GPU Training with Unsloth
Unsloth enables seamless multi-GPU fine-tuning without requiring changes to your training script, automatically detecting distributed launches via torchrun or accelerate and handling per-process device mapping internally to prevent memory relocation errors with quantized models.
Unsloth is an open-source library that accelerates large language model fine-tuning with optimized kernels and automatic optimizations. Setting up multi-GPU training with Unsloth requires no additional code modifications—the library handles distributed device mapping automatically when launched with standard multi-process tools. This guide walks through the complete process using the actual implementation in the unslothai/unsloth repository.
How Multi-GPU Training Works in Unsloth
Unsloth detects distributed training environments by checking environment variables set by standard launchers. When LOCAL_RANK, RANK, and WORLD_SIZE are present, the library automatically creates a per-process device map that pins each process to its assigned GPU. This approach prevents the "Accelerate device relocation" error that typically occurs when quantized weights are moved between devices after loading.
The workflow relies on two core components:
prepare_device_map()inunsloth/models/loader_utils.pyhandles rank detection and CUDA device configurationFastLanguageModel.from_pretrained()inunsloth/models/loader.pyintegrates this mapping when loading quantized models
Environment Prerequisites
Before launching multi-GPU training, ensure your environment meets these requirements:
- Install Unsloth and PyTorch 2.0+ with CUDA support
- Verify multiple GPUs are visible to PyTorch
- Use a launcher that sets distributed environment variables (
torchrunoraccelerate)
The launcher automatically exports LOCAL_RANK (process-specific GPU ID), RANK (global process ID), and WORLD_SIZE (total process count), which Unsloth reads to configure the device map.
Launching Multi-GPU Training
You can start distributed training using three equivalent approaches. Each automatically triggers Unsloth's multi-GPU path because the device_map parameter defaults to "sequential", causing prepare_device_map() to detect the distributed environment.
Using torchrun
The torchrun utility is the standard PyTorch distributed launcher. Specify the number of GPUs with --nproc_per_node:
torchrun --nproc_per_node=2 unsloth-cli.py \
--model_name "unsloth/Llama-3.2-1B-Instruct" \
--max_seq_length 8192 \
--load_in_4bit true \
--per_device_train_batch_size 4 \
--gradient_accumulation_steps 8 \
--learning_rate 2e-6 \
--max_steps 400 \
--output_dir ./outputs
Using the Direct Python API
For custom training loops, call prepare_device_map() before loading the model. This function returns a dictionary mapping the model to the current process's GPU:
from unsloth import FastLanguageModel
from unsloth.models.loader_utils import prepare_device_map
# Environment variables set automatically by torchrun
device_map, _ = prepare_device_map() # Returns {"": "cuda:<local_rank>"}
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Llama-3.2-1B-Instruct",
max_seq_length=8192,
load_in_4bit=True,
device_map=device_map, # Multi-GPU aware mapping
)
# Apply LoRA configuration (identical to single-GPU)
model = FastLanguageModel.get_peft_model(
model,
r=64,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
lora_alpha=32,
lora_dropout=0.1,
)
Using Hugging Face Accelerate
If you prefer the Hugging Face ecosystem, accelerate launch provides equivalent functionality:
accelerate launch unsloth-cli.py \
--model_name "unsloth/Llama-3.2-1B-Instruct" \
--load_in_4bit true \
--per_device_train_batch_size 2 \
--gradient_accumulation_steps 4
Internal Implementation Details
Understanding the source code reveals how Unsloth eliminates manual configuration.
In unsloth/models/loader_utils.py, the prepare_device_map() function (lines 84-99) reads the distributed environment variables and sets the CUDA device for the current process. According to the Unsloth source code, this function checks for LOCAL_RANK in the environment, calls torch.cuda.set_device(local_rank), and returns a device map formatted as {"": "cuda:0"} for the local rank.
The FastLanguageModel.from_pretrained() method in unsloth/models/loader.py (lines 91-99) automatically invokes this function when it detects a string device_map argument and a quantized model request. The method passes the generated device map to the underlying model loader, ensuring each process initializes its model shard on the correct device. This per-process device map is critical because it ensures each rank loads quantized weights—whether 4-bit, 8-bit, or FP8—directly onto its assigned GPU rather than loading on CPU and relocating, which causes errors with quantized tensors.
Summary
- Unsloth automatically detects multi-GPU environments via
torchrunoraccelerateenvironment variables - The
prepare_device_map()function inunsloth/models/loader_utils.pycreates per-process GPU assignments - Quantized models (4-bit, 8-bit, FP8) load directly on target devices, avoiding relocation errors
- No code changes are required between single-GPU and multi-GPU training scripts
- Both CLI (
unsloth-cli.py) and Python API support distributed training with identical syntax
Frequently Asked Questions
Does Unsloth require code changes to switch from single-GPU to multi-GPU?
No. Unsloth detects the distributed launch environment automatically through environment variables like LOCAL_RANK and WORLD_SIZE. The same training script works for both single-GPU and multi-GPU configurations because prepare_device_map() handles the transition internally.
Why does Unsloth use per-process device maps instead of letting Accelerate handle device placement?
Standard Accelerate device relocation fails with quantized weights because moving 4-bit or 8-bit tensors between GPUs after loading causes memory errors. As implemented in unslothai/unsloth, the prepare_device_map() function ensures each process loads quantized weights directly onto its assigned GPU, eliminating cross-device memory moves that break quantized model training.
Can I use DeepSpeed or FSDP with Unsloth's multi-GPU training?
While Unsloth handles the initial model loading and device mapping via prepare_device_map(), you can integrate DeepSpeed or FSDP through the Hugging Face Trainer. However, Unsloth's optimizations primarily target the data parallelism approach used with torchrun and standard distributed data parallel (DDP) configurations.
What happens if I run a multi-GPU script without a launcher like torchrun?
Without torchrun or accelerate launch, the LOCAL_RANK and WORLD_SIZE environment variables remain unset. In this case, prepare_device_map() falls back to single-GPU mode, loading the model on cuda:0 only, effectively running single-GPU training even if multiple GPUs are physically present.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →