# How to Set Up and Run Multi-GPU Training with Unsloth

> Learn how to set up and run multi-GPU training with Unsloth. Achieve seamless distributed fine-tuning without script changes for faster model training.

- Repository: [Unsloth AI/unsloth](https://github.com/unslothai/unsloth)
- Tags: how-to-guide
- Published: 2026-03-20

---

**Unsloth enables seamless multi-GPU fine-tuning without requiring changes to your training script, automatically detecting distributed launches via `torchrun` or `accelerate` and handling per-process device mapping internally to prevent memory relocation errors with quantized models.**

Unsloth is an open-source library that accelerates large language model fine-tuning with optimized kernels and automatic optimizations. Setting up **multi-GPU training with Unsloth** requires no additional code modifications—the library handles distributed device mapping automatically when launched with standard multi-process tools. This guide walks through the complete process using the actual implementation in the `unslothai/unsloth` repository.

## How Multi-GPU Training Works in Unsloth

Unsloth detects distributed training environments by checking environment variables set by standard launchers. When `LOCAL_RANK`, `RANK`, and `WORLD_SIZE` are present, the library automatically creates a per-process **device map** that pins each process to its assigned GPU. This approach prevents the "Accelerate device relocation" error that typically occurs when quantized weights are moved between devices after loading.

The workflow relies on two core components:

- `prepare_device_map()` in [`unsloth/models/loader_utils.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader_utils.py) handles rank detection and CUDA device configuration
- `FastLanguageModel.from_pretrained()` in [`unsloth/models/loader.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader.py) integrates this mapping when loading quantized models

## Environment Prerequisites

Before launching multi-GPU training, ensure your environment meets these requirements:

1. Install Unsloth and PyTorch 2.0+ with CUDA support
2. Verify multiple GPUs are visible to PyTorch
3. Use a launcher that sets distributed environment variables (`torchrun` or `accelerate`)

The launcher automatically exports `LOCAL_RANK` (process-specific GPU ID), `RANK` (global process ID), and `WORLD_SIZE` (total process count), which Unsloth reads to configure the device map.

## Launching Multi-GPU Training

You can start distributed training using three equivalent approaches. Each automatically triggers Unsloth's multi-GPU path because the `device_map` parameter defaults to `"sequential"`, causing `prepare_device_map()` to detect the distributed environment.

### Using torchrun

The `torchrun` utility is the standard PyTorch distributed launcher. Specify the number of GPUs with `--nproc_per_node`:

```bash
torchrun --nproc_per_node=2 unsloth-cli.py \
    --model_name "unsloth/Llama-3.2-1B-Instruct" \
    --max_seq_length 8192 \
    --load_in_4bit true \
    --per_device_train_batch_size 4 \
    --gradient_accumulation_steps 8 \
    --learning_rate 2e-6 \
    --max_steps 400 \
    --output_dir ./outputs

```

### Using the Direct Python API

For custom training loops, call `prepare_device_map()` before loading the model. This function returns a dictionary mapping the model to the current process's GPU:

```python
from unsloth import FastLanguageModel
from unsloth.models.loader_utils import prepare_device_map

# Environment variables set automatically by torchrun

device_map, _ = prepare_device_map()  # Returns {"": "cuda:<local_rank>"}

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Llama-3.2-1B-Instruct",
    max_seq_length=8192,
    load_in_4bit=True,
    device_map=device_map,  # Multi-GPU aware mapping

)

# Apply LoRA configuration (identical to single-GPU)

model = FastLanguageModel.get_peft_model(
    model,
    r=64,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    lora_alpha=32,
    lora_dropout=0.1,
)

```

### Using Hugging Face Accelerate

If you prefer the Hugging Face ecosystem, `accelerate launch` provides equivalent functionality:

```bash
accelerate launch unsloth-cli.py \
    --model_name "unsloth/Llama-3.2-1B-Instruct" \
    --load_in_4bit true \
    --per_device_train_batch_size 2 \
    --gradient_accumulation_steps 4

```

## Internal Implementation Details

Understanding the source code reveals how Unsloth eliminates manual configuration.

In [`unsloth/models/loader_utils.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader_utils.py), the `prepare_device_map()` function (lines 84-99) reads the distributed environment variables and sets the CUDA device for the current process. According to the Unsloth source code, this function checks for `LOCAL_RANK` in the environment, calls `torch.cuda.set_device(local_rank)`, and returns a device map formatted as `{"": "cuda:0"}` for the local rank.

The `FastLanguageModel.from_pretrained()` method in [`unsloth/models/loader.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader.py) (lines 91-99) automatically invokes this function when it detects a string `device_map` argument and a quantized model request. The method passes the generated device map to the underlying model loader, ensuring each process initializes its model shard on the correct device. This per-process device map is critical because it ensures each rank loads quantized weights—whether 4-bit, 8-bit, or FP8—directly onto its assigned GPU rather than loading on CPU and relocating, which causes errors with quantized tensors.

## Summary

- Unsloth automatically detects multi-GPU environments via `torchrun` or `accelerate` environment variables
- The `prepare_device_map()` function in [`unsloth/models/loader_utils.py`](https://github.com/unslothai/unsloth/blob/main/unsloth/models/loader_utils.py) creates per-process GPU assignments
- Quantized models (4-bit, 8-bit, FP8) load directly on target devices, avoiding relocation errors
- No code changes are required between single-GPU and multi-GPU training scripts
- Both CLI ([`unsloth-cli.py`](https://github.com/unslothai/unsloth/blob/main/unsloth-cli.py)) and Python API support distributed training with identical syntax

## Frequently Asked Questions

### Does Unsloth require code changes to switch from single-GPU to multi-GPU?

No. Unsloth detects the distributed launch environment automatically through environment variables like `LOCAL_RANK` and `WORLD_SIZE`. The same training script works for both single-GPU and multi-GPU configurations because `prepare_device_map()` handles the transition internally.

### Why does Unsloth use per-process device maps instead of letting Accelerate handle device placement?

Standard Accelerate device relocation fails with quantized weights because moving 4-bit or 8-bit tensors between GPUs after loading causes memory errors. As implemented in `unslothai/unsloth`, the `prepare_device_map()` function ensures each process loads quantized weights directly onto its assigned GPU, eliminating cross-device memory moves that break quantized model training.

### Can I use DeepSpeed or FSDP with Unsloth's multi-GPU training?

While Unsloth handles the initial model loading and device mapping via `prepare_device_map()`, you can integrate DeepSpeed or FSDP through the Hugging Face Trainer. However, Unsloth's optimizations primarily target the data parallelism approach used with `torchrun` and standard distributed data parallel (DDP) configurations.

### What happens if I run a multi-GPU script without a launcher like torchrun?

Without `torchrun` or `accelerate launch`, the `LOCAL_RANK` and `WORLD_SIZE` environment variables remain unset. In this case, `prepare_device_map()` falls back to single-GPU mode, loading the model on `cuda:0` only, effectively running single-GPU training even if multiple GPUs are physically present.