# MiniMind Pre-Training Resume Across Different GPU Configurations: Automatic Step Scaling Explained

> Learn how MiniMind resumes pre-training across different GPU configurations using automatic step scaling. Discover the lm_checkpoint utility and seamless distributed world size adjustments.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: deep-dive
- Published: 2026-03-24

---

**MiniMind automatically resumes pre-training across different GPU counts by proportionally scaling the step counter in the `lm_checkpoint` utility when the distributed world size changes between saves.**

MiniMind is an open-source, lightweight language model implementation designed for efficient distributed training and experimentation. When working with dynamic hardware environments, GPU availability frequently changes between training sessions, making seamless checkpoint resumption essential. The framework handles MiniMind pre-training resume across different GPU configurations through intelligent step scaling logic embedded directly in its checkpointing system.

## How World Size Tracking Enables Cross-GPU Resume

MiniMind's checkpointing mechanism, implemented in [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py), captures the distributed training context at save time. When `lm_checkpoint` saves a model state, it records the current **world size** (total GPU count) alongside model weights and optimizer states:

```python
ckp_data = {
    'model': state_dict,
    'optimizer': optimizer.state_dict(),
    'epoch': epoch,
    'step': step,
    'world_size': dist.get_world_size() if dist.is_initialized() else 1,
    ...
}

```

This world size metadata allows the training script to detect configuration changes during resumption. The checkpoint file `pretrain_512_resume.pth` stores this value persistently, enabling the framework to recognize transitions between single-GPU and multi-GPU setups, or between different multi-GPU topologies.

## The Step Scaling Mechanism in trainer_utils.py

When loading a checkpoint via `lm_checkpoint` in resume mode, MiniMind compares the saved world size against the current distributed environment. According to the source code in [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) (lines 111-115), if the GPU count differs between the saved checkpoint and the current run, the framework automatically adjusts the step counter to maintain training consistency:

```python
saved_ws = ckp_data.get('world_size', 1)
current_ws = dist.get_world_size() if dist.is_initialized() else 1
if saved_ws != current_ws:
    ckp_data['step'] = ckp_data['step'] * saved_ws // current_ws
    Logger(f'GPU数量变化({saved_ws}→{current_ws})，step已自动转换为{ckp_data["step"]}')

```

**Integer scaling** ensures that the logical training progress remains consistent regardless of hardware changes. For example, when reducing from 4 GPUs to 2 GPUs, the step count doubles to account for the halved gradient accumulation frequency per process.

## Practical Examples: Resuming on Different GPU Counts

### Resuming from 4 GPUs to 2 GPUs

Start initial training on 4 GPUs with the `--from_resume 0` flag:

```bash
torchrun --nproc_per_node=4 trainer/train_pretrain.py \
    --epochs 10 \
    --batch_size 32 \
    --learning_rate 5e-4 \
    --use_wandb \
    --save_dir ../out \
    --from_resume 0

```

Later, resume on 2 GPUs using `--from_resume 1`:

```bash
torchrun --nproc_per_node=2 trainer/train_pretrain.py \
    --epochs 10 \
    --batch_size 32 \
    --learning_rate 5e-4 \
    --use_wandb \
    --save_dir ../out \
    --from_resume 1

```

The `lm_checkpoint` function detects the world size change (4 → 2) and applies the transformation `step = original_step * 4 // 2`, effectively doubling the step counter to maintain equivalent data throughput.

### Scaling Up from 2 GPUs to 4 GPUs

The reverse operation works identically. When resuming with increased GPU resources, the step count decreases proportionally:

```bash
torchrun --nproc_per_node=4 trainer/train_pretrain.py \
    --from_resume 1

```

This triggers the calculation `step = original_step * 2 // 4`, ensuring the training schedule aligns with the increased parallel processing capacity.

## Checkpoint Loading in train_pretrain.py

The resume workflow is orchestrated in [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py) (lines 73-99), where the script initializes training states based on checkpoint data. When `--from_resume 1` is specified, the loader reconstructs the full training context:

```python
ckp_data = lm_checkpoint(lm_config, weight=args.save_weight,
                         save_dir='../checkpoints') if args.from_resume==1 else None

if ckp_data:
    model.load_state_dict(ckp_data['model'])
    optimizer.load_state_dict(ckp_data['optimizer'])
    scaler.load_state_dict(ckp_data['scaler'])
    start_epoch = ckp_data['epoch']
    start_step = ckp_data.get('step', 0)

```

The scaled step value propagates through the training loop, ensuring that learning rate schedules and logging metrics remain accurate despite the hardware configuration change.

## Summary

- **Automatic detection**: MiniMind stores `world_size` in checkpoints via `lm_checkpoint` to track the GPU count used during saving.
- **Proportional scaling**: The framework adjusts step counters using integer division (`step * saved_ws // current_ws`) when GPU counts differ between sessions.
- **Seamless resumption**: Use `--from_resume 1` in [`train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/train_pretrain.py) to trigger automatic loading and scaling without manual intervention.
- **Bidirectional support**: The scaling logic works correctly when increasing or decreasing GPU resources (e.g., 4→2 or 2→4 GPUs).
- **Implementation location**: Critical logic resides in [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) (lines 111-115) with orchestration in [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py) (lines 73-99).

## Frequently Asked Questions

### Does MiniMind require manual step adjustment when changing GPUs?

No. MiniMind automatically handles step adjustment through the `lm_checkpoint` utility in [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py). When the saved world size differs from the current GPU count, the framework applies integer scaling to the step counter without requiring user intervention.

### What happens if I resume with more GPUs than before?

When resuming with increased GPU resources (for example, moving from 2 GPUs to 4 GPUs), MiniMind divides the step count proportionally using the formula `step = saved_step * saved_ws // current_ws`. This ensures the training schedule accounts for the increased parallel gradient computation.

### Where is the GPU count stored in MiniMind checkpoints?

The GPU count (world size) is stored in the `pretrain_512_resume.pth` checkpoint file under the key `'world_size'`. This value is captured via `dist.get_world_size()` during the save operation in [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) and persists alongside model weights and optimizer states.

### Can I resume training on a single GPU after training on multiple GPUs?

Yes. MiniMind supports resuming on any GPU configuration, including single GPU resumption after multi-GPU training. The scaling logic treats single GPU as world size 1, automatically adjusting the step count to reflect the change from distributed to non-distributed training.