MiniMind Pre-Training Resume Across Different GPU Configurations: Automatic Step Scaling Explained
MiniMind automatically resumes pre-training across different GPU counts by proportionally scaling the step counter in the lm_checkpoint utility when the distributed world size changes between saves.
MiniMind is an open-source, lightweight language model implementation designed for efficient distributed training and experimentation. When working with dynamic hardware environments, GPU availability frequently changes between training sessions, making seamless checkpoint resumption essential. The framework handles MiniMind pre-training resume across different GPU configurations through intelligent step scaling logic embedded directly in its checkpointing system.
How World Size Tracking Enables Cross-GPU Resume
MiniMind's checkpointing mechanism, implemented in trainer/trainer_utils.py, captures the distributed training context at save time. When lm_checkpoint saves a model state, it records the current world size (total GPU count) alongside model weights and optimizer states:
ckp_data = {
'model': state_dict,
'optimizer': optimizer.state_dict(),
'epoch': epoch,
'step': step,
'world_size': dist.get_world_size() if dist.is_initialized() else 1,
...
}
This world size metadata allows the training script to detect configuration changes during resumption. The checkpoint file pretrain_512_resume.pth stores this value persistently, enabling the framework to recognize transitions between single-GPU and multi-GPU setups, or between different multi-GPU topologies.
The Step Scaling Mechanism in trainer_utils.py
When loading a checkpoint via lm_checkpoint in resume mode, MiniMind compares the saved world size against the current distributed environment. According to the source code in trainer/trainer_utils.py (lines 111-115), if the GPU count differs between the saved checkpoint and the current run, the framework automatically adjusts the step counter to maintain training consistency:
saved_ws = ckp_data.get('world_size', 1)
current_ws = dist.get_world_size() if dist.is_initialized() else 1
if saved_ws != current_ws:
ckp_data['step'] = ckp_data['step'] * saved_ws // current_ws
Logger(f'GPU数量变化({saved_ws}→{current_ws}),step已自动转换为{ckp_data["step"]}')
Integer scaling ensures that the logical training progress remains consistent regardless of hardware changes. For example, when reducing from 4 GPUs to 2 GPUs, the step count doubles to account for the halved gradient accumulation frequency per process.
Practical Examples: Resuming on Different GPU Counts
Resuming from 4 GPUs to 2 GPUs
Start initial training on 4 GPUs with the --from_resume 0 flag:
torchrun --nproc_per_node=4 trainer/train_pretrain.py \
--epochs 10 \
--batch_size 32 \
--learning_rate 5e-4 \
--use_wandb \
--save_dir ../out \
--from_resume 0
Later, resume on 2 GPUs using --from_resume 1:
torchrun --nproc_per_node=2 trainer/train_pretrain.py \
--epochs 10 \
--batch_size 32 \
--learning_rate 5e-4 \
--use_wandb \
--save_dir ../out \
--from_resume 1
The lm_checkpoint function detects the world size change (4 → 2) and applies the transformation step = original_step * 4 // 2, effectively doubling the step counter to maintain equivalent data throughput.
Scaling Up from 2 GPUs to 4 GPUs
The reverse operation works identically. When resuming with increased GPU resources, the step count decreases proportionally:
torchrun --nproc_per_node=4 trainer/train_pretrain.py \
--from_resume 1
This triggers the calculation step = original_step * 2 // 4, ensuring the training schedule aligns with the increased parallel processing capacity.
Checkpoint Loading in train_pretrain.py
The resume workflow is orchestrated in trainer/train_pretrain.py (lines 73-99), where the script initializes training states based on checkpoint data. When --from_resume 1 is specified, the loader reconstructs the full training context:
ckp_data = lm_checkpoint(lm_config, weight=args.save_weight,
save_dir='../checkpoints') if args.from_resume==1 else None
if ckp_data:
model.load_state_dict(ckp_data['model'])
optimizer.load_state_dict(ckp_data['optimizer'])
scaler.load_state_dict(ckp_data['scaler'])
start_epoch = ckp_data['epoch']
start_step = ckp_data.get('step', 0)
The scaled step value propagates through the training loop, ensuring that learning rate schedules and logging metrics remain accurate despite the hardware configuration change.
Summary
- Automatic detection: MiniMind stores
world_sizein checkpoints vialm_checkpointto track the GPU count used during saving. - Proportional scaling: The framework adjusts step counters using integer division (
step * saved_ws // current_ws) when GPU counts differ between sessions. - Seamless resumption: Use
--from_resume 1intrain_pretrain.pyto trigger automatic loading and scaling without manual intervention. - Bidirectional support: The scaling logic works correctly when increasing or decreasing GPU resources (e.g., 4→2 or 2→4 GPUs).
- Implementation location: Critical logic resides in
trainer/trainer_utils.py(lines 111-115) with orchestration intrainer/train_pretrain.py(lines 73-99).
Frequently Asked Questions
Does MiniMind require manual step adjustment when changing GPUs?
No. MiniMind automatically handles step adjustment through the lm_checkpoint utility in trainer/trainer_utils.py. When the saved world size differs from the current GPU count, the framework applies integer scaling to the step counter without requiring user intervention.
What happens if I resume with more GPUs than before?
When resuming with increased GPU resources (for example, moving from 2 GPUs to 4 GPUs), MiniMind divides the step count proportionally using the formula step = saved_step * saved_ws // current_ws. This ensures the training schedule accounts for the increased parallel gradient computation.
Where is the GPU count stored in MiniMind checkpoints?
The GPU count (world size) is stored in the pretrain_512_resume.pth checkpoint file under the key 'world_size'. This value is captured via dist.get_world_size() during the save operation in trainer/trainer_utils.py and persists alongside model weights and optimizer states.
Can I resume training on a single GPU after training on multiple GPUs?
Yes. MiniMind supports resuming on any GPU configuration, including single GPU resumption after multi-GPU training. The scaling logic treats single GPU as world size 1, automatically adjusting the step count to reflect the change from distributed to non-distributed training.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →