How to Resume Pre-Training from a Checkpoint in MiniMind

MiniMind automatically resumes pre-training from the latest checkpoint when you pass --from_resume 1 to trainer/train_pretrain.py, restoring the model state, optimizer, scaler, and exact training step. The checkpoint system saves both model weights and resume metadata to ../checkpoints, allowing seamless continuation even when switching GPU counts.

The MiniMind repository provides a built-in checkpoint recovery mechanism in its pre-training pipeline. According to the source code in jingyaogong/minimind, the trainer/train_pretrain.py script handles automatic detection and restoration of training state without requiring manual file manipulation.

Understanding MiniMind's Checkpoint System

MiniMind saves two distinct files during training via the lm_checkpoint function in trainer/trainer_utils.py.

Model Weights and Resume Data

Every checkpoint creates:

  • Model weights: <save_dir>/<save_weight>_<hidden_size>{_moe}.pth (e.g., ../checkpoints/pretrain_512.pth)
  • Resume metadata: <save_dir>/<save_weight>_<hidden_size>{_moe}_resume.pth containing the model state dictionary, optimizer state, scaler state, current epoch, global step, world size, and optional Weights & Biases run ID

These files are written to the ../checkpoints directory by default, with the resume file enabling complete state restoration.

Resuming Pre-Training with the --from_resume Flag

To resume training, invoke the pre-training script with the resume flag enabled.

Checkpoint Detection and Loading

When args.from_resume == 1, the script calls lm_checkpoint to probe for existing resume data:

ckp_data = lm_checkpoint(lm_config, weight=args.save_weight,
                         save_dir='../checkpoints') if args.from_resume==1 else None

If the resume file exists, lm_checkpoint returns a dictionary loaded from the _resume.pth file. If no checkpoint is found, training starts from scratch.

Restoring Training State

The script restores the complete training context before entering the loop:

model.load_state_dict(ckp_data['model'])
optimizer.load_state_dict(ckp_data['optimizer'])
scaler.load_state_dict(ckp_data['scaler'])
start_epoch = ckp_data['epoch']
start_step  = ckp_data.get('step', 0)

The training loop then initializes from start_epoch and start_step, with the SkipBatchSampler class (defined in trainer/trainer_utils.py) skipping already-processed batches to prevent duplicate training on seen data.

Handling GPU Configuration Changes

MiniMind automatically adjusts training state when the GPU count differs from the original checkpoint. In trainer/trainer_utils.py, the lm_checkpoint function rescales the step count to maintain consistent learning rate scheduling:

ckp_data['step'] = ckp_data['step'] * saved_ws // current_ws

This conversion ensures that optimizer states and batch sampling remain valid when moving from single-GPU to multi-GPU setups (or vice versa).

Step-by-Step Resume Examples

Command Line Resume

Start a fresh pre-training run:

python -m trainer.train_pretrain \
    --save_dir ../out \
    --save_weight pretrain \
    --epochs 4 \
    --batch_size 32 \
    --learning_rate 5e-4 \
    --hidden_size 512 \
    --data_path ../dataset/pretrain_hq.jsonl

Resume from the latest checkpoint:

python -m trainer.train_pretrain \
    --save_dir ../out \
    --save_weight pretrain \
    --epochs 4 \
    --batch_size 32 \
    --learning_rate 5e-4 \
    --hidden_size 512 \
    --data_path ../dataset/pretrain_hq.jsonl \
    --from_resume 1

Programmatic Resume

You can also invoke the training loop programmatically:

import argparse
from trainer.train_pretrain import main

parser = argparse.ArgumentParser()
parser.add_argument('--from_resume', type=int, default=1)
parser.add_argument('--save_dir', type=str, default='../out')
parser.add_argument('--save_weight', type=str, default='pretrain')
parser.add_argument('--epochs', type=int, default=4)
parser.add_argument('--batch_size', type=int, default=32)
parser.add_argument('--learning_rate', type=float, default=5e-4)
parser.add_argument('--hidden_size', type=int, default=512)
parser.add_argument('--data_path', type=str, default='../dataset/pretrain_hq.jsonl')

args = parser.parse_args()
main(args)  # Automatically loads checkpoint and continues training

Summary

  • MiniMind saves paired checkpoint files: model weights (.pth) and resume metadata (_resume.pth) in ../checkpoints.
  • Pass --from_resume 1 to trainer/train_pretrain.py to automatically detect and load the latest checkpoint.
  • The resume process restores model parameters, optimizer state, gradient scaler, epoch counter, and global step via lm_checkpoint in trainer/trainer_utils.py.
  • GPU count mismatches are handled automatically by rescaling the step count in the checkpoint loader.
  • The SkipBatchSampler ensures no duplicate training by skipping batches processed in previous steps.

Frequently Asked Questions

Where does MiniMind store pre-training checkpoints?

MiniMind writes checkpoints to the ../checkpoints directory relative to the execution path, creating files named <save_weight>_<hidden_size>.pth for weights and <save_weight>_<hidden_size>_resume.pth for training state metadata.

Can I resume training on a different number of GPUs?

Yes. The lm_checkpoint function in trainer/trainer_utils.py automatically rescales the global step count based on the ratio of saved world size to current world size, ensuring optimizer states and learning rate schedules remain consistent across hardware configurations.

What happens if I use --from_resume 1 but no checkpoint exists?

The script treats missing checkpoints as a fresh start and initializes training from scratch without error, since lm_checkpoint returns None when the resume file is not found.

Should I use --from_resume with --from_weight?

Set --from_weight none when resuming to load only the checkpoint data stored in the resume file. Using a specific weight path with --from_weight while also setting --from_resume 1 may cause conflicts; the resume flag prioritizes the checkpoint's saved model state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →