Where Are MiniMind Model Checkpoints Saved? File Paths and Management Guide

MiniMind model checkpoints are saved to a checkpoints directory at the repository root by default, managed through the lm_checkpoint utility function in trainer/trainer_utils.py.

The MiniMind repository (jingyaogong/minimind) provides a streamlined checkpointing system for its compact language models. Knowing exactly where MiniMind model checkpoints are stored and how to manipulate their paths is critical for resuming interrupted training, evaluating intermediate models, or migrating weights across environments.

Default Checkpoint Location in MiniMind

MiniMind defines its checkpoint storage location in the lm_checkpoint function inside trainer/trainer_utils.py. The default save_dir parameter is set to '../checkpoints':

def lm_checkpoint(lm_config, weight='full_sft', model=None,
                  optimizer=None, epoch=0, step=0,
                  wandb=None, save_dir='../checkpoints', **kwargs):

When you launch training from any script within the trainer/ directory (such as train_pretrain.py or train_spo.py), this relative path resolves to a folder named checkpoints/ that sits at the repository root. This design ensures consistency across all training pipelines—whether you are pre-training, running SFT, or applying RLHF via PPO.

All official training scripts invoke lm_checkpoint without overriding save_dir, so weights are automatically written to this single centralized location.

Checkpoint File Naming Convention

MiniMind generates two distinct files per checkpoint save operation, constructed using the weight prefix, model configuration, and MoE status:

  • Model weights only: {weight}_{hidden_size}{moe_path}.pth
  • Full training state: {weight}_{hidden_size}{moe_path}_resume.pth

The code that constructs these paths appears in trainer/trainer_utils.py:

moe_path = '_moe' if lm_config.use_moe else ''
ckp_path = f'{save_dir}/{weight}_{lm_config.hidden_size}{moe_path}.pth'
resume_path = f'{save_dir}/{weight}_{lm_config.hidden_size}{moe_path}_resume.pth'

For example, a 512-hidden-unit model saved during supervised fine-tuning produces:

  • checkpoints/full_sft_512.pth
  • checkpoints/full_sft_512_resume.pth

If the model uses Mixture of Experts (MoE), the filename includes the _moe suffix (e.g., full_sft_512_moe.pth).

How to Save and Load MiniMind Model Checkpoints

The lm_checkpoint function handles both persistence and restoration. When model is provided, it saves; when model=None, it loads the resume file.

Saving Checkpoints During Training

To write a checkpoint during training, call lm_checkpoint with the model and optimizer objects:

from trainer.trainer_utils import lm_checkpoint

# Inside your training loop

lm_checkpoint(
    lm_config=config,
    weight='spo',
    model=model,
    optimizer=optimizer,
    epoch=current_epoch,
    step=global_step,
    wandb=wandb_run,  # Optional: logs checkpoint to WandB

    save_dir='../checkpoints'  # Default location

)

This creates both the .pth file (containing half-precision model weights) and the _resume.pth file (containing optimizer state, epoch count, and step counter).

Resuming Training from a Checkpoint

To resume, invoke lm_checkpoint without passing a model. The function detects model=None and returns the deserialized state dictionary:

from trainer.trainer_utils import lm_checkpoint

# Load the resume file

ckpt_data = lm_checkpoint(
    lm_config=config,
    weight='spo',
    save_dir='../checkpoints'
)

if ckpt_data is not None:
    model.load_state_dict(ckpt_data['model'])
    optimizer.load_state_dict(ckpt_data['optimizer'])
    start_epoch = ckpt_data['epoch']
    start_step = ckpt_data['step']

The function specifically checks for the existence of resume_path using os.path.exists(resume_path) before loading with torch.load(resume_path, map_location='cpu').

Customizing the Checkpoint Directory

You can override the default location by passing a custom save_dir argument:


# Save to a custom absolute or relative path

lm_checkpoint(
    lm_config=config,
    weight='full_sft',
    model=model,
    optimizer=optimizer,
    epoch=epoch,
    step=step,
    save_dir='/path/to/my_experiment_checkpoints'
)

This flexibility allows you to organize checkpoints by experiment date, hyperparameter set, or storage volume without modifying the source code in trainer_utils.py.

What Gets Stored in MiniMind Checkpoints

Understanding the difference between the two generated files helps you choose the right one for your use case:

  • .pth files: Contain only the model's state_dict (saved in half-precision via state_dict_half). These are suitable for inference or weight transfer but cannot resume training.
  • _resume.pth files: Contain a dictionary with keys for 'model', 'optimizer', 'epoch', and 'step'. These are required to resume training exactly where it left off, including optimizer momentum and learning rate schedule state.

Summary

  • MiniMind stores checkpoints in checkpoints/ at the repository root by default, defined by save_dir='../checkpoints' in trainer/trainer_utils.py.
  • The lm_checkpoint function in trainer/trainer_utils.py manages all checkpoint I/O for scripts like train_pretrain.py, train_spo.py, and train_ppo.py.
  • Each save operation produces two files: a .pth file (weights only) and a _resume.pth file (full training state).
  • File names follow the pattern {weight}_{hidden_size}{_moe}.pth, where weight is the training stage identifier (e.g., pretrain, full_sft) and _moe is appended for Mixture of Experts models.
  • Pass a custom save_dir argument to lm_checkpoint to change the storage location without altering the library code.

Frequently Asked Questions

What is the default directory for MiniMind model checkpoints?

By default, MiniMind saves checkpoints to a directory named checkpoints located at the repository root. This is determined by the save_dir='../checkpoints' default parameter in the lm_checkpoint function within trainer/trainer_utils.py. When running training scripts from the trainer/ folder, the relative path ../checkpoints resolves correctly to the root level.

How do I resume training from a MiniMind checkpoint?

To resume training, call lm_checkpoint with model=None and the same weight identifier used during saving. The function loads the *_resume.pth file (not the regular .pth file) and returns a dictionary containing the model state, optimizer state, epoch, and step count. You must manually load these into your model and optimizer objects before continuing the training loop.

Can I change where MiniMind saves its checkpoints?

Yes. Simply pass a different path string to the save_dir parameter when calling lm_checkpoint from your training script. For example, setting save_dir='./experiment_1/checkpoints' directs all output to that custom folder. This override works in all training scripts including train_pretrain.py, train_full_sft.py, and train_lora.py.

What is the difference between .pth and _resume.pth files in MiniMind?

The .pth file contains only the model weights (in half-precision), making it ideal for inference or model distribution. The _resume.pth file contains the complete training state including optimizer statistics, current epoch, and global step count, which is necessary to resume training without loss of progress. Always use the _resume.pth file when continuing training, and use the .pth file for standalone model evaluation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →