How to Start Pre‑Training a MiniMind Model: Complete Guide

To start pre-training MiniMind, run the train_pretrain.py script with your data path and hyperparameters, which orchestrates distributed training, mixed-precision optimization, and automatic checkpointing via the modular components defined in model/model_minimind.py and trainer/trainer_utils.py.

MiniMind is a compact, open-source large language model (LLM) implemented in PyTorch that supports training from scratch or continuing from existing checkpoints. The pre-training pipeline in the jingyaogong/minimind repository provides a streamlined interface for causal language modeling using transformer architectures with optional Mixture of Experts (MoE) layers. This guide explains how to initialize and execute pre-training using the official scripts and core library components.

Core Pre‑Training Architecture

The pre-training system consists of five interconnected components that handle configuration, data processing, and distributed optimization:

  • MiniMindConfig – Defined in model/model_minimind.py (lines 8‑27), this class specifies model dimensions, MoE settings, RoPE parameters, and architectural hyperparameters required to instantiate the transformer.
  • MiniMindForCausalLM – The main model class (starting at line 426 in model/model_minimind.py) implements the forward pass, producing logits and auxiliary loss during pre-training.
  • PreTrainedTokenizerFast – Loaded via AutoTokenizer.from_pretrained in trainer/trainer_utils.py (lines 19‑22), handling BOS/EOS token insertion and vocabulary mapping.
  • PretrainDataset – Located in dataset/lm_dataset.py (lines 31‑48), this dataset class reads JSONL files, tokenizes raw text, pads sequences to fixed length, and masks padding tokens with -100 for loss computation.
  • train_pretrain.py – The main orchestration script that manages argument parsing, distributed setup, optimization loops, and checkpoint persistence.

Step‑by‑Step Pre‑Training Execution

CLI Arguments and Configuration

The entry point accepts comprehensive hyperparameters to control model architecture and training dynamics. In trainer/train_pretrain.py (lines 73‑82), the argument parser initializes default values for batch size, learning rate, and model dimensions:

parser = argparse.ArgumentParser(description="MiniMind Pretraining")
parser.add_argument("--batch_size", type=int, default=32, help="batch size")
parser.add_argument("--learning_rate", type=float, default=5e-4, help="initial LR")
parser.add_argument("--hidden_size", default=512, type=int, help="model hidden dimension")
parser.add_argument("--use_moe", default=0, type=int, choices=[0,1], help="enable MoE")
parser.add_argument("--data_path", type=str, default="../dataset/pretrain_hq.jsonl", help="pre‑training data")

Distributed Training Initialization

For multi-GPU setups, the script automatically detects environment variables set by torchrun. The init_distributed_mode() helper configures process groups and assigns devices:

local_rank = init_distributed_mode()
if dist.is_initialized(): args.device = f"cuda:{local_rank}"

This initialization occurs at lines 99‑102 in trainer/train_pretrain.py, enabling seamless Distributed Data Parallel (DDP) training without code modifications.

Model and Tokenizer Instantiation

The training script creates a MiniMindConfig instance and optionally loads existing weights for resumption. At lines 104‑108, the configuration object is constructed from CLI arguments:

lm_config = MiniMindConfig(
    hidden_size=args.hidden_size,
    num_hidden_layers=args.num_hidden_layers,
    use_moe=bool(args.use_moe)
)
ckp_data = lm_checkpoint(lm_config, weight=args.save_weight,
                        save_dir='../checkpoints') if args.from_resume==1 else None

The init_model utility in trainer/trainer_utils.py (lines 19‑31) then instantiates the model and tokenizer, moving parameters to the target device and reporting trainable parameter counts.

Dataset Preparation and Sampling

Raw text data must be formatted as JSONL with {"text": "..."} structures per line. The PretrainDataset class handles tokenization and batching, while a custom SkipBatchSampler enables step-level resumption:

train_ds = PretrainDataset(args.data_path, tokenizer, max_length=args.max_seq_len)
train_sampler = DistributedSampler(train_ds) if dist.is_initialized() else None
batch_sampler = SkipBatchSampler(train_sampler or indices,
                                args.batch_size, skip)
loader = DataLoader(train_ds, batch_sampler=batch_sampler,
                  num_workers=args.num_workers, pin_memory=True)

This implementation appears at lines 128‑154 in trainer/train_pretrain.py, supporting both single-process and distributed sampling modes.

Optimizer and Mixed Precision Setup

The script configures AdamW optimization with automatic mixed precision (AMP) using bfloat16 by default (lines 110‑113 and 130‑132):

dtype = torch.bfloat16 if args.dtype == "bfloat16" else torch.float16
autocast_ctx = nullcontext() if device_type == "cpu" else torch.cuda.amp.autocast(dtype=dtype)
scaler = torch.cuda.amp.GradScaler(enabled=(args.dtype == 'float16'))
optimizer = optim.AdamW(model.parameters(), lr=args.learning_rate)

Training Loop and Checkpointing

The main training epoch executes forward passes, calculates combined loss (including MoE auxiliary losses), and performs gradient scaling. At lines 143‑146, the model wraps with DDP while excluding rotary embedding buffers from synchronization:

if dist.is_initialized():
    model._ddp_params_and_buffers_to_ignore = {"freqs_cos", "freqs_sin"}
    model = DistributedDataParallel(model, device_ids=[local_rank])

The loop (lines 148‑158) handles gradient accumulation, optimizer stepping at specified intervals, and periodic checkpoint saving to both full-weight files (pretrain_512.pth) and resume states (pretrain_512_resume.pth).

Practical Code Examples

Single‑GPU Training Command

Execute pre-training on a single device using the following bash command, specifying your JSONL data path and model dimensions:

python trainer/train_pretrain.py \
  --data_path data/pretrain_hq.jsonl \
  --save_dir checkpoints \
  --save_weight pretrain \
  --epochs 2 \
  --batch_size 64 \
  --learning_rate 3e-4 \
  --hidden_size 512 \
  --num_hidden_layers 8 \
  --max_seq_len 340 \
  --use_moe 0 \
  --use_wandb

Multi‑GPU Distributed Launch

For scaled training across multiple GPUs, use torchrun to automatically set RANK, LOCAL_RANK, and WORLD_SIZE environment variables:

torchrun --nproc_per_node=4 trainer/train_pretrain.py \
  --data_path data/pretrain_hq.jsonl \
  --epochs 4 \
  --batch_size 32 \
  --learning_rate 5e-4 \
  --use_moe 1 \
  --from_resume 1

The --nproc_per_node flag should match your available GPU count.

Minimal Python API Usage

For programmatic control in notebooks or custom pipelines, import the core components directly:

from trainer.trainer_utils import init_model
from dataset.lm_dataset import PretrainDataset
from model.model_minimind import MiniMindConfig
from torch.utils.data import DataLoader
import torch

# Configure architecture

cfg = MiniMindConfig(hidden_size=512, num_hidden_layers=8, use_moe=False)

# Initialize components

model, tokenizer = init_model(cfg, from_weight='none', device='cpu')
train_ds = PretrainDataset('data/pretrain_hq.jsonl', tokenizer, max_length=340)
loader = DataLoader(train_ds, batch_size=32, shuffle=True)

# Single training step

opt = torch.optim.AdamW(model.parameters(), lr=5e-4)
model.train()
for input_ids, labels in loader:
    opt.zero_grad()
    out = model(input_ids, labels=labels)
    loss = (out.loss + out.aux_loss) / 1  # Adjust for gradient accumulation

    loss.backward()
    opt.step()
    print(f'loss={loss.item():.4f}')
    break

Resuming From Checkpoints

To continue training from an interrupted session, ensure your checkpoint files exist in the save directory and enable the resume flag:

python trainer/train_pretrain.py \
  --from_resume 1 \
  --save_weight pretrain \
  --hidden_size 512 \
  --use_moe 0 \
  --epochs 5

The script automatically detects pretrain_512_resume.pth and restores model weights, optimizer states, and step counters.

Summary

  • Primary Entry Point: Use trainer/train_pretrain.py as the unified interface for launching pre-training jobs with configurable hyperparameters.
  • Data Format: Prepare training data as JSONL files containing {"text": "..."} records, processed by PretrainDataset in dataset/lm_dataset.py.
  • Distributed Support: Launch multi-GPU training via torchrun without code changes; the script handles DDP initialization and sampler coordination internally.
  • Checkpoint Management: The lm_checkpoint utility in trainer/trainer_utils.py manages both full model weights and resumable training states.
  • MoE Capability: Enable Mixture of Experts by setting --use_moe 1 in the configuration, which modifies the MiniMindConfig instantiation at runtime.

Frequently Asked Questions

What data format does MiniMind pre‑training require?

MiniMind expects pre-training data in JSONL format where each line contains a JSON object with a "text" field containing raw text content. The PretrainDataset class in dataset/lm_dataset.py reads these files, tokenizes the content using the MiniMind tokenizer, and yields tensors with BOS/EOS tokens added and padding masked with -100.

How do I enable Mixture of Experts (MoE) during pre‑training?

Set the --use_moe 1 command-line argument when running train_pretrain.py. This boolean flag is passed to MiniMindConfig (defined in model/model_minimind.py), which instantiates the transformer with sparse expert layers instead of standard feed-forward networks. The MoE auxiliary loss is automatically added to the total loss during the forward pass.

Can I resume training if my process is interrupted?

Yes. Specify --from_resume 1 when launching the training script. The system checks for existing checkpoint files in your --save_dir directory (specifically {save_weight}_resume.pth), loads the model state, optimizer state, and training step counter, and continues from the exact batch where training stopped using the SkipBatchSampler implementation.

What hardware requirements are needed for MiniMind pre‑training?

MiniMind supports training on CPU (for debugging) and CUDA GPUs. For practical pre-training, a CUDA-capable GPU is recommended. The repository supports mixed precision training (bfloat16 or float16) to reduce memory consumption, and the DDP implementation allows scaling across multiple GPUs by adjusting --nproc_per_node in the torchrun command.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →