How to Start Pre‑Training a MiniMind Model: Complete Guide
To start pre-training MiniMind, run the train_pretrain.py script with your data path and hyperparameters, which orchestrates distributed training, mixed-precision optimization, and automatic checkpointing via the modular components defined in model/model_minimind.py and trainer/trainer_utils.py.
MiniMind is a compact, open-source large language model (LLM) implemented in PyTorch that supports training from scratch or continuing from existing checkpoints. The pre-training pipeline in the jingyaogong/minimind repository provides a streamlined interface for causal language modeling using transformer architectures with optional Mixture of Experts (MoE) layers. This guide explains how to initialize and execute pre-training using the official scripts and core library components.
Core Pre‑Training Architecture
The pre-training system consists of five interconnected components that handle configuration, data processing, and distributed optimization:
MiniMindConfig– Defined inmodel/model_minimind.py(lines 8‑27), this class specifies model dimensions, MoE settings, RoPE parameters, and architectural hyperparameters required to instantiate the transformer.MiniMindForCausalLM– The main model class (starting at line 426 inmodel/model_minimind.py) implements the forward pass, producing logits and auxiliary loss during pre-training.PreTrainedTokenizerFast– Loaded viaAutoTokenizer.from_pretrainedintrainer/trainer_utils.py(lines 19‑22), handling BOS/EOS token insertion and vocabulary mapping.PretrainDataset– Located indataset/lm_dataset.py(lines 31‑48), this dataset class reads JSONL files, tokenizes raw text, pads sequences to fixed length, and masks padding tokens with-100for loss computation.train_pretrain.py– The main orchestration script that manages argument parsing, distributed setup, optimization loops, and checkpoint persistence.
Step‑by‑Step Pre‑Training Execution
CLI Arguments and Configuration
The entry point accepts comprehensive hyperparameters to control model architecture and training dynamics. In trainer/train_pretrain.py (lines 73‑82), the argument parser initializes default values for batch size, learning rate, and model dimensions:
parser = argparse.ArgumentParser(description="MiniMind Pretraining")
parser.add_argument("--batch_size", type=int, default=32, help="batch size")
parser.add_argument("--learning_rate", type=float, default=5e-4, help="initial LR")
parser.add_argument("--hidden_size", default=512, type=int, help="model hidden dimension")
parser.add_argument("--use_moe", default=0, type=int, choices=[0,1], help="enable MoE")
parser.add_argument("--data_path", type=str, default="../dataset/pretrain_hq.jsonl", help="pre‑training data")
Distributed Training Initialization
For multi-GPU setups, the script automatically detects environment variables set by torchrun. The init_distributed_mode() helper configures process groups and assigns devices:
local_rank = init_distributed_mode()
if dist.is_initialized(): args.device = f"cuda:{local_rank}"
This initialization occurs at lines 99‑102 in trainer/train_pretrain.py, enabling seamless Distributed Data Parallel (DDP) training without code modifications.
Model and Tokenizer Instantiation
The training script creates a MiniMindConfig instance and optionally loads existing weights for resumption. At lines 104‑108, the configuration object is constructed from CLI arguments:
lm_config = MiniMindConfig(
hidden_size=args.hidden_size,
num_hidden_layers=args.num_hidden_layers,
use_moe=bool(args.use_moe)
)
ckp_data = lm_checkpoint(lm_config, weight=args.save_weight,
save_dir='../checkpoints') if args.from_resume==1 else None
The init_model utility in trainer/trainer_utils.py (lines 19‑31) then instantiates the model and tokenizer, moving parameters to the target device and reporting trainable parameter counts.
Dataset Preparation and Sampling
Raw text data must be formatted as JSONL with {"text": "..."} structures per line. The PretrainDataset class handles tokenization and batching, while a custom SkipBatchSampler enables step-level resumption:
train_ds = PretrainDataset(args.data_path, tokenizer, max_length=args.max_seq_len)
train_sampler = DistributedSampler(train_ds) if dist.is_initialized() else None
batch_sampler = SkipBatchSampler(train_sampler or indices,
args.batch_size, skip)
loader = DataLoader(train_ds, batch_sampler=batch_sampler,
num_workers=args.num_workers, pin_memory=True)
This implementation appears at lines 128‑154 in trainer/train_pretrain.py, supporting both single-process and distributed sampling modes.
Optimizer and Mixed Precision Setup
The script configures AdamW optimization with automatic mixed precision (AMP) using bfloat16 by default (lines 110‑113 and 130‑132):
dtype = torch.bfloat16 if args.dtype == "bfloat16" else torch.float16
autocast_ctx = nullcontext() if device_type == "cpu" else torch.cuda.amp.autocast(dtype=dtype)
scaler = torch.cuda.amp.GradScaler(enabled=(args.dtype == 'float16'))
optimizer = optim.AdamW(model.parameters(), lr=args.learning_rate)
Training Loop and Checkpointing
The main training epoch executes forward passes, calculates combined loss (including MoE auxiliary losses), and performs gradient scaling. At lines 143‑146, the model wraps with DDP while excluding rotary embedding buffers from synchronization:
if dist.is_initialized():
model._ddp_params_and_buffers_to_ignore = {"freqs_cos", "freqs_sin"}
model = DistributedDataParallel(model, device_ids=[local_rank])
The loop (lines 148‑158) handles gradient accumulation, optimizer stepping at specified intervals, and periodic checkpoint saving to both full-weight files (pretrain_512.pth) and resume states (pretrain_512_resume.pth).
Practical Code Examples
Single‑GPU Training Command
Execute pre-training on a single device using the following bash command, specifying your JSONL data path and model dimensions:
python trainer/train_pretrain.py \
--data_path data/pretrain_hq.jsonl \
--save_dir checkpoints \
--save_weight pretrain \
--epochs 2 \
--batch_size 64 \
--learning_rate 3e-4 \
--hidden_size 512 \
--num_hidden_layers 8 \
--max_seq_len 340 \
--use_moe 0 \
--use_wandb
Multi‑GPU Distributed Launch
For scaled training across multiple GPUs, use torchrun to automatically set RANK, LOCAL_RANK, and WORLD_SIZE environment variables:
torchrun --nproc_per_node=4 trainer/train_pretrain.py \
--data_path data/pretrain_hq.jsonl \
--epochs 4 \
--batch_size 32 \
--learning_rate 5e-4 \
--use_moe 1 \
--from_resume 1
The --nproc_per_node flag should match your available GPU count.
Minimal Python API Usage
For programmatic control in notebooks or custom pipelines, import the core components directly:
from trainer.trainer_utils import init_model
from dataset.lm_dataset import PretrainDataset
from model.model_minimind import MiniMindConfig
from torch.utils.data import DataLoader
import torch
# Configure architecture
cfg = MiniMindConfig(hidden_size=512, num_hidden_layers=8, use_moe=False)
# Initialize components
model, tokenizer = init_model(cfg, from_weight='none', device='cpu')
train_ds = PretrainDataset('data/pretrain_hq.jsonl', tokenizer, max_length=340)
loader = DataLoader(train_ds, batch_size=32, shuffle=True)
# Single training step
opt = torch.optim.AdamW(model.parameters(), lr=5e-4)
model.train()
for input_ids, labels in loader:
opt.zero_grad()
out = model(input_ids, labels=labels)
loss = (out.loss + out.aux_loss) / 1 # Adjust for gradient accumulation
loss.backward()
opt.step()
print(f'loss={loss.item():.4f}')
break
Resuming From Checkpoints
To continue training from an interrupted session, ensure your checkpoint files exist in the save directory and enable the resume flag:
python trainer/train_pretrain.py \
--from_resume 1 \
--save_weight pretrain \
--hidden_size 512 \
--use_moe 0 \
--epochs 5
The script automatically detects pretrain_512_resume.pth and restores model weights, optimizer states, and step counters.
Summary
- Primary Entry Point: Use
trainer/train_pretrain.pyas the unified interface for launching pre-training jobs with configurable hyperparameters. - Data Format: Prepare training data as JSONL files containing
{"text": "..."}records, processed byPretrainDatasetindataset/lm_dataset.py. - Distributed Support: Launch multi-GPU training via
torchrunwithout code changes; the script handles DDP initialization and sampler coordination internally. - Checkpoint Management: The
lm_checkpointutility intrainer/trainer_utils.pymanages both full model weights and resumable training states. - MoE Capability: Enable Mixture of Experts by setting
--use_moe 1in the configuration, which modifies theMiniMindConfiginstantiation at runtime.
Frequently Asked Questions
What data format does MiniMind pre‑training require?
MiniMind expects pre-training data in JSONL format where each line contains a JSON object with a "text" field containing raw text content. The PretrainDataset class in dataset/lm_dataset.py reads these files, tokenizes the content using the MiniMind tokenizer, and yields tensors with BOS/EOS tokens added and padding masked with -100.
How do I enable Mixture of Experts (MoE) during pre‑training?
Set the --use_moe 1 command-line argument when running train_pretrain.py. This boolean flag is passed to MiniMindConfig (defined in model/model_minimind.py), which instantiates the transformer with sparse expert layers instead of standard feed-forward networks. The MoE auxiliary loss is automatically added to the total loss during the forward pass.
Can I resume training if my process is interrupted?
Yes. Specify --from_resume 1 when launching the training script. The system checks for existing checkpoint files in your --save_dir directory (specifically {save_weight}_resume.pth), loads the model state, optimizer state, and training step counter, and continues from the exact batch where training stopped using the SkipBatchSampler implementation.
What hardware requirements are needed for MiniMind pre‑training?
MiniMind supports training on CPU (for debugging) and CUDA GPUs. For practical pre-training, a CUDA-capable GPU is recommended. The repository supports mixed precision training (bfloat16 or float16) to reduce memory consumption, and the DDP implementation allows scaling across multiple GPUs by adjusting --nproc_per_node in the torchrun command.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →