MiniMind Full LLM Training Pipeline: SFT, LoRA, DPO, and RLHF Explained
MiniMind provides a complete end-to-end LLM training pipeline that includes supervised fine-tuning (SFT), LoRA parameter-efficient fine-tuning, Direct Preference Optimization (DPO), and RLHF via Proximal Policy Optimization (PPO).
The jingyaogong/minimind repository implements a lightweight yet comprehensive training stack for large language models. Whether you need full-model supervised fine-tuning or advanced preference alignment, the MiniMind training pipeline covers every major stage of modern LLM development with dedicated scripts and shared infrastructure.
Supervised Fine-Tuning (SFT) Implementation
MiniMind implements full-model supervised fine-tuning in trainer/train_full_sft.py. This script loads a base model via init_model from model/model_minimind.py, constructs an SFTDataset for tokenized instruction data, and executes a standard cross-entropy training loop.
The SFT trainer supports production-grade features including gradient accumulation, automatic mixed-precision training via torch.cuda.amp.autocast, and distributed data parallel (DDP) initialization through utilities in trainer/trainer_utils.py. Checkpoints are saved periodically based on the --save_interval parameter.
python trainer/train_full_sft.py \
--data_path data/sft_dataset.jsonl \
--save_dir out/sft \
--epochs 3 \
--batch_size 32 \
--learning_rate 1e-6 \
--use_wandb
LoRA Parameter-Efficient Fine-Tuning
For scenarios requiring efficient adaptation, trainer/train_lora.py implements Low-Rank Adaptation (LoRA) fine-tuning. The script calls apply_lora from model/model_lora.py to inject trainable adapter layers into the frozen base model, then explicitly freezes all non-LoRA parameters by checking for 'lora' in parameter names.
This approach updates only the low-rank adapter weights while preserving the base model, significantly reducing memory requirements. The save_lora function stores adapter weights separately as {lora_name}_{hidden}.pth files for later merging or inference.
python trainer/train_lora.py \
--data_path data/lora_identity.jsonl \
--save_dir out/lora \
--lora_name my_lora \
--epochs 10 \
--learning_rate 1e-4 \
--use_moe 0
model, tokenizer = init_model(lm_config, args.from_weight, device=args.device)
apply_lora(model)
# Freeze all parameters except LoRA adapters
for name, p in model.named_parameters():
p.requires_grad = 'lora' in name
Direct Preference Optimization (DPO)
MiniMind supports preference learning without reward models through trainer/train_dpo.py. This script implements Direct Preference Optimization using paired chosen and rejected responses loaded via DPODataset from dataset/lm_dataset.py.
The training loop maintains a frozen reference model (ref_model) alongside the trainable policy model. For each batch, the script computes log probabilities for both models and optimizes the DPO loss defined in dpo_loss, using the --beta parameter to control the divergence penalty temperature.
python trainer/train_dpo.py \
--data_path data/dpo_pairs.jsonl \
--save_dir out/dpo \
--epochs 2 \
--learning_rate 4e-8 \
--beta 0.1
ref_outputs = ref_model(x) # Frozen reference model
policy_outputs = model(x) # Learnable policy model
loss = dpo_loss(ref_log_probs, policy_log_probs, mask, beta=args.beta)
RLHF via Proximal Policy Optimization (PPO)
The repository implements full RLHF training in trainer/train_ppo.py using Proximal Policy Optimization. This script orchestrates three distinct models: an actor (the policy being trained), a critic (value function derived from MiniMindForCausalLM), and a reward model (any Hugging Face compatible model such as internlm2-1_8b-reward).
The PPO loop calculates advantages by comparing reward model outputs against critic value estimates, applies a KL penalty to prevent policy drift, and optimizes the clipped surrogate objective. The script supports specialized "reasoning" models that incorporate format-based rewards via the --reasoning flag.
python trainer/train_ppo.py \
--data_path data/rlaif-mini.jsonl \
--save_dir out/ppo \
--epochs 1 \
--learning_rate 8e-8 \
--critic_learning_rate 8e-8 \
--reward_model_path path/to/reward-model \
--reasoning 1
responses = actor_model.generate(...)
rewards = calculate_rewards(prompts, responses, reward_model, reward_tokenizer)
values = critic_model(input_ids=gen_out, attention_mask=mask)
advantages = rewards - values.detach()
# PPO surrogate loss
policy_loss = -torch.min(ratio * advantages,
torch.clamp(ratio, 1 - eps, 1 + eps) * advantages).mean()
Shared Training Infrastructure
All training stages share common infrastructure defined in trainer/trainer_utils.py and model/model_minimind.py. The MiniMindConfig class provides unified hyperparameter management, while the trainer utilities handle learning rate scheduling (cosine with warmup), Weights & Biases integration via SwanLab, deterministic seed control, and DDP initialization.
This modular architecture ensures that optimization techniques, checkpointing logic, and logging conventions remain consistent across SFT, LoRA, DPO, and PPO training stages.
Summary
- Full SFT support:
trainer/train_full_sft.pyprovides distributed, mixed-precision supervised fine-tuning withSFTDatasetand gradient accumulation. - LoRA fine-tuning:
trainer/train_lora.pyandmodel/model_lora.pyenable parameter-efficient training viaapply_loraandsave_lorautilities. - DPO alignment:
trainer/train_dpo.pyimplements reference-model-based preference optimization using thedpo_lossfunction. - RLHF with PPO:
trainer/train_ppo.pyoffers complete actor-critic training with external reward model support andcalculate_rewardsintegration. - Unified infrastructure: Shared
trainer/trainer_utils.pyandMiniMindConfigprovide consistent distributed training, checkpointing, and logging across all pipeline stages.
Frequently Asked Questions
Does MiniMind support distributed training for the full LLM training pipeline?
Yes. According to the MiniMind source code, all training scripts—including train_full_sft.py, train_lora.py, train_dpo.py, and train_ppo.py—import DDP initialization and distributed sampling utilities from trainer/trainer_utils.py. The scripts support multi-GPU training via PyTorch DDP and include automatic mixed-precision support through torch.cuda.amp.autocast.
What is the difference between MiniMind's DPO and RLHF implementations?
MiniMind's DPO (train_dpo.py) performs offline preference learning using a frozen reference model and paired chosen/rejected datasets without requiring a separate reward model. The RLHF implementation (train_ppo.py) uses online Proximal Policy Optimization with three distinct models: an actor (policy), a critic (value function), and an external reward model loaded from Hugging Face (e.g., internlm2-1_8b-reward), enabling more complex advantage estimation through the calculate_rewards function.
Can LoRA adapters be used with DPO or RLHF training in MiniMind?
While MiniMind provides separate dedicated scripts for each stage, the modular architecture allows sequential composition. You first train LoRA adapters using train_lora.py, which saves weights via save_lora in model/model_lora.py. These adapter weights can then be loaded into the policy model before initiating DPO training in train_dpo.py or PPO training in train_ppo.py, effectively combining parameter-efficient fine-tuning with preference alignment.
Which reward models are compatible with MiniMind's RLHF pipeline?
The PPO script (train_ppo.py) accepts any Hugging Face Transformers-compatible reward model via the --reward_model_path argument. The implementation calls calculate_rewards to compute scores from the reward model's logits, making it compatible with standard classification-style reward models like internlm2-1_8b-reward or custom-trained MiniMind-based reward models.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →