# MiniMind Full LLM Training Pipeline: SFT, LoRA, DPO, and RLHF Explained

> Explore the MiniMind LLM training pipeline covering SFT, LoRA, DPO, and RLHF. Train powerful language models efficiently with this comprehensive toolkit.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: deep-dive
- Published: 2026-03-24

---

**MiniMind provides a complete end-to-end LLM training pipeline that includes supervised fine-tuning (SFT), LoRA parameter-efficient fine-tuning, Direct Preference Optimization (DPO), and RLHF via Proximal Policy Optimization (PPO).**

The jingyaogong/minimind repository implements a lightweight yet comprehensive training stack for large language models. Whether you need full-model supervised fine-tuning or advanced preference alignment, the MiniMind training pipeline covers every major stage of modern LLM development with dedicated scripts and shared infrastructure.

## Supervised Fine-Tuning (SFT) Implementation

MiniMind implements full-model supervised fine-tuning in [`trainer/train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_full_sft.py). This script loads a base model via `init_model` from [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py), constructs an `SFTDataset` for tokenized instruction data, and executes a standard cross-entropy training loop.

The SFT trainer supports production-grade features including gradient accumulation, automatic mixed-precision training via `torch.cuda.amp.autocast`, and distributed data parallel (DDP) initialization through utilities in [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py). Checkpoints are saved periodically based on the `--save_interval` parameter.

```bash
python trainer/train_full_sft.py \
  --data_path data/sft_dataset.jsonl \
  --save_dir out/sft \
  --epochs 3 \
  --batch_size 32 \
  --learning_rate 1e-6 \
  --use_wandb

```

## LoRA Parameter-Efficient Fine-Tuning

For scenarios requiring efficient adaptation, [`trainer/train_lora.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_lora.py) implements Low-Rank Adaptation (LoRA) fine-tuning. The script calls `apply_lora` from [`model/model_lora.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_lora.py) to inject trainable adapter layers into the frozen base model, then explicitly freezes all non-LoRA parameters by checking for `'lora'` in parameter names.

This approach updates only the low-rank adapter weights while preserving the base model, significantly reducing memory requirements. The `save_lora` function stores adapter weights separately as `{lora_name}_{hidden}.pth` files for later merging or inference.

```bash
python trainer/train_lora.py \
  --data_path data/lora_identity.jsonl \
  --save_dir out/lora \
  --lora_name my_lora \
  --epochs 10 \
  --learning_rate 1e-4 \
  --use_moe 0

```

```python
model, tokenizer = init_model(lm_config, args.from_weight, device=args.device)
apply_lora(model)

# Freeze all parameters except LoRA adapters

for name, p in model.named_parameters():
    p.requires_grad = 'lora' in name

```

## Direct Preference Optimization (DPO)

MiniMind supports preference learning without reward models through [`trainer/train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_dpo.py). This script implements Direct Preference Optimization using paired chosen and rejected responses loaded via `DPODataset` from [`dataset/lm_dataset.py`](https://github.com/jingyaogong/minimind/blob/main/dataset/lm_dataset.py).

The training loop maintains a frozen **reference model** (`ref_model`) alongside the trainable policy model. For each batch, the script computes log probabilities for both models and optimizes the DPO loss defined in `dpo_loss`, using the `--beta` parameter to control the divergence penalty temperature.

```bash
python trainer/train_dpo.py \
  --data_path data/dpo_pairs.jsonl \
  --save_dir out/dpo \
  --epochs 2 \
  --learning_rate 4e-8 \
  --beta 0.1

```

```python
ref_outputs = ref_model(x)                     # Frozen reference model

policy_outputs = model(x)                      # Learnable policy model

loss = dpo_loss(ref_log_probs, policy_log_probs, mask, beta=args.beta)

```

## RLHF via Proximal Policy Optimization (PPO)

The repository implements full RLHF training in [`trainer/train_ppo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_ppo.py) using Proximal Policy Optimization. This script orchestrates three distinct models: an **actor** (the policy being trained), a **critic** (value function derived from `MiniMindForCausalLM`), and a **reward model** (any Hugging Face compatible model such as `internlm2-1_8b-reward`).

The PPO loop calculates advantages by comparing reward model outputs against critic value estimates, applies a KL penalty to prevent policy drift, and optimizes the clipped surrogate objective. The script supports specialized "reasoning" models that incorporate format-based rewards via the `--reasoning` flag.

```bash
python trainer/train_ppo.py \
  --data_path data/rlaif-mini.jsonl \
  --save_dir out/ppo \
  --epochs 1 \
  --learning_rate 8e-8 \
  --critic_learning_rate 8e-8 \
  --reward_model_path path/to/reward-model \
  --reasoning 1

```

```python
responses = actor_model.generate(...)
rewards = calculate_rewards(prompts, responses, reward_model, reward_tokenizer)
values = critic_model(input_ids=gen_out, attention_mask=mask)
advantages = rewards - values.detach()

# PPO surrogate loss

policy_loss = -torch.min(ratio * advantages,
                         torch.clamp(ratio, 1 - eps, 1 + eps) * advantages).mean()

```

## Shared Training Infrastructure

All training stages share common infrastructure defined in [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) and [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py). The `MiniMindConfig` class provides unified hyperparameter management, while the trainer utilities handle learning rate scheduling (cosine with warmup), Weights & Biases integration via SwanLab, deterministic seed control, and DDP initialization.

This modular architecture ensures that optimization techniques, checkpointing logic, and logging conventions remain consistent across SFT, LoRA, DPO, and PPO training stages.

## Summary

- **Full SFT support**: [`trainer/train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_full_sft.py) provides distributed, mixed-precision supervised fine-tuning with `SFTDataset` and gradient accumulation.
- **LoRA fine-tuning**: [`trainer/train_lora.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_lora.py) and [`model/model_lora.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_lora.py) enable parameter-efficient training via `apply_lora` and `save_lora` utilities.
- **DPO alignment**: [`trainer/train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_dpo.py) implements reference-model-based preference optimization using the `dpo_loss` function.
- **RLHF with PPO**: [`trainer/train_ppo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_ppo.py) offers complete actor-critic training with external reward model support and `calculate_rewards` integration.
- **Unified infrastructure**: Shared [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) and `MiniMindConfig` provide consistent distributed training, checkpointing, and logging across all pipeline stages.

## Frequently Asked Questions

### Does MiniMind support distributed training for the full LLM training pipeline?

Yes. According to the MiniMind source code, all training scripts—including [`train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/train_full_sft.py), [`train_lora.py`](https://github.com/jingyaogong/minimind/blob/main/train_lora.py), [`train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/train_dpo.py), and [`train_ppo.py`](https://github.com/jingyaogong/minimind/blob/main/train_ppo.py)—import DDP initialization and distributed sampling utilities from [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py). The scripts support multi-GPU training via PyTorch DDP and include automatic mixed-precision support through `torch.cuda.amp.autocast`.

### What is the difference between MiniMind's DPO and RLHF implementations?

MiniMind's **DPO** ([`train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/train_dpo.py)) performs offline preference learning using a frozen reference model and paired chosen/rejected datasets without requiring a separate reward model. The **RLHF** implementation ([`train_ppo.py`](https://github.com/jingyaogong/minimind/blob/main/train_ppo.py)) uses online Proximal Policy Optimization with three distinct models: an actor (policy), a critic (value function), and an external reward model loaded from Hugging Face (e.g., `internlm2-1_8b-reward`), enabling more complex advantage estimation through the `calculate_rewards` function.

### Can LoRA adapters be used with DPO or RLHF training in MiniMind?

While MiniMind provides separate dedicated scripts for each stage, the modular architecture allows sequential composition. You first train LoRA adapters using [`train_lora.py`](https://github.com/jingyaogong/minimind/blob/main/train_lora.py), which saves weights via `save_lora` in [`model/model_lora.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_lora.py). These adapter weights can then be loaded into the policy model before initiating DPO training in [`train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/train_dpo.py) or PPO training in [`train_ppo.py`](https://github.com/jingyaogong/minimind/blob/main/train_ppo.py), effectively combining parameter-efficient fine-tuning with preference alignment.

### Which reward models are compatible with MiniMind's RLHF pipeline?

The PPO script ([`train_ppo.py`](https://github.com/jingyaogong/minimind/blob/main/train_ppo.py)) accepts any Hugging Face Transformers-compatible reward model via the `--reward_model_path` argument. The implementation calls `calculate_rewards` to compute scores from the reward model's logits, making it compatible with standard classification-style reward models like `internlm2-1_8b-reward` or custom-trained MiniMind-based reward models.