# How to Train Your Own Small Language Model from Scratch with MiniMind

> Train your own small language model from scratch with MiniMind. Follow sequential steps from tokenizer prep to RL alignment using the jingyaogong/minimind PyTorch components.

- Repository: [jingyaogong/minimind](https://github.com/jingyaogong/minimind)
- Tags: how-to-guide
- Published: 2026-03-24

---

**You can train a complete small language model from scratch using MiniMind by running the provided training scripts in sequence: tokenizer preparation (optional), pre-training on raw text, supervised fine-tuning (SFT), and optional RL alignment, all supported by modular PyTorch components in the jingyaogong/minimind repository.**

MiniMind is a lightweight, open-source Transformer implementation designed for training compact language models on modest hardware. The repository provides a full-stack training framework that includes a custom tokenizer, configurable model architecture with optional Mixture-of-Experts (MoE), and scripts for every stage from pre-training to deployment.

## Overview of the MiniMind Training Pipeline

MiniMind organizes the training lifecycle into distinct stages, each handled by a dedicated script in the `trainer/` directory. The pipeline supports both full-parameter training and parameter-efficient methods like LoRA.

The complete workflow includes:

- **Tokenizer preparation** – Train a Byte-Level BPE tokenizer tailored to your domain
- **Pre-training** – Language modeling objective on large raw-text corpora
- **Supervised fine-tuning** – Instruction-following training using ChatML formatted data
- **Parameter-efficient tuning** – LoRA adapters for low-resource fine-tuning
- **Reinforcement learning** – PPO, DPO, or GRPO alignment with preference data
- **Distillation** – Transfer knowledge from large teacher models to MiniMind students

## Architecture and Key Components

Understanding the core codebase helps you customize training effectively. The implementation resides in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) and follows standard Transformer conventions with modern optimizations.

### Model Definition (`MiniMindConfig` and `MiniMindModel`)

The configuration class `MiniMindConfig` inherits from `transformers.PretrainedConfig` and declares all hyperparameters including optional MoE settings. The network architecture in `MiniMindModel` stacks multiple `MiniMindBlock` layers, each containing:

- **RMSNorm** for stable layer normalization
- Rotary positional embeddings (`precompute_freqs_cis`)
- Standard **Attention** mechanism with optional FlashAttention support
- **FeedForward** or **MOEFeedForward** when `use_moe=True`

The causal language modeling head shares weights with the token embedding matrix, reducing parameter count.

### Tokenizer System

MiniMind ships with a pre-trained 6,400-token BPE tokenizer stored in [`model/tokenizer.json`](https://github.com/jingyaogong/minimind/blob/main/model/tokenizer.json) and [`model/tokenizer_config.json`](https://github.com/jingyaogong/minimind/blob/main/model/tokenizer_config.json). For domain-specific vocabularies, [`trainer/train_tokenizer.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_tokenizer.py) uses the HuggingFace `tokenizers` library to rebuild the vocabulary and outputs the same JSON format expected by the training scripts.

### Dataset Classes

The [`dataset/lm_dataset.py`](https://github.com/jingyaogong/minimind/blob/main/dataset/lm_dataset.py) file provides PyTorch Dataset implementations for each training stage:

- `PretrainDataset` – Handles raw text with BOS/EOS tokens and label masking
- `SFTDataset` – Applies ChatML templates via `apply_chat_template` for instruction data
- `DPODataset` and `RLAIFDataset` – Structure preference pairs for RL training

All datasets use `datasets.load_dataset`, accepting JSONL files with specific schema requirements.

### Training Utilities

[`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py) contains essential infrastructure code including:

- `init_distributed_mode` for PyTorch DDP setup
- `get_lr` for learning rate scheduling
- `lm_checkpoint` for model saving and resumption
- `get_model_params` for counting trainable parameters

## Step-by-Step Training Guide

Follow these stages sequentially to build your model. All scripts support distributed training via the `RANK` and `LOCAL_RANK` environment variables.

### (Optional) Train a Custom Tokenizer

If your domain requires specialized vocabulary, train a custom tokenizer before pre-training:

```bash
python -m trainer.train_tokenizer \
    --data_path data/my_corpus.jsonl \
    --tokenizer_dir model/my_tokenizer \
    --vocab_size 6400

```

This generates [`tokenizer.json`](https://github.com/jingyaogong/minimind/blob/main/tokenizer.json) and [`tokenizer_config.json`](https://github.com/jingyaogong/minimind/blob/main/tokenizer_config.json) in the specified directory. Reference this directory in subsequent training scripts using `--tokenizer_path model/my_tokenizer`.

### Pre-training on Raw Text

Run [`trainer/train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_pretrain.py) to train the base model on a language modeling objective:

```bash
python -m trainer.train_pretrain \
    --save_dir out/checkpoints \
    --save_weight pretrain \
    --epochs 4 \
    --batch_size 32 \
    --learning_rate 5e-4 \
    --hidden_size 512 \
    --num_hidden_layers 8 \
    --max_seq_len 340 \
    --data_path data/pretrain_hq.jsonl \
    --use_wandb \
    --wandb_project MiniMind-Pretrain

```

The script constructs a `MiniMindConfig` from CLI arguments, loads the tokenizer, initializes a `PretrainDataset`, and executes the training loop. Checkpoints are saved as `pretrain_512.pth` (or your specified hidden size) at intervals defined by `--save_interval`.

### Supervised Fine-Tuning (SFT)

Convert your pre-trained checkpoint into a chat model using [`trainer/train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_full_sft.py):

```bash
python -m trainer.train_full_sft \
    --save_dir out/checkpoints_sft \
    --save_weight sft \
    --epochs 2 \
    --batch_size 16 \
    --learning_rate 2e-4 \
    --hidden_size 512 \
    --data_path data/sft_dataset.jsonl \
    --use_lora 0

```

The `SFTDataset` class processes conversation data in ChatML format, generating labels that mask the system and user turns while supervising only the assistant responses. Set `--use_lora 1` to switch to adapter-based training (see next section).

### Parameter-Efficient Fine-Tuning with LoRA

For hardware-constrained environments, [`trainer/train_lora.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_lora.py) inserts low-rank adapters into attention and feed-forward projections:

```bash
python -m trainer.train_lora \
    --save_dir out/checkpoints_lora \
    --save_weight lora \
    --epochs 3 \
    --batch_size 32 \
    --learning_rate 1e-4 \
    --hidden_size 512 \
    --lora_rank 8 \
    --data_path data/sft_dataset.jsonl

```

This freezes the base model weights and updates only the adapter parameters, producing checkpoints as small as 2 MB for a 6M-parameter MiniMind model.

### Reinforcement Learning Alignment

Align your model with human preferences using PPO or DPO. First train a reward model, then optimize the policy:

```bash

# Train reward model

python -m trainer.train_ppo \
    --stage reward \
    --save_dir out/ppo \
    --data_path data/ppo_data.jsonl \
    --epochs 1

# Train policy with PPO

python -m trainer.train_ppo \
    --stage policy \
    --save_dir out/ppo \
    --data_path data/ppo_data.jsonl \
    --epochs 2 \
    --kl_coef 0.1

```

Alternatively, use [`trainer/train_dpo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_dpo.py) for Direct Preference Optimization without a separate reward model, or [`trainer/train_grpo.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/train_grpo.py) for gradient-based RPO training. These scripts utilize `DPODataset` to process chosen versus rejected response pairs.

## Deploying Your Model

After training, serve your model via an OpenAI-compatible API or interactive web interface.

Start the API server using [`scripts/serve_openai_api.py`](https://github.com/jingyaogong/minimind/blob/main/scripts/serve_openai_api.py):

```bash
python -m scripts.serve_openai_api \
    --model_path out/checkpoints/pretrain_512.pth \
    --tokenizer_path model \
    --port 8000

```

Access the endpoint with any OpenAI client:

```python
import openai
openai.api_base = "http://localhost:8000/v1"
openai.api_key = "dummy"

resp = openai.ChatCompletion.create(
    model="mini-mind",
    messages=[{"role": "user", "content": "Hello, how are you?"}],
    max_tokens=128
)
print(resp["choices"][0]["message"]["content"])

```

For interactive demos, use [`scripts/web_demo.py`](https://github.com/jingyaogong/minimind/blob/main/scripts/web_demo.py) to launch a Gradio interface.

## Summary

- **MiniMind** provides a complete training stack in `jingyaogong/minimind` with scripts for tokenizer training, pre-training, SFT, LoRA, and RL alignment.
- The architecture in [`model/model_minimind.py`](https://github.com/jingyaogong/minimind/blob/main/model/model_minimind.py) supports configurable sizes, FlashAttention, and optional MoE layers via `MiniMindConfig`.
- **Pre-training** uses [`train_pretrain.py`](https://github.com/jingyaogong/minimind/blob/main/train_pretrain.py) with `PretrainDataset` for raw text corpora.
- **Supervised fine-tuning** applies ChatML formatting through [`train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/train_full_sft.py) and `SFTDataset`.
- **LoRA training** via [`train_lora.py`](https://github.com/jingyaogong/minimind/blob/main/train_lora.py) enables efficient fine-tuning with minimal GPU memory.
- **Deployment** options include [`scripts/serve_openai_api.py`](https://github.com/jingyaogong/minimind/blob/main/scripts/serve_openai_api.py) for API serving and [`scripts/web_demo.py`](https://github.com/jingyaogong/minimind/blob/main/scripts/web_demo.py) for web demos.

## Frequently Asked Questions

### What hardware is required to train a MiniMind model from scratch?

You can train small MiniMind variants (512 hidden size, 8 layers) on a single consumer GPU with 8-12 GB VRAM. For larger configurations or full pre-training on extensive corpora, multi-GPU setups using PyTorch DistributedDataParallel (via `RANK` environment variables) are supported through [`trainer/trainer_utils.py`](https://github.com/jingyaogong/minimind/blob/main/trainer/trainer_utils.py).

### How do I format data for supervised fine-tuning?

SFT data should be JSONL files where each line contains a `conversations` list with ChatML-formatted messages (system, user, assistant roles). The `SFTDataset` class in [`dataset/lm_dataset.py`](https://github.com/jingyaogong/minimind/blob/main/dataset/lm_dataset.py) automatically applies the chat template using the tokenizer's `apply_chat_template` method and creates labels that only supervise the assistant's responses.

### Can I use a custom tokenizer instead of the default 6,400-token vocabulary?

Yes. Run `python -m trainer.train_tokenizer` with your domain-specific corpus to generate a new BPE tokenizer. The script outputs [`tokenizer.json`](https://github.com/jingyaogong/minimind/blob/main/tokenizer.json) and [`tokenizer_config.json`](https://github.com/jingyaogong/minimind/blob/main/tokenizer_config.json) compatible with all training scripts. Specify your custom tokenizer path using the `--tokenizer_path` argument in any training script.

### What is the difference between full SFT and LoRA training?

Full SFT ([`train_full_sft.py`](https://github.com/jingyaogong/minimind/blob/main/train_full_sft.py) with `--use_lora 0`) updates all model parameters and produces larger checkpoints (hundreds of MB). LoRA training ([`train_lora.py`](https://github.com/jingyaogong/minimind/blob/main/train_lora.py)) inserts low-rank adapter matrices into the attention and feed-forward layers, freezing base weights and producing small checkpoints (~2 MB) suitable for storage-constrained environments or rapid task switching.