How to Train Your Own Small Language Model from Scratch with MiniMind
You can train a complete small language model from scratch using MiniMind by running the provided training scripts in sequence: tokenizer preparation (optional), pre-training on raw text, supervised fine-tuning (SFT), and optional RL alignment, all supported by modular PyTorch components in the jingyaogong/minimind repository.
MiniMind is a lightweight, open-source Transformer implementation designed for training compact language models on modest hardware. The repository provides a full-stack training framework that includes a custom tokenizer, configurable model architecture with optional Mixture-of-Experts (MoE), and scripts for every stage from pre-training to deployment.
Overview of the MiniMind Training Pipeline
MiniMind organizes the training lifecycle into distinct stages, each handled by a dedicated script in the trainer/ directory. The pipeline supports both full-parameter training and parameter-efficient methods like LoRA.
The complete workflow includes:
- Tokenizer preparation – Train a Byte-Level BPE tokenizer tailored to your domain
- Pre-training – Language modeling objective on large raw-text corpora
- Supervised fine-tuning – Instruction-following training using ChatML formatted data
- Parameter-efficient tuning – LoRA adapters for low-resource fine-tuning
- Reinforcement learning – PPO, DPO, or GRPO alignment with preference data
- Distillation – Transfer knowledge from large teacher models to MiniMind students
Architecture and Key Components
Understanding the core codebase helps you customize training effectively. The implementation resides in model/model_minimind.py and follows standard Transformer conventions with modern optimizations.
Model Definition (MiniMindConfig and MiniMindModel)
The configuration class MiniMindConfig inherits from transformers.PretrainedConfig and declares all hyperparameters including optional MoE settings. The network architecture in MiniMindModel stacks multiple MiniMindBlock layers, each containing:
- RMSNorm for stable layer normalization
- Rotary positional embeddings (
precompute_freqs_cis) - Standard Attention mechanism with optional FlashAttention support
- FeedForward or MOEFeedForward when
use_moe=True
The causal language modeling head shares weights with the token embedding matrix, reducing parameter count.
Tokenizer System
MiniMind ships with a pre-trained 6,400-token BPE tokenizer stored in model/tokenizer.json and model/tokenizer_config.json. For domain-specific vocabularies, trainer/train_tokenizer.py uses the HuggingFace tokenizers library to rebuild the vocabulary and outputs the same JSON format expected by the training scripts.
Dataset Classes
The dataset/lm_dataset.py file provides PyTorch Dataset implementations for each training stage:
PretrainDataset– Handles raw text with BOS/EOS tokens and label maskingSFTDataset– Applies ChatML templates viaapply_chat_templatefor instruction dataDPODatasetandRLAIFDataset– Structure preference pairs for RL training
All datasets use datasets.load_dataset, accepting JSONL files with specific schema requirements.
Training Utilities
trainer/trainer_utils.py contains essential infrastructure code including:
init_distributed_modefor PyTorch DDP setupget_lrfor learning rate schedulinglm_checkpointfor model saving and resumptionget_model_paramsfor counting trainable parameters
Step-by-Step Training Guide
Follow these stages sequentially to build your model. All scripts support distributed training via the RANK and LOCAL_RANK environment variables.
(Optional) Train a Custom Tokenizer
If your domain requires specialized vocabulary, train a custom tokenizer before pre-training:
python -m trainer.train_tokenizer \
--data_path data/my_corpus.jsonl \
--tokenizer_dir model/my_tokenizer \
--vocab_size 6400
This generates tokenizer.json and tokenizer_config.json in the specified directory. Reference this directory in subsequent training scripts using --tokenizer_path model/my_tokenizer.
Pre-training on Raw Text
Run trainer/train_pretrain.py to train the base model on a language modeling objective:
python -m trainer.train_pretrain \
--save_dir out/checkpoints \
--save_weight pretrain \
--epochs 4 \
--batch_size 32 \
--learning_rate 5e-4 \
--hidden_size 512 \
--num_hidden_layers 8 \
--max_seq_len 340 \
--data_path data/pretrain_hq.jsonl \
--use_wandb \
--wandb_project MiniMind-Pretrain
The script constructs a MiniMindConfig from CLI arguments, loads the tokenizer, initializes a PretrainDataset, and executes the training loop. Checkpoints are saved as pretrain_512.pth (or your specified hidden size) at intervals defined by --save_interval.
Supervised Fine-Tuning (SFT)
Convert your pre-trained checkpoint into a chat model using trainer/train_full_sft.py:
python -m trainer.train_full_sft \
--save_dir out/checkpoints_sft \
--save_weight sft \
--epochs 2 \
--batch_size 16 \
--learning_rate 2e-4 \
--hidden_size 512 \
--data_path data/sft_dataset.jsonl \
--use_lora 0
The SFTDataset class processes conversation data in ChatML format, generating labels that mask the system and user turns while supervising only the assistant responses. Set --use_lora 1 to switch to adapter-based training (see next section).
Parameter-Efficient Fine-Tuning with LoRA
For hardware-constrained environments, trainer/train_lora.py inserts low-rank adapters into attention and feed-forward projections:
python -m trainer.train_lora \
--save_dir out/checkpoints_lora \
--save_weight lora \
--epochs 3 \
--batch_size 32 \
--learning_rate 1e-4 \
--hidden_size 512 \
--lora_rank 8 \
--data_path data/sft_dataset.jsonl
This freezes the base model weights and updates only the adapter parameters, producing checkpoints as small as 2 MB for a 6M-parameter MiniMind model.
Reinforcement Learning Alignment
Align your model with human preferences using PPO or DPO. First train a reward model, then optimize the policy:
# Train reward model
python -m trainer.train_ppo \
--stage reward \
--save_dir out/ppo \
--data_path data/ppo_data.jsonl \
--epochs 1
# Train policy with PPO
python -m trainer.train_ppo \
--stage policy \
--save_dir out/ppo \
--data_path data/ppo_data.jsonl \
--epochs 2 \
--kl_coef 0.1
Alternatively, use trainer/train_dpo.py for Direct Preference Optimization without a separate reward model, or trainer/train_grpo.py for gradient-based RPO training. These scripts utilize DPODataset to process chosen versus rejected response pairs.
Deploying Your Model
After training, serve your model via an OpenAI-compatible API or interactive web interface.
Start the API server using scripts/serve_openai_api.py:
python -m scripts.serve_openai_api \
--model_path out/checkpoints/pretrain_512.pth \
--tokenizer_path model \
--port 8000
Access the endpoint with any OpenAI client:
import openai
openai.api_base = "http://localhost:8000/v1"
openai.api_key = "dummy"
resp = openai.ChatCompletion.create(
model="mini-mind",
messages=[{"role": "user", "content": "Hello, how are you?"}],
max_tokens=128
)
print(resp["choices"][0]["message"]["content"])
For interactive demos, use scripts/web_demo.py to launch a Gradio interface.
Summary
- MiniMind provides a complete training stack in
jingyaogong/minimindwith scripts for tokenizer training, pre-training, SFT, LoRA, and RL alignment. - The architecture in
model/model_minimind.pysupports configurable sizes, FlashAttention, and optional MoE layers viaMiniMindConfig. - Pre-training uses
train_pretrain.pywithPretrainDatasetfor raw text corpora. - Supervised fine-tuning applies ChatML formatting through
train_full_sft.pyandSFTDataset. - LoRA training via
train_lora.pyenables efficient fine-tuning with minimal GPU memory. - Deployment options include
scripts/serve_openai_api.pyfor API serving andscripts/web_demo.pyfor web demos.
Frequently Asked Questions
What hardware is required to train a MiniMind model from scratch?
You can train small MiniMind variants (512 hidden size, 8 layers) on a single consumer GPU with 8-12 GB VRAM. For larger configurations or full pre-training on extensive corpora, multi-GPU setups using PyTorch DistributedDataParallel (via RANK environment variables) are supported through trainer/trainer_utils.py.
How do I format data for supervised fine-tuning?
SFT data should be JSONL files where each line contains a conversations list with ChatML-formatted messages (system, user, assistant roles). The SFTDataset class in dataset/lm_dataset.py automatically applies the chat template using the tokenizer's apply_chat_template method and creates labels that only supervise the assistant's responses.
Can I use a custom tokenizer instead of the default 6,400-token vocabulary?
Yes. Run python -m trainer.train_tokenizer with your domain-specific corpus to generate a new BPE tokenizer. The script outputs tokenizer.json and tokenizer_config.json compatible with all training scripts. Specify your custom tokenizer path using the --tokenizer_path argument in any training script.
What is the difference between full SFT and LoRA training?
Full SFT (train_full_sft.py with --use_lora 0) updates all model parameters and produces larger checkpoints (hundreds of MB). LoRA training (train_lora.py) inserts low-rank adapter matrices into the attention and feed-forward layers, freezing base weights and producing small checkpoints (~2 MB) suitable for storage-constrained environments or rapid task switching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →