train-llm-from-scratch
A straightforward method for training your LLM, from downloading data to generating text.
Explore the RLHF value head architecture: a two-layer MLP projecting transformer states to scalar values. Learn how zero-initialization ensures stable training. Understand the critic network in LLM training from scratch.
How Configuration Is Managed for LLM Training Jobs in train-llm-from-scratchDiscover how LLM training jobs use a layered JSON config system with dataclass defaults, file merging, and CLI overrides. Learn configuration management in train-llm-from-scratch.
How to Use Streamlit for Controlling LLM Training: A Complete Guide to train-llm-from-scratchControl LLM training from scratch with an intuitive Streamlit interface. Manage data prep, pretraining, fine-tuning, and RLHF from your browser. Perfect for train-llm-from-scratch.
How to Evaluate an LLM's Performance on the GSM8K Benchmark: A Complete Implementation GuideLearn how to evaluate an LLM's performance on the GSM8K benchmark with this complete implementation guide. Discover the steps to accurately measure your model's math problem-solving capabilities.
How to Perform Inference and Generate Text with a Trained LLM: A Complete GuideEasily perform inference and generate text with a trained LLM. Learn how to load your model and use generate_reply for seamless text generation with custom controls.
How to Format Instructions and Use Chat Templates for LLM InputLearn to format instructions and use chat templates for LLM input with our tokenizer-agnostic implementation. Convert messages to token sequences and masks efficiently.
How to Process and Tokenize Data for LLM Training: A Complete Pipeline GuideLearn to process and tokenize data for LLM training with a complete pipeline guide. Convert raw text to efficient HDF5 token files using tiktoken and specialized loss masks.
How to Implement Proximal Policy Optimization (PPO) for LLM Fine-TuningLearn how to implement Proximal Policy Optimization PPO for LLM fine-tuning. Stabilize RLHF training with trust regions and advantage estimation.
Direct Preference Optimization (DPO) Implementation Guide: From Theory to Code in train-llm-from-scratchLearn how to implement Direct Preference Optimization DPO in train-llm-from-scratch. Understand DPO theory and see code examples for aligning language models with human preferences easily.
How to Train a Reward Model for LLM Alignment: A Complete Implementation GuideLearn to train a reward model for LLM alignment. This guide details implementation using a scalar reward head and Bradley-Terry loss on human preference data.
How to Perform Supervised Fine-Tuning (SFT) on a Pretrained LLM: A Complete Implementation GuideLearn to perform Supervised Fine-Tuning SFT on a pretrained LLM with this implementation guide. Discover prompt masking and masked cross-entropy loss for efficient training.
Understanding the Post-Training Pipeline for LLMs: From SFT to GRPOExplore the LLM post-training pipeline from SFT to GRPO. Learn about reward modeling, preference alignment, and reinforcement learning in the FareedKhan-dev/train-llm-from-scratch repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →