train-llm-from-scratch

A straightforward method for training your LLM, from downloading data to generating text.

37 articles 2.6k View on GitHub ↗
37 articles
Value Head Architecture in RLHF Training: A Deep Dive into the Critic Network

Explore the RLHF value head architecture: a two-layer MLP projecting transformer states to scalar values. Learn how zero-initialization ensures stable training. Understand the critic network in LLM training from scratch.

deep-dive
Jun 11, 2026
How Configuration Is Managed for LLM Training Jobs in train-llm-from-scratch

Discover how LLM training jobs use a layered JSON config system with dataclass defaults, file merging, and CLI overrides. Learn configuration management in train-llm-from-scratch.

internals
Jun 11, 2026
How to Use Streamlit for Controlling LLM Training: A Complete Guide to train-llm-from-scratch

Control LLM training from scratch with an intuitive Streamlit interface. Manage data prep, pretraining, fine-tuning, and RLHF from your browser. Perfect for train-llm-from-scratch.

how-to-guide
Jun 11, 2026
How to Evaluate an LLM's Performance on the GSM8K Benchmark: A Complete Implementation Guide

Learn how to evaluate an LLM's performance on the GSM8K benchmark with this complete implementation guide. Discover the steps to accurately measure your model's math problem-solving capabilities.

how-to-guide
Jun 11, 2026
How to Perform Inference and Generate Text with a Trained LLM: A Complete Guide

Easily perform inference and generate text with a trained LLM. Learn how to load your model and use generate_reply for seamless text generation with custom controls.

tutorial
Jun 11, 2026
How to Format Instructions and Use Chat Templates for LLM Input

Learn to format instructions and use chat templates for LLM input with our tokenizer-agnostic implementation. Convert messages to token sequences and masks efficiently.

how-to-guide
Jun 11, 2026
How to Process and Tokenize Data for LLM Training: A Complete Pipeline Guide

Learn to process and tokenize data for LLM training with a complete pipeline guide. Convert raw text to efficient HDF5 token files using tiktoken and specialized loss masks.

tutorial
Jun 11, 2026
How to Implement Proximal Policy Optimization (PPO) for LLM Fine-Tuning

Learn how to implement Proximal Policy Optimization PPO for LLM fine-tuning. Stabilize RLHF training with trust regions and advantage estimation.

tutorial
Jun 11, 2026
Direct Preference Optimization (DPO) Implementation Guide: From Theory to Code in train-llm-from-scratch

Learn how to implement Direct Preference Optimization DPO in train-llm-from-scratch. Understand DPO theory and see code examples for aligning language models with human preferences easily.

tutorial
Jun 11, 2026
How to Train a Reward Model for LLM Alignment: A Complete Implementation Guide

Learn to train a reward model for LLM alignment. This guide details implementation using a scalar reward head and Bradley-Terry loss on human preference data.

how-to-guide
Jun 11, 2026
How to Perform Supervised Fine-Tuning (SFT) on a Pretrained LLM: A Complete Implementation Guide

Learn to perform Supervised Fine-Tuning SFT on a pretrained LLM with this implementation guide. Discover prompt masking and masked cross-entropy loss for efficient training.

how-to-guide
Jun 11, 2026
Understanding the Post-Training Pipeline for LLMs: From SFT to GRPO

Explore the LLM post-training pipeline from SFT to GRPO. Learn about reward modeling, preference alignment, and reinforcement learning in the FareedKhan-dev/train-llm-from-scratch repository.

deep-dive
Jun 11, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →