LLMs-from-scratch

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

20 articles 93.4k View on GitHub ↗
20 articles
Understanding the Difference Between Pretraining and Instruction Finetuning

Learn the crucial difference between pretraining and instruction finetuning. Discover how models learn from raw text and adapt to user commands. LLMs from scratch explained.

deep-dive
May 12, 2026
Training Speed Optimization Techniques for LLM Pretraining: A PyTorch Performance Guide

Optimize LLM pretraining speed with PyTorch. Discover techniques like torch.compile, bfloat16, DDP, and throughput aggregation for faster model training.

performance
May 12, 2026
Understanding Multi-Head Latent Attention (MLA) Optimizations: DeepSeek-Inspired KV Cache Compression

Learn how Multi-Head Latent Attention MLA optimizes GPU memory with KV cache compression inspired by DeepSeek. Understand on-the-fly up-projection for efficient attention computation.

deep-dive
May 12, 2026
Implementing Mixed Expert (MoE) Layers in Transformers: A Complete Guide to the LLMs-from-Scratch Implementation

Learn to implement Mixed Expert MoE layers in transformers with this LLMs-from-scratch guide. Understand gating networks and expert routing to optimize models efficiently.

tutorial
May 12, 2026
Handling Near-Duplicate Detection in Instruction Datasets: A Practical Guide

Efficiently handle near-duplicate detection in instruction datasets with this practical guide. Learn to remove redundant entries using TF-IDF and cosine similarity.

how-to-guide
May 12, 2026
Converting GPT Checkpoints to Other Model Architectures: A Complete Guide to LLMs-from-scratch

Convert GPT checkpoints to PyTorch with LLMs-from-scratch. Explore original GPT-2, LLaMA-style, and Qwen-3-style model architectures easily.

how-to-guide
May 12, 2026
Embedding Layers vs Linear Projections in PyTorch: Understanding the Difference in LLMs-from-scratch

Understand the difference between embedding layers and linear projections in PyTorch for LLMs. Learn how these distinct components transform data within transformer architectures.

deep-dive
May 12, 2026
Building a User Interface for Interacting with Finetuned LLMs

Create a user interface for interacting with finetuned LLMs. Explore the rasbt/LLMs-from-scratch repository for a Chainlit-based web interface to chat with GPT-2 models.

how-to-guide
May 12, 2026
Evaluating Instruction-Following Models with LLM-as-a-Judge: Implementation Guide

Implement LLM-as-a-judge to evaluate instruction-following models. Learn how to score responses using Llama 3 and Ollama with the LLMs-from-scratch framework. Get started today!

how-to-guide
May 12, 2026
Creating Instruction Finetuning Datasets from Scratch: A Complete Guide to the LLMs-from-scratch Pipeline

Learn to create instruction finetuning datasets from scratch for LLMs. This guide covers data formatting with JSON, duplicate cleaning, and GPT training for custom models.

how-to-guide
May 12, 2026
Implementing Learning Rate Schedulers for LLM Pretraining: Warm-Up and Cosine Decay in PyTorch

Implement learning rate schedulers for LLM pretraining with PyTorch. Combine linear warm-up and cosine decay for stable convergence in transformer models. Learn how from LLMs-from-scratch.

how-to-guide
May 12, 2026
Setting up Multi-GPU Training with PyTorch DDP: A Complete Guide from LLMs-from-scratch

Learn to set up multi-GPU training with PyTorch DDP. This guide simplifies scaling your LLM training across multiple GPUs with minimal code changes.

how-to-guide
May 12, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →