# LLMs-from-scratch | Sebastian Raschka | Knowledge Base | Instagit

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

GitHub Stars: 93.4k

Repository: https://github.com/rasbt/LLMs-from-scratch

---

## Articles

### [Understanding the Difference Between Pretraining and Instruction Finetuning](/rasbt/LLMs-from-scratch/difference-between-pretraining-and-instruction-finetuning)

Learn the crucial difference between pretraining and instruction finetuning. Discover how models learn from raw text and adapt to user commands. LLMs from scratch explained.

- Tags: deep-dive
- Published: 2026-05-12

### [Training Speed Optimization Techniques for LLM Pretraining: A PyTorch Performance Guide](/rasbt/LLMs-from-scratch/training-speed-optimization-techniques-for-llm-pretraining)

Optimize LLM pretraining speed with PyTorch. Discover techniques like torch.compile, bfloat16, DDP, and throughput aggregation for faster model training.

- Tags: performance
- Published: 2026-05-12

### [Understanding Multi-Head Latent Attention (MLA) Optimizations: DeepSeek-Inspired KV Cache Compression](/rasbt/LLMs-from-scratch/understanding-multi-head-latent-attention-mla-optimizations)

Learn how Multi-Head Latent Attention MLA optimizes GPU memory with KV cache compression inspired by DeepSeek. Understand on-the-fly up-projection for efficient attention computation.

- Tags: deep-dive
- Published: 2026-05-12

### [Implementing Mixed Expert (MoE) Layers in Transformers: A Complete Guide to the LLMs-from-Scratch Implementation](/rasbt/LLMs-from-scratch/implementing-mixed-expert-moe-layers-in-transformers)

Learn to implement Mixed Expert MoE layers in transformers with this LLMs-from-scratch guide. Understand gating networks and expert routing to optimize models efficiently.

- Tags: tutorial
- Published: 2026-05-12

### [Handling Near-Duplicate Detection in Instruction Datasets: A Practical Guide](/rasbt/LLMs-from-scratch/handling-near-duplicate-detection-in-instruction-datasets)

Efficiently handle near-duplicate detection in instruction datasets with this practical guide. Learn to remove redundant entries using TF-IDF and cosine similarity.

- Tags: how-to-guide
- Published: 2026-05-12

### [Converting GPT Checkpoints to Other Model Architectures: A Complete Guide to LLMs-from-scratch](/rasbt/LLMs-from-scratch/converting-gpt-checkpoints-to-other-model-architectures)

Convert GPT checkpoints to PyTorch with LLMs-from-scratch. Explore original GPT-2, LLaMA-style, and Qwen-3-style model architectures easily.

- Tags: how-to-guide
- Published: 2026-05-12

### [Embedding Layers vs Linear Projections in PyTorch: Understanding the Difference in LLMs-from-scratch](/rasbt/LLMs-from-scratch/difference-between-embedding-layers-and-linear-projections)

Understand the difference between embedding layers and linear projections in PyTorch for LLMs. Learn how these distinct components transform data within transformer architectures.

- Tags: deep-dive
- Published: 2026-05-12

### [Building a User Interface for Interacting with Finetuned LLMs](/rasbt/LLMs-from-scratch/building-ui-for-interacting-with-finetuned-llms)

Create a user interface for interacting with finetuned LLMs. Explore the rasbt/LLMs-from-scratch repository for a Chainlit-based web interface to chat with GPT-2 models.

- Tags: how-to-guide
- Published: 2026-05-12

### [Evaluating Instruction-Following Models with LLM-as-a-Judge: Implementation Guide](/rasbt/LLMs-from-scratch/evaluating-instruction-following-models-with-llm-as-a-judge)

Implement LLM-as-a-judge to evaluate instruction-following models. Learn how to score responses using Llama 3 and Ollama with the LLMs-from-scratch framework. Get started today!

- Tags: how-to-guide
- Published: 2026-05-12

### [Creating Instruction Finetuning Datasets from Scratch: A Complete Guide to the LLMs-from-scratch Pipeline](/rasbt/LLMs-from-scratch/creating-instruction-finetuning-datasets-from-scratch)

Learn to create instruction finetuning datasets from scratch for LLMs. This guide covers data formatting with JSON, duplicate cleaning, and GPT training for custom models.

- Tags: how-to-guide
- Published: 2026-05-12

### [Implementing Learning Rate Schedulers for LLM Pretraining: Warm-Up and Cosine Decay in PyTorch](/rasbt/LLMs-from-scratch/implementing-learning-rate-schedulers-for-llm-pretraining)

Implement learning rate schedulers for LLM pretraining with PyTorch. Combine linear warm-up and cosine decay for stable convergence in transformer models. Learn how from LLMs-from-scratch.

- Tags: how-to-guide
- Published: 2026-05-12

### [Setting up Multi-GPU Training with PyTorch DDP: A Complete Guide from LLMs-from-scratch](/rasbt/LLMs-from-scratch/setting-up-multi-gpu-training-with-pytorch-ddp)

Learn to set up multi-GPU training with PyTorch DDP. This guide simplifies scaling your LLM training across multiple GPUs with minimal code changes.

- Tags: how-to-guide
- Published: 2026-05-12

### [Memory-Efficient Model Weight Loading Strategies in LLMs-from-Scratch](/rasbt/LLMs-from-scratch/memory-efficient-model-weight-loading-strategies)

Discover memory-efficient model weight loading strategies for LLMs. Learn to use map_location, safetensors, and sharded checkpoints to keep RAM usage low when initializing large language models.

- Tags: performance
- Published: 2026-05-12

### [How to Extend the Tiktoken BPE Tokenizer with New Tokens: A Complete Guide](/rasbt/LLMs-from-scratch/how-to-extend-tiktoken-bpe-tokenizer-with-new-tokens)

Learn to extend the Tiktoken BPE tokenizer with new tokens. Load custom vocabularies, augment special tokens, and wrap in a helper class for your language model. Complete guide.

- Tags: how-to-guide
- Published: 2026-05-12

### [Implementing Direct Preference Optimization (DPO) for LLM Alignment: A Complete Guide](/rasbt/LLMs-from-scratch/implementing-direct-preference-optimization-for-llm-alignment)

Implement Direct Preference Optimization DPO for LLM alignment. This guide simplifies RL pipelines by directly fine-tuning LLMs on human preferences using binary classification loss.

- Tags: tutorial
- Published: 2026-05-12

### [Grouped-Query Attention vs Multi-Head Attention: Understanding GQA Implementation](/rasbt/LLMs-from-scratch/understanding-grouped-query-attention-vs-multi-head-attention)

Understand Grouped-Query Attention (GQA) vs Multi-Head Attention (MHA). Learn how GQA optimizes memory and KV-cache for LLMs, offering efficiency with comparable performance.

- Tags: deep-dive
- Published: 2026-05-12

### [How to Convert a GPT Architecture to Llama 3.2: RoPE, Scaling, and KV-Cache Implementation](/rasbt/LLMs-from-scratch/converting-gpt-architecture-to-llama-3-2)

Convert GPT to Llama 3.2 with our LLMs-from-scratch repository. Swap embeddings, adjust scaling, and update KV-cache, preserving all pretrained weights for seamless integration.

- Tags: deep-dive
- Published: 2026-05-12

### [Implementing Parameter-Efficient Finetuning with LoRA from Scratch: A Complete Guide](/rasbt/LLMs-from-scratch/implementing-parameter-efficient-finetuning-with-lora-from-scratch)

Learn to implement parameter-efficient finetuning with LoRA from scratch. This guide explains LoRA's low-rank matrices for efficient LLM fine-tuning with minimal memory.

- Tags: how-to-guide
- Published: 2026-05-12

### [How to Implement KV Cache for Efficient LLM Inference: A Complete Guide to the LLMs-from-Scratch Implementation](/rasbt/LLMs-from-scratch/how-to-implement-kv-cache-for-efficient-llm-inference)

Learn to implement KV cache for efficient LLM inference and reduce complexity from O(T·d) to O(d). This guide provides a complete LLMs-from-scratch implementation.

- Tags: how-to-guide
- Published: 2026-05-12

### [How Causal Masking Works in Multi-Head Attention: Implementation Deep Dive](/rasbt/LLMs-from-scratch/how-does-causal-masking-work-in-multi-head-attention)

Understand causal masking in multi-head attention. Learn how LLMs-from-scratch implements this autoregressive generation technique by blocking future tokens. Deep dive into the implementation.

- Tags: deep-dive
- Published: 2026-05-12

