LLMs-from-scratch
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
Learn the crucial difference between pretraining and instruction finetuning. Discover how models learn from raw text and adapt to user commands. LLMs from scratch explained.
Training Speed Optimization Techniques for LLM Pretraining: A PyTorch Performance GuideOptimize LLM pretraining speed with PyTorch. Discover techniques like torch.compile, bfloat16, DDP, and throughput aggregation for faster model training.
Understanding Multi-Head Latent Attention (MLA) Optimizations: DeepSeek-Inspired KV Cache CompressionLearn how Multi-Head Latent Attention MLA optimizes GPU memory with KV cache compression inspired by DeepSeek. Understand on-the-fly up-projection for efficient attention computation.
Implementing Mixed Expert (MoE) Layers in Transformers: A Complete Guide to the LLMs-from-Scratch ImplementationLearn to implement Mixed Expert MoE layers in transformers with this LLMs-from-scratch guide. Understand gating networks and expert routing to optimize models efficiently.
Handling Near-Duplicate Detection in Instruction Datasets: A Practical GuideEfficiently handle near-duplicate detection in instruction datasets with this practical guide. Learn to remove redundant entries using TF-IDF and cosine similarity.
Converting GPT Checkpoints to Other Model Architectures: A Complete Guide to LLMs-from-scratchConvert GPT checkpoints to PyTorch with LLMs-from-scratch. Explore original GPT-2, LLaMA-style, and Qwen-3-style model architectures easily.
Embedding Layers vs Linear Projections in PyTorch: Understanding the Difference in LLMs-from-scratchUnderstand the difference between embedding layers and linear projections in PyTorch for LLMs. Learn how these distinct components transform data within transformer architectures.
Building a User Interface for Interacting with Finetuned LLMsCreate a user interface for interacting with finetuned LLMs. Explore the rasbt/LLMs-from-scratch repository for a Chainlit-based web interface to chat with GPT-2 models.
Evaluating Instruction-Following Models with LLM-as-a-Judge: Implementation GuideImplement LLM-as-a-judge to evaluate instruction-following models. Learn how to score responses using Llama 3 and Ollama with the LLMs-from-scratch framework. Get started today!
Creating Instruction Finetuning Datasets from Scratch: A Complete Guide to the LLMs-from-scratch PipelineLearn to create instruction finetuning datasets from scratch for LLMs. This guide covers data formatting with JSON, duplicate cleaning, and GPT training for custom models.
Implementing Learning Rate Schedulers for LLM Pretraining: Warm-Up and Cosine Decay in PyTorchImplement learning rate schedulers for LLM pretraining with PyTorch. Combine linear warm-up and cosine decay for stable convergence in transformer models. Learn how from LLMs-from-scratch.
Setting up Multi-GPU Training with PyTorch DDP: A Complete Guide from LLMs-from-scratchLearn to set up multi-GPU training with PyTorch DDP. This guide simplifies scaling your LLM training across multiple GPUs with minimal code changes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →