nn-zero-to-hero
Neural Networks: Zero to Hero
Master PyTorch tensor shapes using explicit contracts and reusable classes. Prevent bugs and write cleaner code with these nn-zero-to-hero patterns.
What Is the Adam Optimizer and How Does It Differ From SGD?Understand the Adam optimizer and how it differs from SGD. Learn about adaptive learning rates and gradient-based optimization for faster model training.
How Residual Connections Enable Training of Very Deep Neural NetworksDiscover how residual connections enable training of very deep neural networks by preserving gradient flow and preventing the vanishing gradient problem. Learn about identity shortcut paths and residual learning.
The Role of Activation Functions Like tanh and ReLU in Neural Networks: A Deep Dive into karpathy/nn-zero-to-heroExplore the crucial role of activation functions like tanh and ReLU in neural networks. Understand how they introduce non-linearity, manage gradients, and enable complex function approximation in deep learning architectures.
How to Diagnose Overfitting vs Underfitting in Neural Networks: A Practical Guide with nn-zero-to-heroEasily diagnose overfitting vs underfitting in neural networks by comparing training and validation loss. Learn practical tips from nn-zero-to-hero to build better models.
Common Pitfalls When Training Deep Neural Networks: Lessons from nn-zero-to-heroAvoid common pitfalls when training deep neural networks. Learn how vanishing/exploding gradients, poor initialization, and learning rate issues destabilize training and hinder convergence. Master nn-zero-to-hero lessons for ef...
Why Tokenization Causes Issues in LLMs and How to Debug ThemDiscover why tokenization causes LLM failures and learn effective debugging strategies. Address issues with vocabularies, BPE merges, and non-bijective mappings to improve your models.
Byte Pair Encoding (BPE) Tokenization: How GPT Tokenizers WorkLearn how Byte Pair Encoding BPE tokenization works. Discover this sub-word segmentation algorithm for efficient text to integer token conversion in GPT models.
How GPT Architecture Performs Autoregressive Language Modeling: Inside nn-zero-to-heroDiscover how GPT architecture performs autoregressive language modeling. Learn about causal attention, next-token prediction, and sequence generation in this deep dive.
How the Attention Mechanism and Self-Attention Work: A Deep Dive into nn-zero-to-heroUnderstand the attention mechanism and self-attention in neural networks. Learn how they dynamically weigh input importance and allow tokens to attend to each other.
How to Implement a Transformer Architecture from Scratch in PyTorchImplement a Transformer architecture from scratch in PyTorch. Learn to build a GPT-style model using token embeddings, attention, and feed-forward blocks with the nn-zero-to-hero repository.
Understanding the Internals of PyTorch torch.nn Modules: A Deep Dive into nn.ModuleExplore PyTorch nnModule internals. Discover how it registers parameters, handles forward passes via __call__, and manages state with state_dict.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →