# nn-zero-to-hero | Andrej | Knowledge Base | Instagit

Neural Networks: Zero to Hero

GitHub Stars: 22.6k

Repository: https://github.com/karpathy/nn-zero-to-hero

---

## Articles

### [How to Efficiently Manage Tensor Shapes in PyTorch: Patterns from nn-zero-to-hero](/karpathy/nn-zero-to-hero/how-to-efficiently-manage-tensor-shapes-pytorch)

Master PyTorch tensor shapes using explicit contracts and reusable classes. Prevent bugs and write cleaner code with these nn-zero-to-hero patterns.

- Tags: tutorial
- Published: 2026-05-23

### [What Is the Adam Optimizer and How Does It Differ From SGD?](/karpathy/nn-zero-to-hero/what-is-adam-optimizer-how-it-differs-sgd)

Understand the Adam optimizer and how it differs from SGD. Learn about adaptive learning rates and gradient-based optimization for faster model training.

- Tags: deep-dive
- Published: 2026-05-23

### [How Residual Connections Enable Training of Very Deep Neural Networks](/karpathy/nn-zero-to-hero/how-residual-connections-help-train-very-deep-networks)

Discover how residual connections enable training of very deep neural networks by preserving gradient flow and preventing the vanishing gradient problem. Learn about identity shortcut paths and residual learning.

- Tags: deep-dive
- Published: 2026-05-23

### [The Role of Activation Functions Like tanh and ReLU in Neural Networks: A Deep Dive into karpathy/nn-zero-to-hero](/karpathy/nn-zero-to-hero/role-activation-functions-tanh-relu-neural-networks)

Explore the crucial role of activation functions like tanh and ReLU in neural networks. Understand how they introduce non-linearity, manage gradients, and enable complex function approximation in deep learning architectures.

- Tags: deep-dive
- Published: 2026-05-23

### [How to Diagnose Overfitting vs Underfitting in Neural Networks: A Practical Guide with nn-zero-to-hero](/karpathy/nn-zero-to-hero/how-to-diagnose-overfitting-vs-underfitting-neural-networks)

Easily diagnose overfitting vs underfitting in neural networks by comparing training and validation loss. Learn practical tips from nn-zero-to-hero to build better models.

- Tags: how-to-guide
- Published: 2026-05-23

### [Common Pitfalls When Training Deep Neural Networks: Lessons from nn-zero-to-hero](/karpathy/nn-zero-to-hero/what-are-common-pitfalls-training-deep-neural-networks)

Avoid common pitfalls when training deep neural networks. Learn how vanishing/exploding gradients, poor initialization, and learning rate issues destabilize training and hinder convergence. Master nn-zero-to-hero lessons for ef...

- Tags: deep-dive
- Published: 2026-05-23

### [Why Tokenization Causes Issues in LLMs and How to Debug Them](/karpathy/nn-zero-to-hero/why-tokenization-cause-issues-llms-how-debug)

Discover why tokenization causes LLM failures and learn effective debugging strategies. Address issues with vocabularies, BPE merges, and non-bijective mappings to improve your models.

- Tags: deep-dive
- Published: 2026-05-23

### [Byte Pair Encoding (BPE) Tokenization: How GPT Tokenizers Work](/karpathy/nn-zero-to-hero/what-is-byte-pair-encoding-bpe-tokenization-implementation)

Learn how Byte Pair Encoding BPE tokenization works. Discover this sub-word segmentation algorithm for efficient text to integer token conversion in GPT models.

- Tags: deep-dive
- Published: 2026-05-23

### [How GPT Architecture Performs Autoregressive Language Modeling: Inside nn-zero-to-hero](/karpathy/nn-zero-to-hero/how-gpt-architecture-perform-autoregressive-language-modeling)

Discover how GPT architecture performs autoregressive language modeling. Learn about causal attention, next-token prediction, and sequence generation in this deep dive.

- Tags: deep-dive
- Published: 2026-05-23

### [How the Attention Mechanism and Self-Attention Work: A Deep Dive into nn-zero-to-hero](/karpathy/nn-zero-to-hero/what-is-attention-mechanism-how-self-attention-work)

Understand the attention mechanism and self-attention in neural networks. Learn how they dynamically weigh input importance and allow tokens to attend to each other.

- Tags: deep-dive
- Published: 2026-05-23

### [How to Implement a Transformer Architecture from Scratch in PyTorch](/karpathy/nn-zero-to-hero/how-to-implement-transformer-architecture-from-scratch)

Implement a Transformer architecture from scratch in PyTorch. Learn to build a GPT-style model using token embeddings, attention, and feed-forward blocks with the nn-zero-to-hero repository.

- Tags: how-to-guide
- Published: 2026-05-23

### [Understanding the Internals of PyTorch torch.nn Modules: A Deep Dive into nn.Module](/karpathy/nn-zero-to-hero/what-are-internals-pytorch-torch-nn-modules)

Explore PyTorch nnModule internals. Discover how it registers parameters, handles forward passes via __call__, and manages state with state_dict.

- Tags: deep-dive
- Published: 2026-05-23

### [How WaveNet-Style CNNs Work for Sequence Modeling: Causal Dilated Convolutions Explained](/karpathy/nn-zero-to-hero/how-wavenet-style-cnns-work-sequence-modeling)

Discover how WaveNet-style CNNs use causal dilated convolutions to model sequences. Learn how massive receptive fields capture long-range dependencies without recurrence.

- Tags: deep-dive
- Published: 2026-05-23

### [Learning Rate Tuning: Why It Matters and How to Find Optimal Hyperparameters in Neural Networks](/karpathy/nn-zero-to-hero/why-learning-rate-tuning-important-how-find-optimal-hyperparameters)

Master learning rate tuning in neural networks. Discover why optimal hyperparameters are crucial for fast yet stable convergence and prevent common optimization pitfalls.

- Tags: deep-dive
- Published: 2026-05-23

### [How to Implement Manual Backpropagation Without PyTorch Autograd](/karpathy/nn-zero-to-hero/how-to-implement-manual-backpropagation-without-pytorch-autograd)

Learn to implement manual backpropagation without PyTorch autograd. Compute gradients using the chain rule and apply SGD updates for deeper neural network understanding.

- Tags: how-to-guide
- Published: 2026-05-23

### [Bigram vs MLP Language Models: From Counting to Neural Networks](/karpathy/nn-zero-to-hero/difference-between-bigram-mlp-language-models)

Explore the difference between Bigram and MLP language models. Understand how Bigram uses counts while MLP leverages neural networks and context for character prediction.

- Tags: deep-dive
- Published: 2026-05-23

### [How to Implement a Character-Level Language Model from Scratch: A Complete Guide](/karpathy/nn-zero-to-hero/how-to-implement-character-level-language-model-from-scratch)

Master character-level language models from scratch. See the nn-zero-to-hero repo's step-by-step guide from bigrams to transformers for your next NLP project.

- Tags: how-to-guide
- Published: 2026-05-23

### [Training, Validation, and Test Data Splits for Neural Networks: The 80/10/10 Guide](/karpathy/nn-zero-to-hero/what-are-training-validation-test-data-splits)

Understand training, validation, and test data splits for neural networks. Learn the 80/10/10 method from karpathy nn-zero-to-hero for optimal model performance.

- Tags: deep-dive
- Published: 2026-05-23

### [How Cross Entropy Loss Works for Classification: Inside Karpathy’s nn-zero-to-hero Implementation](/karpathy/nn-zero-to-hero/how-cross-entropy-loss-function-work-classification)

Understand cross entropy loss for classification. Learn how it measures discrepancies by converting logits to probabilities and calculating negative log-likelihood using Karpathy's nn-zero-to-hero implementation.

- Tags: deep-dive
- Published: 2026-05-23

### [What Is Batch Normalization and How Does It Stabilize Neural Network Training?](/karpathy/nn-zero-to-hero/what-is-batch-normalization-how-it-stabilizes-training)

Discover Batch Normalization and how it stabilizes neural network training by normalizing layer activations, enabling faster convergence and higher learning rates.

- Tags: deep-dive
- Published: 2026-05-23

### [How Backpropagation Is Implemented in Micrograd: A Deep Dive into the Value Class](/karpathy/nn-zero-to-hero/how-is-backpropagation-implemented-micrograd)

Discover how micrograd implements backpropagation using its Value class. Learn about reverse-mode automatic differentiation and computation history in this deep dive.

- Tags: deep-dive
- Published: 2026-05-23

### [Understanding the Computational Graph Structure in micrograd Value Objects](/karpathy/nn-zero-to-hero/what-is-computational-graph-structure-micrograd-value-objects)

Explore the computational graph structure in micrograd Value objects. Learn how parent references, operations, and backward closures enable reverse-mode automatic differentiation for efficient deep learning.

- Tags: internals
- Published: 2026-05-23

### [How micrograd's Automatic Differentiation Engine Works: A Deep Dive into Reverse-Mode Autograd](/karpathy/nn-zero-to-hero/how-does-micrograd-automatic-differentiation-engine-work)

Explore micrograd's reverse-mode autograd engine. Learn how its Value class builds a dynamic graph and backpropagates gradients for efficient deep learning.

- Tags: deep-dive
- Published: 2026-05-23

