nn-zero-to-hero

Neural Networks: Zero to Hero

23 articles 22.6k View on GitHub ↗
23 articles
How to Efficiently Manage Tensor Shapes in PyTorch: Patterns from nn-zero-to-hero

Master PyTorch tensor shapes using explicit contracts and reusable classes. Prevent bugs and write cleaner code with these nn-zero-to-hero patterns.

tutorial
May 23, 2026
What Is the Adam Optimizer and How Does It Differ From SGD?

Understand the Adam optimizer and how it differs from SGD. Learn about adaptive learning rates and gradient-based optimization for faster model training.

deep-dive
May 23, 2026
How Residual Connections Enable Training of Very Deep Neural Networks

Discover how residual connections enable training of very deep neural networks by preserving gradient flow and preventing the vanishing gradient problem. Learn about identity shortcut paths and residual learning.

deep-dive
May 23, 2026
The Role of Activation Functions Like tanh and ReLU in Neural Networks: A Deep Dive into karpathy/nn-zero-to-hero

Explore the crucial role of activation functions like tanh and ReLU in neural networks. Understand how they introduce non-linearity, manage gradients, and enable complex function approximation in deep learning architectures.

deep-dive
May 23, 2026
How to Diagnose Overfitting vs Underfitting in Neural Networks: A Practical Guide with nn-zero-to-hero

Easily diagnose overfitting vs underfitting in neural networks by comparing training and validation loss. Learn practical tips from nn-zero-to-hero to build better models.

how-to-guide
May 23, 2026
Common Pitfalls When Training Deep Neural Networks: Lessons from nn-zero-to-hero

Avoid common pitfalls when training deep neural networks. Learn how vanishing/exploding gradients, poor initialization, and learning rate issues destabilize training and hinder convergence. Master nn-zero-to-hero lessons for ef...

deep-dive
May 23, 2026
Why Tokenization Causes Issues in LLMs and How to Debug Them

Discover why tokenization causes LLM failures and learn effective debugging strategies. Address issues with vocabularies, BPE merges, and non-bijective mappings to improve your models.

deep-dive
May 23, 2026
Byte Pair Encoding (BPE) Tokenization: How GPT Tokenizers Work

Learn how Byte Pair Encoding BPE tokenization works. Discover this sub-word segmentation algorithm for efficient text to integer token conversion in GPT models.

deep-dive
May 23, 2026
How GPT Architecture Performs Autoregressive Language Modeling: Inside nn-zero-to-hero

Discover how GPT architecture performs autoregressive language modeling. Learn about causal attention, next-token prediction, and sequence generation in this deep dive.

deep-dive
May 23, 2026
How the Attention Mechanism and Self-Attention Work: A Deep Dive into nn-zero-to-hero

Understand the attention mechanism and self-attention in neural networks. Learn how they dynamically weigh input importance and allow tokens to attend to each other.

deep-dive
May 23, 2026
How to Implement a Transformer Architecture from Scratch in PyTorch

Implement a Transformer architecture from scratch in PyTorch. Learn to build a GPT-style model using token embeddings, attention, and feed-forward blocks with the nn-zero-to-hero repository.

how-to-guide
May 23, 2026
Understanding the Internals of PyTorch torch.nn Modules: A Deep Dive into nn.Module

Explore PyTorch nnModule internals. Discover how it registers parameters, handles forward passes via __call__, and manages state with state_dict.

deep-dive
May 23, 2026

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →