# annotated_deep_learning_paper_implementations | labml.ai | Knowledge Base | Instagit

🧑‍🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adabelief, sophia, ...), gans(cyclegan, stylegan2, ...), 🎮 reinforcement learning (ppo, dqn), capsnet, distillation, ... 🧠

GitHub Stars: 65.8k

Repository: https://github.com/labmlai/annotated_deep_learning_paper_implementations

---

## Articles

### [Evidential Deep Learning for Classification Uncertainty: A Complete Implementation Guide](/labmlai/annotated_deep_learning_paper_implementations/how-is-evidential-deep-learning-implemented-for-classification-uncertainty)

Learn to implement evidential deep learning for classification uncertainty. Transform neural network outputs into evidence parameters defining a Dirichlet distribution for explicit uncertainty quantification with labmlai.

- Tags: how-to-guide
- Published: 2026-03-04

### [Pre-LN and Post-LN Transformer Architectures: Key Differences and Implementation Guide](/labmlai/annotated_deep_learning_paper_implementations/what-is-the-difference-between-pre-ln-and-post-ln-transformer-architectures)

Understand the key differences between Pre-LN and Post-LN transformer architectures. Learn how Pre-LN improves stable training of deep models and explore implementation guides.

- Tags: deep-dive
- Published: 2026-03-04

### [How FNet Is Different From Standard Transformer Attention: A Deep Dive into Fourier-Based Token Mixing](/labmlai/annotated_deep_learning_paper_implementations/how-is-fnet-different-from-standard-transformer-attention)

Discover how FNet revolutionizes Transformers by replacing attention with Fourier transforms. Experience linear complexity, reduced computation, and faster training. Learn the key differences today.

- Tags: deep-dive
- Published: 2026-03-04

### [How Capsule Networks Implement Dynamic Routing in PyTorch](/labmlai/annotated_deep_learning_paper_implementations/how-does-capsule-networks-implement-dynamic-routing)

Learn how Capsule Networks use dynamic routing via routing by agreement. Explore PyTorch implementation details and the Router class at labmlai annotated deep learning paper implementations.

- Tags: deep-dive
- Published: 2026-03-04

### [Feedforward Transformer vs Standard Attention: Architectural Differences Explained](/labmlai/annotated_deep_learning_paper_implementations/what-is-the-difference-between-feedforward-transformer-and-standard-attention)

Discover the key differences between Feedforward Transformer and standard attention. Learn how Feedforward Transformers use FFNs solely, contrasting with attention's O(L²) complexity.

- Tags: deep-dive
- Published: 2026-03-04

### [How Stable Diffusion Uses Latent Space for Image Generation: Technical Architecture Explained](/labmlai/annotated_deep_learning_paper_implementations/how-is-the-stable-diffusion-latent-space-used-for-image-generation)

Discover how Stable Diffusion leverages its latent space architecture for image generation. Learn about the text-conditioned U-Net and VAE encoding/decoding process.

- Tags: architecture
- Published: 2026-03-04

### [How the Switch Transformer Routing Mechanism Works: Token-to-Expert Selection Explained](/labmlai/annotated_deep_learning_paper_implementations/how-does-the-switch-transformer-routing-mechanism-work)

Learn how the Switch Transformer routing mechanism assigns tokens to experts with softmax, enforces capacity limits, and scales output by confidence for efficient sparse training.

- Tags: deep-dive
- Published: 2026-03-04

### [How ALiBi Removes Positional Embeddings While Preserving Attention Patterns](/labmlai/annotated_deep_learning_paper_implementations/how-does-alibi-remove-positional-embeddings-while-preserving-attention-patterns)

Discover how ALiBi removes positional embeddings by adding linear biases to attention logits. Learn how this preserves attention patterns & relative ordering without learned vectors.

- Tags: deep-dive
- Published: 2026-03-04

### [RAdam vs Adam Optimizer: Key Differences for Stable Deep Learning Training](/labmlai/annotated_deep_learning_paper_implementations/what-is-the-difference-between-radam-and-standard-adam-optimizer-for-deep-learning)

Discover the key differences between RAdam and Adam optimizers for stable deep learning training. RAdam offers automatic warm-up and faster convergence.

- Tags: deep-dive
- Published: 2026-03-04

### [How the Dueling Network Architecture Is Implemented in DQN](/labmlai/annotated_deep_learning_paper_implementations/how-is-the-dueling-network-architecture-implemented-in-dqn)

Learn how the dueling network architecture is implemented in DQN. Discover how state-value and action-advantage streams combine for stable learning in this detailed guide.

- Tags: deep-dive
- Published: 2026-03-04

### [How PPO with GAE Improves Reinforcement Learning Sample Efficiency](/labmlai/annotated_deep_learning_paper_implementations/how-does-ppo-with-gae-improve-reinforcement-learning-sample-efficiency)

Discover how PPO with GAE enhances reinforcement learning sample efficiency. Learn to update policies stably and informatively with low-variance advantage estimates from limited data.

- Tags: tutorial
- Published: 2026-03-04

### [StyleGAN 2 vs Original GAN: 8 Key Architectural Differences](/labmlai/annotated_deep_learning_paper_implementations/what-makes-stylegan-2-different-from-the-original-gan-implementation)

Discover 8 key architectural differences between StyleGAN 2 and original GANs. Learn how StyleGAN 2 achieves high-fidelity image synthesis with its improved generator-discriminator pipeline.

- Tags: deep-dive
- Published: 2026-03-04

### [How kNN-LM Combines Retrieval and Language Modeling: Implementation Guide](/labmlai/annotated_deep_learning_paper_implementations/how-does-knn-lm-combine-retrieval-and-language-modeling)

Learn how kNN-LM combines retrieval and language modeling by interpolating next-token and k-nearest neighbor distributions. Implement this powerful technique with our guide.

- Tags: how-to-guide
- Published: 2026-03-04

### [LayerNorm vs GroupNorm for Transformer Training: Key Differences and Implementation Guide](/labmlai/annotated_deep_learning_paper_implementations/what-are-the-key-differences-between-layernorm-and-groupnorm-for-transformer-training)

Discover the key differences between LayerNorm and GroupNorm for transformer training. Learn which is best for NLP and vision tasks with this implementation guide.

- Tags: deep-dive
- Published: 2026-03-04

### [How Rotary Positional Embedding Is Implemented in PyTorch: A Deep Dive into labmlai/annotated_deep_learning_paper_implementations](/labmlai/annotated_deep_learning_paper_implementations/how-is-rotary-positional-embedding-implemented-in-pytorch)

Learn how Rotary Positional Embedding is implemented in PyTorch using labmlai. Discover efficient rotation techniques with cached sine cosine matrices.

- Tags: deep-dive
- Published: 2026-03-04

### [How the Attention Mechanism Works in Transformer XL with Relative Positional Embeddings](/labmlai/annotated_deep_learning_paper_implementations/how-does-the-attention-mechanism-work-in-transformer-xl-with-relative-positional-embeddings)

Explore how Transformer XL uses the attention mechanism with relative positional embeddings for advanced context reuse and temporal coherence. Learn about segment-level recurrence.

- Tags: deep-dive
- Published: 2026-03-04

### [GAT vs GATv2: Key Differences in Graph Attention Networks Explained](/labmlai/annotated_deep_learning_paper_implementations/what-are-the-differences-between-gat-and-gatv2-for-graph-neural-networks)

GATv2 enhances GAT with dynamic attention, learning query-dependent neighbor rankings impossible for original GAT. Explore key GAT vs GATv2 differences.

- Tags: deep-dive
- Published: 2026-03-04

### [How WGAN-GP Addresses Mode Collapse in GAN Training: A Code-First Guide](/labmlai/annotated_deep_learning_paper_implementations/how-does-wgan-gp-address-mode-collapse-in-gan-training)

Learn how WGAN-GP prevents mode collapse in GANs. Discover its gradient penalty approach for stable critic training and better mode coverage. Code first guide.

- Tags: deep-dive
- Published: 2026-03-04

### [PonderNet Adaptive Computation in Transformers: Implementation Guide](/labmlai/annotated_deep_learning_paper_implementations/how-is-pondernet-implemented-for-adaptive-computation-in-transformers)

Learn how PonderNet implements adaptive computation in transformers. Discover dynamic step adjustment with halting probabilities and reconstruction loss optimization. Get the guide now.

- Tags: implementation-guide
- Published: 2026-03-04

### [Nucleus Sampling vs Top-k Sampling for Language Model Decoding: Implementation Differences](/labmlai/annotated_deep_learning_paper_implementations/how-do-nucleus-sampling-and-top-k-sampling-differ-for-language-model-decoding)

Discover the implementation differences between nucleus sampling and top-k sampling for language model decoding. Learn how nucleus sampling adapts to probability distributions for better text generation.

- Tags: deep-dive
- Published: 2026-03-04

### [Difference Between DDPM and DDIM Sampling in Diffusion Models: Implementation Guide](/labmlai/annotated_deep_learning_paper_implementations/what-is-the-difference-between-ddpm-and-ddim-sampling-in-diffusion-models)

Discover the core differences between DDPM and DDIM sampling in diffusion models. Learn how DDIM offers faster, high-quality generation with fewer steps in this implementation guide.

- Tags: tutorial
- Published: 2026-03-04

### [How Zero3 Enables Large Model Training on Limited GPU Memory: A Deep Dive into Sharded Data Parallelism](/labmlai/annotated_deep_learning_paper_implementations/how-does-zero3-enable-large-model-training-on-limited-gpu-memory)

Discover how Zero3 uses sharded data parallelism to train large models on limited GPU memory by fetching active parameter shards on demand. Learn more today.

- Tags: deep-dive
- Published: 2026-03-04

### [How LoRA is Implemented for Efficient Transformer Fine-Tuning in PyTorch](/labmlai/annotated_deep_learning_paper_implementations/how-is-lora-implemented-for-efficient-transformer-fine-tuning-in-pytorch)

Learn how LoRA fine-tuning efficiently adapts transformers in PyTorch. Freeze pre-trained weights and inject low-rank matrices drastically reducing parameters.

- Tags: how-to-guide
- Published: 2026-03-04

### [How Flash Attention Improves Transformer Memory Efficiency Compared to Standard Attention](/labmlai/annotated_deep_learning_paper_implementations/how-does-flash-attention-improve-transformer-memory-efficiency-compared-to-standard-attention)

Discover how Flash Attention slashes transformer memory use from O(L²) to O(L) by processing attention block-wise sans full matrix materialization. Learn the technique.

- Tags: deep-dive
- Published: 2026-03-04

