# d2l-zh | Dive into Deep Learning (D2L.ai) | Knowledge Base | Instagit

《动手学深度学习》：面向中文读者、能运行、可讨论。中英文版被70多个国家的500多所大学用于教学。

GitHub Stars: 75.9k

Repository: https://github.com/d2l-ai/d2l-zh

---

## Articles

### [How d2l-zh Explains Learning Rate Optimization for Model Convergence](/d2l-ai/d2l-zh/learning-rate-optimization-model-convergence-d2l-zh)

d2l-zh explains learning rate optimization for model convergence detailing optimal ranges and dynamic schedules like cosine decay with warm-up for stable, fast deep network training.

- Tags: tutorial
- Published: 2026-03-01

### [Strategies for Handling Imbalanced Datasets in Deep Learning: A Practical Guide from d2l-zh](/d2l-ai/d2l-zh/handling-imbalanced-datasets-deep-learning-d2l-zh)

Discover d2l-zh strategies for imbalanced datasets in deep learning: reweighting loss, focal loss, and resampling. Improve your model accuracy now.

- Tags: how-to-guide
- Published: 2026-03-01

### [How d2l-zh Addresses the Challenges of Training Very Deep Neural Networks](/d2l-ai/d2l-zh/challenges-training-very-deep-neural-networks-d2l-zh)

Learn how d2l-zh tackles deep neural network training challenges using residual connections, batch normalization, and more for stable, multi-layer network training.

- Tags: deep-dive
- Published: 2026-03-01

### [Activation Functions in d2l-zh: Types, Implementation, and Usage Across Deep Learning Frameworks](/d2l-ai/d2l-zh/activation-functions-types-usage-d2l-zh)

Explore activation functions in d2l-zh: ReLU, tanh, Sigmoid, and masked Softmax. Understand their implementation and usage across PyTorch TensorFlow MXNet and Paddle for deep learning.

- Tags: deep-dive
- Published: 2026-03-01

### [How d2l-zh Explains Gradient Descent and Its Variants: From First Principles to Adaptive Optimization](/d2l-ai/d2l-zh/gradient-descent-concepts-variants-d2l-zh)

Explore how d2l-zh explains gradient descent and its variants using Taylor expansion, stochastic sampling, momentum, and adaptive learning rates for efficient optimization.

- Tags: deep-dive
- Published: 2026-03-01

### [Hyperparameter Tuning in d2l-zh: Best Practices from Dive into Deep Learning](/d2l-ai/d2l-zh/hyperparameter-tuning-best-practices-d2l-zh)

Master hyperparameter tuning in d2l-zh Learn best practices like validation sets and K-fold cross-validation to optimize your deep learning models.

- Tags: best-practices
- Published: 2026-03-01

### [How d2l-zh Explains the Encoder-Decoder Architecture for Sequence-to-Sequence Tasks](/d2l-ai/d2l-zh/encoder-decoder-architecture-seq2seq-d2l-zh)

Understand the encoder-decoder architecture for sequence-to-sequence tasks with d2l-zh. Learn how GRU-based recurrent networks implement this powerful framework for machine translation.

- Tags: deep-dive
- Published: 2026-03-01

### [The Role and Impact of Batch Normalization in Model Training According to d2l-zh](/d2l-ai/d2l-zh/batch-normalization-role-impact-d2l-zh)

Discover how batch normalization stabilizes hidden activations and accelerates deep neural network training according to d2l-zh. Learn its impact and role.

- Tags: deep-dive
- Published: 2026-03-01

### [Preventing Overfitting in Deep Learning Models: Techniques from d2l-zh](/d2l-ai/d2l-zh/preventing-overfitting-techniques-d2l-zh)

Learn how to prevent overfitting in deep learning models with d2l-zh. Explore techniques like controlling model capacity, L₂ regularization, and dropout. Improve your model's generalization.

- Tags: tutorial
- Published: 2026-03-01

### [How Word Embeddings Are Implemented and Utilized in d2l-zh](/d2l-ai/d2l-zh/word-embeddings-implementation-utilization-d2l-zh)

Discover how word embeddings are implemented and utilized in d2l-zh. Explore static pre-trained vectors like GloVe and fastText, and learn about learned embeddings within models like BERT.

- Tags: how-to-guide
- Published: 2026-03-01

### [Differences Between LSTM and GRU in d2l-zh: Architectural and Performance Comparison](/d2l-ai/d2l-zh/lstm-vs-gru-differences-d2l-zh)

Explore LSTM vs GRU differences in d2l-zh. Understand how LSTM's three gates and cell state contrast with GRU's two gates for efficient long-term dependency management.

- Tags: deep-dive
- Published: 2026-03-01

### [How d2l-zh Explains the Backpropagation Algorithm for Training Neural Networks](/d2l-ai/d2l-zh/backpropagation-algorithm-explanation-d2l-zh)

Learn how d2l-zh explains backpropagation using computational graphs and the chain rule for efficient neural network training. Understand gradient computation from output to input.

- Tags: deep-dive
- Published: 2026-03-01

### [Vanishing Gradient Problem in Deep Learning: How d2l-zh Mitigates It](/d2l-ai/d2l-zh/vanishing-gradient-problem-addressing-d2l-zh)

Understand the vanishing gradient problem in deep learning and discover how d2l-zh effectively mitigates it using ReLU, Xavier initialization, and LSTM.

- Tags: deep-dive
- Published: 2026-03-01

### [How d2l-zh Approaches Natural Language Processing (NLP) Tasks with Deep Learning](/d2l-ai/d2l-zh/nlp-tasks-deep-learning-approaches-d2l-zh)

Explore how d2l-zh leverages deep learning for NLP tasks. Discover its three stage pipeline static word embeddings contextual Transformer pre-training and task specific fine-tuning with multiple back-ends.

- Tags: how-to-guide
- Published: 2026-03-01

### [Practical Applications of Deep Learning in Computer Vision: The d2l-zh Implementation Guide](/d2l-ai/d2l-zh/computer-vision-applications-d2l-zh)

Explore d2l-zh's practical deep learning applications in computer vision including image classification object detection semantic segmentation and more Learn to implement these models with framework agnostic code

- Tags: how-to-guide
- Published: 2026-03-01

### [How d2l-zh Explains Optimization Strategies for Deep Learning Models: From Gradient Descent to Adam](/d2l-ai/d2l-zh/optimization-strategies-deep-learning-d2l-zh)

Explore deep learning optimization strategies from Gradient Descent to Adam with d2l-zh. Learn theory and code in chapter_optimization.

- Tags: deep-dive
- Published: 2026-03-01

### [Advanced Concepts of Attention Mechanisms in Dive into Deep Learning (d2l-zh)](/d2l-ai/d2l-zh/advanced-attention-mechanisms-d2l-zh)

Explore advanced attention mechanisms like Scaled Dot-Product, Additive, and Multi-Head Attention in d2l-zh. Dive deep into self-attention with positional encoding and masked softmax utilities for cutting-edge AI.

- Tags: deep-dive
- Published: 2026-03-01

### [How d2l-zh Details the Functioning of Recurrent Neural Networks (RNNs): From Mathematical Foundations to Multi-Framework Implementation](/d2l-ai/d2l-zh/rnn-functioning-details-d2l-zh)

Explore how d2l-zh explains Recurrent Neural Networks RNNs. Learn mathematical foundations and see PyTorch TensorFlow MXNet and PaddlePaddle implementations.

- Tags: deep-dive
- Published: 2026-03-01

### [Key Architectural Components of Convolutional Neural Networks (CNNs) According to d2l-zh](/d2l-ai/d2l-zh/cnn-architecture-components-d2l-zh)

Discover the 8 key architectural components of CNNs explained by d2l-zh including convolutional layers pooling layers batch normalization and more Learn how CNNs achieve spatial invariance and efficiency

- Tags: architecture
- Published: 2026-03-01

### [How d2l-zh Introduces and Explains Multilayer Perceptrons (MLPs): Theory, Math, and Code](/d2l-ai/d2l-zh/multilayer-perceptrons-introduction-d2l-zh)

Explore how d2l-zh explains multilayer perceptrons (MLPs) with theory, math, and code. Understand architecture activation functions and get hands-on implementations.

- Tags: deep-dive
- Published: 2026-03-01

### [Mathematical Basis for Linear Networks in d2l-zh: Affine Transformations and Maximum Likelihood](/d2l-ai/d2l-zh/mathematical-basis-linear-networks-d2l-zh)

Understand the math behind d2l-zh linear networks. Explore affine transformations and maximum likelihood estimation for MSE and cross-entropy loss functions.

- Tags: deep-dive
- Published: 2026-03-01

### [How d2l-zh Explains Machine Learning Fundamentals for Practitioners: A Code-First Approach](/d2l-ai/d2l-zh/machine-learning-fundamentals-d2l-zh-explanation)

d2l-zh explains machine learning fundamentals with runnable code examples and a seven-step workflow. Learn core concepts through practical application and a focused utility library.

- Tags: deep-dive
- Published: 2026-03-01

### [Core Concepts of Deep Learning in d2l-zh: The Four-Component Framework Explained](/d2l-ai/d2l-zh/core-concepts-deep-learning-d2l-zh)

Explore deep learning core concepts in d2l-zh. Understand the four components data model objective function and optimization algorithm plus layered architectures and back-propagation with this guide.

- Tags: deep-dive
- Published: 2026-03-01

