L1 vs L2 vs Dropout Regularization: Key Differences in Deep Learning

L1 regularization drives weights to exact zero for feature selection, L2 shrinks all weights smoothly toward zero for stable optimization, and Dropout prevents neuron co-adaptation by randomly zeroing activations during training without modifying the loss function.

Regularization techniques are essential for preventing overfitting in neural networks. According to the scutan90/DeepLearning-500-questions repository, understanding the distinctions between L1, L2, and Dropout regularization is fundamental to building generalizable deep learning models. These three approaches differ fundamentally in their mathematical formulation, effect on network parameters, and implementation strategies.

Mathematical Foundations

L1 Regularization (Lasso)

L1 regularization adds the L1-norm of the weight vector to the loss function:


L_total = L_data + λ·∑|w_i|

As documented in ch13_优化算法/第十三章_优化算法.md, this linear penalty creates a sharp corner at zero in the optimization landscape. The absolute value function is non-differentiable at zero, which can lead to slower convergence but drives many weights to exactly zero.

L2 Regularization (Ridge)

L2 regularization adds the squared L2-norm of the weights to the loss:


L_total = L_data + λ·∑w_i²

This quadratic penalty is smooth and differentiable everywhere, making gradient-based optimization more stable. According to the repository's optimization algorithms chapter, L2 provides consistent gradients that shrink all weights proportionally toward zero without inducing exact sparsity.

Dropout

Dropout does not add an explicit term to the loss function. During each forward pass, neurons are retained with probability p and set to zero otherwise. At test time, activations are scaled by 1/(1-p) to maintain expected values.

As explained in ch03_深度学习基础/第三章_深度学习基础.md, this creates an implicit ensemble of sub-networks. Each training iteration samples a different "thinned" network, preventing neurons from co-adapting to specific features.

Effects on Network Weights

Sparsity vs. Shrinkage

  • L1 regularization encourages sparsity—many weights become exactly zero, effectively performing feature selection. The sharp corner at zero in the L1 penalty drives irrelevant weights to extinction, making the model more interpretable and compact.
  • L2 regularization encourages small but dense weights—all parameters shrink toward zero but rarely become exactly zero. This distributes importance across all features rather than selecting a subset, which often improves generalization when all features contribute weakly to the prediction.
  • Dropout forces robustness—neurons cannot rely on specific other neurons because any unit may be dropped during training. Each neuron must learn features that are useful independently, reducing complex co-adaptations that lead to overfitting.

Optimization Characteristics

According to ch13_优化算法/第十三章_优化算法.md:

  • L1 creates a non-differentiable point at zero, which can cause optimization difficulties with standard SGD. Subgradient methods or coordinate descent are often more effective for L1 problems.
  • L2 provides smooth, consistent gradients that work well with standard stochastic gradient descent and its variants (Adam, RMSprop).
  • Dropout introduces training stochasticity—each mini-batch sees a different network architecture. This can smooth the loss surface but typically requires more training epochs to converge.

Implementation in PyTorch

The scutan90/DeepLearning-500-questions repository emphasizes practical implementation. Below are canonical PyTorch implementations for each technique.

L2 Regularization (Weight Decay)

PyTorch optimizers implement L2 via the weight_decay parameter:

import torch.optim as optim

# L2 regularization via weight_decay parameter

optimizer = optim.SGD(model.parameters(), 
                      lr=0.01, 
                      momentum=0.9,
                      weight_decay=1e-4)  # λ = 0.0001

The optimizer automatically adds λ·∑w_i² to the loss during the backward pass.

L1 Regularization (Manual Implementation)

L1 requires manual computation as PyTorch does not provide a built-in parameter:

import torch
import torch.nn as nn

def l1_penalty(model, lambda_l1=1e-5):
    """Compute L1 norm of all model parameters."""
    return sum(p.abs().sum() for p in model.parameters()) * lambda_l1

criterion = nn.CrossEntropyLoss()
optimizer = optim.SGD(model.parameters(), lr=0.01)

# Training loop

for inputs, targets in dataloader:
    optimizer.zero_grad()
    outputs = model(inputs)
    loss = criterion(outputs, targets)
    
    # Add L1 penalty manually

    loss = loss + l1_penalty(model, lambda_l1=1e-5)
    
    loss.backward()
    optimizer.step()

Dropout

Dropout is implemented as a layer that automatically handles training and evaluation modes:

import torch.nn as nn

class Net(nn.Module):
    def __init__(self, dropout_p=0.5):
        super().__init__()
        self.fc1 = nn.Linear(784, 256)
        self.relu = nn.ReLU()
        self.dropout = nn.Dropout(p=dropout_p)  # Dropout probability

        self.fc2 = nn.Linear(256, 10)

    def forward(self, x):
        x = self.relu(self.fc1(x))
        x = self.dropout(x)  # Active only during training

        return self.fc2(x)

model = Net(dropout_p=0.3)

# Dropout automatically disabled during model.eval()

When to Use Each Technique

According to ch14_超参数调整/第十四章_超参数调整.md and ch03_深度学习基础/第三章_深度学习基础.md:

Choose L1 when:

  • You need feature selection or model interpretability
  • Deploying to resource-constrained environments where sparse weights reduce memory footprint
  • Working with high-dimensional data where many features are irrelevant

Choose L2 when:

  • You need a baseline regularization for general deep learning pipelines
  • Training CNNs, RNNs, or Transformers where smooth optimization is critical
  • You want a single hyperparameter (weight_decay) that is easy to tune

Choose Dropout when:

  • You have large fully-connected layers prone to co-adaptation
  • You need strong regularization for very deep networks
  • You want an implicit ensemble effect without training multiple models

Combining Techniques: As noted in the repository, L2 and Dropout are frequently used together—L2 prevents weight growth while Dropout prevents complex co-adaptations. However, placing Dropout directly after Batch Normalization layers should be avoided as it can interfere with batch statistics.

Summary

  • L1 regularization adds λ·∑|w_i| to the loss, driving weights to exactly zero and creating sparse models suitable for feature selection.
  • L2 regularization adds λ·∑w_i² to the loss, shrinking all weights smoothly toward zero without inducing sparsity, providing stable gradient flow.
  • Dropout stochastically zeroes neurons during training (probability p) and scales outputs at test time, preventing co-adaptation and effectively ensembling sub-networks.
  • According to scutan90/DeepLearning-500-questions, L1 is preferred for sparsity, L2 for general weight decay, and Dropout for deep fully-connected networks; they can be combined but require careful placement relative to BatchNorm.

Frequently Asked Questions

Can I use L1 and L2 regularization together?

Yes, you can combine L1 and L2 regularization in a technique called Elastic Net. In PyTorch, you would apply L2 via the optimizer's weight_decay parameter and manually add the L1 penalty to the loss function during the training loop, as shown in the implementation examples above.

Why does Dropout require scaling at test time?

During training, Dropout retains each neuron with probability p, meaning the expected output is scaled by p compared to using all neurons. To ensure that the expected output remains consistent at test time (when all neurons are active), activations must be scaled by 1/(1-p) (inverted dropout) or weights must be scaled at test time. This maintains the same expected magnitude for downstream layers.

Is Dropout still necessary when using Batch Normalization?

According to ch14_超参数调整/第十四章_超参数调整.md, Dropout and Batch Normalization serve different purposes and can be used together, but placement matters. Dropout should not be placed directly after BatchNorm because it disrupts the batch statistics calculated during training. Modern architectures often use one or the other, with BatchNorm providing regularization through noise in batch statistics and Dropout providing explicit stochasticity.

How do I choose the regularization strength λ for L1 or L2?

The regularization coefficient λ is typically chosen through cross-validation or grid search. As documented in ch14_超参数调整/第十四章_超参数调整.md, common starting values for L2 weight decay in deep learning range from 1e-4 to 1e-2. For L1, smaller values (e.g., 1e-5 to 1e-3) are often used because the absolute value penalty can drive weights to zero more aggressively than the squared penalty. You should monitor validation loss and model sparsity (for L1) to determine the optimal trade-off between fitting the training data and maintaining generalization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →