# L1 vs L2 vs Dropout Regularization: Key Differences in Deep Learning

> Understand L1 L2 and Dropout regularization key differences in deep learning. Learn how L1 zeros weights L2 shrinks them and Dropout prevents co-adaptation for better model performance.

- Repository: [scutan90/DeepLearning-500-questions](https://github.com/scutan90/DeepLearning-500-questions)
- Tags: deep-dive
- Published: 2026-03-06

---

**L1 regularization drives weights to exact zero for feature selection, L2 shrinks all weights smoothly toward zero for stable optimization, and Dropout prevents neuron co-adaptation by randomly zeroing activations during training without modifying the loss function.**

Regularization techniques are essential for preventing overfitting in neural networks. According to the `scutan90/DeepLearning-500-questions` repository, understanding the distinctions between **L1, L2, and Dropout regularization** is fundamental to building generalizable deep learning models. These three approaches differ fundamentally in their mathematical formulation, effect on network parameters, and implementation strategies.

## Mathematical Foundations

### L1 Regularization (Lasso)

L1 regularization adds the L1-norm of the weight vector to the loss function:

```

L_total = L_data + λ·∑|w_i|

```

As documented in `ch13_优化算法/第十三章_优化算法.md`, this linear penalty creates a sharp corner at zero in the optimization landscape. The absolute value function is non-differentiable at zero, which can lead to slower convergence but drives many weights to exactly zero.

### L2 Regularization (Ridge)

L2 regularization adds the squared L2-norm of the weights to the loss:

```

L_total = L_data + λ·∑w_i²

```

This quadratic penalty is smooth and differentiable everywhere, making gradient-based optimization more stable. According to the repository's optimization algorithms chapter, L2 provides consistent gradients that shrink all weights proportionally toward zero without inducing exact sparsity.

### Dropout

Dropout does not add an explicit term to the loss function. During each forward pass, neurons are retained with probability *p* and set to zero otherwise. At test time, activations are scaled by `1/(1-p)` to maintain expected values.

As explained in `ch03_深度学习基础/第三章_深度学习基础.md`, this creates an implicit ensemble of sub-networks. Each training iteration samples a different "thinned" network, preventing neurons from co-adapting to specific features.

## Effects on Network Weights

### Sparsity vs. Shrinkage

- **L1 regularization** encourages **sparsity**—many weights become exactly zero, effectively performing feature selection. The sharp corner at zero in the L1 penalty drives irrelevant weights to extinction, making the model more interpretable and compact.
- **L2 regularization** encourages **small but dense** weights—all parameters shrink toward zero but rarely become exactly zero. This distributes importance across all features rather than selecting a subset, which often improves generalization when all features contribute weakly to the prediction.
- **Dropout** forces **robustness**—neurons cannot rely on specific other neurons because any unit may be dropped during training. Each neuron must learn features that are useful independently, reducing complex co-adaptations that lead to overfitting.

### Optimization Characteristics

According to `ch13_优化算法/第十三章_优化算法.md`:

- **L1** creates a non-differentiable point at zero, which can cause optimization difficulties with standard SGD. Subgradient methods or coordinate descent are often more effective for L1 problems.
- **L2** provides smooth, consistent gradients that work well with standard stochastic gradient descent and its variants (Adam, RMSprop).
- **Dropout** introduces training stochasticity—each mini-batch sees a different network architecture. This can smooth the loss surface but typically requires more training epochs to converge.

## Implementation in PyTorch

The `scutan90/DeepLearning-500-questions` repository emphasizes practical implementation. Below are canonical PyTorch implementations for each technique.

### L2 Regularization (Weight Decay)

PyTorch optimizers implement L2 via the `weight_decay` parameter:

```python
import torch.optim as optim

# L2 regularization via weight_decay parameter

optimizer = optim.SGD(model.parameters(), 
                      lr=0.01, 
                      momentum=0.9,
                      weight_decay=1e-4)  # λ = 0.0001

```

The optimizer automatically adds `λ·∑w_i²` to the loss during the backward pass.

### L1 Regularization (Manual Implementation)

L1 requires manual computation as PyTorch does not provide a built-in parameter:

```python
import torch
import torch.nn as nn

def l1_penalty(model, lambda_l1=1e-5):
    """Compute L1 norm of all model parameters."""
    return sum(p.abs().sum() for p in model.parameters()) * lambda_l1

criterion = nn.CrossEntropyLoss()
optimizer = optim.SGD(model.parameters(), lr=0.01)

# Training loop

for inputs, targets in dataloader:
    optimizer.zero_grad()
    outputs = model(inputs)
    loss = criterion(outputs, targets)
    
    # Add L1 penalty manually

    loss = loss + l1_penalty(model, lambda_l1=1e-5)
    
    loss.backward()
    optimizer.step()

```

### Dropout

Dropout is implemented as a layer that automatically handles training and evaluation modes:

```python
import torch.nn as nn

class Net(nn.Module):
    def __init__(self, dropout_p=0.5):
        super().__init__()
        self.fc1 = nn.Linear(784, 256)
        self.relu = nn.ReLU()
        self.dropout = nn.Dropout(p=dropout_p)  # Dropout probability

        self.fc2 = nn.Linear(256, 10)

    def forward(self, x):
        x = self.relu(self.fc1(x))
        x = self.dropout(x)  # Active only during training

        return self.fc2(x)

model = Net(dropout_p=0.3)

# Dropout automatically disabled during model.eval()

```

## When to Use Each Technique

According to `ch14_超参数调整/第十四章_超参数调整.md` and `ch03_深度学习基础/第三章_深度学习基础.md`:

**Choose L1 when:**
- You need **feature selection** or model interpretability
- Deploying to resource-constrained environments where sparse weights reduce memory footprint
- Working with high-dimensional data where many features are irrelevant

**Choose L2 when:**
- You need a **baseline regularization** for general deep learning pipelines
- Training CNNs, RNNs, or Transformers where smooth optimization is critical
- You want a single hyperparameter (`weight_decay`) that is easy to tune

**Choose Dropout when:**
- You have **large fully-connected layers** prone to co-adaptation
- You need **strong regularization** for very deep networks
- You want an implicit ensemble effect without training multiple models

**Combining Techniques:**
As noted in the repository, L2 and Dropout are frequently used together—L2 prevents weight growth while Dropout prevents complex co-adaptations. However, placing Dropout directly after Batch Normalization layers should be avoided as it can interfere with batch statistics.

## Summary

- **L1 regularization** adds `λ·∑|w_i|` to the loss, driving weights to exactly zero and creating sparse models suitable for feature selection.
- **L2 regularization** adds `λ·∑w_i²` to the loss, shrinking all weights smoothly toward zero without inducing sparsity, providing stable gradient flow.
- **Dropout** stochastically zeroes neurons during training (probability *p*) and scales outputs at test time, preventing co-adaptation and effectively ensembling sub-networks.
- According to `scutan90/DeepLearning-500-questions`, L1 is preferred for sparsity, L2 for general weight decay, and Dropout for deep fully-connected networks; they can be combined but require careful placement relative to BatchNorm.

## Frequently Asked Questions

### Can I use L1 and L2 regularization together?

Yes, you can combine L1 and L2 regularization in a technique called **Elastic Net**. In PyTorch, you would apply L2 via the optimizer's `weight_decay` parameter and manually add the L1 penalty to the loss function during the training loop, as shown in the implementation examples above.

### Why does Dropout require scaling at test time?

During training, Dropout retains each neuron with probability *p*, meaning the expected output is scaled by *p* compared to using all neurons. To ensure that the expected output remains consistent at test time (when all neurons are active), activations must be scaled by `1/(1-p)` (inverted dropout) or weights must be scaled at test time. This maintains the same expected magnitude for downstream layers.

### Is Dropout still necessary when using Batch Normalization?

According to `ch14_超参数调整/第十四章_超参数调整.md`, Dropout and Batch Normalization serve different purposes and can be used together, but placement matters. Dropout should not be placed directly after BatchNorm because it disrupts the batch statistics calculated during training. Modern architectures often use one or the other, with BatchNorm providing regularization through noise in batch statistics and Dropout providing explicit stochasticity.

### How do I choose the regularization strength λ for L1 or L2?

The regularization coefficient λ is typically chosen through cross-validation or grid search. As documented in `ch14_超参数调整/第十四章_超参数调整.md`, common starting values for L2 weight decay in deep learning range from `1e-4` to `1e-2`. For L1, smaller values (e.g., `1e-5` to `1e-3`) are often used because the absolute value penalty can drive weights to zero more aggressively than the squared penalty. You should monitor validation loss and model sparsity (for L1) to determine the optimal trade-off between fitting the training data and maintaining generalization.