# Activation Functions in d2l-zh: Types, Implementation, and Usage Across Deep Learning Frameworks

> Explore activation functions in d2l-zh: ReLU, tanh, Sigmoid, and masked Softmax. Understand their implementation and usage across PyTorch TensorFlow MXNet and Paddle for deep learning.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: deep-dive
- Published: 2026-03-01

---

**The d2l-zh educational framework implements four primary activation functions—ReLU, tanh, Sigmoid, and masked Softmax—across PyTorch, TensorFlow, MXNet, and Paddle backends, with ReLU serving as the default non-linearity for modern architectures while Sigmoid remains primarily for legacy educational examples.**

The d2l-zh repository (d2l-ai/d2l-zh) provides a unified interface for deep learning education across multiple frameworks. Understanding the **activation functions** available in this codebase is essential for implementing neural networks effectively. This guide examines the four core activation types used throughout the library, their specific backend implementations, and practical usage patterns in both modern and classical architectures.

## Core Activation Functions in d2l-zh

### ReLU (Rectified Linear Unit)

**ReLU** is the default activation function in d2l-zh for modern architectures. In [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py) (lines 1294-1296), the implementation wraps `nn.ReLU()` for PyTorch. The function appears in position-wise feed-forward networks within Transformer blocks, where the code pattern `self.dense2(self.relu(self.dense1(X)))` applies the non-linearity between two linear transformations.

### Tanh (Hyperbolic Tangent)

The **tanh** activation appears frequently in attention mechanisms and classic RNN cells throughout d2l-zh. According to [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py) (lines 1163-1164), the implementation uses `torch.tanh` directly. In additive attention implementations, intermediate feature maps pass through tanh to produce bounded output values that stabilize similarity scores.

### Sigmoid

**Sigmoid** functions in d2l-zh serve primarily educational purposes, appearing in legacy architectures like LeNet-5. The [`chapter_convolutional-neural-networks/lenet.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_convolutional-neural-networks/lenet.md) file demonstrates `nn.Sigmoid()` in PyTorch (lines 61-68) and `tf.keras.layers.Activation('sigmoid')` in TensorFlow (lines 78-80). These implementations follow classical convolutional network designs where sigmoid activations follow each convolutional or dense layer.

### Masked Softmax

Unlike standard activations, **masked Softmax** in d2l-zh specifically handles variable-length sequences in language models and machine translation. The `masked_softmax` function in [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py) (lines 1127-1134) masks padding tokens before applying `nn.functional.softmax`, preventing invalid positions from affecting probability distributions. TensorFlow implements analogous logic in [`d2l/tensorflow.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/tensorflow.py) (lines 1061-1068).

## Implementation Across Deep Learning Frameworks

### PyTorch Implementation

In the PyTorch backend ([`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py)), activation functions appear as explicit layer objects or functional calls. ReLU instantiates as `self.relu = nn.ReLU()` within module definitions, while tanh operates as `torch.tanh(X)` for tensor transformations. The masked softmax implementation handles batch matrix operations with explicit padding masks.

### TensorFlow and Keras

The TensorFlow backend ([`d2l/tensorflow.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/tensorflow.py)) utilizes Keras layer activations, particularly for Sigmoid in legacy models. The LeNet implementation demonstrates `tf.keras.layers.Activation('sigmoid')` following convolutional layers. Modern architectures prefer `tf.keras.layers.ReLU()` for consistency with current best practices.

### MXNet and Gluon

MXNet support in [`d2l/mxnet.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/mxnet.py) employs activation keywords in layer constructors. Dense layers specify activation via the `activation='relu'` or `activation='tanh'` parameters, allowing the backend to handle non-linearity application internally. This approach reduces explicit activation layer definitions in model code.

### PaddlePaddle

The Paddle backend ([`d2l/paddle.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/paddle.py)) follows PyTorch patterns with explicit activation layers. The `PositionWiseFFN` class demonstrates `self.relu = nn.ReLU()` initialization and application within feed-forward networks, maintaining API consistency across the d2l-zh framework.

## Practical Usage Examples

The following snippets demonstrate activation function usage across different backends in d2l-zh.

PyTorch implementation with ReLU and explicit layer definition:

```python
import torch
from torch import nn
from d2l import torch as d2l

class SimpleFFN(nn.Module):
    def __init__(self, in_dim, hidden_dim, out_dim):
        super().__init__()
        self.dense1 = nn.Linear(in_dim, hidden_dim)
        self.relu = nn.ReLU()               # ReLU activation layer

        self.dense2 = nn.Linear(hidden_dim, out_dim)

    def forward(self, X):
        X = self.relu(self.dense1(X))       # Apply ReLU

        return self.dense2(X)

```

TensorFlow implementation using Sigmoid in the classic LeNet architecture:

```python
import tensorflow as tf
from d2l import tensorflow as d2l

def lenet():
    return tf.keras.models.Sequential([
        tf.keras.layers.Conv2D(6, 5, padding='same', activation='sigmoid'),
        tf.keras.layers.AvgPool2D(2, 2),
        tf.keras.layers.Conv2D(16, 5, activation='sigmoid'),
        tf.keras.layers.AvgPool2D(2, 2),
        tf.keras.layers.Flatten(),
        tf.keras.layers.Dense(120, activation='sigmoid'),
        tf.keras.layers.Dense(84, activation='sigmoid'),
        tf.keras.layers.Dense(10)  # logits

    ])

```

MXNet implementation using keyword-based activation arguments:

```python
from mxnet import gluon, nd
from d2l import mxnet as d2l

net = gluon.nn.Sequential()
with net.name_scope():
    net.add(gluon.nn.Dense(256, activation='relu'))   # ReLU via keyword

    net.add(gluon.nn.Dense(128, activation='tanh'))   # tanh via keyword

    net.add(gluon.nn.Dense(10))                      # logits

net.initialize()

```

PaddlePaddle implementation with explicit ReLU layers in Transformer blocks:

```python
import paddle
from paddle import nn
from d2l import paddle as d2l

class PositionWiseFFN(nn.Layer):
    def __init__(self, ffn_num_input, ffn_num_hiddens, ffn_num_outputs):
        super().__init__()
        self.dense1 = nn.Linear(ffn_num_input, ffn_num_hiddens)
        self.relu = nn.ReLU()                     # Explicit ReLU layer

        self.dense2 = nn.Linear(ffn_num_hiddens, ffn_num_outputs)

    def forward(self, X):
        return self.dense2(self.relu(self.dense1(X)))

```

## Summary

- **ReLU** serves as the default activation for modern architectures in d2l-zh, implemented via `nn.ReLU()` in PyTorch and Paddle or `activation='relu'` in MXNet.
- **Tanh** remains essential for attention mechanisms and RNN cells, providing bounded outputs through `torch.tanh` or `activation='tanh'`.
- **Sigmoid** persists primarily in educational contexts like LeNet-5, demonstrating historical deep learning patterns rather than modern best practices.
- **Masked Softmax** handles variable-length sequences by filtering padding tokens before probability normalization in sequence modeling tasks.

## Frequently Asked Questions

### What is the default activation function used in d2l-zh modern neural networks?

d2l-zh uses **ReLU** as the default activation function for modern architectures. In PyTorch implementations, this appears as `nn.ReLU()` in modules like `PositionWiseFFN`, while MXNet uses the `activation='relu'` keyword in dense layer constructors. ReLU's computational efficiency and gradient-preserving properties make it the standard choice for MLPs, Transformers, and CNNs throughout the codebase.

### Why does d2l-zh include Sigmoid activation in its tutorials?

The repository retains **Sigmoid** primarily for educational purposes, specifically in the LeNet-5 example found in [`chapter_convolutional-neural-networks/lenet.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_convolutional-neural-networks/lenet.md). This demonstrates historical deep learning patterns where sigmoid activations followed each convolutional layer, helping learners understand the evolution from classical to modern architectures. The implementation uses `nn.Sigmoid()` in PyTorch and `tf.keras.layers.Activation('sigmoid')` in TensorFlow.

### How does d2l-zh handle activation functions in sequence models with padding?

For sequence models, d2l-zh implements **masked Softmax** rather than standard hidden-layer activations. The `masked_softmax` function in [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py) (lines 1127-1134) masks padding tokens before applying softmax, ensuring invalid positions do not affect probability distributions in machine translation and language modeling tasks. TensorFlow implements analogous logic in [`d2l/tensorflow.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/tensorflow.py) (lines 1061-1068).

### What is the difference between PyTorch and MXNet activation implementations in d2l-zh?

In **PyTorch**, d2l-zh uses explicit activation layer objects like `self.relu = nn.ReLU()` that are called in the forward pass, or functional calls like `torch.tanh(X)`. In **MXNet**, activations are specified via string keywords (`activation='relu'` or `activation='tanh'`) in layer constructors such as `gluon.nn.Dense`, allowing the backend to handle non-linearity internally without explicit layer definitions in model code.