Activation Functions in d2l-zh: Types, Implementation, and Usage Across Deep Learning Frameworks

The d2l-zh educational framework implements four primary activation functions—ReLU, tanh, Sigmoid, and masked Softmax—across PyTorch, TensorFlow, MXNet, and Paddle backends, with ReLU serving as the default non-linearity for modern architectures while Sigmoid remains primarily for legacy educational examples.

The d2l-zh repository (d2l-ai/d2l-zh) provides a unified interface for deep learning education across multiple frameworks. Understanding the activation functions available in this codebase is essential for implementing neural networks effectively. This guide examines the four core activation types used throughout the library, their specific backend implementations, and practical usage patterns in both modern and classical architectures.

Core Activation Functions in d2l-zh

ReLU (Rectified Linear Unit)

ReLU is the default activation function in d2l-zh for modern architectures. In d2l/torch.py (lines 1294-1296), the implementation wraps nn.ReLU() for PyTorch. The function appears in position-wise feed-forward networks within Transformer blocks, where the code pattern self.dense2(self.relu(self.dense1(X))) applies the non-linearity between two linear transformations.

Tanh (Hyperbolic Tangent)

The tanh activation appears frequently in attention mechanisms and classic RNN cells throughout d2l-zh. According to d2l/torch.py (lines 1163-1164), the implementation uses torch.tanh directly. In additive attention implementations, intermediate feature maps pass through tanh to produce bounded output values that stabilize similarity scores.

Sigmoid

Sigmoid functions in d2l-zh serve primarily educational purposes, appearing in legacy architectures like LeNet-5. The chapter_convolutional-neural-networks/lenet.md file demonstrates nn.Sigmoid() in PyTorch (lines 61-68) and tf.keras.layers.Activation('sigmoid') in TensorFlow (lines 78-80). These implementations follow classical convolutional network designs where sigmoid activations follow each convolutional or dense layer.

Masked Softmax

Unlike standard activations, masked Softmax in d2l-zh specifically handles variable-length sequences in language models and machine translation. The masked_softmax function in d2l/torch.py (lines 1127-1134) masks padding tokens before applying nn.functional.softmax, preventing invalid positions from affecting probability distributions. TensorFlow implements analogous logic in d2l/tensorflow.py (lines 1061-1068).

Implementation Across Deep Learning Frameworks

PyTorch Implementation

In the PyTorch backend (d2l/torch.py), activation functions appear as explicit layer objects or functional calls. ReLU instantiates as self.relu = nn.ReLU() within module definitions, while tanh operates as torch.tanh(X) for tensor transformations. The masked softmax implementation handles batch matrix operations with explicit padding masks.

TensorFlow and Keras

The TensorFlow backend (d2l/tensorflow.py) utilizes Keras layer activations, particularly for Sigmoid in legacy models. The LeNet implementation demonstrates tf.keras.layers.Activation('sigmoid') following convolutional layers. Modern architectures prefer tf.keras.layers.ReLU() for consistency with current best practices.

MXNet and Gluon

MXNet support in d2l/mxnet.py employs activation keywords in layer constructors. Dense layers specify activation via the activation='relu' or activation='tanh' parameters, allowing the backend to handle non-linearity application internally. This approach reduces explicit activation layer definitions in model code.

PaddlePaddle

The Paddle backend (d2l/paddle.py) follows PyTorch patterns with explicit activation layers. The PositionWiseFFN class demonstrates self.relu = nn.ReLU() initialization and application within feed-forward networks, maintaining API consistency across the d2l-zh framework.

Practical Usage Examples

The following snippets demonstrate activation function usage across different backends in d2l-zh.

PyTorch implementation with ReLU and explicit layer definition:

import torch
from torch import nn
from d2l import torch as d2l

class SimpleFFN(nn.Module):
    def __init__(self, in_dim, hidden_dim, out_dim):
        super().__init__()
        self.dense1 = nn.Linear(in_dim, hidden_dim)
        self.relu = nn.ReLU()               # ReLU activation layer

        self.dense2 = nn.Linear(hidden_dim, out_dim)

    def forward(self, X):
        X = self.relu(self.dense1(X))       # Apply ReLU

        return self.dense2(X)

TensorFlow implementation using Sigmoid in the classic LeNet architecture:

import tensorflow as tf
from d2l import tensorflow as d2l

def lenet():
    return tf.keras.models.Sequential([
        tf.keras.layers.Conv2D(6, 5, padding='same', activation='sigmoid'),
        tf.keras.layers.AvgPool2D(2, 2),
        tf.keras.layers.Conv2D(16, 5, activation='sigmoid'),
        tf.keras.layers.AvgPool2D(2, 2),
        tf.keras.layers.Flatten(),
        tf.keras.layers.Dense(120, activation='sigmoid'),
        tf.keras.layers.Dense(84, activation='sigmoid'),
        tf.keras.layers.Dense(10)  # logits

    ])

MXNet implementation using keyword-based activation arguments:

from mxnet import gluon, nd
from d2l import mxnet as d2l

net = gluon.nn.Sequential()
with net.name_scope():
    net.add(gluon.nn.Dense(256, activation='relu'))   # ReLU via keyword

    net.add(gluon.nn.Dense(128, activation='tanh'))   # tanh via keyword

    net.add(gluon.nn.Dense(10))                      # logits

net.initialize()

PaddlePaddle implementation with explicit ReLU layers in Transformer blocks:

import paddle
from paddle import nn
from d2l import paddle as d2l

class PositionWiseFFN(nn.Layer):
    def __init__(self, ffn_num_input, ffn_num_hiddens, ffn_num_outputs):
        super().__init__()
        self.dense1 = nn.Linear(ffn_num_input, ffn_num_hiddens)
        self.relu = nn.ReLU()                     # Explicit ReLU layer

        self.dense2 = nn.Linear(ffn_num_hiddens, ffn_num_outputs)

    def forward(self, X):
        return self.dense2(self.relu(self.dense1(X)))

Summary

  • ReLU serves as the default activation for modern architectures in d2l-zh, implemented via nn.ReLU() in PyTorch and Paddle or activation='relu' in MXNet.
  • Tanh remains essential for attention mechanisms and RNN cells, providing bounded outputs through torch.tanh or activation='tanh'.
  • Sigmoid persists primarily in educational contexts like LeNet-5, demonstrating historical deep learning patterns rather than modern best practices.
  • Masked Softmax handles variable-length sequences by filtering padding tokens before probability normalization in sequence modeling tasks.

Frequently Asked Questions

What is the default activation function used in d2l-zh modern neural networks?

d2l-zh uses ReLU as the default activation function for modern architectures. In PyTorch implementations, this appears as nn.ReLU() in modules like PositionWiseFFN, while MXNet uses the activation='relu' keyword in dense layer constructors. ReLU's computational efficiency and gradient-preserving properties make it the standard choice for MLPs, Transformers, and CNNs throughout the codebase.

Why does d2l-zh include Sigmoid activation in its tutorials?

The repository retains Sigmoid primarily for educational purposes, specifically in the LeNet-5 example found in chapter_convolutional-neural-networks/lenet.md. This demonstrates historical deep learning patterns where sigmoid activations followed each convolutional layer, helping learners understand the evolution from classical to modern architectures. The implementation uses nn.Sigmoid() in PyTorch and tf.keras.layers.Activation('sigmoid') in TensorFlow.

How does d2l-zh handle activation functions in sequence models with padding?

For sequence models, d2l-zh implements masked Softmax rather than standard hidden-layer activations. The masked_softmax function in d2l/torch.py (lines 1127-1134) masks padding tokens before applying softmax, ensuring invalid positions do not affect probability distributions in machine translation and language modeling tasks. TensorFlow implements analogous logic in d2l/tensorflow.py (lines 1061-1068).

What is the difference between PyTorch and MXNet activation implementations in d2l-zh?

In PyTorch, d2l-zh uses explicit activation layer objects like self.relu = nn.ReLU() that are called in the forward pass, or functional calls like torch.tanh(X). In MXNet, activations are specified via string keywords (activation='relu' or activation='tanh') in layer constructors such as gluon.nn.Dense, allowing the backend to handle non-linearity internally without explicit layer definitions in model code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →