How d2l-zh Details the Functioning of Recurrent Neural Networks (RNNs): From Mathematical Foundations to Multi-Framework Implementation

d2l-zh (Dive into Deep Learning Chinese edition) explains RNNs by first formalizing hidden-state recurrence and parameter sharing mathematically, then providing both scratch implementations and high-level API wrappers for PyTorch, TensorFlow, MXNet, and PaddlePaddle.

The d2l-zh repository delivers a comprehensive treatment of Recurrent Neural Networks (RNNs) that bridges theoretical machine learning concepts with production-ready code. This article examines how d2l-zh details the functioning of Recurrent Neural Networks (RNNs) through mathematical derivations in chapter_recurrent-neural-networks/rnn_origin.md and hands-on implementations across four major deep learning frameworks.

Mathematical Formalization of Hidden-State Recurrence

According to the source code in chapter_recurrent-neural-networks/rnn_origin.md, the core mechanism of RNNs relies on hidden-state recurrence, where the model maintains a hidden state vector H that captures temporal dependencies across sequence steps.

At time step t, the hidden state H_t is computed by combining the current input X_t with the previous hidden state H_{t-1} through a fully connected layer with activation function φ:


H_t = φ(X_t W_xh + H_{t-1} W_hh + b_h)

Here, W_xh represents the input-to-hidden weight matrix, W_hh denotes the hidden-to-hidden recurrent weight matrix, and b_h is the bias vector. This recurrence relation, corresponding to Equation (6) in the source text, creates a computational graph that chains operations across the sequence length.

The output O_t at each time step is generated by projecting the hidden state through an output layer:


O_t = H_t W_hq + b_q

Where W_hq and b_q constitute the hidden-to-output transformation parameters, as specified in Equation (8) of rnn_origin.md.

Parameter Sharing and Training Objectives

A critical characteristic detailed in chapter_recurrent-neural-networks/rnn_origin.md is parameter sharing across time steps. The weight matrices W_xh, W_hh, W_hq and biases b_h, b_q remain identical for every time step, ensuring the total parameter count remains constant regardless of sequence length.

For character-level language modeling tasks, the RNN outputs a probability distribution via softmax activation. The training objective minimizes cross-entropy loss summed across all time steps. Model performance is evaluated using perplexity, calculated as the exponent of the average cross-entropy loss, as detailed in the "Perplexity" section of rnn_origin.md.

Implementation Patterns: From Scratch to High-Level APIs

The d2l-zh codebase provides dual implementation strategies in d2l/torch.py, d2l/tensorflow.py, d2l/paddle.py, and d2l/mxnet.py, allowing learners to understand RNN mechanics at different abstraction levels.

Scratch Implementation (RNNModelScratch)

The RNNModelScratch class in d2l/torch.py implements the recurrence mechanism explicitly without relying on framework-specific recurrent layers. This approach manually handles one-hot encoding, parameter initialization, and the forward pass loop over time steps.

class RNNModelScratch:
    """RNN model implemented from scratch."""
    def __init__(self, vocab_size, num_hiddens, device, get_params, init_state, forward):
        self.vocab_size, self.num_hiddens = vocab_size, num_hiddens
        self.params = get_params(vocab_size, num_hiddens, device)
        self.init_state, self.forward = init_state, forward

    def __call__(self, X, state):
        X = self._one_hot(X)  # X: (batch, seq_len, vocab)

        outputs, state = self.forward(X, state, self.params)
        return outputs.reshape((-1, self.vocab_size)), state

Key implementation details include explicit one-hot encoding of inputs, manual parameter management through get_params, and stateful training where the hidden state propagates between mini-batches.

High-Level API Implementation (RNNModel)

For production use, d2l-zh provides the RNNModel class that delegates recurrent computation to optimized framework implementations such as nn.LSTM or nn.GRU.

class RNNModel(nn.Module):
    """RNN model using built-in recurrent layers."""
    def __init__(self, rnn_layer, vocab_size, num_hiddens, **kwargs):
        super(RNNModel, self).__init__(**kwargs)
        self.rnn = rnn_layer(input_size=vocab_size,
                             hidden_size=num_hiddens,
                             num_layers=2)
        self.dense = nn.Linear(num_hiddens, vocab_size)

    def forward(self, inputs, state):
        X = inputs.permute(1, 0, 2)  # (batch, seq_len, vocab) -> (seq_len, batch, vocab)

        Y, state = self.rnn(X, state)
        output = self.dense(Y.reshape(-1, Y.shape[-1]))
        return output, state

This implementation accepts a pre-configured recurrent layer (e.g., torch.nn.LSTM) and applies a final dense projection to map hidden states to vocabulary logits. The architectural pattern remains consistent across TensorFlow (tf.keras.layers.LSTM), PaddlePaddle (paddle.nn.LSTM), and MXNet (mxnet.gluon.rnn.LSTM) implementations.

Character-Level Language Modeling Training Pipeline

The repository includes complete training workflows demonstrating how d2l-zh applies RNN theory to character-level language modeling using the "Time Machine" dataset.

import d2l.torch as d2l
import torch
from d2l.torch import RNNModel, try_gpu

# Hyper-parameters

batch_size, num_steps = 32, 35
num_hiddens = 256
vocab = d2l.load_data_time_machine('torch')
vocab_size = len(vocab)

# Model initialization

rnn_layer = torch.nn.LSTM(input_size=vocab_size,
                          hidden_size=num_hiddens,
                          num_layers=2)
net = RNNModel(rnn_layer, vocab_size, num_hiddens)

# Training configuration

loss = torch.nn.CrossEntropyLoss()
trainer = torch.optim.SGD(net.parameters(), lr=1)
device = try_gpu()

# Training loop with state truncation

def train_epoch(net, train_iter, loss, trainer, device):
    net.train()
    state = None
    for X, Y in train_iter:
        X = d2l.to_one_hot(X, vocab_size).to(device)
        Y = Y.t().reshape(-1).to(device)
        
        y_hat, state = net(X, state)
        
        # Detach hidden state to prevent backprop through entire history

        if isinstance(state, tuple):
            state = (state[0].detach(), state[1].detach())
        else:
            state = state.detach()
            
        l = loss(y_hat, Y)
        trainer.zero_grad()
        l.backward()
        trainer.step()

This example illustrates critical practical considerations: input transposition for batch-first versus time-first conventions, explicit hidden state detachment to manage gradient flow through time, and one-hot encoding of character indices.

Summary

  • d2l-zh explains RNNs through a progression from mathematical formalization (H_t = φ(X_t W_xh + H_{t-1} W_hh + b_h)) to multi-framework code implementations.
  • The repository enforces parameter sharing across time steps, keeping the model size constant regardless of sequence length.
  • Two implementation modes are provided: RNNModelScratch for educational transparency and RNNModel for leveraging optimized framework primitives.
  • Training pipelines demonstrate truncated backpropagation through time via explicit state detachment in the training loop.
  • Complete implementations exist for PyTorch, TensorFlow, PaddlePaddle, and MXNet in their respective d2l/ package files.

Frequently Asked Questions

What mathematical formalism does d2l-zh use to explain RNN hidden states?

According to chapter_recurrent-neural-networks/rnn_origin.md, d2l-zh defines the hidden state update as H_t = φ(X_t W_xh + H_{t-1} W_hh + b_h), where φ is the activation function, W_xh and W_hh are shared weight matrices, and b_h is the bias term. This recurrence relation captures temporal dependencies by combining current inputs with previous hidden states through learned linear transformations.

How does d2l-zh handle backpropagation through time (BPTT) in its scratch implementation?

In the RNNModelScratch implementation within d2l/torch.py, the hidden state is explicitly detached after each forward pass using state.detach() (or state[0].detach(), state[1].detach() for LSTM tuples). This truncation prevents gradients from flowing back through the entire sequence history while maintaining stateful training across mini-batches.

Can I use d2l-zh's RNN implementations with frameworks other than PyTorch?

Yes, d2l-zh provides parallel implementations in d2l/tensorflow.py, d2l/paddle.py, and d2l/mxnet.py. Each file contains equivalent RNNModelScratch and RNNModel classes adapted to their respective framework APIs, such as tf.keras.layers.LSTM for TensorFlow and paddle.nn.LSTM for PaddlePaddle.

Where does d2l-zh explain bidirectional RNNs or advanced architectures?

Bidirectional RNNs are demonstrated in contrib/to-rm-mx-contrib-text/chapter_natural-language-processing/sentiment-analysis-rnn.md, which applies stacked and bidirectional recurrent layers to sentiment analysis tasks. This file illustrates how to extend the basic RNNModel pattern for sequence classification rather than language modeling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →