# How d2l-zh Details the Functioning of Recurrent Neural Networks (RNNs): From Mathematical Foundations to Multi-Framework Implementation

> Explore how d2l-zh explains Recurrent Neural Networks RNNs. Learn mathematical foundations and see PyTorch TensorFlow MXNet and PaddlePaddle implementations.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: deep-dive
- Published: 2026-03-01

---

**d2l-zh (Dive into Deep Learning Chinese edition) explains RNNs by first formalizing hidden-state recurrence and parameter sharing mathematically, then providing both scratch implementations and high-level API wrappers for PyTorch, TensorFlow, MXNet, and PaddlePaddle.**

The d2l-zh repository delivers a comprehensive treatment of Recurrent Neural Networks (RNNs) that bridges theoretical machine learning concepts with production-ready code. This article examines how d2l-zh details the functioning of Recurrent Neural Networks (RNNs) through mathematical derivations in [`chapter_recurrent-neural-networks/rnn_origin.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_recurrent-neural-networks/rnn_origin.md) and hands-on implementations across four major deep learning frameworks.

## Mathematical Formalization of Hidden-State Recurrence

According to the source code in [`chapter_recurrent-neural-networks/rnn_origin.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_recurrent-neural-networks/rnn_origin.md), the core mechanism of RNNs relies on **hidden-state recurrence**, where the model maintains a hidden state vector **H** that captures temporal dependencies across sequence steps.

At time step **t**, the hidden state **H_t** is computed by combining the current input **X_t** with the previous hidden state **H_{t-1}** through a fully connected layer with activation function **φ**:

```

H_t = φ(X_t W_xh + H_{t-1} W_hh + b_h)

```

Here, **W_xh** represents the input-to-hidden weight matrix, **W_hh** denotes the hidden-to-hidden recurrent weight matrix, and **b_h** is the bias vector. This recurrence relation, corresponding to Equation (6) in the source text, creates a computational graph that chains operations across the sequence length.

The output **O_t** at each time step is generated by projecting the hidden state through an output layer:

```

O_t = H_t W_hq + b_q

```

Where **W_hq** and **b_q** constitute the hidden-to-output transformation parameters, as specified in Equation (8) of [`rnn_origin.md`](https://github.com/d2l-ai/d2l-zh/blob/main/rnn_origin.md).

## Parameter Sharing and Training Objectives

A critical characteristic detailed in [`chapter_recurrent-neural-networks/rnn_origin.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_recurrent-neural-networks/rnn_origin.md) is **parameter sharing** across time steps. The weight matrices **W_xh**, **W_hh**, **W_hq** and biases **b_h**, **b_q** remain identical for every time step, ensuring the total parameter count remains constant regardless of sequence length.

For character-level language modeling tasks, the RNN outputs a probability distribution via softmax activation. The training objective minimizes **cross-entropy loss** summed across all time steps. Model performance is evaluated using **perplexity**, calculated as the exponent of the average cross-entropy loss, as detailed in the "Perplexity" section of [`rnn_origin.md`](https://github.com/d2l-ai/d2l-zh/blob/main/rnn_origin.md).

## Implementation Patterns: From Scratch to High-Level APIs

The d2l-zh codebase provides dual implementation strategies in [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py), [`d2l/tensorflow.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/tensorflow.py), [`d2l/paddle.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/paddle.py), and [`d2l/mxnet.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/mxnet.py), allowing learners to understand RNN mechanics at different abstraction levels.

### Scratch Implementation (RNNModelScratch)

The `RNNModelScratch` class in [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py) implements the recurrence mechanism explicitly without relying on framework-specific recurrent layers. This approach manually handles one-hot encoding, parameter initialization, and the forward pass loop over time steps.

```python
class RNNModelScratch:
    """RNN model implemented from scratch."""
    def __init__(self, vocab_size, num_hiddens, device, get_params, init_state, forward):
        self.vocab_size, self.num_hiddens = vocab_size, num_hiddens
        self.params = get_params(vocab_size, num_hiddens, device)
        self.init_state, self.forward = init_state, forward

    def __call__(self, X, state):
        X = self._one_hot(X)  # X: (batch, seq_len, vocab)

        outputs, state = self.forward(X, state, self.params)
        return outputs.reshape((-1, self.vocab_size)), state

```

Key implementation details include explicit one-hot encoding of inputs, manual parameter management through `get_params`, and stateful training where the hidden state propagates between mini-batches.

### High-Level API Implementation (RNNModel)

For production use, d2l-zh provides the `RNNModel` class that delegates recurrent computation to optimized framework implementations such as `nn.LSTM` or `nn.GRU`.

```python
class RNNModel(nn.Module):
    """RNN model using built-in recurrent layers."""
    def __init__(self, rnn_layer, vocab_size, num_hiddens, **kwargs):
        super(RNNModel, self).__init__(**kwargs)
        self.rnn = rnn_layer(input_size=vocab_size,
                             hidden_size=num_hiddens,
                             num_layers=2)
        self.dense = nn.Linear(num_hiddens, vocab_size)

    def forward(self, inputs, state):
        X = inputs.permute(1, 0, 2)  # (batch, seq_len, vocab) -> (seq_len, batch, vocab)

        Y, state = self.rnn(X, state)
        output = self.dense(Y.reshape(-1, Y.shape[-1]))
        return output, state

```

This implementation accepts a pre-configured recurrent layer (e.g., `torch.nn.LSTM`) and applies a final dense projection to map hidden states to vocabulary logits. The architectural pattern remains consistent across TensorFlow (`tf.keras.layers.LSTM`), PaddlePaddle (`paddle.nn.LSTM`), and MXNet (`mxnet.gluon.rnn.LSTM`) implementations.

## Character-Level Language Modeling Training Pipeline

The repository includes complete training workflows demonstrating how d2l-zh applies RNN theory to character-level language modeling using the "Time Machine" dataset.

```python
import d2l.torch as d2l
import torch
from d2l.torch import RNNModel, try_gpu

# Hyper-parameters

batch_size, num_steps = 32, 35
num_hiddens = 256
vocab = d2l.load_data_time_machine('torch')
vocab_size = len(vocab)

# Model initialization

rnn_layer = torch.nn.LSTM(input_size=vocab_size,
                          hidden_size=num_hiddens,
                          num_layers=2)
net = RNNModel(rnn_layer, vocab_size, num_hiddens)

# Training configuration

loss = torch.nn.CrossEntropyLoss()
trainer = torch.optim.SGD(net.parameters(), lr=1)
device = try_gpu()

# Training loop with state truncation

def train_epoch(net, train_iter, loss, trainer, device):
    net.train()
    state = None
    for X, Y in train_iter:
        X = d2l.to_one_hot(X, vocab_size).to(device)
        Y = Y.t().reshape(-1).to(device)
        
        y_hat, state = net(X, state)
        
        # Detach hidden state to prevent backprop through entire history

        if isinstance(state, tuple):
            state = (state[0].detach(), state[1].detach())
        else:
            state = state.detach()
            
        l = loss(y_hat, Y)
        trainer.zero_grad()
        l.backward()
        trainer.step()

```

This example illustrates critical practical considerations: input transposition for batch-first versus time-first conventions, explicit hidden state detachment to manage gradient flow through time, and one-hot encoding of character indices.

## Summary

- **d2l-zh** explains RNNs through a progression from mathematical formalization (`H_t = φ(X_t W_xh + H_{t-1} W_hh + b_h)`) to multi-framework code implementations.
- The repository enforces **parameter sharing** across time steps, keeping the model size constant regardless of sequence length.
- **Two implementation modes** are provided: `RNNModelScratch` for educational transparency and `RNNModel` for leveraging optimized framework primitives.
- Training pipelines demonstrate **truncated backpropagation through time** via explicit state detachment in the training loop.
- Complete implementations exist for **PyTorch, TensorFlow, PaddlePaddle, and MXNet** in their respective `d2l/` package files.

## Frequently Asked Questions

### What mathematical formalism does d2l-zh use to explain RNN hidden states?

According to [`chapter_recurrent-neural-networks/rnn_origin.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_recurrent-neural-networks/rnn_origin.md), d2l-zh defines the hidden state update as `H_t = φ(X_t W_xh + H_{t-1} W_hh + b_h)`, where **φ** is the activation function, **W_xh** and **W_hh** are shared weight matrices, and **b_h** is the bias term. This recurrence relation captures temporal dependencies by combining current inputs with previous hidden states through learned linear transformations.

### How does d2l-zh handle backpropagation through time (BPTT) in its scratch implementation?

In the `RNNModelScratch` implementation within [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py), the hidden state is explicitly detached after each forward pass using `state.detach()` (or `state[0].detach(), state[1].detach()` for LSTM tuples). This truncation prevents gradients from flowing back through the entire sequence history while maintaining stateful training across mini-batches.

### Can I use d2l-zh's RNN implementations with frameworks other than PyTorch?

Yes, d2l-zh provides parallel implementations in [`d2l/tensorflow.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/tensorflow.py), [`d2l/paddle.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/paddle.py), and [`d2l/mxnet.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/mxnet.py). Each file contains equivalent `RNNModelScratch` and `RNNModel` classes adapted to their respective framework APIs, such as `tf.keras.layers.LSTM` for TensorFlow and `paddle.nn.LSTM` for PaddlePaddle.

### Where does d2l-zh explain bidirectional RNNs or advanced architectures?

Bidirectional RNNs are demonstrated in [`contrib/to-rm-mx-contrib-text/chapter_natural-language-processing/sentiment-analysis-rnn.md`](https://github.com/d2l-ai/d2l-zh/blob/main/contrib/to-rm-mx-contrib-text/chapter_natural-language-processing/sentiment-analysis-rnn.md), which applies stacked and bidirectional recurrent layers to sentiment analysis tasks. This file illustrates how to extend the basic `RNNModel` pattern for sequence classification rather than language modeling.