# How d2l-zh Approaches Natural Language Processing (NLP) Tasks with Deep Learning

> Explore how d2l-zh leverages deep learning for NLP tasks. Discover its three stage pipeline static word embeddings contextual Transformer pre-training and task specific fine-tuning with multiple back-ends.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: how-to-guide
- Published: 2026-03-01

---

**d2l-zh structures its NLP curriculum as a three-stage pipeline: static word embeddings → contextual Transformer pre-training → task-specific fine-tuning, implemented with interchangeable MXNet, PyTorch, and Paddle back-ends.**

The d2l-zh repository (the Chinese edition of *Dive into Deep Learning*) provides a comprehensive, code-first guide to modern **Natural Language Processing (NLP) tasks with deep learning**. Its pedagogical architecture moves from classical shallow embeddings to state-of-the-art BERT fine-tuning, using self-contained Jupyter notebooks that execute across three major deep-learning frameworks.

## Word-Level Embedding Foundations

The curriculum begins with **context-independent word embeddings** to establish baseline representations. These chapters demonstrate how to train lookup tables from co-occurrence statistics and highlight their limitations on polysemous words.

### Word2Vec and GloVe Implementations

The [`word2vec.md`](https://github.com/d2l-ai/d2l-zh/blob/main/word2vec.md) and [`glove.md`](https://github.com/d2l-ai/d2l-zh/blob/main/glove.md) notebooks in `chapter_natural-language-processing-pretraining/` implement shallow embedding pipelines using `nn.Embedding(vocab_size, embed_dim)`. These modules train embeddings from scratch on corpus statistics, storing parameters in a dense matrix that maps discrete tokens to continuous vectors.

- Source: [`chapter_natural-language-processing-pretraining/word2vec.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_natural-language-processing-pretraining/word2vec.md)
- Source: [`chapter_natural-language-processing-pretraining/glove.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_natural-language-processing-pretraining/glove.md)

These implementations emphasize why static vectors fail to disambiguate words like "crane" (bird vs. machine), motivating the shift to contextual models.

## Contextual Pre-training with BERT

The core of d2l-zh’s modern NLP pipeline is a **Transformer-based encoder** trained on two self-supervised tasks: Masked Language Modeling (MLM) and Next Sentence Prediction (NSP).

### BERT Input Representation

The `get_tokens_and_segments` function constructs BERT-style inputs by concatenating token lists with special delimiters:

```python
def get_tokens_and_segments(tokens_a, tokens_b=None):
    tokens = ['<cls>'] + tokens_a + ['<sep>']
    segments = [0] * (len(tokens_a) + 2)
    if tokens_b is not None:
        tokens += tokens_b + ['<sep>']
        segments += [1] * (len(tokens_b) + 1)
    return tokens, segments

```

This generates the `[CLS] … [SEP] … [SEP]` pattern and corresponding segment IDs (0 for the first sentence, 1 for the second).

### Model Architecture

The `BERTEncoder` class stacks `d2l.EncoderBlock` layers (multi-head self-attention + feed-forward) and adds **learnable position embeddings**:

- `BERTEncoder`: Lines 86-108 in [`chapter_natural-language-processing-pretraining/bert.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_natural-language-processing-pretraining/bert.md)
- Implements the Transformer encoder stack with configurable `num_layers`, `num_hiddens`, and `num_heads`

### Pre-training Objectives

Two heads attach to the encoder output:

1. **`MaskLM`**: Implements **Masked Language Modeling** (lines 38-60). It predicts randomly masked tokens using a small MLP head attached to the final hidden states.
2. **`NextSentencePred`**: Implements **Next-Sentence Prediction** (lines 73-81). It performs binary classification on the `[CLS]` representation to determine if sentence B follows sentence A.

The `BERTModel` class (lines 70-92) wraps the encoder and both heads, exposing a `forward(tokens, segments, ...)` method returning encoder output, MLM logits, and NSP logits.

### Loading Pretrained Weights

The `load_pretrained_model` helper downloads a zipped checkpoint, builds a `Vocab` from [`vocab.json`](https://github.com/d2l-ai/d2l-zh/blob/main/vocab.json), instantiates `BERTModel`, and loads saved parameters:

```python
devices = d2l.try_all_gpus()
bert, vocab = load_pretrained_model(
    'bert.small', num_hiddens=256, ffn_num_hiddens=512,
    num_heads=4, num_layers=2, dropout=0.1, max_len=512,
    devices=devices)

```

This enables immediate fine-tuning on downstream tasks without training the encoder from scratch.

## Downstream NLP Applications

### Sentiment Analysis with RNNs and CNNs

Before Transformer fine-tuning, the curriculum covers **recurrent and convolutional architectures** for text classification:

- **RNN-based classifier**: [`chapter_natural-language-processing-applications/sentiment-analysis-rnn.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_natural-language-processing-applications/sentiment-analysis-rnn.md) demonstrates feeding word embeddings into a bidirectional LSTM with a final linear classifier.
- **CNN-based classifier**: [`chapter_natural-language-processing-applications/sentiment-analysis-cnn.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_natural-language-processing-applications/sentiment-analysis-cnn.md) implements 1-D convolutions over embeddings, max-pooling, and dense output layers.

Both notebooks reuse the `d2l.Tokenizer` and `d2l.Vocab` utilities, illustrating the framework-agnostic `@tab` syntax that generates code for MXNet, PyTorch, and Paddle simultaneously.

### Natural Language Inference with BERT Fine-Tuning

The [`natural-language-inference-bert.md`](https://github.com/d2l-ai/d2l-zh/blob/main/natural-language-inference-bert.md) chapter provides an end-to-end example of **fine-tuning pretrained BERT** on the SNLI dataset:

1. **Dataset preparation**: `SNLIBERTDataset` tokenizes premise/hypothesis pairs, constructs token/segment IDs using `get_tokens_and_segments`, and pads sequences to `max_len`.

2. **Model definition**: `BERTClassifier` reuses the encoder from the pretrained BERT, adds a hidden projection (`bert.hidden`), and an output layer for three classes (entailment, contradiction, neutral).

   ```python
   class BERTClassifier(nn.Block):
       def __init__(self, bert):
           super(BERTClassifier, self).__init__()
           self.encoder = bert.encoder
           self.hidden = bert.hidden
           self.output = nn.Dense(3)  # 3 classes for SNLI

   
       def forward(self, inputs):
           tokens_X, segments_X, valid_lens_X = inputs
           encoded_X = self.encoder(tokens_X, segments_X, valid_lens_X)
           return self.output(self.hidden(encoded_X[:, 0, :]))
   ```

3. **Training loop**: Only the parameters of the added MLP (`net.output`) are initialized from scratch; the encoder parameters are **fine-tuned** using a small learning rate.

   ```python
   net = BERTClassifier(bert)               # attach pretrained encoder

   net.output.initialize(ctx=devices)       # new classifier head

   trainer = gluon.Trainer(net.collect_params(),
                           'adam', {'learning_rate': 1e-4})
   d2l.train_ch13(net, train_iter, test_iter,
                  gluon.loss.SoftmaxCrossEntropyLoss(),
                  trainer, num_epochs=5, devices=devices,
                  split_fn=d2l.split_batch_multi_inputs)
   ```

This pattern—*pre-train → attach → fine-tune*—represents the canonical deep-learning approach to modern NLP tasks in d2l-zh.

## Unified API Across Back-ends

All notebooks rely on the high-level `d2l` package (`d2l.mxnet`, `d2l.torch`, `d2l.paddle`). The same logical flow appears in each language-specific block (`#@tab mxnet, pytorch, paddle`), allowing readers to concentrate on **architectural ideas** rather than framework syntax.

Key utilities defined in [`d2l/__init__.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/__init__.py) and sub-modules include:

- `d2l.Vocab`: Bidirectional token↔index mapping with reserved symbols (`<unk>`, `<pad>`, `<cls>`, `<sep>`).
- `d2l.download_extract`: Fetches datasets and pretrained checkpoints from the centralized `DATA_HUB`.
- `d2l.get_tokens_and_segments`: Constructs BERT-style input sequences with segment IDs.
- `d2l.train_ch13`: Generic training loop supporting multi-input batches via `split_fn`.

These utilities guarantee reproducibility and hide boilerplate I/O, enabling focus on model architecture.

## Summary

- **d2l-zh** structures NLP education as a progression from static embeddings to Transformer fine-tuning, implemented in `chapter_natural-language-processing-pretraining/` and `chapter_natural-language-processing-applications/`.
- **Static embeddings** (Word2Vec, GloVe) provide baseline representations via `nn.Embedding`, but lack context sensitivity.
- **BERT pre-training** involves `BERTEncoder`, `MaskLM`, and `NextSentencePred` heads, trained on masked language modeling and next-sentence prediction tasks.
- **Fine-tuning** follows the *pre-train → attach → fine-tune* pattern, where `BERTClassifier` adds task-specific heads (e.g., for SNLI) and updates all parameters with differential learning rates.
- **Multi-framework support** via the `d2l` package allows identical code to run on MXNet, PyTorch, and Paddle using `#@tab` directives.

## Frequently Asked Questions

### What is the difference between static embeddings and BERT in d2l-zh?

Static embeddings (Word2Vec and GloVe) map each word to a single fixed vector regardless of context, implemented via `nn.Embedding` in [`word2vec.md`](https://github.com/d2l-ai/d2l-zh/blob/main/word2vec.md) and [`glove.md`](https://github.com/d2l-ai/d2l-zh/blob/main/glove.md). BERT, defined in [`bert.md`](https://github.com/d2l-ai/d2l-zh/blob/main/bert.md), uses a `BERTEncoder` to generate **contextualized** representations where the same word receives different vectors based on surrounding tokens, enabling disambiguation of polysemous words like "bank."

### How does d2l-zh support multiple deep learning frameworks for NLP?

The repository uses a **framework-agnostic** `@tab` syntax (e.g., `#@tab mxnet, pytorch, paddle`) within Markdown cells. The `d2l` package provides unified APIs (`d2l.mxnet`, `d2l.torch`, `d2l.paddle`) that implement identical functions like `Vocab`, `get_tokens_and_segments`, and `train_ch13`, allowing learners to focus on architectural concepts rather than framework-specific syntax.

### What are the two pre-training tasks used in d2l-zh's BERT implementation?

According to [`chapter_natural-language-processing-pretraining/bert.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_natural-language-processing-pretraining/bert.md), BERT pre-trains on **Masked Language Modeling (MLM)** and **Next Sentence Prediction (NSP)**. The `MaskLM` class (lines 38-60) predicts randomly masked tokens from corrupted inputs, while `NextSentencePred` (lines 73-81) classifies whether sentence B logically follows sentence A using the `[CLS]` token representation.

### How is BERT fine-tuned for Natural Language Inference in d2l-zh?

Fine-tuning follows the pattern defined in [`chapter_natural-language-processing-applications/natural-language-inference-bert.md`](https://github.com/d2l-ai/d2l-zh/blob/main/chapter_natural-language-processing-applications/natural-language-inference-bert.md). First, `SNLIBERTDataset` prepares premise-hypothesis pairs with `get_tokens_and_segments`. Then, `BERTClassifier` attaches a new output layer to the pretrained `BERTEncoder`. During training, `d2l.train_ch13` updates all parameters but initializes only the new classifier head from scratch, using `split_fn=d2l.split_batch_multi_inputs` to handle multi-input batches (tokens, segments, valid lengths).