How d2l-zh Approaches Natural Language Processing (NLP) Tasks with Deep Learning

d2l-zh structures its NLP curriculum as a three-stage pipeline: static word embeddings → contextual Transformer pre-training → task-specific fine-tuning, implemented with interchangeable MXNet, PyTorch, and Paddle back-ends.

The d2l-zh repository (the Chinese edition of Dive into Deep Learning) provides a comprehensive, code-first guide to modern Natural Language Processing (NLP) tasks with deep learning. Its pedagogical architecture moves from classical shallow embeddings to state-of-the-art BERT fine-tuning, using self-contained Jupyter notebooks that execute across three major deep-learning frameworks.

Word-Level Embedding Foundations

The curriculum begins with context-independent word embeddings to establish baseline representations. These chapters demonstrate how to train lookup tables from co-occurrence statistics and highlight their limitations on polysemous words.

Word2Vec and GloVe Implementations

The word2vec.md and glove.md notebooks in chapter_natural-language-processing-pretraining/ implement shallow embedding pipelines using nn.Embedding(vocab_size, embed_dim). These modules train embeddings from scratch on corpus statistics, storing parameters in a dense matrix that maps discrete tokens to continuous vectors.

These implementations emphasize why static vectors fail to disambiguate words like "crane" (bird vs. machine), motivating the shift to contextual models.

Contextual Pre-training with BERT

The core of d2l-zh’s modern NLP pipeline is a Transformer-based encoder trained on two self-supervised tasks: Masked Language Modeling (MLM) and Next Sentence Prediction (NSP).

BERT Input Representation

The get_tokens_and_segments function constructs BERT-style inputs by concatenating token lists with special delimiters:

def get_tokens_and_segments(tokens_a, tokens_b=None):
    tokens = ['<cls>'] + tokens_a + ['<sep>']
    segments = [0] * (len(tokens_a) + 2)
    if tokens_b is not None:
        tokens += tokens_b + ['<sep>']
        segments += [1] * (len(tokens_b) + 1)
    return tokens, segments

This generates the [CLS] … [SEP] … [SEP] pattern and corresponding segment IDs (0 for the first sentence, 1 for the second).

Model Architecture

The BERTEncoder class stacks d2l.EncoderBlock layers (multi-head self-attention + feed-forward) and adds learnable position embeddings:

Pre-training Objectives

Two heads attach to the encoder output:

  1. MaskLM: Implements Masked Language Modeling (lines 38-60). It predicts randomly masked tokens using a small MLP head attached to the final hidden states.
  2. NextSentencePred: Implements Next-Sentence Prediction (lines 73-81). It performs binary classification on the [CLS] representation to determine if sentence B follows sentence A.

The BERTModel class (lines 70-92) wraps the encoder and both heads, exposing a forward(tokens, segments, ...) method returning encoder output, MLM logits, and NSP logits.

Loading Pretrained Weights

The load_pretrained_model helper downloads a zipped checkpoint, builds a Vocab from vocab.json, instantiates BERTModel, and loads saved parameters:

devices = d2l.try_all_gpus()
bert, vocab = load_pretrained_model(
    'bert.small', num_hiddens=256, ffn_num_hiddens=512,
    num_heads=4, num_layers=2, dropout=0.1, max_len=512,
    devices=devices)

This enables immediate fine-tuning on downstream tasks without training the encoder from scratch.

Downstream NLP Applications

Sentiment Analysis with RNNs and CNNs

Before Transformer fine-tuning, the curriculum covers recurrent and convolutional architectures for text classification:

Both notebooks reuse the d2l.Tokenizer and d2l.Vocab utilities, illustrating the framework-agnostic @tab syntax that generates code for MXNet, PyTorch, and Paddle simultaneously.

Natural Language Inference with BERT Fine-Tuning

The natural-language-inference-bert.md chapter provides an end-to-end example of fine-tuning pretrained BERT on the SNLI dataset:

  1. Dataset preparation: SNLIBERTDataset tokenizes premise/hypothesis pairs, constructs token/segment IDs using get_tokens_and_segments, and pads sequences to max_len.

  2. Model definition: BERTClassifier reuses the encoder from the pretrained BERT, adds a hidden projection (bert.hidden), and an output layer for three classes (entailment, contradiction, neutral).

    class BERTClassifier(nn.Block):
        def __init__(self, bert):
            super(BERTClassifier, self).__init__()
            self.encoder = bert.encoder
            self.hidden = bert.hidden
            self.output = nn.Dense(3)  # 3 classes for SNLI
    
    
        def forward(self, inputs):
            tokens_X, segments_X, valid_lens_X = inputs
            encoded_X = self.encoder(tokens_X, segments_X, valid_lens_X)
            return self.output(self.hidden(encoded_X[:, 0, :]))
  3. Training loop: Only the parameters of the added MLP (net.output) are initialized from scratch; the encoder parameters are fine-tuned using a small learning rate.

    net = BERTClassifier(bert)               # attach pretrained encoder
    
    net.output.initialize(ctx=devices)       # new classifier head
    
    trainer = gluon.Trainer(net.collect_params(),
                            'adam', {'learning_rate': 1e-4})
    d2l.train_ch13(net, train_iter, test_iter,
                   gluon.loss.SoftmaxCrossEntropyLoss(),
                   trainer, num_epochs=5, devices=devices,
                   split_fn=d2l.split_batch_multi_inputs)

This pattern—pre-train → attach → fine-tune—represents the canonical deep-learning approach to modern NLP tasks in d2l-zh.

Unified API Across Back-ends

All notebooks rely on the high-level d2l package (d2l.mxnet, d2l.torch, d2l.paddle). The same logical flow appears in each language-specific block (#@tab mxnet, pytorch, paddle), allowing readers to concentrate on architectural ideas rather than framework syntax.

Key utilities defined in d2l/__init__.py and sub-modules include:

  • d2l.Vocab: Bidirectional token↔index mapping with reserved symbols (<unk>, <pad>, <cls>, <sep>).
  • d2l.download_extract: Fetches datasets and pretrained checkpoints from the centralized DATA_HUB.
  • d2l.get_tokens_and_segments: Constructs BERT-style input sequences with segment IDs.
  • d2l.train_ch13: Generic training loop supporting multi-input batches via split_fn.

These utilities guarantee reproducibility and hide boilerplate I/O, enabling focus on model architecture.

Summary

  • d2l-zh structures NLP education as a progression from static embeddings to Transformer fine-tuning, implemented in chapter_natural-language-processing-pretraining/ and chapter_natural-language-processing-applications/.
  • Static embeddings (Word2Vec, GloVe) provide baseline representations via nn.Embedding, but lack context sensitivity.
  • BERT pre-training involves BERTEncoder, MaskLM, and NextSentencePred heads, trained on masked language modeling and next-sentence prediction tasks.
  • Fine-tuning follows the pre-train → attach → fine-tune pattern, where BERTClassifier adds task-specific heads (e.g., for SNLI) and updates all parameters with differential learning rates.
  • Multi-framework support via the d2l package allows identical code to run on MXNet, PyTorch, and Paddle using #@tab directives.

Frequently Asked Questions

What is the difference between static embeddings and BERT in d2l-zh?

Static embeddings (Word2Vec and GloVe) map each word to a single fixed vector regardless of context, implemented via nn.Embedding in word2vec.md and glove.md. BERT, defined in bert.md, uses a BERTEncoder to generate contextualized representations where the same word receives different vectors based on surrounding tokens, enabling disambiguation of polysemous words like "bank."

How does d2l-zh support multiple deep learning frameworks for NLP?

The repository uses a framework-agnostic @tab syntax (e.g., #@tab mxnet, pytorch, paddle) within Markdown cells. The d2l package provides unified APIs (d2l.mxnet, d2l.torch, d2l.paddle) that implement identical functions like Vocab, get_tokens_and_segments, and train_ch13, allowing learners to focus on architectural concepts rather than framework-specific syntax.

What are the two pre-training tasks used in d2l-zh's BERT implementation?

According to chapter_natural-language-processing-pretraining/bert.md, BERT pre-trains on Masked Language Modeling (MLM) and Next Sentence Prediction (NSP). The MaskLM class (lines 38-60) predicts randomly masked tokens from corrupted inputs, while NextSentencePred (lines 73-81) classifies whether sentence B logically follows sentence A using the [CLS] token representation.

How is BERT fine-tuned for Natural Language Inference in d2l-zh?

Fine-tuning follows the pattern defined in chapter_natural-language-processing-applications/natural-language-inference-bert.md. First, SNLIBERTDataset prepares premise-hypothesis pairs with get_tokens_and_segments. Then, BERTClassifier attaches a new output layer to the pretrained BERTEncoder. During training, d2l.train_ch13 updates all parameters but initializes only the new classifier head from scratch, using split_fn=d2l.split_batch_multi_inputs to handle multi-input batches (tokens, segments, valid lengths).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →