How Word Embeddings Are Implemented and Utilized in d2l-zh

The d2l-zh library implements word embeddings through two primary mechanisms: static pre-trained vectors via the TokenEmbedding class for GloVe and fastText, and learned embeddings via framework-specific nn.Embedding layers inside models like BERT.

The d2l-zh repository provides educational implementations of deep learning algorithms for Chinese readers, with comprehensive support for natural language processing tasks. Understanding how word embeddings are implemented and utilized in d2l-zh is essential for leveraging pre-trained semantic representations or training custom embedding layers from scratch across PyTorch, Paddle, and MXNet backends.

Static Pre-trained Embeddings

The TokenEmbedding class serves as the primary interface for loading pre-trained word embeddings such as GloVe and fastText. This class handles download, extraction, and vector lookup operations across all supported deep learning frameworks.

Architecture

The implementation follows a consistent pipeline across d2l/torch.py, d2l/paddle.py, and d2l/mxnet.py. The TokenEmbedding class performs the following operations:

  • Download and extraction: The download() and download_extract() functions fetch archives from the official data hub and unpack them.
  • File parsing: The _load_embedding() method processes vec.txt files where each line contains a token followed by its vector dimensions.
  • Index construction: The method builds idx_to_token (a list mapping indices to tokens) and idx_to_vec (an NDArray storing vector values, with index 0 reserved for unknown tokens).
  • Vocabulary alignment: When a Vocabulary instance is provided, the embedding vectors are filtered and reordered to match the vocabulary's token set.

Retrieval APIs

The class provides efficient vector retrieval through:

  • __getitem__(tokens): Returns NDArray vectors for a list of tokens, mapping unknown words to the zero vector at index 0.
  • get_vecs_by_tokens(tokens): Explicit method for vector retrieval in contrib implementations.
from d2l import torch as d2l

# Load GloVe 100-dimensional vectors

embed = d2l.TokenEmbedding('glove', 'glove.6b.100d.txt')
vectors = embed['deep', 'learning', 'unknown_word']  # Returns NDArray

Learned Embeddings Inside Neural Models

For end-to-end training, d2l-zh utilizes framework-native embedding layers within model architectures. The BERT implementation demonstrates this pattern.

In d2l/torch.py, the BERTEncoder initializes learned embeddings:


# Inside BERTEncoder.__init__

self.token_embedding = nn.Embedding(vocab_size, num_hiddens)
self.segment_embedding = nn.Embedding(2, num_hiddens)

During the forward pass in BERTEncoder.forward:

X = self.token_embedding(tokens) + self.segment_embedding(segments)

These embeddings start from random initialization and update via backpropagation, capturing task-specific semantic relationships. The same architecture appears in d2l/paddle.py and d2l/mxnet.py with their respective embedding implementations.

Practical Usage

Loading Pre-trained Embeddings with Vocabulary Alignment

from d2l import mxnet as d2l
from d2lzh.text import vocab, embedding

# Build vocabulary from corpus

tokenized_corpus = [['deep', 'learning'], ['natural', 'language', 'processing']]
my_vocab = vocab.Vocabulary(
    counter={t: i for i, t in enumerate(set(sum(tokenized_corpus, [])))},
    min_freq=1,
    reserved_tokens=['<pad>', '<unk>']
)

# Create aligned embedding

glove = embedding.create('glove', 'glove.6b.100d.txt', vocabulary=my_vocab)
vecs = glove[['deep', 'language', 'unknown']]

Using Learned Embeddings in BERT

import torch
from d2l import torch as d2l

# Initialize BERT with learned embeddings

vocab_size, num_hiddens = 10000, 768
net = d2l.BERTModel(
    vocab_size, num_hiddens,
    norm_shape=[num_hiddens],
    ffn_num_input=4*num_hiddens,
    ffn_num_hiddens=4*num_hiddens,
    num_heads=8, num_layers=6, dropout=0.1
)

# Forward pass

tokens = torch.tensor([[1, 2, 3, 4]])
encoded, _, _ = net(tokens, segments=None, valid_lens=None)

Key Source Files

  • d2l/torch.py – PyTorch implementations of TokenEmbedding and BERT embedding layers
  • d2l/paddle.py – PaddlePaddle equivalents
  • d2l/mxnet.py – MXNet/Gluon equivalents
  • contrib/to-rm-mx-contrib-text/d2lzh/text/embedding.py – High-level embedding creation utilities
  • contrib/to-rm-mx-contrib-text/d2lzh/text/vocab.py – Vocabulary construction and alignment

Summary

  • The TokenEmbedding class provides static pre-trained vectors from GloVe and fastText with O(1) lookup via dictionary mapping.
  • Framework-specific nn.Embedding layers enable learned embeddings in models like BERT, starting from random initialization and updating via backpropagation.
  • Vocabulary alignment mechanisms allow pre-trained embeddings to be filtered and reordered to match custom vocabularies.
  • The d2l-zh library abstracts backend differences across PyTorch, Paddle, and MXNet while maintaining consistent APIs for word embeddings.

Frequently Asked Questions

How do I load pre-trained GloVe embeddings in d2l-zh?

Use the TokenEmbedding class from your preferred backend module. For example, in PyTorch: from d2l import torch as d2l; embed = d2l.TokenEmbedding('glove', 'glove.6b.100d.txt'). This downloads the vectors automatically and provides dictionary-based lookup for tokens.

What is the difference between TokenEmbedding and nn.Embedding in d2l-zh?

TokenEmbedding loads static, pre-trained vectors from files like GloVe or fastText and provides read-only lookup for semantic similarity tasks. In contrast, nn.Embedding (used inside models like BERT) initializes random vectors that are updated during training to capture task-specific patterns.

How does d2l-zh handle unknown tokens in word embeddings?

The TokenEmbedding class reserves index 0 for unknown tokens, assigning a zero vector by default. When looking up tokens not present in the pre-trained vocabulary, the class returns the vector at index 0, ensuring graceful handling of out-of-vocabulary words.

Can I align pre-trained embeddings with a custom vocabulary in d2l-zh?

Yes. Pass your Vocabulary instance to the embedding.create() function or TokenEmbedding constructor. The implementation filters the pre-trained vectors to include only tokens present in your vocabulary and reorders them to match your vocabulary's index mapping, enabling seamless integration with downstream models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →