How Word Embeddings Are Implemented and Utilized in d2l-zh
The d2l-zh library implements word embeddings through two primary mechanisms: static pre-trained vectors via the TokenEmbedding class for GloVe and fastText, and learned embeddings via framework-specific nn.Embedding layers inside models like BERT.
The d2l-zh repository provides educational implementations of deep learning algorithms for Chinese readers, with comprehensive support for natural language processing tasks. Understanding how word embeddings are implemented and utilized in d2l-zh is essential for leveraging pre-trained semantic representations or training custom embedding layers from scratch across PyTorch, Paddle, and MXNet backends.
Static Pre-trained Embeddings
The TokenEmbedding class serves as the primary interface for loading pre-trained word embeddings such as GloVe and fastText. This class handles download, extraction, and vector lookup operations across all supported deep learning frameworks.
Architecture
The implementation follows a consistent pipeline across d2l/torch.py, d2l/paddle.py, and d2l/mxnet.py. The TokenEmbedding class performs the following operations:
- Download and extraction: The
download()anddownload_extract()functions fetch archives from the official data hub and unpack them. - File parsing: The
_load_embedding()method processesvec.txtfiles where each line contains a token followed by its vector dimensions. - Index construction: The method builds
idx_to_token(a list mapping indices to tokens) andidx_to_vec(an NDArray storing vector values, with index 0 reserved for unknown tokens). - Vocabulary alignment: When a
Vocabularyinstance is provided, the embedding vectors are filtered and reordered to match the vocabulary's token set.
Retrieval APIs
The class provides efficient vector retrieval through:
__getitem__(tokens): Returns NDArray vectors for a list of tokens, mapping unknown words to the zero vector at index 0.get_vecs_by_tokens(tokens): Explicit method for vector retrieval in contrib implementations.
from d2l import torch as d2l
# Load GloVe 100-dimensional vectors
embed = d2l.TokenEmbedding('glove', 'glove.6b.100d.txt')
vectors = embed['deep', 'learning', 'unknown_word'] # Returns NDArray
Learned Embeddings Inside Neural Models
For end-to-end training, d2l-zh utilizes framework-native embedding layers within model architectures. The BERT implementation demonstrates this pattern.
In d2l/torch.py, the BERTEncoder initializes learned embeddings:
# Inside BERTEncoder.__init__
self.token_embedding = nn.Embedding(vocab_size, num_hiddens)
self.segment_embedding = nn.Embedding(2, num_hiddens)
During the forward pass in BERTEncoder.forward:
X = self.token_embedding(tokens) + self.segment_embedding(segments)
These embeddings start from random initialization and update via backpropagation, capturing task-specific semantic relationships. The same architecture appears in d2l/paddle.py and d2l/mxnet.py with their respective embedding implementations.
Practical Usage
Loading Pre-trained Embeddings with Vocabulary Alignment
from d2l import mxnet as d2l
from d2lzh.text import vocab, embedding
# Build vocabulary from corpus
tokenized_corpus = [['deep', 'learning'], ['natural', 'language', 'processing']]
my_vocab = vocab.Vocabulary(
counter={t: i for i, t in enumerate(set(sum(tokenized_corpus, [])))},
min_freq=1,
reserved_tokens=['<pad>', '<unk>']
)
# Create aligned embedding
glove = embedding.create('glove', 'glove.6b.100d.txt', vocabulary=my_vocab)
vecs = glove[['deep', 'language', 'unknown']]
Using Learned Embeddings in BERT
import torch
from d2l import torch as d2l
# Initialize BERT with learned embeddings
vocab_size, num_hiddens = 10000, 768
net = d2l.BERTModel(
vocab_size, num_hiddens,
norm_shape=[num_hiddens],
ffn_num_input=4*num_hiddens,
ffn_num_hiddens=4*num_hiddens,
num_heads=8, num_layers=6, dropout=0.1
)
# Forward pass
tokens = torch.tensor([[1, 2, 3, 4]])
encoded, _, _ = net(tokens, segments=None, valid_lens=None)
Key Source Files
- d2l/torch.py – PyTorch implementations of
TokenEmbeddingand BERT embedding layers - d2l/paddle.py – PaddlePaddle equivalents
- d2l/mxnet.py – MXNet/Gluon equivalents
- contrib/to-rm-mx-contrib-text/d2lzh/text/embedding.py – High-level embedding creation utilities
- contrib/to-rm-mx-contrib-text/d2lzh/text/vocab.py – Vocabulary construction and alignment
Summary
- The
TokenEmbeddingclass provides static pre-trained vectors from GloVe and fastText with O(1) lookup via dictionary mapping. - Framework-specific
nn.Embeddinglayers enable learned embeddings in models like BERT, starting from random initialization and updating via backpropagation. - Vocabulary alignment mechanisms allow pre-trained embeddings to be filtered and reordered to match custom vocabularies.
- The d2l-zh library abstracts backend differences across PyTorch, Paddle, and MXNet while maintaining consistent APIs for word embeddings.
Frequently Asked Questions
How do I load pre-trained GloVe embeddings in d2l-zh?
Use the TokenEmbedding class from your preferred backend module. For example, in PyTorch: from d2l import torch as d2l; embed = d2l.TokenEmbedding('glove', 'glove.6b.100d.txt'). This downloads the vectors automatically and provides dictionary-based lookup for tokens.
What is the difference between TokenEmbedding and nn.Embedding in d2l-zh?
TokenEmbedding loads static, pre-trained vectors from files like GloVe or fastText and provides read-only lookup for semantic similarity tasks. In contrast, nn.Embedding (used inside models like BERT) initializes random vectors that are updated during training to capture task-specific patterns.
How does d2l-zh handle unknown tokens in word embeddings?
The TokenEmbedding class reserves index 0 for unknown tokens, assigning a zero vector by default. When looking up tokens not present in the pre-trained vocabulary, the class returns the vector at index 0, ensuring graceful handling of out-of-vocabulary words.
Can I align pre-trained embeddings with a custom vocabulary in d2l-zh?
Yes. Pass your Vocabulary instance to the embedding.create() function or TokenEmbedding constructor. The implementation filters the pre-trained vectors to include only tokens present in your vocabulary and reorders them to match your vocabulary's index mapping, enabling seamless integration with downstream models.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →