# How Word Embeddings Are Implemented and Utilized in d2l-zh

> Discover how word embeddings are implemented and utilized in d2l-zh. Explore static pre-trained vectors like GloVe and fastText, and learn about learned embeddings within models like BERT.

- Repository: [Dive into Deep Learning (D2L.ai)/d2l-zh](https://github.com/d2l-ai/d2l-zh)
- Tags: how-to-guide
- Published: 2026-03-01

---

**The d2l-zh library implements word embeddings through two primary mechanisms: static pre-trained vectors via the `TokenEmbedding` class for GloVe and fastText, and learned embeddings via framework-specific `nn.Embedding` layers inside models like BERT.**

The d2l-zh repository provides educational implementations of deep learning algorithms for Chinese readers, with comprehensive support for natural language processing tasks. Understanding how word embeddings are implemented and utilized in d2l-zh is essential for leveraging pre-trained semantic representations or training custom embedding layers from scratch across PyTorch, Paddle, and MXNet backends.

## Static Pre-trained Embeddings

The `TokenEmbedding` class serves as the primary interface for loading pre-trained word embeddings such as GloVe and fastText. This class handles download, extraction, and vector lookup operations across all supported deep learning frameworks.

### Architecture

The implementation follows a consistent pipeline across [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py), [`d2l/paddle.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/paddle.py), and [`d2l/mxnet.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/mxnet.py). The `TokenEmbedding` class performs the following operations:

- **Download and extraction**: The `download()` and `download_extract()` functions fetch archives from the official data hub and unpack them.
- **File parsing**: The `_load_embedding()` method processes [`vec.txt`](https://github.com/d2l-ai/d2l-zh/blob/main/vec.txt) files where each line contains a token followed by its vector dimensions.
- **Index construction**: The method builds `idx_to_token` (a list mapping indices to tokens) and `idx_to_vec` (an NDArray storing vector values, with index 0 reserved for unknown tokens).
- **Vocabulary alignment**: When a `Vocabulary` instance is provided, the embedding vectors are filtered and reordered to match the vocabulary's token set.

### Retrieval APIs

The class provides efficient vector retrieval through:

- `__getitem__(tokens)`: Returns NDArray vectors for a list of tokens, mapping unknown words to the zero vector at index 0.
- `get_vecs_by_tokens(tokens)`: Explicit method for vector retrieval in contrib implementations.

```python
from d2l import torch as d2l

# Load GloVe 100-dimensional vectors

embed = d2l.TokenEmbedding('glove', 'glove.6b.100d.txt')
vectors = embed['deep', 'learning', 'unknown_word']  # Returns NDArray

```

## Learned Embeddings Inside Neural Models

For end-to-end training, d2l-zh utilizes framework-native embedding layers within model architectures. The BERT implementation demonstrates this pattern.

In [`d2l/torch.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/torch.py), the `BERTEncoder` initializes learned embeddings:

```python

# Inside BERTEncoder.__init__

self.token_embedding = nn.Embedding(vocab_size, num_hiddens)
self.segment_embedding = nn.Embedding(2, num_hiddens)

```

During the forward pass in `BERTEncoder.forward`:

```python
X = self.token_embedding(tokens) + self.segment_embedding(segments)

```

These embeddings start from random initialization and update via backpropagation, capturing task-specific semantic relationships. The same architecture appears in [`d2l/paddle.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/paddle.py) and [`d2l/mxnet.py`](https://github.com/d2l-ai/d2l-zh/blob/main/d2l/mxnet.py) with their respective embedding implementations.

## Practical Usage

### Loading Pre-trained Embeddings with Vocabulary Alignment

```python
from d2l import mxnet as d2l
from d2lzh.text import vocab, embedding

# Build vocabulary from corpus

tokenized_corpus = [['deep', 'learning'], ['natural', 'language', 'processing']]
my_vocab = vocab.Vocabulary(
    counter={t: i for i, t in enumerate(set(sum(tokenized_corpus, [])))},
    min_freq=1,
    reserved_tokens=['<pad>', '<unk>']
)

# Create aligned embedding

glove = embedding.create('glove', 'glove.6b.100d.txt', vocabulary=my_vocab)
vecs = glove[['deep', 'language', 'unknown']]

```

### Using Learned Embeddings in BERT

```python
import torch
from d2l import torch as d2l

# Initialize BERT with learned embeddings

vocab_size, num_hiddens = 10000, 768
net = d2l.BERTModel(
    vocab_size, num_hiddens,
    norm_shape=[num_hiddens],
    ffn_num_input=4*num_hiddens,
    ffn_num_hiddens=4*num_hiddens,
    num_heads=8, num_layers=6, dropout=0.1
)

# Forward pass

tokens = torch.tensor([[1, 2, 3, 4]])
encoded, _, _ = net(tokens, segments=None, valid_lens=None)

```

## Key Source Files

- **d2l/torch.py** – PyTorch implementations of `TokenEmbedding` and BERT embedding layers
- **d2l/paddle.py** – PaddlePaddle equivalents
- **d2l/mxnet.py** – MXNet/Gluon equivalents
- **contrib/to-rm-mx-contrib-text/d2lzh/text/embedding.py** – High-level embedding creation utilities
- **contrib/to-rm-mx-contrib-text/d2lzh/text/vocab.py** – Vocabulary construction and alignment

## Summary

- The `TokenEmbedding` class provides static pre-trained vectors from GloVe and fastText with O(1) lookup via dictionary mapping.
- Framework-specific `nn.Embedding` layers enable learned embeddings in models like BERT, starting from random initialization and updating via backpropagation.
- Vocabulary alignment mechanisms allow pre-trained embeddings to be filtered and reordered to match custom vocabularies.
- The d2l-zh library abstracts backend differences across PyTorch, Paddle, and MXNet while maintaining consistent APIs for word embeddings.

## Frequently Asked Questions

### How do I load pre-trained GloVe embeddings in d2l-zh?

Use the `TokenEmbedding` class from your preferred backend module. For example, in PyTorch: `from d2l import torch as d2l; embed = d2l.TokenEmbedding('glove', 'glove.6b.100d.txt')`. This downloads the vectors automatically and provides dictionary-based lookup for tokens.

### What is the difference between TokenEmbedding and nn.Embedding in d2l-zh?

`TokenEmbedding` loads static, pre-trained vectors from files like GloVe or fastText and provides read-only lookup for semantic similarity tasks. In contrast, `nn.Embedding` (used inside models like BERT) initializes random vectors that are updated during training to capture task-specific patterns.

### How does d2l-zh handle unknown tokens in word embeddings?

The `TokenEmbedding` class reserves index 0 for unknown tokens, assigning a zero vector by default. When looking up tokens not present in the pre-trained vocabulary, the class returns the vector at index 0, ensuring graceful handling of out-of-vocabulary words.

### Can I align pre-trained embeddings with a custom vocabulary in d2l-zh?

Yes. Pass your `Vocabulary` instance to the `embedding.create()` function or `TokenEmbedding` constructor. The implementation filters the pre-trained vectors to include only tokens present in your vocabulary and reorders them to match your vocabulary's index mapping, enabling seamless integration with downstream models.