How to Implement Word2Vec and GloVe Semantic Embeddings with PyTorch Embeddings

The Microsoft AI-For-Beginners curriculum demonstrates loading pretrained Word2Vec vectors via gensim and GloVe vectors via torchtext, then injecting them into trainable PyTorch models using torch.nn.Embedding.from_pretrained() to enrich text classification with semantic knowledge.

The microsoft/AI-For-Beginners repository provides a complete walkthrough of semantic embedding integration in lesson 5-NLP → 14-Embeddings. Instead of training embeddings from scratch, you can implement Word2Vec and GloVe semantic embeddings with PyTorch by transferring weights from pretrained models into torch.nn.Embedding layers, allowing your network to leverage billions of tokens of pre-computed linguistic context.

Loading Pretrained Word2Vec Vectors with Gensim

The Word2Vec implementation in EmbeddingsPyTorch.ipynb uses the gensim library to parse Google's binary Word2Vec format and align it with your task vocabulary. Because the pretrained vocabulary rarely matches your corpus exactly, the code handles out-of-vocabulary (OOV) tokens by initializing them with random Gaussian noise while copying known vectors directly from the pretrained matrix.

import torch
import gensim
import numpy as np

# Load GoogleNews vectors (300-dimensional)

wv = gensim.models.KeyedVectors.load_word2vec_format(
    "GoogleNews-vectors-negative300.bin.gz", binary=True
)

# Build embedding matrix aligned to your vocab

embed_dim = wv.vector_size
matrix = np.zeros((len(vocab), embed_dim))

for i, token in enumerate(vocab):
    if token in wv:
        matrix[i] = wv[token]
    else:
        matrix[i] = np.random.normal(scale=0.6, size=(embed_dim,))

# Create PyTorch embedding layer (fine-tuning enabled)

embedding = torch.nn.Embedding.from_pretrained(
    torch.tensor(matrix, dtype=torch.float), freeze=False
)

Setting freeze=False allows gradient updates to adapt the semantic vectors to your specific domain during downstream training.

Integrating GloVe Embeddings with Torchtext

For GloVe integration, the lesson leverages torchtext's built-in downloader to fetch Common Crawl or Wikipedia-derived vectors without manual file management. The torchtext.vocab.GloVe class returns a vocabulary object containing the tensor of pretrained vectors, which integrates seamlessly with PyTorch's embedding constructor.

import torch
import torchtext

# Download 50-dimensional GloVe vectors trained on 6B tokens

glove = torchtext.vocab.GloVe(name="6B", dim=50)

# Build embedding layer directly from pretrained weights

embedding = torch.nn.Embedding.from_pretrained(
    glove.vectors, freeze=False
)

When your custom vocabulary indices differ from GloVe's internal mapping, you must remap indices or reindex your vocabulary to align with the glove.stoi dictionary before constructing the layer.

Building a Classifier with Semantic Embeddings

Once the embedding layer is initialized with Word2Vec or GloVe weights, the lesson implements an EmbedClassifier that aggregates token representations into a fixed-length sentence vector. The architecture looks up embeddings for input token indices, performs mean pooling across the sequence dimension, and passes the result through a linear classifier.

class EmbedClassifier(torch.nn.Module):
    def __init__(self, vocab_size, embed_dim, n_classes):
        super().__init__()
        self.embedding = torch.nn.Embedding(vocab_size, embed_dim)
        self.fc = torch.nn.Linear(embed_dim, n_classes)

    def forward(self, x):
        # x: [batch_size, seq_len]

        emb = self.embedding(x)               # [batch, seq_len, embed_dim]

        pooled = torch.mean(emb, dim=1)        # [batch, embed_dim]

        return self.fc(pooled)                 # [batch, n_classes]

Replace the randomly initialized self.embedding with your pretrained embedding layer from the previous sections to inject semantic knowledge.

Handling Variable-Length Sequences with EmbeddingBag

The notebook demonstrates two strategies for batching sequences of different lengths. The first pads every sequence to the maximum length in the batch using a padify helper function. The second, more memory-efficient approach uses torch.nn.EmbeddingBag, which processes concatenated sequences and a separate offsets tensor indicating where each example begins.

class BagClassifier(torch.nn.Module):
    def __init__(self, vocab_size, embed_dim, n_classes):
        super().__init__()
        self.embedding = torch.nn.EmbeddingBag(
            vocab_size, embed_dim, mode='mean'
        )
        self.fc = torch.nn.Linear(embed_dim, n_classes)

    def forward(self, text, offsets):
        # text: 1-D tensor [total_tokens]

        # offsets: 1-D tensor [batch_size] marking sequence starts

        emb = self.embedding(text, offsets)   # [batch, embed_dim]

        return self.fc(emb)

EmbeddingBag eliminates the need for padding tokens and computes the mean (or sum) of embeddings internally, making it ideal for large-vocabulary text classification tasks.

Training Custom Embeddings from Scratch

If your domain vocabulary diverges significantly from general English, the companion notebook lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb demonstrates training Word2Vec-style embeddings using the Continuous Bag-of-Words (CBoW) architecture. This approach learns embeddings directly from your corpus co-occurrence statistics rather than loading pretrained weights.

Summary

  • Word2Vec integration requires loading binary vectors via gensim.models.KeyedVectors, aligning them with your vocabulary, and passing the resulting NumPy matrix to torch.nn.Embedding.from_pretrained().
  • GloVe integration uses torchtext.vocab.GloVe to download vectors and constructs the embedding layer directly from the returned tensor.
  • Fine-tuning control is managed via the freeze parameter; set freeze=False to adapt semantic vectors during training or freeze=True to keep them static.
  • Variable-length handling can use either padding with padify or EmbeddingBag with offset indices for efficient aggregation of token embeddings.

Frequently Asked Questions

How do I handle out-of-vocabulary words when using pretrained Word2Vec vectors?

Iterate through your vocabulary and check membership in the pretrained model's key index. For tokens missing from the GoogleNews vectors, initialize their embedding vector with small random values drawn from a normal distribution (e.g., scale=0.6) to prevent zero-gradient issues during backpropagation.

Can I use higher-dimensional GloVe vectors like 300d instead of 50d?

Yes. Change the dim parameter in torchtext.vocab.GloVe(name="6B", dim=300) to match the desired vector size, ensuring your EmbedClassifier initializes its linear layer with the matching embed_dim. The 6B token corpus provides 50d, 100d, 200d, and 300d variants.

Should I freeze pretrained embeddings or allow fine-tuning?

Freeze embeddings (freeze=True) when working with small datasets to prevent overfitting and preserve general semantic relationships. Enable fine-tuning (freeze=False) when you have substantial domain-specific data that could benefit from adapting the vector space to task-specific nuances, such as medical or legal terminology.

What is the difference between using Embedding with padding versus EmbeddingBag?

torch.nn.Embedding requires fixed-length inputs (achieved via padding) and returns a 3-D tensor [batch, seq_len, embed_dim] that you manually pool. torch.nn.EmbeddingBag accepts variable-length sequences as a flattened 1-D tensor with offsets, computes the mean or sum internally, and returns a 2-D tensor [batch, embed_dim], reducing memory overhead and computation time for classification tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →