How to Train Custom Word Embeddings Using CBoW in PyTorch: AI-For-Beginners Guide

You can train custom word embeddings using the Continuous Bag-of-Words (CBoW) architecture by implementing an embedding layer to learn vector representations, averaging context word vectors to predict target words, and optimizing with cross-entropy loss as demonstrated in the Microsoft AI-For-Beginners repository.

The Continuous Bag-of-Words (CBoW) model learns dense vector representations by predicting a target word from its surrounding context words. In the microsoft/AI-For-Beginners curriculum, this implementation is contained within the Natural Language Processing (NLP) lesson track, providing a hands-on approach to training domain-specific embeddings from scratch using PyTorch rather than relying on pre-trained vectors.

Understanding the CBoW Architecture

CBoW operates on a simple yet powerful premise: given a sequence of words, the model uses the context words surrounding a target word to predict the target itself. For a context window of size C, the model takes the 2C surrounding words (C words to the left and C to the right), converts them to embedding vectors, aggregates them (typically via averaging), and projects the result through a linear layer to produce a probability distribution over the entire vocabulary.

This architecture forces the model to learn semantic relationships in the data, placing words with similar contexts close together in the embedding space. Unlike Skip-gram, which predicts context from target words, CBoW is computationally efficient and works well with frequent words, making it ideal for beginners learning neural language models.

Step-by-Step Implementation

Vocabulary Construction and Dataset Preparation

Training begins with corpus tokenization and vocabulary mapping. You must create a word-to-index dictionary to convert text into integer IDs that torch.nn.Embedding can process. For each target word in your corpus, you generate training pairs consisting of the context word indices (the sliding window) and the target word index.

In lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb, the dataset preparation involves:

  1. Tokenizing the input corpus into individual words
  2. Building a vocabulary dictionary mapping unique words to integer indices
  3. Creating (context, target) pairs using a sliding window approach
  4. Converting these pairs into PyTorch tensors for batch processing

Model Architecture with torch.nn.Embedding

The CBoW model requires two core components: an embedding layer and a projection layer. The embedding layer stores the trainable word vectors, while the projection layer maps the aggregated context representation to vocabulary logits.

The forward pass follows this computational graph:

  • Lookup embeddings for each context word using self.emb(context)
  • Aggregate context embeddings via mean() to create a single context vector
  • Pass the vector through a linear layer: self.linear(avg_emb)
  • Return raw logits for loss calculation

Training Loop and Loss Optimization

The training process uses CrossEntropyLoss, which combines LogSoftmax and Negative Log-Likelihood loss into a single efficient operation. The optimizer (typically SGD or Adam) adjusts the embedding weights and linear layer parameters to minimize the difference between predicted and actual target words.

The standard training loop iterates over epochs, processes batches of (context, target) pairs, computes the loss via criterion(logits, target), performs backpropagation with loss.backward(), and updates parameters using optimizer.step().

Complete Training Example

The following self-contained example demonstrates the full CBoW workflow, from data preparation to embedding extraction, mirroring the implementation found in the AI-For-Beginners notebook.

import torch
import torch.nn as nn
from torch.utils.data import DataLoader, Dataset

# ---- 1. Corpus and Vocabulary Construction ---------------------------------

corpus = [
    "the quick brown fox jumps over the lazy dog",
    "the fox is quick and the dog is lazy",
    "brown fox brown dog"
]

tokenized = [sentence.split() for sentence in corpus]
vocab = {word: idx for idx, word in enumerate(
    sorted({w for s in tokenized for w in s})
)}
vocab_size = len(vocab)
idx2word = {i: w for w, i in vocab.items()}

# ---- 2. CBOW Dataset Generation --------------------------------------------

class CBOWDataset(Dataset):
    def __init__(self, tokenized, vocab, window=2):
        self.samples = []
        for sentence in tokenized:
            ids = [vocab[w] for w in sentence]
            for i in range(window, len(ids) - window):
                context = ids[i-window:i] + ids[i+1:i+window+1]
                target = ids[i]
                self.samples.append((context, target))
    
    def __len__(self): 
        return len(self.samples)
    
    def __getitem__(self, idx):
        ctx, tgt = self.samples[idx]
        return torch.tensor(ctx, dtype=torch.long), torch.tensor(tgt, dtype=torch.long)

dataset = CBOWDataset(tokenized, vocab, window=2)
loader = DataLoader(dataset, batch_size=4, shuffle=True)

# ---- 3. CBOW Model Definition ------------------------------------------------

class CBOW(nn.Module):
    def __init__(self, vocab_size, embed_dim):
        super().__init__()
        self.emb = nn.Embedding(vocab_size, embed_dim)
        self.linear = nn.Linear(embed_dim, vocab_size)
    
    def forward(self, context):
        # context shape: (batch, 2*window)

        embeds = self.emb(context)          # (batch, context_size, embed_dim)

        avg_emb = embeds.mean(dim=1)        # (batch, embed_dim)

        out = self.linear(avg_emb)          # (batch, vocab_size)

        return out

embed_dim = 50
model = CBOW(vocab_size, embed_dim)

# ---- 4. Training Configuration and Loop --------------------------------------

criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.05)

for epoch in range(30):
    total_loss = 0.0
    for ctx, tgt in loader:
        optimizer.zero_grad()
        logits = model(ctx)
        loss = criterion(logits, tgt)
        loss.backward()
        optimizer.step()
        total_loss += loss.item()
    print(f'Epoch {epoch+1:02d} – Loss: {total_loss:.4f}')

# ---- 5. Extracting Custom Embeddings ----------------------------------------

with torch.no_grad():
    word = "fox"
    vector = model.emb.weight[vocab[word]]
    print(f'Embedding for "{word}": {vector[:5]} ...')

This implementation creates custom word embeddings specific to your corpus. After training, model.emb.weight contains the learned vectors, where semantically similar words (like "fox" and "dog" in the example) will have higher cosine similarity than unrelated terms.

Key Files and Resources

The AI-For-Beginners repository provides the following resources for CBoW implementation:

  • lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb – The complete notebook containing data preprocessing, model definition, training visualization, and embedding analysis (GitHub source)

  • lessons/5-NLP/15-LanguageModeling/README.md – Lesson overview explaining the theoretical foundations of CBoW and its relationship to other language modeling techniques (GitHub source)

  • environment.yml – Specification of required dependencies including PyTorch and supporting libraries needed to run the notebook (GitHub source)

Summary

Training custom word embeddings with CBoW in PyTorch involves these essential steps:

  • Corpus preparation: Tokenize text and build a vocabulary mapping to convert words into integer indices
  • Context window generation: Create (context, target) training pairs using a sliding window across your text
  • Model architecture: Implement torch.nn.Embedding for vector storage and a linear projection layer for vocabulary prediction
  • Aggregation strategy: Average context word embeddings to create a single representation vector
  • Optimization: Use CrossEntropyLoss and SGD/Adam to minimize prediction error and update embedding weights
  • Extraction: Retrieve trained vectors from the embedding layer weight matrix for downstream NLP tasks

Frequently Asked Questions

What is the difference between CBoW and Skip-gram?

CBoW predicts the target word from context words, averaging their embeddings to make a prediction, while Skip-gram predicts context words from the target word. CBoW is computationally faster and performs better on frequent words, whereas Skip-gram works better for rare words and produces richer semantic relationships for smaller datasets.

What context window size should I use for CBoW training?

A context window of 2 to 5 words on each side (total context size of 4-10) is standard. Smaller windows (1-2) capture syntactic relationships better, while larger windows (5-10) capture more general semantic associations. The optimal size depends on your corpus size and the specific semantic relationships you want to capture.

How do I evaluate the quality of custom word embeddings?

Evaluate embeddings through intrinsic tasks like analogy completion (king - man + woman ≈ queen) or word similarity benchmarks (SimLex-999), or through extrinsic tasks where you use the vectors as features in downstream models for text classification or sentiment analysis. Visualizing embeddings with t-SNE or UMAP can also reveal semantic clustering quality.

Can I use GPU acceleration for training CBoW models in PyTorch?

Yes, simply move your model and data to the GPU using .to(device) where device = torch.device("cuda" if torch.cuda.is_available() else "cpu"). Both the embedding lookup and linear layer operations support CUDA acceleration, significantly speeding up training for large vocabularies or extensive corpora.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →