# How to Train Custom Word Embeddings Using CBoW in PyTorch: AI-For-Beginners Guide

> Learn to train custom word embeddings with CBoW in PyTorch. This AI for Beginners guide shows how to create vector representations for your NLP projects.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: how-to-guide
- Published: 2026-08-29

---

**You can train custom word embeddings using the Continuous Bag-of-Words (CBoW) architecture by implementing an embedding layer to learn vector representations, averaging context word vectors to predict target words, and optimizing with cross-entropy loss as demonstrated in the Microsoft AI-For-Beginners repository.**

The **Continuous Bag-of-Words (CBoW)** model learns dense vector representations by predicting a target word from its surrounding context words. In the **microsoft/AI-For-Beginners** curriculum, this implementation is contained within the Natural Language Processing (NLP) lesson track, providing a hands-on approach to training domain-specific embeddings from scratch using **PyTorch** rather than relying on pre-trained vectors.

## Understanding the CBoW Architecture

CBoW operates on a simple yet powerful premise: given a sequence of words, the model uses the context words surrounding a target word to predict the target itself. For a context window of size *C*, the model takes the 2*C* surrounding words (C words to the left and C to the right), converts them to embedding vectors, aggregates them (typically via averaging), and projects the result through a linear layer to produce a probability distribution over the entire vocabulary.

This architecture forces the model to learn **semantic relationships** in the data, placing words with similar contexts close together in the embedding space. Unlike Skip-gram, which predicts context from target words, CBoW is computationally efficient and works well with frequent words, making it ideal for beginners learning neural language models.

## Step-by-Step Implementation

### Vocabulary Construction and Dataset Preparation

Training begins with corpus tokenization and vocabulary mapping. You must create a word-to-index dictionary to convert text into integer IDs that **torch.nn.Embedding** can process. For each target word in your corpus, you generate training pairs consisting of the context word indices (the sliding window) and the target word index.

In `lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb`, the dataset preparation involves:

1. Tokenizing the input corpus into individual words
2. Building a vocabulary dictionary mapping unique words to integer indices
3. Creating (context, target) pairs using a sliding window approach
4. Converting these pairs into PyTorch tensors for batch processing

### Model Architecture with torch.nn.Embedding

The CBoW model requires two core components: an embedding layer and a projection layer. The embedding layer stores the trainable word vectors, while the projection layer maps the aggregated context representation to vocabulary logits.

The forward pass follows this computational graph:
- Lookup embeddings for each context word using `self.emb(context)`
- Aggregate context embeddings via `mean()` to create a single context vector
- Pass the vector through a linear layer: `self.linear(avg_emb)`
- Return raw logits for loss calculation

### Training Loop and Loss Optimization

The training process uses **CrossEntropyLoss**, which combines LogSoftmax and Negative Log-Likelihood loss into a single efficient operation. The optimizer (typically **SGD** or **Adam**) adjusts the embedding weights and linear layer parameters to minimize the difference between predicted and actual target words.

The standard training loop iterates over epochs, processes batches of (context, target) pairs, computes the loss via `criterion(logits, target)`, performs backpropagation with `loss.backward()`, and updates parameters using `optimizer.step()`.

## Complete Training Example

The following self-contained example demonstrates the full CBoW workflow, from data preparation to embedding extraction, mirroring the implementation found in the AI-For-Beginners notebook.

```python
import torch
import torch.nn as nn
from torch.utils.data import DataLoader, Dataset

# ---- 1. Corpus and Vocabulary Construction ---------------------------------

corpus = [
    "the quick brown fox jumps over the lazy dog",
    "the fox is quick and the dog is lazy",
    "brown fox brown dog"
]

tokenized = [sentence.split() for sentence in corpus]
vocab = {word: idx for idx, word in enumerate(
    sorted({w for s in tokenized for w in s})
)}
vocab_size = len(vocab)
idx2word = {i: w for w, i in vocab.items()}

# ---- 2. CBOW Dataset Generation --------------------------------------------

class CBOWDataset(Dataset):
    def __init__(self, tokenized, vocab, window=2):
        self.samples = []
        for sentence in tokenized:
            ids = [vocab[w] for w in sentence]
            for i in range(window, len(ids) - window):
                context = ids[i-window:i] + ids[i+1:i+window+1]
                target = ids[i]
                self.samples.append((context, target))
    
    def __len__(self): 
        return len(self.samples)
    
    def __getitem__(self, idx):
        ctx, tgt = self.samples[idx]
        return torch.tensor(ctx, dtype=torch.long), torch.tensor(tgt, dtype=torch.long)

dataset = CBOWDataset(tokenized, vocab, window=2)
loader = DataLoader(dataset, batch_size=4, shuffle=True)

# ---- 3. CBOW Model Definition ------------------------------------------------

class CBOW(nn.Module):
    def __init__(self, vocab_size, embed_dim):
        super().__init__()
        self.emb = nn.Embedding(vocab_size, embed_dim)
        self.linear = nn.Linear(embed_dim, vocab_size)
    
    def forward(self, context):
        # context shape: (batch, 2*window)

        embeds = self.emb(context)          # (batch, context_size, embed_dim)

        avg_emb = embeds.mean(dim=1)        # (batch, embed_dim)

        out = self.linear(avg_emb)          # (batch, vocab_size)

        return out

embed_dim = 50
model = CBOW(vocab_size, embed_dim)

# ---- 4. Training Configuration and Loop --------------------------------------

criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.05)

for epoch in range(30):
    total_loss = 0.0
    for ctx, tgt in loader:
        optimizer.zero_grad()
        logits = model(ctx)
        loss = criterion(logits, tgt)
        loss.backward()
        optimizer.step()
        total_loss += loss.item()
    print(f'Epoch {epoch+1:02d} – Loss: {total_loss:.4f}')

# ---- 5. Extracting Custom Embeddings ----------------------------------------

with torch.no_grad():
    word = "fox"
    vector = model.emb.weight[vocab[word]]
    print(f'Embedding for "{word}": {vector[:5]} ...')

```

This implementation creates **custom word embeddings** specific to your corpus. After training, `model.emb.weight` contains the learned vectors, where semantically similar words (like "fox" and "dog" in the example) will have higher cosine similarity than unrelated terms.

## Key Files and Resources

The AI-For-Beginners repository provides the following resources for CBoW implementation:

- **`lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb`** – The complete notebook containing data preprocessing, model definition, training visualization, and embedding analysis ([GitHub source](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/15-LanguageModeling/CBoW-PyTorch.ipynb))

- **[`lessons/5-NLP/15-LanguageModeling/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/15-LanguageModeling/README.md)** – Lesson overview explaining the theoretical foundations of CBoW and its relationship to other language modeling techniques ([GitHub source](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/15-LanguageModeling/README.md))

- **[`environment.yml`](https://github.com/microsoft/AI-For-Beginners/blob/main/environment.yml)** – Specification of required dependencies including PyTorch and supporting libraries needed to run the notebook ([GitHub source](https://github.com/microsoft/AI-For-Beginners/blob/main/environment.yml))

## Summary

Training custom word embeddings with CBoW in PyTorch involves these essential steps:

- **Corpus preparation**: Tokenize text and build a vocabulary mapping to convert words into integer indices
- **Context window generation**: Create (context, target) training pairs using a sliding window across your text
- **Model architecture**: Implement `torch.nn.Embedding` for vector storage and a linear projection layer for vocabulary prediction
- **Aggregation strategy**: Average context word embeddings to create a single representation vector
- **Optimization**: Use CrossEntropyLoss and SGD/Adam to minimize prediction error and update embedding weights
- **Extraction**: Retrieve trained vectors from the embedding layer weight matrix for downstream NLP tasks

## Frequently Asked Questions

### What is the difference between CBoW and Skip-gram?

**CBoW predicts the target word from context words**, averaging their embeddings to make a prediction, while **Skip-gram predicts context words from the target word**. CBoW is computationally faster and performs better on frequent words, whereas Skip-gram works better for rare words and produces richer semantic relationships for smaller datasets.

### What context window size should I use for CBoW training?

A **context window of 2 to 5 words** on each side (total context size of 4-10) is standard. Smaller windows (1-2) capture syntactic relationships better, while larger windows (5-10) capture more general semantic associations. The optimal size depends on your corpus size and the specific semantic relationships you want to capture.

### How do I evaluate the quality of custom word embeddings?

Evaluate embeddings through **intrinsic tasks** like analogy completion (king - man + woman ≈ queen) or word similarity benchmarks (SimLex-999), or through **extrinsic tasks** where you use the vectors as features in downstream models for text classification or sentiment analysis. Visualizing embeddings with t-SNE or UMAP can also reveal semantic clustering quality.

### Can I use GPU acceleration for training CBoW models in PyTorch?

Yes, simply move your model and data to the GPU using `.to(device)` where `device = torch.device("cuda" if torch.cuda.is_available() else "cpu")`. Both the embedding lookup and linear layer operations support CUDA acceleration, significantly speeding up training for large vocabularies or extensive corpora.