# How Word2Vec and GloVe Embeddings Are Used in NLP: Microsoft AI-For-Beginners Guide

> Learn how Word2Vec and GloVe embeddings power NLP with the Microsoft AI-For-Beginners guide. Explore CBoW Skip-Gram and pre-trained vector integration.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-29

---

**The Microsoft AI-For-Beginners curriculum implements Word2Vec embeddings through CBoW and Skip-Gram architectures in TensorFlow and PyTorch, while demonstrating GloVe integration via pre-trained vector loading and vocabulary alignment utilities.**

The microsoft/AI-For-Beginners repository provides hands-on instruction for semantic embeddings in Lesson 5-NLP (14-Embeddings and 15-LanguageModeling). This curriculum bridges theory and practice by teaching learners how to train predictive **Word2Vec** models from scratch and inject static **GloVe** vectors into neural networks.

## Word2Vec Architectures in the Curriculum

The repository covers two predictive architectures that learn dense vector representations by optimizing context-word relationships.

### Continuous Bag-of-Words (CBoW) Implementation

The CBoW model predicts a target word from its surrounding context window. In `lessons/5-NLP/15-LanguageModeling/CBoW-TF.ipynb`, the curriculum implements this using a standard **Embedding layer** initialized with random weights and trained on the AG News dataset.

The TensorFlow implementation defines the embedding layer as:

```python
import tensorflow as tf
from tensorflow import keras

vocab_size = 5000
embed_dim = 30
vectorizer = keras.layers.experimental.preprocessing.TextVectorization(
    max_tokens=vocab_size, input_shape=(1,))

# Word2Vec embedding layer

embedder = keras.layers.Embedding(vocab_size, embed_dim, input_length=1)

model = keras.Sequential([
    embedder,
    keras.layers.Dense(vocab_size, activation='softmax')
])
model.compile(optimizer=keras.optimizers.SGD(learning_rate=0.1),
              loss='sparse_categorical_crossentropy')
model.fit(ds, epochs=200)

```

The PyTorch counterpart in `CBoW-PyTorch.ipynb` utilizes `nn.Embedding(vocab_size, embed_dim)` to achieve the same functionality. Both implementations learn a matrix that maps each vocabulary token to a low-dimensional vector capable of capturing semantic relationships.

### Skip-Gram Model Overview

While the practical notebooks focus on CBoW implementations, the curriculum documentation explains **Skip-Gram** as the inverse architecture: predicting surrounding context words from a central target word. Both approaches generate the same Word2Vec representation format, allowing learners to interchange architectures based on dataset characteristics.

## GloVe Embeddings Integration

Unlike Word2Vec's predictive training, **GloVe** (Global Vectors) utilizes a count-based approach that factorizes word-context co-occurrence matrices.

### Loading Pre-trained Vectors

The curriculum demonstrates loading static 300-dimensional GloVe vectors using the `gensim` library. These pre-trained embeddings require no additional training and can be directly injected into model layers:

```python
from gensim.models import KeyedVectors

glove_path = 'glove.6B.300d.txt'
glove = KeyedVectors.load_word2vec_format(glove_path, binary=False)

# Retrieve vector for specific terms

neural_vec = glove['neural']

```

### Handling Vocabulary Mismatches with torchnlp.py

When integrating external embeddings, vocabulary mismatches between the source corpus and model tokenizer require resolution. The utility module [`lessons/5-NLP/14-Embeddings/torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/14-Embeddings/torchnlp.py) contains the `encode()` function that abstracts token-to-index conversion.

According to lines 33-38 of [`torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/torchnlp.py), the implementation supports both TorchText vocabulary objects (`vocab.get_stoi()`) and GloVe objects (`vocab.stoi`) through unified mapping logic:

```python
from lessons.5_NLP.14_Embeddings.torchnlp import encode

sentence = "AI for beginners"
indices = encode(sentence)  # Returns list of token ids aligned to embedding matrix

```

This ensures seamless integration regardless of whether embeddings are trained or pre-trained.

## Practical Usage and Downstream Applications

After training Word2Vec models, the curriculum demonstrates how to extract the learned embedding matrix and perform semantic similarity queries. The embedding vectors can be retrieved and inspected using:

```python
import numpy as np

# Extract full embedding matrix

vectors = embedder(vectorizer(vocab))  # Shape: (vocab_size, embed_dim)

def close_words(word, n=5):
    vec = embedder(vectorizer(word))[0]
    # Euclidean distance search

    distances = np.linalg.norm(vectors - vec, axis=1)
    idx = distances.argsort()[:n]
    return [vocab[i] for i in idx]

# Find semantically similar terms

print(close_words('paris'))  # ['paris', 'philippines', 'seoul', ...]

```

This functionality enables downstream tasks such as analogical reasoning, document similarity, and feature initialization for classification networks.

## Summary

- **Word2Vec Implementation**: The curriculum provides complete CBoW training pipelines in both TensorFlow and PyTorch, utilizing `keras.layers.Embedding` and `nn.Embedding` layers with configurable dimensions (typically 30-300).
- **GloVe Integration**: Pre-trained 300-dimensional vectors are loaded via `gensim`, offering static embeddings that require no additional training overhead.
- **Vocabulary Alignment**: The [`torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/torchnlp.py) utility handles `stoi` (string-to-index) mappings across different vocabulary types, ensuring compatibility between custom tokenizers and external embedding sources.
- **Semantic Querying**: Trained embeddings support nearest-neighbor search using Euclidean distance, enabling practical similarity tasks directly within the notebook environment.

## Frequently Asked Questions

### What is the difference between Word2Vec and GloVe in the AI-For-Beginners curriculum?

**Word2Vec** uses predictive neural architectures (CBoW or Skip-Gram) to learn embeddings dynamically from a specific dataset, while **GloVe** employs count-based matrix factorization on global co-occurrence statistics. The curriculum teaches Word2Vec training from scratch on AG News and treats GloVe as pre-trained static vectors ready for immediate use.

### How does the curriculum handle vocabulary mismatches when using pre-trained embeddings?

The [`torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/torchnlp.py) module provides an `encode()` function that normalizes access to `stoi` mappings. According to the source code at [`lessons/5-NLP/14-Embeddings/torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/14-Embeddings/torchnlp.py), it detects whether the vocabulary object uses `get_stoi()` (TorchText) or direct `stoi` attribute access (GloVe), ensuring tokens map to correct indices regardless of embedding source.

### Can the learned Word2Vec embeddings be used for tasks other than text classification?

Yes. The CBoW notebooks demonstrate extracting the embedding matrix to perform semantic similarity searches. Learners can use these vectors for clustering, analogy completion, or as pre-initialized weights in downstream models such as LSTMs or Transformers.

### What datasets are used to train the Word2Vec models in the repository?

The curriculum utilizes the **AG News** dataset for training CBoW implementations. This dataset provides sufficient text volume to demonstrate semantic relationships while remaining computationally manageable for educational purposes.