How to Represent Text Using Bag-of-Words and TF-IDF in PyTorch
Microsoft's AI for Beginners curriculum demonstrates how to convert raw text into numerical tensors using Bag-of-Words and TF-IDF representations that integrate directly with PyTorch DataLoaders and linear classifiers.
The Microsoft AI-For-Beginners repository provides a practical introduction to natural language processing fundamentals in lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb. This guide walks through the complete pipeline for representing text using Bag-of-Words and TF-IDF in PyTorch, from tokenization with torchtext to training a neural classifier on the AG-NEWS dataset.
Loading the AG-NEWS Dataset with torchtext
The implementation uses the AG-NEWS dataset provided by torchtext.datasets, which contains news headlines categorized into four classes: World, Sports, Business, and Sci/Tech. Each sample returns a tuple of (label, text) where labels are 1-indexed integers.
The dataset is loaded with a specified root directory for caching:
import torchtext
import os
os.makedirs('./data', exist_ok=True)
train_dataset, test_dataset = torchtext.datasets.AG_NEWS(root='./data')
classes = ['World', 'Sports', 'Business', 'Sci/Tech']
Building a Vocabulary and Tokenizer
Before creating vector representations, you must tokenize raw strings and map tokens to integer indices. The implementation uses the basic_english tokenizer from torchtext.data.utils and builds a vocabulary using Python's collections.Counter.
import collections
from torchtext.data.utils import get_tokenizer
tokenizer = get_tokenizer('basic_english')
counter = collections.Counter()
# Count all tokens in the training set
for _, line in train_dataset:
counter.update(tokenizer(line))
# Create vocabulary with minimum frequency of 1
vocab = torchtext.vocab.vocab(counter, min_freq=1)
vocab_size = len(vocab)
stoi = vocab.get_stoi() # String-to-int mapping
def encode(text):
"""Convert text to list of integer indices."""
return [stoi[t] for t in tokenizer(text)]
The stoi dictionary provides the mapping required to convert any token into its corresponding index within the fixed-size feature vector.
Implementing Bag-of-Words (BoW) in PyTorch
Bag-of-Words represents each document as a dense vector of length vocab_size where each entry counts how many times the corresponding token appears. The to_bow function in TextRepresentationPyTorch.ipynb initializes a zero tensor and increments indices based on the encoded tokens.
import torch
def to_bow(text, bow_vocab_size=vocab_size):
vec = torch.zeros(bow_vocab_size, dtype=torch.float32)
for idx in encode(text):
if idx < bow_vocab_size:
vec[idx] += 1
return vec
# Example usage
sample_text = train_dataset[0][1]
bow_vec = to_bow(sample_text)
print(bow_vec[:10]) # Shows counts for first 10 vocabulary tokens
To efficiently batch these vectors during training, the implementation defines a collate function called bowify that processes raw batches from the DataLoader:
from torch.utils.data import DataLoader
def bowify(batch):
# Convert 1-indexed labels to 0-indexed for PyTorch
labels = torch.LongTensor([lbl - 1 for lbl, _ in batch])
features = torch.stack([to_bow(txt) for _, txt in batch])
return labels, features
train_loader = DataLoader(
train_dataset,
batch_size=16,
collate_fn=bowify,
shuffle=True
)
Computing TF-IDF Vectors with Scikit-Learn
While Bag-of-Words counts raw frequencies, TF-IDF (Term Frequency-Inverse Document Frequency) statistically weights terms to emphasize discriminative words and downweight common terms. Since torchtext does not natively provide TF-IDF, the notebook integrates Scikit-Learn's TfidfVectorizer and converts the output to PyTorch tensors.
from sklearn.feature_extraction.text import TfidfVectorizer
# Initialize vectorizer with unigrams and bigrams
tfidf = TfidfVectorizer(ngram_range=(1, 2))
# Fit on a corpus of raw text strings
corpus = [
'I like hot dogs.',
'The dog ran fast.',
'Its hot outside.'
]
tfidf.fit(corpus)
# Transform new text and convert to PyTorch tensor
new_sent = ['My dog likes hot dogs on a hot day.']
tfidf_vec = tfidf.transform(new_sent).toarray()
tfidf_tensor = torch.tensor(tfidf_vec, dtype=torch.float32)
print(tfidf_tensor.shape) # torch.Size([1, vocab_dim])
The toarray() method converts the sparse matrix output to a dense numpy array, which torch.tensor() then transforms into a PyTorch-compatible format.
Training a Linear Text Classifier
Once text is represented as fixed-length vectors—whether via BoW or TF-IDF—you can feed them into a standard PyTorch classifier. The implementation uses a simple linear layer followed by LogSoftmax for multi-class classification across the four AG-NEWS categories.
# Define the network architecture
net = torch.nn.Sequential(
torch.nn.Linear(vocab_size, 4),
torch.nn.LogSoftmax(dim=1)
)
# Training loop configuration
optimizer = torch.optim.Adam(net.parameters(), lr=0.01)
loss_fn = torch.nn.NLLLoss()
def train_one_epoch(net, loader):
net.train()
for labels, feats in loader:
optimizer.zero_grad()
out = net(feats)
loss = loss_fn(out, labels)
loss.backward()
optimizer.step()
Because the Bag-of-Words and TF-IDF representations produce tensors of shape (batch_size, vocab_size), they interface directly with torch.nn.Linear layers without requiring embedding lookups or recurrent architectures.
Summary
- Bag-of-Words creates fixed-size count vectors by summing token occurrences, implemented manually using
torch.zerosand index incrementing. - TF-IDF provides statistically weighted features using
sklearn.feature_extraction.text.TfidfVectorizer, converted to PyTorch tensors via.toarray()andtorch.tensor(). - The bowify collate function enables efficient batching within
torch.utils.data.DataLoaderby stacking individual BoW vectors. - Both representations integrate with standard PyTorch linear classifiers (
torch.nn.Linear+LogSoftmax) for text classification tasks on the AG-NEWS dataset.
Frequently Asked Questions
What is the difference between Bag-of-Words and TF-IDF in text representation?
Bag-of-Words represents documents as raw frequency counts of vocabulary tokens, treating each word's occurrence equally. TF-IDF weights these frequencies by how unique a term is across the entire corpus, reducing the impact of common words like "the" or "and" while emphasizing rare, discriminative terms. According to the Microsoft AI-For-Beginners curriculum, TF-IDF typically improves classification performance on noisy corpora compared to raw BoW counts.
How do I convert Scikit-Learn TF-IDF vectors to PyTorch tensors?
First, ensure your TF-IDF vectorizer has been fitted on your training corpus using TfidfVectorizer.fit(). When transforming new text, call .toarray() on the resulting sparse matrix to convert it to a dense numpy array, then wrap it with torch.tensor(array, dtype=torch.float32) to create a PyTorch-compatible tensor suitable for model input.
Why does the bowify function subtract 1 from the labels?
The AG-NEWS dataset provided by torchtext uses 1-based indexing for labels (1 through 4), while PyTorch's loss functions like NLLLoss and CrossEntropyLoss expect 0-based class indices (0 through 3). The line lbl - 1 in the bowify collate function remaps the labels to the correct range for neural network training.
Can I use N-grams with the Bag-of-Words implementation shown?
The manual BoW implementation in TextRepresentationPyTorch.ipynb handles unigrams (single tokens) based on the vocabulary built from the tokenizer. To use N-grams (pairs or triplets of words), you would need to either modify the tokenizer to return N-gram tokens or use Scikit-Learn's CountVectorizer with the ngram_range parameter, then convert the output to PyTorch tensors following the same pattern as the TF-IDF example.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →