# How to Represent Text Using Bag-of-Words and TF-IDF in PyTorch

> Learn to represent text with Bag-of-Words and TF-IDF in PyTorch using Microsoft's AI for Beginners. Convert text to numerical tensors for AI models.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: tutorial
- Published: 2026-08-29

---

**Microsoft's AI for Beginners curriculum demonstrates how to convert raw text into numerical tensors using Bag-of-Words and TF-IDF representations that integrate directly with PyTorch DataLoaders and linear classifiers.**

The Microsoft AI-For-Beginners repository provides a practical introduction to natural language processing fundamentals in `lessons/5-NLP/13-TextRep/TextRepresentationPyTorch.ipynb`. This guide walks through the complete pipeline for representing text using Bag-of-Words and TF-IDF in PyTorch, from tokenization with `torchtext` to training a neural classifier on the AG-NEWS dataset.

## Loading the AG-NEWS Dataset with torchtext

The implementation uses the **AG-NEWS** dataset provided by `torchtext.datasets`, which contains news headlines categorized into four classes: World, Sports, Business, and Sci/Tech. Each sample returns a tuple of `(label, text)` where labels are 1-indexed integers.

The dataset is loaded with a specified root directory for caching:

```python
import torchtext
import os

os.makedirs('./data', exist_ok=True)
train_dataset, test_dataset = torchtext.datasets.AG_NEWS(root='./data')
classes = ['World', 'Sports', 'Business', 'Sci/Tech']

```

## Building a Vocabulary and Tokenizer

Before creating vector representations, you must tokenize raw strings and map tokens to integer indices. The implementation uses the **basic_english** tokenizer from `torchtext.data.utils` and builds a vocabulary using Python's `collections.Counter`.

```python
import collections
from torchtext.data.utils import get_tokenizer

tokenizer = get_tokenizer('basic_english')
counter = collections.Counter()

# Count all tokens in the training set

for _, line in train_dataset:
    counter.update(tokenizer(line))

# Create vocabulary with minimum frequency of 1

vocab = torchtext.vocab.vocab(counter, min_freq=1)
vocab_size = len(vocab)
stoi = vocab.get_stoi()  # String-to-int mapping

def encode(text):
    """Convert text to list of integer indices."""
    return [stoi[t] for t in tokenizer(text)]

```

The `stoi` dictionary provides the mapping required to convert any token into its corresponding index within the fixed-size feature vector.

## Implementing Bag-of-Words (BoW) in PyTorch

**Bag-of-Words** represents each document as a dense vector of length `vocab_size` where each entry counts how many times the corresponding token appears. The `to_bow` function in `TextRepresentationPyTorch.ipynb` initializes a zero tensor and increments indices based on the encoded tokens.

```python
import torch

def to_bow(text, bow_vocab_size=vocab_size):
    vec = torch.zeros(bow_vocab_size, dtype=torch.float32)
    for idx in encode(text):
        if idx < bow_vocab_size:
            vec[idx] += 1
    return vec

# Example usage

sample_text = train_dataset[0][1]
bow_vec = to_bow(sample_text)
print(bow_vec[:10])  # Shows counts for first 10 vocabulary tokens

```

To efficiently batch these vectors during training, the implementation defines a **collate function** called `bowify` that processes raw batches from the DataLoader:

```python
from torch.utils.data import DataLoader

def bowify(batch):
    # Convert 1-indexed labels to 0-indexed for PyTorch

    labels = torch.LongTensor([lbl - 1 for lbl, _ in batch])
    features = torch.stack([to_bow(txt) for _, txt in batch])
    return labels, features

train_loader = DataLoader(
    train_dataset, 
    batch_size=16,
    collate_fn=bowify, 
    shuffle=True
)

```

## Computing TF-IDF Vectors with Scikit-Learn

While Bag-of-Words counts raw frequencies, **TF-IDF** (Term Frequency-Inverse Document Frequency) statistically weights terms to emphasize discriminative words and downweight common terms. Since `torchtext` does not natively provide TF-IDF, the notebook integrates **Scikit-Learn's** `TfidfVectorizer` and converts the output to PyTorch tensors.

```python
from sklearn.feature_extraction.text import TfidfVectorizer

# Initialize vectorizer with unigrams and bigrams

tfidf = TfidfVectorizer(ngram_range=(1, 2))

# Fit on a corpus of raw text strings

corpus = [
    'I like hot dogs.',
    'The dog ran fast.',
    'Its hot outside.'
]
tfidf.fit(corpus)

# Transform new text and convert to PyTorch tensor

new_sent = ['My dog likes hot dogs on a hot day.']
tfidf_vec = tfidf.transform(new_sent).toarray()
tfidf_tensor = torch.tensor(tfidf_vec, dtype=torch.float32)

print(tfidf_tensor.shape)  # torch.Size([1, vocab_dim])

```

The `toarray()` method converts the sparse matrix output to a dense numpy array, which `torch.tensor()` then transforms into a PyTorch-compatible format.

## Training a Linear Text Classifier

Once text is represented as fixed-length vectors—whether via BoW or TF-IDF—you can feed them into a standard PyTorch classifier. The implementation uses a simple linear layer followed by **LogSoftmax** for multi-class classification across the four AG-NEWS categories.

```python

# Define the network architecture

net = torch.nn.Sequential(
    torch.nn.Linear(vocab_size, 4),
    torch.nn.LogSoftmax(dim=1)
)

# Training loop configuration

optimizer = torch.optim.Adam(net.parameters(), lr=0.01)
loss_fn = torch.nn.NLLLoss()

def train_one_epoch(net, loader):
    net.train()
    for labels, feats in loader:
        optimizer.zero_grad()
        out = net(feats)
        loss = loss_fn(out, labels)
        loss.backward()
        optimizer.step()

```

Because the Bag-of-Words and TF-IDF representations produce tensors of shape `(batch_size, vocab_size)`, they interface directly with `torch.nn.Linear` layers without requiring embedding lookups or recurrent architectures.

## Summary

- **Bag-of-Words** creates fixed-size count vectors by summing token occurrences, implemented manually using `torch.zeros` and index incrementing.
- **TF-IDF** provides statistically weighted features using `sklearn.feature_extraction.text.TfidfVectorizer`, converted to PyTorch tensors via `.toarray()` and `torch.tensor()`.
- The **bowify** collate function enables efficient batching within `torch.utils.data.DataLoader` by stacking individual BoW vectors.
- Both representations integrate with standard PyTorch linear classifiers (`torch.nn.Linear` + `LogSoftmax`) for text classification tasks on the AG-NEWS dataset.

## Frequently Asked Questions

### What is the difference between Bag-of-Words and TF-IDF in text representation?

Bag-of-Words represents documents as raw frequency counts of vocabulary tokens, treating each word's occurrence equally. TF-IDF weights these frequencies by how unique a term is across the entire corpus, reducing the impact of common words like "the" or "and" while emphasizing rare, discriminative terms. According to the Microsoft AI-For-Beginners curriculum, TF-IDF typically improves classification performance on noisy corpora compared to raw BoW counts.

### How do I convert Scikit-Learn TF-IDF vectors to PyTorch tensors?

First, ensure your TF-IDF vectorizer has been fitted on your training corpus using `TfidfVectorizer.fit()`. When transforming new text, call `.toarray()` on the resulting sparse matrix to convert it to a dense numpy array, then wrap it with `torch.tensor(array, dtype=torch.float32)` to create a PyTorch-compatible tensor suitable for model input.

### Why does the bowify function subtract 1 from the labels?

The AG-NEWS dataset provided by `torchtext` uses 1-based indexing for labels (1 through 4), while PyTorch's loss functions like `NLLLoss` and `CrossEntropyLoss` expect 0-based class indices (0 through 3). The line `lbl - 1` in the `bowify` collate function remaps the labels to the correct range for neural network training.

### Can I use N-grams with the Bag-of-Words implementation shown?

The manual BoW implementation in `TextRepresentationPyTorch.ipynb` handles unigrams (single tokens) based on the vocabulary built from the tokenizer. To use N-grams (pairs or triplets of words), you would need to either modify the tokenizer to return N-gram tokens or use Scikit-Learn's `CountVectorizer` with the `ngram_range` parameter, then convert the output to PyTorch tensors following the same pattern as the TF-IDF example.