# Python Libraries Used for NLP in AI for Beginners: PyTorch and TorchText Deep Dive

> Discover PyTorch and TorchText, the Python libraries for NLP in AI for Beginners. Learn how they handle tokenization, vocabulary, and dataset loading for your AI projects.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-28

---

**The Microsoft AI for Beginners curriculum uses PyTorch and TorchText as its exclusive Python libraries for natural language processing, leveraging torch for neural network operations and torchtext for tokenization, vocabulary management, and dataset loading.**

The microsoft/AI-For-Beginners repository provides a comprehensive introduction to artificial intelligence concepts, with its natural language processing (NLP) modules built entirely around this streamlined two-library ecosystem. Unlike many educational resources that require multiple specialized NLP packages, this curriculum demonstrates end-to-end text processing and deep learning using only PyTorch's native text utilities, allowing learners to master NLP fundamentals without managing complex third-party dependencies.

## Core Python Libraries for NLP in AI for Beginners

### PyTorch (torch) — The Deep Learning Engine

The repository utilizes the **torch** library as its foundational framework for building and training neural networks on text data. As implemented in [`lessons/5-NLP/18-Transformers/torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/torchnlp.py), PyTorch handles all tensor operations, automatic differentiation, GPU acceleration, and optimization loops required for NLP model training across the curriculum's RNN, Transformer, and Generative Network lessons.

### TorchText — Text Processing Utilities

Complementing PyTorch, the **torchtext** library provides specialized preprocessing utilities that bridge raw text and numerical tensors. According to the source code in [`lessons/5-NLP/18-Transformers/torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/torchnlp.py), torchtext supplies essential tools including the `get_tokenizer` function, `ngrams_iterator` for n-gram generation, built-in datasets such as `AG_NEWS`, and vocabulary construction through `torchtext.vocab.vocab`.

## Implementation in the Source Code

The NLP lessons share a unified utility architecture across multiple modules. The file [`torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/torchnlp.py) appears consistently in `lessons/5-NLP/14-Embeddings/`, `lessons/5-NLP/16-RNN/`, `lessons/5-NLP/17-GenerativeNetworks/`, and `lessons/5-NLP/18-Transformers/`, providing standardized text processing functions.

### Tokenization and Dataset Loading

The curriculum loads the AG_NEWS dataset and constructs vocabularies using generator-based tokenization to minimize memory overhead:

```python
import torch
import torchtext
from collections import Counter

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = torchtext.data.utils.get_tokenizer('basic_english')

def load_dataset(ngrams=1, min_freq=1):
    train_dataset, test_dataset = torchtext.datasets.AG_NEWS(root='./data')
    train_dataset = list(train_dataset)
    test_dataset = list(test_dataset)

    # Build vocab from training data

    counter = Counter()
    for _, line in train_dataset:
        counter.update(
            torchtext.data.utils.ngrams_iterator(tokenizer(line), ngrams=ngrams)
        )
    vocab = torchtext.vocab.vocab(counter, min_freq=min_freq)
    return train_dataset, test_dataset, vocab

```

### Text Encoding Pipeline

Raw strings are converted to integer indices via vocabulary lookups, handling unknown tokens gracefully:

```python
def encode(text, vocab, unk=0):
    stoi = vocab.get_stoi() if hasattr(vocab, "get_stoi") else vocab.stoi
    return [stoi.get(tok, unk) for tok in tokenizer(text)]

# Example usage

_, _, vocab = load_dataset()
encoded = encode("AI for beginners is awesome!", vocab)
print(encoded)   # → [124, 58, 342, …]

```

### Training Loop Structure

The same PyTorch patterns are reused across lesson types, abstracting device placement and optimization:

```python
def train_epoch(net, dataloader, lr=0.01):
    optimizer = torch.optim.Adam(net.parameters(), lr=lr)
    loss_fn = torch.nn.CrossEntropyLoss().to(device)

    net.train()
    for labels, features in dataloader:
        optimizer.zero_grad()
        features, labels = features.to(device), labels.to(device)
        outputs = net(features)
        loss = loss_fn(outputs, labels)
        loss.backward()
        optimizer.step()

```

## Standard Library Alternative

While the primary NLP curriculum depends on torchtext and PyTorch, the repository includes [`examples/04-text-sentiment.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/examples/04-text-sentiment.py) demonstrating rudimentary sentiment analysis using only Python standard library components. This file implements manual text cleaning with regular expressions and frequency counting via `collections.Counter`, requiring no external NLP dependencies.

## Summary

- **PyTorch (torch)** provides all tensor operations, neural network layers, and automatic differentiation for deep learning NLP models in the curriculum.
- **TorchText** exclusively handles text tokenization, vocabulary generation via `torchtext.vocab.vocab`, and dataset retrieval through `torchtext.datasets.AG_NEWS`.
- The shared [`torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/torchnlp.py) utility file ensures consistent implementation across the Embeddings, RNN, Generative Networks, and Transformers lessons.
- Source code searches confirm that third-party NLP libraries such as NLTK, spaCy, and Hugging Face Transformers are not used in the NLP modules.
- CPU fallback support is built into all examples via `torch.device("cuda" if torch.cuda.is_available() else "cpu")`.

## Frequently Asked Questions

### Does AI for Beginners use NLTK or spaCy for text processing?

No, the curriculum does not import NLTK, spaCy, or other specialized NLP libraries. According to the source code analysis, all tokenization is performed using `torchtext.data.utils.get_tokenizer('basic_english')`, and vocabulary management is handled by `torchtext.vocab` classes.

### What dataset does the curriculum use for text classification examples?

The repository uses the **AG_NEWS** dataset loaded via `torchtext.datasets.AG_NEWS`. This built-in dataset is accessed in [`lessons/5-NLP/18-Transformers/torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/torchnlp.py) and reused across the RNN and Transformers modules for news categorization tasks.

### Is Hugging Face Transformers used in the Transformers lesson?

No, despite the lesson title referring to the Transformer architecture, the code does not utilize the Hugging Face **transformers** library. Instead, [`lessons/5-NLP/18-Transformers/torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/18-Transformers/torchnlp.py) implements transformer models using native PyTorch `nn.Module` classes, constructing attention mechanisms from scratch for educational clarity.

### Can the NLP examples run without a GPU?

Yes, all NLP examples include automatic CPU fallback. As implemented in the [`torchnlp.py`](https://github.com/microsoft/AI-For-Beginners/blob/main/torchnlp.py) utility files, the code defines `device = torch.device("cuda" if torch.cuda.is_available() else "cpu")`, allowing text processing and model training to execute on standard CPU hardware when CUDA is unavailable.