# Transformer and BERT Model Architectures in Microsoft AI For Beginners: Complete NLP Guide

> Explore Transformer and BERT model architectures in Microsoft AI For Beginners. Learn about multi-head self-attention, Masked Language Modeling, and Next Sentence Prediction in this comprehensive NLP guide.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-25

---

**The Microsoft AI For Beginners curriculum covers the original Transformer encoder-decoder stack with multi-head self-attention and BERT, a pure encoder variant pretrained with Masked Language Modeling and Next Sentence Prediction.**

The **Transformer** and **BERT** architectures form the backbone of modern natural language processing, and the `microsoft/AI-For-Beginners` repository provides comprehensive, hands-on lessons covering both theory and implementation. These notebooks in `lessons/5-NLP/18-Transformers/` walk you through building and fine-tuning these models using TensorFlow and PyTorch.

## Transformer Architecture Foundations

The original Transformer, introduced in "Attention Is All You Need," eliminates recurrent connections entirely in favor of **self-attention mechanisms**. According to the curriculum materials, this design "gave rise to the now-familiar Transformer models" that power systems like BERT, DistilBERT, BigBird, and OpenGPT-3.

### Core Encoder Components

Each Transformer encoder layer in `TransformersTF.ipynb` contains three sub-layers:

1. **Multi-head self-attention** — the input sequence attends to itself across multiple representation sub-spaces in parallel
2. **Position-wise feed-forward network** — a two-layer fully-connected MLP applied independently to each token position
3. **Add & Norm** — residual connections followed by layer normalization for stable deep-network training

### Decoder Extensions

The decoder mirrors the encoder structure but adds two critical attention variants:

- **Masked self-attention** — prevents tokens from attending to future positions during generation
- **Encoder-decoder attention** — allows the decoder to attend to the full encoder output

**Positional encodings** inject sequence order information, compensating for the otherwise permutation-invariant attention mechanism.

## BERT Architecture: Bidirectional Encoder Design

BERT (Bidirectional Encoder Representations from Transformers) is implemented in the curriculum as a **pure encoder** Transformer that processes text bidirectionally. Unlike the original Transformer, BERT lacks a decoder and is designed specifically for transfer learning.

### Model Variants and Specifications

The `TransformersTF.ipynb` notebook documents two standard configurations:

| Variant | Layers | Hidden Size | Attention Heads |
|---------|--------|-------------|-----------------|
| **BERT-base** | 12 | 768 | 12 |
| **BERT-large** | 24 | 1024 | 16 |

These specifications correspond to the `L-12_H-768_A-12` and `L-24_H-1024_A-16` model names you'll encounter in TensorFlow Hub and Hugging Face repositories.

### Pretraining Tasks

BERT learns deep contextual representations through two unsupervised objectives:

- **Masked Language Modeling (MLM)** — randomly masks input tokens and trains the model to reconstruct them, forcing bidirectional context understanding
- **Next Sentence Prediction (NSP)** — binary classification of whether two sentences are contiguous, teaching inter-sentence relationships

The curriculum notes BERT was pretrained on **Wikipedia and BooksCorpus** before fine-tuning for downstream tasks.

## Practical Implementation in TensorFlow

The file `lessons/5-NLP/18-Transformers/TransformersTF.ipynb` demonstrates loading and using BERT through TensorFlow Hub with frozen weights for demonstration:

```python
import tensorflow_hub as hub
import tensorflow_text as text  # Required for preprocessing ops

# Load BERT preprocessing and encoder modules from TF Hub

bert_preprocess = hub.KerasLayer(
    "https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3")
bert_encoder = hub.KerasLayer(
    "https://tfhub.dev/tensorflow/bert_en_uncased_L-12_H-768_A-12/3",
    trainable=False)  # Freeze encoder for quick demos

# Build classification model on top of [CLS] pooled output

inputs = tf.keras.Input(shape=(), dtype=tf.string, name="text")
preprocessed = bert_preprocess(inputs)
outputs = bert_encoder(preprocessed)["pooled_output"]   # Shape: [batch, 768]

logits = tf.keras.layers.Dense(num_classes, activation="softmax")(outputs)

model = tf.keras.Model(inputs, logits)
model.compile(optimizer="adam", loss="sparse_categorical_crossentropy",
              metrics=["accuracy"])

```

The `trainable=False` setting freezes the 110 million parameters of BERT-base, enabling fast feature extraction. For production fine-tuning, set `trainable=True` and use learning rates around **2e-5 to 5e-5**.

## PyTorch Implementation with Hugging Face

The PyTorch notebook at `lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb` uses the Hugging Face `transformers` library for a more flexible training loop:

```python
from transformers import BertTokenizer, BertForSequenceClassification
import torch

# Load pretrained tokenizer and model with classification head

tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
model = BertForSequenceClassification.from_pretrained(
    "bert-base-uncased", num_labels=num_classes)

def encode(texts):
    return tokenizer(texts, padding=True, truncation=True,
                    return_tensors="pt")

# Forward pass with labels for loss computation

texts = ["Example sentence 1", "Another example"]
inputs = encode(texts)
labels = torch.tensor([0, 1])

outputs = model(**inputs, labels=labels)
loss = outputs.loss
logits = outputs.logits
loss.backward()  # Backpropagate for fine-tuning

```

The `BertForSequenceClassification` class automatically attaches a classification head to the [CLS] token representation, outputting task-specific logits.

## BERT for Named Entity Recognition

The curriculum extends BERT applications in `lessons/5-NLP/19-NER/NER-TF.ipynb` and `NER-PyTorch.ipynb`, demonstrating why BERT's **contextualized token representations** outperform traditional sequence labeling approaches. Each token output from BERT's final encoder layer serves as input to a token-level classifier, enabling fine-grained entity extraction.

## Key Lesson Files

| File Path | Description |
|-----------|-------------|
| `lessons/5-NLP/18-Transformers/TransformersTF.ipynb` | TensorFlow Transformer encoder implementation and detailed BERT walkthrough |
| `lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb` | Hugging Face Transformers BERT fine-tuning for text classification |
| `lessons/5-NLP/19-NER/NER-TF.ipynb` | BERT application to named entity recognition with TensorFlow |
| `lessons/5-NLP/19-NER/NER-PyTorch.ipynb` | PyTorch NER implementation building on earlier BERT encoder lessons |

## Summary

- **Transformer architecture** uses stacked encoder-decoder layers with multi-head self-attention, feed-forward networks, and residual connections instead of recurrence
- **BERT-base** (12 layers, 768 hidden, 12 heads) and **BERT-large** (24 layers, 1024 hidden, 16 heads) are pure encoder variants pretrained with MLM and NSP objectives
- The Microsoft AI For Beginners curriculum provides runnable TensorFlow and PyTorch implementations in `lessons/5-NLP/18-Transformers/`
- **Fine-tuning** requires small learning rates (~2e-5) and typically freezes or minimally updates the pretrained encoder weights
- BERT's bidirectional context understanding makes it particularly effective for token-level tasks like named entity recognition covered in lesson 19

## Frequently Asked Questions

### What is the difference between the original Transformer and BERT?

The original Transformer is an **encoder-decoder architecture** designed for sequence-to-sequence tasks like machine translation. **BERT uses only the encoder stack** and is trained bidirectionally with masked language modeling, making it suitable for classification and token-level understanding tasks rather than generation.

### How does multi-head self-attention work in the Transformer?

Multi-head self-attention runs **multiple attention computations in parallel** with different learned projections, allowing the model to jointly attend to information from different representation subspaces. As implemented in the curriculum notebooks, each head learns distinct syntactic and semantic relationships before outputs are concatenated and projected.

### Why does BERT use [CLS] and [SEP] special tokens?

The **[CLS] token** at position zero serves as a sentence-level representation for classification tasks, while **[SEP] tokens** mark sentence boundaries and separate paired inputs for next-sentence prediction. These design choices enable BERT to handle both single-sentence and sentence-pair tasks with a unified architecture.

### When should I freeze versus fine-tune BERT weights?

**Freeze BERT** (`trainable=False`) for quick prototyping, limited compute, or when your downstream dataset is small—this treats BERT as a feature extractor. **Fine-tune end-to-end** when you have sufficient data and compute, as this adapts the pretrained representations to your specific domain and typically yields higher accuracy on the target task.