Transformer and BERT Model Architectures in Microsoft AI For Beginners: Complete NLP Guide

The Microsoft AI For Beginners curriculum covers the original Transformer encoder-decoder stack with multi-head self-attention and BERT, a pure encoder variant pretrained with Masked Language Modeling and Next Sentence Prediction.

The Transformer and BERT architectures form the backbone of modern natural language processing, and the microsoft/AI-For-Beginners repository provides comprehensive, hands-on lessons covering both theory and implementation. These notebooks in lessons/5-NLP/18-Transformers/ walk you through building and fine-tuning these models using TensorFlow and PyTorch.

Transformer Architecture Foundations

The original Transformer, introduced in "Attention Is All You Need," eliminates recurrent connections entirely in favor of self-attention mechanisms. According to the curriculum materials, this design "gave rise to the now-familiar Transformer models" that power systems like BERT, DistilBERT, BigBird, and OpenGPT-3.

Core Encoder Components

Each Transformer encoder layer in TransformersTF.ipynb contains three sub-layers:

  1. Multi-head self-attention — the input sequence attends to itself across multiple representation sub-spaces in parallel
  2. Position-wise feed-forward network — a two-layer fully-connected MLP applied independently to each token position
  3. Add & Norm — residual connections followed by layer normalization for stable deep-network training

Decoder Extensions

The decoder mirrors the encoder structure but adds two critical attention variants:

  • Masked self-attention — prevents tokens from attending to future positions during generation
  • Encoder-decoder attention — allows the decoder to attend to the full encoder output

Positional encodings inject sequence order information, compensating for the otherwise permutation-invariant attention mechanism.

BERT Architecture: Bidirectional Encoder Design

BERT (Bidirectional Encoder Representations from Transformers) is implemented in the curriculum as a pure encoder Transformer that processes text bidirectionally. Unlike the original Transformer, BERT lacks a decoder and is designed specifically for transfer learning.

Model Variants and Specifications

The TransformersTF.ipynb notebook documents two standard configurations:

Variant Layers Hidden Size Attention Heads
BERT-base 12 768 12
BERT-large 24 1024 16

These specifications correspond to the L-12_H-768_A-12 and L-24_H-1024_A-16 model names you'll encounter in TensorFlow Hub and Hugging Face repositories.

Pretraining Tasks

BERT learns deep contextual representations through two unsupervised objectives:

  • Masked Language Modeling (MLM) — randomly masks input tokens and trains the model to reconstruct them, forcing bidirectional context understanding
  • Next Sentence Prediction (NSP) — binary classification of whether two sentences are contiguous, teaching inter-sentence relationships

The curriculum notes BERT was pretrained on Wikipedia and BooksCorpus before fine-tuning for downstream tasks.

Practical Implementation in TensorFlow

The file lessons/5-NLP/18-Transformers/TransformersTF.ipynb demonstrates loading and using BERT through TensorFlow Hub with frozen weights for demonstration:

import tensorflow_hub as hub
import tensorflow_text as text  # Required for preprocessing ops

# Load BERT preprocessing and encoder modules from TF Hub

bert_preprocess = hub.KerasLayer(
    "https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3")
bert_encoder = hub.KerasLayer(
    "https://tfhub.dev/tensorflow/bert_en_uncased_L-12_H-768_A-12/3",
    trainable=False)  # Freeze encoder for quick demos

# Build classification model on top of [CLS] pooled output

inputs = tf.keras.Input(shape=(), dtype=tf.string, name="text")
preprocessed = bert_preprocess(inputs)
outputs = bert_encoder(preprocessed)["pooled_output"]   # Shape: [batch, 768]

logits = tf.keras.layers.Dense(num_classes, activation="softmax")(outputs)

model = tf.keras.Model(inputs, logits)
model.compile(optimizer="adam", loss="sparse_categorical_crossentropy",
              metrics=["accuracy"])

The trainable=False setting freezes the 110 million parameters of BERT-base, enabling fast feature extraction. For production fine-tuning, set trainable=True and use learning rates around 2e-5 to 5e-5.

PyTorch Implementation with Hugging Face

The PyTorch notebook at lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb uses the Hugging Face transformers library for a more flexible training loop:

from transformers import BertTokenizer, BertForSequenceClassification
import torch

# Load pretrained tokenizer and model with classification head

tokenizer = BertTokenizer.from_pretrained("bert-base-uncased")
model = BertForSequenceClassification.from_pretrained(
    "bert-base-uncased", num_labels=num_classes)

def encode(texts):
    return tokenizer(texts, padding=True, truncation=True,
                    return_tensors="pt")

# Forward pass with labels for loss computation

texts = ["Example sentence 1", "Another example"]
inputs = encode(texts)
labels = torch.tensor([0, 1])

outputs = model(**inputs, labels=labels)
loss = outputs.loss
logits = outputs.logits
loss.backward()  # Backpropagate for fine-tuning

The BertForSequenceClassification class automatically attaches a classification head to the [CLS] token representation, outputting task-specific logits.

BERT for Named Entity Recognition

The curriculum extends BERT applications in lessons/5-NLP/19-NER/NER-TF.ipynb and NER-PyTorch.ipynb, demonstrating why BERT's contextualized token representations outperform traditional sequence labeling approaches. Each token output from BERT's final encoder layer serves as input to a token-level classifier, enabling fine-grained entity extraction.

Key Lesson Files

File Path Description
lessons/5-NLP/18-Transformers/TransformersTF.ipynb TensorFlow Transformer encoder implementation and detailed BERT walkthrough
lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb Hugging Face Transformers BERT fine-tuning for text classification
lessons/5-NLP/19-NER/NER-TF.ipynb BERT application to named entity recognition with TensorFlow
lessons/5-NLP/19-NER/NER-PyTorch.ipynb PyTorch NER implementation building on earlier BERT encoder lessons

Summary

  • Transformer architecture uses stacked encoder-decoder layers with multi-head self-attention, feed-forward networks, and residual connections instead of recurrence
  • BERT-base (12 layers, 768 hidden, 12 heads) and BERT-large (24 layers, 1024 hidden, 16 heads) are pure encoder variants pretrained with MLM and NSP objectives
  • The Microsoft AI For Beginners curriculum provides runnable TensorFlow and PyTorch implementations in lessons/5-NLP/18-Transformers/
  • Fine-tuning requires small learning rates (~2e-5) and typically freezes or minimally updates the pretrained encoder weights
  • BERT's bidirectional context understanding makes it particularly effective for token-level tasks like named entity recognition covered in lesson 19

Frequently Asked Questions

What is the difference between the original Transformer and BERT?

The original Transformer is an encoder-decoder architecture designed for sequence-to-sequence tasks like machine translation. BERT uses only the encoder stack and is trained bidirectionally with masked language modeling, making it suitable for classification and token-level understanding tasks rather than generation.

How does multi-head self-attention work in the Transformer?

Multi-head self-attention runs multiple attention computations in parallel with different learned projections, allowing the model to jointly attend to information from different representation subspaces. As implemented in the curriculum notebooks, each head learns distinct syntactic and semantic relationships before outputs are concatenated and projected.

Why does BERT use [CLS] and [SEP] special tokens?

The [CLS] token at position zero serves as a sentence-level representation for classification tasks, while [SEP] tokens mark sentence boundaries and separate paired inputs for next-sentence prediction. These design choices enable BERT to handle both single-sentence and sentence-pair tasks with a unified architecture.

When should I freeze versus fine-tune BERT weights?

Freeze BERT (trainable=False) for quick prototyping, limited compute, or when your downstream dataset is small—this treats BERT as a feature extractor. Fine-tune end-to-end when you have sufficient data and compute, as this adapts the pretrained representations to your specific domain and typically yields higher accuracy on the target task.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →