# How Transformers and BERT Work: Understanding Attention Mechanisms in AI for Beginners

> Understand how Transformers and BERT work using attention mechanisms. Explore AI for Beginners and unlock the power of parallel processing and bidirectional encoding for deep language understanding.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-29

---

**Transformers revolutionized natural‑language processing by replacing recurrent connections with self‑attention mechanisms that process every token in parallel, while BERT leverages bidirectional encoding and masked language modeling to capture deep contextual relationships from both directions.**

The `microsoft/AI-For-Beginners` repository provides a comprehensive introduction to these architectures in lesson 18 of the NLP curriculum. This guide explains how Transformers and BERT work by examining the scaled dot‑product attention mechanism, the encoder stack implementation, and the practical fine‑tuning workflows documented in `lessons/5-NLP/18-Transformers/`.

## The Core Transformer Architecture

The Transformer eliminates sequential processing by allowing every token in a sequence to attend to every other token simultaneously. This parallel computation relies on a specific attention mechanism implemented in the encoder layers.

### Scaled Dot‑Product Attention

The attention mechanism computes a **context vector** for each token by weighing the relevance of all other tokens in the sequence. According to the source code analysis, this process follows four distinct steps:

1. **Query, Key, Value Projections** – Each input token is projected into three vectors: **Query (Q)**, **Key (K)**, and **Value (V)**.
2. **Similarity Scoring** – The system calculates attention scores by computing the dot product between a query vector `Qᵢ` and all key vectors `Kⱼ`.
3. **Scaling and Softmax** – Scores are divided by the square root of the vector dimension `√d` to stabilize gradients, then passed through a softmax function to create a probability distribution.
4. **Weighted Aggregation** – The final context vector is the weighted sum of all value vectors `V`, using the softmax outputs as weights.

This **scaled dot‑product attention** ensures that token representations incorporate contextual information from the entire sequence rather than just preceding tokens.

### Multi‑Head Attention

Rather than performing a single attention operation, the Transformer uses **multi‑head attention** to capture different types of relationships in parallel. Multiple attention heads run simultaneously, each learning distinct syntactic and semantic patterns. The outputs from all heads are concatenated and linearly transformed to produce the final representation.

### The Transformer Encoder Stack

The complete encoder consists of a stack of identical layers, each containing two sub‑layers:
- **Multi‑head self‑attention** mechanism
- **Position‑wise feed‑forward network** (applied independently to each token)

Each sub‑layer employs **residual connections** followed by layer normalization, allowing gradients to flow through deep networks during training. The architecture diagram in `lessons/5-NLP/18-Transformers/images/transformer-layer.png` illustrates this structure.

## How BERT Extends the Transformer

**BERT (Bidirectional Encoder Representations from Transformers)** builds exclusively on the Transformer encoder stack, modifying the training process to achieve bidirectional understanding.

### Bidirectional Context Processing

Traditional language models process text left‑to‑right or right‑to‑left, limiting context awareness. BERT processes the **entire sentence simultaneously**, allowing each token to attend to both preceding and following tokens. This bidirectional approach creates richer contextual embeddings than unidirectional alternatives.

### Pre‑Training Objectives

BERT learns language representations through two unsupervised tasks:

- **Masked Language Modeling (MLM)** – The model randomly masks approximately **15% of input tokens** and trains to predict the original vocabulary IDs. This forces the network to learn deep bidirectional semantics rather than simply predicting the next token.
- **Next Sentence Prediction (NSP)** – The model receives pairs of sentences and predicts whether the second sentence logically follows the first. This objective trains BERT to understand inter‑sentence relationships and discourse structure.

Pre‑training occurs on massive corpora including Wikipedia and BookCorpus, creating general‑purpose language representations.

### Fine‑Tuning for Downstream Tasks

After pre‑training, BERT adapts to specific tasks through **fine‑tuning**. This process involves adding a lightweight task‑specific classification head on top of the pooled encoder output and training on a small labeled dataset. The `microsoft/AI-For-Beginners` notebooks demonstrate this using the AG News dataset for text classification.

Critical fine‑tuning parameters include:
- Using a **low learning rate** (typically `2e-5`) to preserve pre‑trained weights
- Applying the **AdamW optimizer** for stable convergence
- Freezing or minimally updating the base BERT layers

## Practical Implementation: PyTorch and TensorFlow

The repository provides complete implementations in both frameworks located at `lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb` and `lessons/5-NLP/18-Transformers/TransformersTF.ipynb`.

### PyTorch BERT Fine‑Tuning

The PyTorch implementation uses the Hugging Face `transformers` library to load pre‑trained weights and attach a classification head:

```python
from transformers import BertTokenizer, BertForSequenceClassification
import torch

model_name = "bert-base-uncased"
tokenizer = BertTokenizer.from_pretrained(model_name)
model = BertForSequenceClassification.from_pretrained(model_name, num_labels=4)

# Prepare input

sentence = "Transformers enable models to attend to all words simultaneously."
inputs = tokenizer(sentence, return_tensors="pt", truncation=True, padding=True)

# Inference

logits = model(**inputs).logits
predicted_class = torch.argmax(logits, dim=-1).item()
print("Predicted class index:", predicted_class)

```

This code loads `bert-base-uncased`, tokenizes input text, and extracts logits for four classification categories. The `BertForSequenceClassification` class automatically attaches a linear layer on top of the pooled output representations.

### TensorFlow BERT Integration

The TensorFlow approach uses TensorFlow Hub to access the BERT encoder:

```python
import tensorflow_hub as hub
import tensorflow as tf

bert_layer = hub.KerasLayer(
    "https://tfhub.dev/google/bert_uncased_L-12_H-768_A-12/3",
    trainable=False,
)

preprocess_layer = hub.KerasLayer(
    "https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3"
)

def build_model():
    input_text = tf.keras.layers.Input(shape=(), dtype=tf.string)
    encoder_inputs = preprocess_layer(input_text)
    outputs = bert_layer(encoder_inputs)
    pooled_output = outputs["pooled_output"]  # [batch, 768]

    
    logits = tf.keras.layers.Dense(4, activation="softmax")(pooled_output)
    return tf.keras.Model(inputs=input_text, outputs=logits)

model = build_model()
model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=2e-5),
              loss="sparse_categorical_crossentropy",
              metrics=["accuracy"])

```

This implementation loads the `bert_uncased_L-12_H-768_A-12` model (12 layers, 768 hidden dimensions, 12 attention heads) from TensorFlow Hub, applies preprocessing operations, and attaches a softmax classification layer to the `pooled_output` representation.

## Summary

- **Transformers** replace recurrent neural networks with **self‑attention mechanisms** that process all tokens in parallel, capturing long‑range dependencies more efficiently than sequential architectures.
- **Scaled dot‑product attention** computes relevance scores between Query and Key vectors, scales them by `√d`, and produces weighted Value vectors to form context‑aware representations.
- **BERT** utilizes the Transformer encoder with **bidirectional training objectives** (Masked Language Modeling and Next Sentence Prediction) to understand context from both directions simultaneously.
- **Fine‑tuning** BERT for specific tasks requires only a lightweight classification head, a low learning rate (`2e-5`), and minimal training data while retaining the pre‑trained knowledge.
- The `microsoft/AI-For-Beginners` repository provides executable reference implementations in `lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb` and `TransformersTF.ipynb`.

## Frequently Asked Questions

### What is the difference between Transformers and BERT?

**Transformers** refer to the general architecture utilizing self‑attention mechanisms, consisting of encoder and decoder stacks. **BERT** specifically implements only the Transformer encoder stack with specialized pre‑training objectives (Masked Language Modeling and Next Sentence Prediction) designed for bidirectional context understanding. While Transformers serve as a broad architectural framework, BERT represents a specific pre‑trained model built upon that framework.

### How does the scaled dot‑product attention mechanism work?

The mechanism projects input tokens into three matrix representations: **Queries (Q)**, **Keys (K)**, and **Values (V)**. It calculates compatibility scores via matrix multiplication between Q and K, divides by the square root of the key dimension to prevent extreme gradient values, applies softmax to normalize scores into probabilities, and multiplies these weights by V to produce context vectors. This computation occurs in parallel across all positions, enabling efficient GPU utilization.

### Why is BERT considered bidirectional while other models are not?

BERT employs **Masked Language Modeling** during pre‑training, which randomly hides 15% of tokens and requires the model to predict them based on surrounding context from both directions. Unlike autoregressive models that only attend to previous tokens (left‑to‑right), BERT's attention mechanism allows every token to directly connect to all other tokens in the sentence simultaneously, creating true bidirectional representations that incorporate both left and right context.

### What learning rate should I use when fine‑tuning BERT?

The `microsoft/AI-For-Beginners` source code recommends using a **low learning rate of 2e-5** when fine‑tuning BERT on downstream tasks. Higher learning rates can catastrophically overwrite the pre‑trained weights that encode general language understanding. The accompanying notebooks demonstrate this configuration using the AdamW optimizer, which decouples weight decay from gradient updates, providing more stable fine‑tuning on small datasets like AG News.