How Transformers and BERT Work: Understanding Attention Mechanisms in AI for Beginners

Transformers revolutionized natural‑language processing by replacing recurrent connections with self‑attention mechanisms that process every token in parallel, while BERT leverages bidirectional encoding and masked language modeling to capture deep contextual relationships from both directions.

The microsoft/AI-For-Beginners repository provides a comprehensive introduction to these architectures in lesson 18 of the NLP curriculum. This guide explains how Transformers and BERT work by examining the scaled dot‑product attention mechanism, the encoder stack implementation, and the practical fine‑tuning workflows documented in lessons/5-NLP/18-Transformers/.

The Core Transformer Architecture

The Transformer eliminates sequential processing by allowing every token in a sequence to attend to every other token simultaneously. This parallel computation relies on a specific attention mechanism implemented in the encoder layers.

Scaled Dot‑Product Attention

The attention mechanism computes a context vector for each token by weighing the relevance of all other tokens in the sequence. According to the source code analysis, this process follows four distinct steps:

  1. Query, Key, Value Projections – Each input token is projected into three vectors: Query (Q), Key (K), and Value (V).
  2. Similarity Scoring – The system calculates attention scores by computing the dot product between a query vector Qᵢ and all key vectors Kⱼ.
  3. Scaling and Softmax – Scores are divided by the square root of the vector dimension √d to stabilize gradients, then passed through a softmax function to create a probability distribution.
  4. Weighted Aggregation – The final context vector is the weighted sum of all value vectors V, using the softmax outputs as weights.

This scaled dot‑product attention ensures that token representations incorporate contextual information from the entire sequence rather than just preceding tokens.

Multi‑Head Attention

Rather than performing a single attention operation, the Transformer uses multi‑head attention to capture different types of relationships in parallel. Multiple attention heads run simultaneously, each learning distinct syntactic and semantic patterns. The outputs from all heads are concatenated and linearly transformed to produce the final representation.

The Transformer Encoder Stack

The complete encoder consists of a stack of identical layers, each containing two sub‑layers:

  • Multi‑head self‑attention mechanism
  • Position‑wise feed‑forward network (applied independently to each token)

Each sub‑layer employs residual connections followed by layer normalization, allowing gradients to flow through deep networks during training. The architecture diagram in lessons/5-NLP/18-Transformers/images/transformer-layer.png illustrates this structure.

How BERT Extends the Transformer

BERT (Bidirectional Encoder Representations from Transformers) builds exclusively on the Transformer encoder stack, modifying the training process to achieve bidirectional understanding.

Bidirectional Context Processing

Traditional language models process text left‑to‑right or right‑to‑left, limiting context awareness. BERT processes the entire sentence simultaneously, allowing each token to attend to both preceding and following tokens. This bidirectional approach creates richer contextual embeddings than unidirectional alternatives.

Pre‑Training Objectives

BERT learns language representations through two unsupervised tasks:

  • Masked Language Modeling (MLM) – The model randomly masks approximately 15% of input tokens and trains to predict the original vocabulary IDs. This forces the network to learn deep bidirectional semantics rather than simply predicting the next token.
  • Next Sentence Prediction (NSP) – The model receives pairs of sentences and predicts whether the second sentence logically follows the first. This objective trains BERT to understand inter‑sentence relationships and discourse structure.

Pre‑training occurs on massive corpora including Wikipedia and BookCorpus, creating general‑purpose language representations.

Fine‑Tuning for Downstream Tasks

After pre‑training, BERT adapts to specific tasks through fine‑tuning. This process involves adding a lightweight task‑specific classification head on top of the pooled encoder output and training on a small labeled dataset. The microsoft/AI-For-Beginners notebooks demonstrate this using the AG News dataset for text classification.

Critical fine‑tuning parameters include:

  • Using a low learning rate (typically 2e-5) to preserve pre‑trained weights
  • Applying the AdamW optimizer for stable convergence
  • Freezing or minimally updating the base BERT layers

Practical Implementation: PyTorch and TensorFlow

The repository provides complete implementations in both frameworks located at lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb and lessons/5-NLP/18-Transformers/TransformersTF.ipynb.

PyTorch BERT Fine‑Tuning

The PyTorch implementation uses the Hugging Face transformers library to load pre‑trained weights and attach a classification head:

from transformers import BertTokenizer, BertForSequenceClassification
import torch

model_name = "bert-base-uncased"
tokenizer = BertTokenizer.from_pretrained(model_name)
model = BertForSequenceClassification.from_pretrained(model_name, num_labels=4)

# Prepare input

sentence = "Transformers enable models to attend to all words simultaneously."
inputs = tokenizer(sentence, return_tensors="pt", truncation=True, padding=True)

# Inference

logits = model(**inputs).logits
predicted_class = torch.argmax(logits, dim=-1).item()
print("Predicted class index:", predicted_class)

This code loads bert-base-uncased, tokenizes input text, and extracts logits for four classification categories. The BertForSequenceClassification class automatically attaches a linear layer on top of the pooled output representations.

TensorFlow BERT Integration

The TensorFlow approach uses TensorFlow Hub to access the BERT encoder:

import tensorflow_hub as hub
import tensorflow as tf

bert_layer = hub.KerasLayer(
    "https://tfhub.dev/google/bert_uncased_L-12_H-768_A-12/3",
    trainable=False,
)

preprocess_layer = hub.KerasLayer(
    "https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3"
)

def build_model():
    input_text = tf.keras.layers.Input(shape=(), dtype=tf.string)
    encoder_inputs = preprocess_layer(input_text)
    outputs = bert_layer(encoder_inputs)
    pooled_output = outputs["pooled_output"]  # [batch, 768]

    
    logits = tf.keras.layers.Dense(4, activation="softmax")(pooled_output)
    return tf.keras.Model(inputs=input_text, outputs=logits)

model = build_model()
model.compile(optimizer=tf.keras.optimizers.Adam(learning_rate=2e-5),
              loss="sparse_categorical_crossentropy",
              metrics=["accuracy"])

This implementation loads the bert_uncased_L-12_H-768_A-12 model (12 layers, 768 hidden dimensions, 12 attention heads) from TensorFlow Hub, applies preprocessing operations, and attaches a softmax classification layer to the pooled_output representation.

Summary

  • Transformers replace recurrent neural networks with self‑attention mechanisms that process all tokens in parallel, capturing long‑range dependencies more efficiently than sequential architectures.
  • Scaled dot‑product attention computes relevance scores between Query and Key vectors, scales them by √d, and produces weighted Value vectors to form context‑aware representations.
  • BERT utilizes the Transformer encoder with bidirectional training objectives (Masked Language Modeling and Next Sentence Prediction) to understand context from both directions simultaneously.
  • Fine‑tuning BERT for specific tasks requires only a lightweight classification head, a low learning rate (2e-5), and minimal training data while retaining the pre‑trained knowledge.
  • The microsoft/AI-For-Beginners repository provides executable reference implementations in lessons/5-NLP/18-Transformers/TransformersPyTorch.ipynb and TransformersTF.ipynb.

Frequently Asked Questions

What is the difference between Transformers and BERT?

Transformers refer to the general architecture utilizing self‑attention mechanisms, consisting of encoder and decoder stacks. BERT specifically implements only the Transformer encoder stack with specialized pre‑training objectives (Masked Language Modeling and Next Sentence Prediction) designed for bidirectional context understanding. While Transformers serve as a broad architectural framework, BERT represents a specific pre‑trained model built upon that framework.

How does the scaled dot‑product attention mechanism work?

The mechanism projects input tokens into three matrix representations: Queries (Q), Keys (K), and Values (V). It calculates compatibility scores via matrix multiplication between Q and K, divides by the square root of the key dimension to prevent extreme gradient values, applies softmax to normalize scores into probabilities, and multiplies these weights by V to produce context vectors. This computation occurs in parallel across all positions, enabling efficient GPU utilization.

Why is BERT considered bidirectional while other models are not?

BERT employs Masked Language Modeling during pre‑training, which randomly hides 15% of tokens and requires the model to predict them based on surrounding context from both directions. Unlike autoregressive models that only attend to previous tokens (left‑to‑right), BERT's attention mechanism allows every token to directly connect to all other tokens in the sentence simultaneously, creating true bidirectional representations that incorporate both left and right context.

What learning rate should I use when fine‑tuning BERT?

The microsoft/AI-For-Beginners source code recommends using a low learning rate of 2e-5 when fine‑tuning BERT on downstream tasks. Higher learning rates can catastrophically overwrite the pre‑trained weights that encode general language understanding. The accompanying notebooks demonstrate this configuration using the AdamW optimizer, which decouples weight decay from gradient updates, providing more stable fine‑tuning on small datasets like AG News.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →