How the Transformer Architecture Works: From Self-Attention to Vision Models
The Transformer architecture replaces sequential recurrence with parallel self-attention, allowing every token to directly attend to every other token while processing sequences in parallel.
The Transformer architecture, introduced in "Attention Is All You Need," has become the foundation of modern deep learning. This guide explains the implementation details found in the rohitg00/ai-engineering-from-scratch repository, providing a code-first walkthrough of how embeddings, attention mechanisms, and feed-forward layers combine to process text and images.
Core Mechanisms of the Transformer Architecture
Token Embeddings and Positional Encoding
The architecture begins by converting discrete tokens into dense vectors. Because self-attention is permutation-invariant, the model requires positional information to understand sequence order. In the repository's mini-GPT implementation, the Embedding class handles this initial projection at lines 45–58 of phases/10-llms-from-scratch/04-pre-training-mini-gpt/code/main.py.
A fixed or learned positional encoding is added to these embeddings, ensuring the model distinguishes between identical tokens appearing at different positions. This combination creates the input representation that flows through subsequent layers.
Multi-Head Self-Attention (MHSA)
The defining innovation of the Transformer architecture is multi-head self-attention, which computes scaled dot-product attention across multiple representation subspaces simultaneously. Each head performs the operation:
Attention(Q, K, V) = softmax(QK^T / sqrt(d_k))V
The MultiHeadAttention class in phases/19-capstone-projects/34-transformer-block/code/main.py (lines 47–84) implements this mechanism. By splitting the embedding dimension across multiple heads, the model can attend to different contextual relationships in parallel—capturing both local syntactic patterns and long-range semantic dependencies without the sequential bottleneck of recurrent networks.
Residual Connections and Layer Normalization
Training stability in deep Transformer models relies on residual connections (skip connections) surrounding each sub-layer. The architecture follows a "pre-norm" or "post-norm" pattern where the input is added to the sub-layer output, preserving gradients during backpropagation.
In phases/19-capstone-projects/34-transformer-block/code/main.py (lines 85–101), the implementation applies LayerNorm before the attention mechanism, followed by the residual addition. This pattern repeats after the feed-forward network, ensuring information flows directly from early layers to deep layers, mitigating the vanishing gradient problem.
Position-Wise Feed-Forward Networks
Following attention, each token passes through a position-wise feed-forward network (FFN)—a two-layer MLP applied independently to each position. This adds non-linearity and depth, allowing the model to transform attention outputs into complex feature representations.
The FeedForward class at lines 103–113 of phases/19-capstone-projects/34-transformer-block/code/main.py typically implements this as:
class FeedForward(nn.Module):
def __init__(self, embed_dim, ff_dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(embed_dim, ff_dim),
nn.GELU(), # or SwiGLU in modern variants
nn.Linear(ff_dim, embed_dim)
)
def forward(self, x):
return self.net(x)
Assembling Deep Models: Stacking Transformer Blocks
A complete Transformer architecture stacks multiple identical blocks to increase model capacity. Each TransformerBlock encapsulates the attention, normalization, and feed-forward operations described above.
The stacking logic appears at lines 133–138 of phases/19-capstone-projects/34-transformer-block/code/main.py, where the model constructs a sequence of blocks:
class TransformerBlock(nn.Module):
def __init__(self, embed_dim, num_heads, ff_dim):
super().__init__()
self.attn = MultiHeadAttention(embed_dim, num_heads)
self.ln1 = LayerNorm(embed_dim)
self.ff = FeedForward(embed_dim, ff_dim)
self.ln2 = LayerNorm(embed_dim)
def forward(self, x):
# Self-attention + residual
x = x + self.attn(self.ln1(x), self.ln1(x), self.ln1(x))
# Feed-forward + residual
x = x + self.ff(self.ln2(x))
return x
Deeper stacks enable the model to learn hierarchical representations, with lower blocks capturing syntactic features and higher blocks modeling abstract semantics.
Architectural Variants: Encoders, Decoders, and Vision Transformers
Encoder vs. Decoder Stacks
The Transformer architecture supports two distinct configurations. The encoder uses bidirectional self-attention, allowing every token to attend to all other tokens in the sequence—ideal for understanding tasks like classification or embedding generation.
The decoder modifies this with causal masking (implemented via an attention mask) to prevent positions from attending to subsequent tokens during training, preserving autoregressive generation properties. Decoders also include cross-attention layers that attend to encoder outputs, enabling sequence-to-sequence translation. The mini-GPT lesson in the repository implements a decoder-style stack for language modeling.
Vision Transformers (ViT)
The Transformer architecture extends beyond NLP to computer vision through Vision Transformers (ViT). Instead of processing text tokens, ViT splits images into fixed-size patches, flattens them, and projects them into embedding vectors treated as sequence tokens.
In phases/07-transformers-deep-dive/09-vision-transformers/code/main.py, the implementation shows how patches are extracted and processed:
# Split image into patches, flatten, project to embedding dim
patches = img.unfold(2, patch_size, patch_size).unfold(3, patch_size, patch_size)
patches = patches.reshape(batch, -1, patch_size*patch_size*channels)
embeds = self.patch_proj(patches) + self.pos_enc[:num_patches]
These patch embeddings then flow through standard TransformerBlock layers, demonstrating the architecture's versatility across modalities. A standalone ViT implementation also exists at phases/19-capstone-projects/59-vit-transformer/code/main.py.
Summary
- Self-attention mechanism: Replaces recurrence with parallel token-to-token interactions, implemented in
MultiHeadAttentionclasses. - Residual connections and LayerNorm: Stabilize training in deep stacks, visible in the
TransformerBlockforward pass at lines 85–101. - Position-wise FFN: Adds non-linear transformation capacity after attention operations.
- Modular stacking: Deep models are built by repeating
TransformerBlockunits, as shown in the layer construction at lines 133–138. - Cross-modal applicability: The same architectural components power both language models (in
phases/10-llms-from-scratch/) and vision models (inphases/07-transformers-deep-dive/09-vision-transformers/).
Frequently Asked Questions
What is the main advantage of the Transformer architecture over recurrent networks?
The Transformer architecture eliminates sequential processing constraints by computing attention in parallel across the entire sequence. While RNNs process tokens one at a time, creating O(n) sequential operations, Transformers process all positions simultaneously, achieving O(1) sequential steps and enabling GPU acceleration across long contexts.
How does multi-head self-attention work in the Transformer architecture?
Multi-head self-attention splits the embedding dimension into multiple heads (typically 8 or 16), each learning different attention patterns. As implemented in phases/19-capstone-projects/34-transformer-block/code/main.py (lines 47–84), each head computes scaled dot-product attention independently, and the results are concatenated and linearly projected, allowing the model to jointly attend to information from different representation subspaces.
What is the difference between the encoder and decoder in the Transformer architecture?
The encoder uses bidirectional self-attention to build contextualized representations of input sequences, suitable for understanding tasks. The decoder employs causal masking to ensure predictions depend only on previous tokens, making it suitable for autoregressive generation. Decoders also include cross-attention layers that query encoder outputs, enabling tasks like machine translation where the model must attend to both the generated sequence and the source input.
Can the Transformer architecture process images, or is it limited to text?
Yes, the Transformer architecture processes images through Vision Transformers (ViT). As shown in phases/07-transformers-deep-dive/09-vision-transformers/code/main.py, images are divided into patches, linearly embedded, and treated as token sequences. These patch embeddings then pass through standard Transformer blocks, proving the architecture's general applicability to any data representable as a sequence of vectors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →