Building a Transformer Model from Scratch in Python: A Complete Implementation Guide
You can build a complete GPT-style transformer using only PyTorch by implementing scaled dot-product attention, multi-head attention, and stacking transformer blocks with pre-normalization and residual connections.
The ai-engineering-from-scratch repository provides a comprehensive curriculum for constructing transformer architectures without relying on high-level libraries like Hugging Face. Building a transformer model from scratch in Python reveals the mathematical foundations of modern large language models while giving you full control over every tensor operation. This guide walks through the exact implementation path found in the repository's deep-dive phase, from basic self-attention to a fully functional decoder-only model.
Core Architecture of a Decoder-Only Transformer
A minimal transformer built following the curriculum implements the classic GPT-style data flow. Input token indices first pass through an embedding layer, combine with positional encodings, then traverse a series of identical transformer blocks before projecting back to vocabulary logits.
As implemented in phases/07-transformers-deep-dive/05-full-transformer/docs/en.md, the architecture follows this pipeline:
input_ids → token embedding → add positional encoding →
N × [LayerNorm → Multi-Head Attention → Residual → LayerNorm → FFN → Residual] →
final LayerNorm → language-model head → logits
Each component uses pre-normalization (LayerNorm before sub-layers), which stabilizes training compared to the original post-norm design.
Implementing Scaled Dot-Product Attention
The foundation of any transformer is the attention mechanism. In phases/07-transformers-deep-dive/02-self-attention-from-scratch/code/self_attention.py, the curriculum derives and implements the scaled dot-product attention formula from first principles.
import torch
import torch.nn.functional as F
def scaled_dot_product_attention(q, k, v, mask=None):
"""q, k, v: (B, T, D) tensors."""
d_k = q.size(-1)
scores = torch.matmul(q, k.transpose(-2, -1)) / d_k**0.5
if mask is not None:
scores = scores.masked_fill(mask == 0, float("-inf"))
attn = F.softmax(scores, dim=-1)
return torch.matmul(attn, v)
This function computes queries, keys, and values, scales the dot-product by $\sqrt{d_k}$ to prevent softmax saturation, and applies an optional causal mask to prevent attending to future positions during training.
Multi-Head Attention with Causal Masking
Single-head attention is wrapped into a multi-head module as detailed in phases/07-transformers-deep-dive/03-multi-head-attention/docs/en.md. The implementation splits the model dimension across multiple heads, computes attention in parallel, and enforces causality through a lower-triangular mask.
class MultiHeadAttention(torch.nn.Module):
def __init__(self, d_model, n_heads):
super().__init__()
assert d_model % n_heads == 0
self.n_heads = n_heads
self.d_head = d_model // n_heads
self.qkv_proj = torch.nn.Linear(d_model, 3 * d_model)
self.out_proj = torch.nn.Linear(d_model, d_model)
def forward(self, x, mask=None):
B, T, D = x.shape
qkv = self.qkv_proj(x) # (B, T, 3*D)
q, k, v = qkv.chunk(3, dim=-1) # three (B, T, D)
# reshape for heads: (B, n_heads, T, d_head)
q = q.view(B, T, self.n_heads, self.d_head).transpose(1, 2)
k = k.view(B, T, self.n_heads, self.d_head).transpose(1, 2)
v = v.view(B, T, self.n_heads, self.d_head).transpose(1, 2)
# causal mask
if mask is None:
mask = torch.tril(torch.ones(T, T, dtype=torch.bool, device=x.device))
mask = mask.unsqueeze(0).unsqueeze(1) # (1, 1, T, T)
attn = scaled_dot_product_attention(q, k, v, mask)
attn = attn.transpose(1, 2).contiguous().view(B, T, D)
return self.out_proj(attn)
The causal mask ensures that when predicting token $i$, the model can only attend to positions $0$ through $i$, maintaining the autoregressive property required for language generation.
Positional Encoding
Since self-attention is permutation-invariant, the model requires positional information to distinguish token order. The curriculum implements sinusoidal positional encodings as described in phases/07-transformers-deep-dive/04-positional-encoding/docs/en.md.
Unlike learned positional embeddings, sinusoidal encodings use deterministic sine and cosine functions of different frequencies, allowing the model to generalize to sequence lengths longer than those seen during training. These encodings are added directly to the token embeddings before entering the transformer blocks.
Composing the Full Transformer Block
A complete transformer block combines attention with a feed-forward network (FFN) and normalization layers. According to the architecture in phases/07-transformers-deep-dive/05-full-transformer/docs/en.md, each block contains:
- LayerNorm applied before the attention sub-layer
- Multi-Head Self-Attention with residual connection
- LayerNorm applied before the FFN sub-layer
- Feed-Forward Network: A two-layer MLP with hidden dimension $4 \times d_{model}$ and GELU activation, followed by a residual connection
This pre-norm configuration prevents gradient vanishing in deep networks and has become the standard for modern LLMs.
Assembling the Complete GPT-Style Model
The final model, located in phases/07-transformers-deep-dive/05-full-transformer/code/main.py, stacks multiple blocks and adds input embeddings plus an output language-modeling head.
class TinyGPT(torch.nn.Module):
def __init__(self, vocab_size=64, d_model=32, n_heads=4, n_layers=2, max_len=16):
super().__init__()
self.token_emb = torch.nn.Embedding(vocab_size, d_model)
self.pos_emb = torch.nn.Embedding(max_len, d_model)
self.blocks = torch.nn.ModuleList([
torch.nn.ModuleDict({
"ln1": torch.nn.LayerNorm(d_model),
"attn": MultiHeadAttention(d_model, n_heads),
"ln2": torch.nn.LayerNorm(d_model),
"ffn": torch.nn.Sequential(
torch.nn.Linear(d_model, 4 * d_model),
torch.nn.GELU(),
torch.nn.Linear(4 * d_model, d_model),
),
})
for _ in range(n_layers)
])
self.ln_final = torch.nn.LayerNorm(d_model)
self.head = torch.nn.Linear(d_model, vocab_size, bias=False)
def forward(self, idx):
B, T = idx.shape
x = self.token_emb(idx) + self.pos_emb(torch.arange(T, device=idx.device))
for block in self.blocks:
x = x + block["attn"](block["ln1"](x))
x = x + block["ffn"](block["ln2"](x))
x = self.ln_final(x)
return self.head(x)
# Sanity check
model = TinyGPT()
logits = model(torch.randint(0, 64, (4, 16)))
print(logits.shape) # torch.Size([4, 16, 64])
This implementation creates a minimal but complete transformer with only PyTorch dependencies, capable of next-token prediction on small vocabularies.
Training Loop and Production Extensions
The curriculum concludes with a minimal training loop implementing masked language modeling on synthetic data, demonstrating the full data-through-model pipeline. Beyond the basic architecture, the repository covers production-grade optimizations:
- KV-Cache and Flash Attention (
phases/07-transformers-deep-dive/12-kv-cache-flash-attention/docs/en.md) for efficient inference - Scaling Laws (
phases/07-transformers-deep-dive/13-scaling-laws/docs/en.md) governing model size and compute-optimal training - Mixture-of-Experts (
phases/07-transformers-deep-dive/11-mixture-of-experts/docs/en.md) for conditional computation
Summary
- Scaled dot-product attention forms the mathematical core of transformers, computing weighted relationships between all token pairs simultaneously.
- Multi-head attention parallelizes multiple attention computations across different representation subspaces, with causal masking ensuring autoregressive generation.
- Pre-normalization (LayerNorm before sub-layers) and residual connections stabilize training in deep transformer stacks.
- Sinusoidal positional encodings inject sequence order information without adding learnable parameters.
- The complete implementation in
phases/07-transformers-deep-dive/05-full-transformer/code/main.pyrequires only PyTorch and produces a functional GPT-style model suitable for educational experimentation.
Frequently Asked Questions
Do I need the Hugging Face Transformers library to build this model?
No. The ai-engineering-from-scratch curriculum explicitly prohibits external transformer libraries; you only need PyTorch and the Python standard library. This constraint ensures you understand every matrix multiplication and tensor reshape rather than treating the model as a black box.
What is the difference between pre-norm and post-norm transformer architectures?
Pre-norm applies Layer normalization before the attention and FFN sub-layers, while post-norm applies it after. The repository uses pre-norm because it prevents gradient vanishing in deep networks and eliminates the need for careful learning rate warm-up schedules that post-norm architectures require.
Why does the implementation use sinusoidal positional encodings instead of learned embeddings?
Sinusoidal encodings allow the model to generalize to sequence lengths longer than those seen during training because the encoding functions are deterministic and continuous. Learned positional embeddings are limited to the maximum sequence length observed during training, though they may offer slightly more flexibility within that range.
How does the causal mask prevent the model from looking at future tokens?
The causal mask is a lower-triangular boolean matrix that multiplies the attention scores. In scaled_dot_product_attention, positions corresponding to future tokens are filled with negative infinity before the softmax, forcing those attention weights to zero. This ensures that prediction of token $i$ depends only on tokens $0$ through $i-1$, maintaining the autoregressive property essential for language modeling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →