# Building a Transformer Model from Scratch in Python: A Complete Implementation Guide

> Build a GPT-style transformer model from scratch using PyTorch. Implement scaled dot-product attention, multi-head attention, and transformer blocks for a complete guide.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: how-to-guide
- Published: 2026-07-26

---

**You can build a complete GPT-style transformer using only PyTorch by implementing scaled dot-product attention, multi-head attention, and stacking transformer blocks with pre-normalization and residual connections.**

The *ai-engineering-from-scratch* repository provides a comprehensive curriculum for constructing transformer architectures without relying on high-level libraries like Hugging Face. Building a transformer model from scratch in Python reveals the mathematical foundations of modern large language models while giving you full control over every tensor operation. This guide walks through the exact implementation path found in the repository's deep-dive phase, from basic self-attention to a fully functional decoder-only model.

## Core Architecture of a Decoder-Only Transformer

A minimal transformer built following the curriculum implements the classic GPT-style data flow. Input token indices first pass through an embedding layer, combine with positional encodings, then traverse a series of identical transformer blocks before projecting back to vocabulary logits.

As implemented in [`phases/07-transformers-deep-dive/05-full-transformer/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/05-full-transformer/docs/en.md), the architecture follows this pipeline:

```

input_ids → token embedding → add positional encoding → 
N × [LayerNorm → Multi-Head Attention → Residual → LayerNorm → FFN → Residual] → 
final LayerNorm → language-model head → logits

```

Each component uses **pre-normalization** (LayerNorm before sub-layers), which stabilizes training compared to the original post-norm design.

## Implementing Scaled Dot-Product Attention

The foundation of any transformer is the attention mechanism. In [`phases/07-transformers-deep-dive/02-self-attention-from-scratch/code/self_attention.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/02-self-attention-from-scratch/code/self_attention.py), the curriculum derives and implements the scaled dot-product attention formula from first principles.

```python
import torch
import torch.nn.functional as F

def scaled_dot_product_attention(q, k, v, mask=None):
    """q, k, v: (B, T, D) tensors."""
    d_k = q.size(-1)
    scores = torch.matmul(q, k.transpose(-2, -1)) / d_k**0.5
    if mask is not None:
        scores = scores.masked_fill(mask == 0, float("-inf"))
    attn = F.softmax(scores, dim=-1)
    return torch.matmul(attn, v)

```

This function computes **queries**, **keys**, and **values**, scales the dot-product by $\sqrt{d_k}$ to prevent softmax saturation, and applies an optional causal mask to prevent attending to future positions during training.

## Multi-Head Attention with Causal Masking

Single-head attention is wrapped into a multi-head module as detailed in [`phases/07-transformers-deep-dive/03-multi-head-attention/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/03-multi-head-attention/docs/en.md). The implementation splits the model dimension across multiple heads, computes attention in parallel, and enforces causality through a lower-triangular mask.

```python
class MultiHeadAttention(torch.nn.Module):
    def __init__(self, d_model, n_heads):
        super().__init__()
        assert d_model % n_heads == 0
        self.n_heads = n_heads
        self.d_head = d_model // n_heads
        self.qkv_proj = torch.nn.Linear(d_model, 3 * d_model)
        self.out_proj = torch.nn.Linear(d_model, d_model)

    def forward(self, x, mask=None):
        B, T, D = x.shape
        qkv = self.qkv_proj(x)                     # (B, T, 3*D)

        q, k, v = qkv.chunk(3, dim=-1)             # three (B, T, D)

        # reshape for heads: (B, n_heads, T, d_head)

        q = q.view(B, T, self.n_heads, self.d_head).transpose(1, 2)
        k = k.view(B, T, self.n_heads, self.d_head).transpose(1, 2)
        v = v.view(B, T, self.n_heads, self.d_head).transpose(1, 2)

        # causal mask

        if mask is None:
            mask = torch.tril(torch.ones(T, T, dtype=torch.bool, device=x.device))
        mask = mask.unsqueeze(0).unsqueeze(1)  # (1, 1, T, T)

        attn = scaled_dot_product_attention(q, k, v, mask)
        attn = attn.transpose(1, 2).contiguous().view(B, T, D)
        return self.out_proj(attn)

```

The **causal mask** ensures that when predicting token $i$, the model can only attend to positions $0$ through $i$, maintaining the autoregressive property required for language generation.

## Positional Encoding

Since self-attention is permutation-invariant, the model requires positional information to distinguish token order. The curriculum implements **sinusoidal positional encodings** as described in [`phases/07-transformers-deep-dive/04-positional-encoding/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/04-positional-encoding/docs/en.md).

Unlike learned positional embeddings, sinusoidal encodings use deterministic sine and cosine functions of different frequencies, allowing the model to generalize to sequence lengths longer than those seen during training. These encodings are added directly to the token embeddings before entering the transformer blocks.

## Composing the Full Transformer Block

A complete transformer block combines attention with a feed-forward network (FFN) and normalization layers. According to the architecture in [`phases/07-transformers-deep-dive/05-full-transformer/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/05-full-transformer/docs/en.md), each block contains:

- **LayerNorm** applied before the attention sub-layer
- **Multi-Head Self-Attention** with residual connection
- **LayerNorm** applied before the FFN sub-layer
- **Feed-Forward Network**: A two-layer MLP with hidden dimension $4 \times d_{model}$ and GELU activation, followed by a residual connection

This **pre-norm** configuration prevents gradient vanishing in deep networks and has become the standard for modern LLMs.

## Assembling the Complete GPT-Style Model

The final model, located in [`phases/07-transformers-deep-dive/05-full-transformer/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/05-full-transformer/code/main.py), stacks multiple blocks and adds input embeddings plus an output language-modeling head.

```python
class TinyGPT(torch.nn.Module):
    def __init__(self, vocab_size=64, d_model=32, n_heads=4, n_layers=2, max_len=16):
        super().__init__()
        self.token_emb = torch.nn.Embedding(vocab_size, d_model)
        self.pos_emb = torch.nn.Embedding(max_len, d_model)
        self.blocks = torch.nn.ModuleList([
            torch.nn.ModuleDict({
                "ln1": torch.nn.LayerNorm(d_model),
                "attn": MultiHeadAttention(d_model, n_heads),
                "ln2": torch.nn.LayerNorm(d_model),
                "ffn": torch.nn.Sequential(
                    torch.nn.Linear(d_model, 4 * d_model),
                    torch.nn.GELU(),
                    torch.nn.Linear(4 * d_model, d_model),
                ),
            })
            for _ in range(n_layers)
        ])
        self.ln_final = torch.nn.LayerNorm(d_model)
        self.head = torch.nn.Linear(d_model, vocab_size, bias=False)

    def forward(self, idx):
        B, T = idx.shape
        x = self.token_emb(idx) + self.pos_emb(torch.arange(T, device=idx.device))
        for block in self.blocks:
            x = x + block["attn"](block["ln1"](x))
            x = x + block["ffn"](block["ln2"](x))
        x = self.ln_final(x)
        return self.head(x)

# Sanity check

model = TinyGPT()
logits = model(torch.randint(0, 64, (4, 16)))
print(logits.shape)   # torch.Size([4, 16, 64])

```

This implementation creates a minimal but complete transformer with only PyTorch dependencies, capable of next-token prediction on small vocabularies.

## Training Loop and Production Extensions

The curriculum concludes with a minimal training loop implementing masked language modeling on synthetic data, demonstrating the full data-through-model pipeline. Beyond the basic architecture, the repository covers production-grade optimizations:

- **KV-Cache and Flash Attention** ([`phases/07-transformers-deep-dive/12-kv-cache-flash-attention/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/12-kv-cache-flash-attention/docs/en.md)) for efficient inference
- **Scaling Laws** ([`phases/07-transformers-deep-dive/13-scaling-laws/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/13-scaling-laws/docs/en.md)) governing model size and compute-optimal training
- **Mixture-of-Experts** ([`phases/07-transformers-deep-dive/11-mixture-of-experts/docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/11-mixture-of-experts/docs/en.md)) for conditional computation

## Summary

- **Scaled dot-product attention** forms the mathematical core of transformers, computing weighted relationships between all token pairs simultaneously.
- **Multi-head attention** parallelizes multiple attention computations across different representation subspaces, with causal masking ensuring autoregressive generation.
- **Pre-normalization** (LayerNorm before sub-layers) and residual connections stabilize training in deep transformer stacks.
- **Sinusoidal positional encodings** inject sequence order information without adding learnable parameters.
- The complete implementation in [`phases/07-transformers-deep-dive/05-full-transformer/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/05-full-transformer/code/main.py) requires only PyTorch and produces a functional GPT-style model suitable for educational experimentation.

## Frequently Asked Questions

### Do I need the Hugging Face Transformers library to build this model?

No. The *ai-engineering-from-scratch* curriculum explicitly prohibits external transformer libraries; you only need PyTorch and the Python standard library. This constraint ensures you understand every matrix multiplication and tensor reshape rather than treating the model as a black box.

### What is the difference between pre-norm and post-norm transformer architectures?

**Pre-norm** applies Layer normalization before the attention and FFN sub-layers, while **post-norm** applies it after. The repository uses pre-norm because it prevents gradient vanishing in deep networks and eliminates the need for careful learning rate warm-up schedules that post-norm architectures require.

### Why does the implementation use sinusoidal positional encodings instead of learned embeddings?

Sinusoidal encodings allow the model to generalize to sequence lengths longer than those seen during training because the encoding functions are deterministic and continuous. Learned positional embeddings are limited to the maximum sequence length observed during training, though they may offer slightly more flexibility within that range.

### How does the causal mask prevent the model from looking at future tokens?

The causal mask is a lower-triangular boolean matrix that multiplies the attention scores. In `scaled_dot_product_attention`, positions corresponding to future tokens are filled with negative infinity before the softmax, forcing those attention weights to zero. This ensures that prediction of token $i$ depends only on tokens $0$ through $i-1$, maintaining the autoregressive property essential for language modeling.