# How to Implement a Transformer Architecture from Scratch in PyTorch

> Implement a Transformer architecture from scratch in PyTorch. Learn to build a GPT-style model using token embeddings, attention, and feed-forward blocks with the nn-zero-to-hero repository.

- Repository: [Andrej/nn-zero-to-hero](https://github.com/karpathy/nn-zero-to-hero)
- Tags: how-to-guide
- Published: 2026-05-23

---

**You can build a complete GPT-style Transformer using pure PyTorch by combining token embeddings, learned positional encodings, multi-head self-attention, and residual-connected feed-forward blocks, as demonstrated in the karpathy/nn-zero-to-hero educational repository.**

Implementing a Transformer architecture from scratch requires understanding how self-attention mechanisms replace recurrent layers to process sequences in parallel. Andrew Karpathy’s **nn-zero-to-hero** repository provides a hands-on roadmap that evolves from simple bigram models to full decoder-only Transformers. This guide breaks down the exact PyTorch implementation found in the `lectures/makemore/` notebooks, referencing specific line numbers and file paths to ensure you can replicate every layer.

## Tokenization and Embeddings

The first step in any Transformer pipeline is converting raw text into integer indices. In `lectures/makemore/makemore_part8_tokenizer.ipynb`, Karpathy builds a character-level tokenizer that provides `encode()` and `decode()` methods. These map each unique character to an integer ID (with 0 reserved as a start/stop token), creating the sparse input representation that the model consumes.

Once tokenized, these integers pass through an **embedding layer** that projects them into dense vectors. As shown in `makemore_part4_backprop.ipynb` (lines 32-40), the `nn.Embedding` class acts as a lookup table with shape `(vocab_size, d_model)`. These embeddings are learned during training and represent the semantic meaning of each token in a continuous vector space.

## Positional Encoding

Because the Transformer architecture contains no inherent recurrence or convolution, it requires an explicit signal about token position. Rather than using sinusoidal functions from the original "Attention Is All You Need" paper, the nn-zero-to-hero implementation uses **learned positional embeddings**.

In `makemore_part7_gpt.ipynb` (lines 112-120), a second `nn.Embedding` layer is instantiated with dimensions `(max_len, d_model)`. During the forward pass, token embeddings and positional embeddings are summed: `x = token_emb(idx) + pos_emb(positions)`. This allows the model to learn positional representations specific to the dataset rather than relying on fixed trigonometric patterns.

## Self-Attention Mechanism

The core innovation of the Transformer is **scaled dot-product self-attention**, which computes a weighted sum of all tokens in a sequence relative to each other. The mathematical operation is `Attention(Q, K, V) = softmax(Q·Kᵀ / √d) · V`.

In `makemore_part7_gpt.ipynb` (lines 139-159), Karpathy implements this using pure PyTorch tensor operations. For a single attention head, the input tensor is projected into three matrices: Queries (`Q`), Keys (`K`), and Values (`V`). The attention weights are computed as the matrix multiplication of `Q` and `K` transposed, scaled by the square root of the head dimension to prevent softmax saturation, then passed through a softmax function. These weights are multiplied by `V` to produce the final attention output.

## Multi-Head Attention

To allow the model to jointly attend to information from different representation subspaces, the repository implements **multi-head attention** by splitting the embedding dimension into parallel heads. Instead of performing one attention calculation with `d_model` dimensions, the model performs `n_head` separate attention calculations with `d_model / n_head` dimensions each.

As shown in `makemore_part7_gpt.ipynb` (lines 167-175), the implementation reshapes the projected `Q`, `K`, and `V` tensors from `(B, T, d_model)` to `(B, n_head, T, head_dim)` using `.view()` and `.transpose()`. Attention is computed independently for each head, and the results are concatenated and passed through a final linear projection to return to the original embedding dimension.

## Feed-Forward Networks and Normalization

Each Transformer block contains a position-wise **feed-forward network (FFN)** consisting of two linear transformations with a GELU activation in between. This structure appears in `makemore_part5_cnn1.ipynb` (lines 70-78) as `Linear → GELU → Linear`, typically expanding the dimension to `4 * d_model` before projecting back down.

**Layer normalization** stabilizes training by normalizing across the feature dimension. The repository transitions from `BatchNorm1d` (used in earlier MLP lectures) to `nn.LayerNorm` specifically for the Transformer architecture, as seen in `makemore_part7_gpt.ipynb` (lines 185-193). Unlike batch normalization, layer norm computes statistics independently for each token, making it suitable for variable-length sequences.

## Residual Connections and Block Stacking

Deep Transformers rely on **residual connections** to propagate gradients through many layers. In `makemore_part7_gpt.ipynb` (lines 199-206), each sub-layer (attention and feed-forward) follows the pattern `x = x + sublayer(layer_norm(x))`. These skip connections add the input of the sub-layer to its output, preventing the vanishing gradient problem in deep networks.

A complete Transformer **stacks** multiple identical blocks. The repository demonstrates a minimal but functional 4-layer stack in `makemore_part7_gpt.ipynb` (lines 210-225). Each block contains:
1. LayerNorm → Multi-head attention → Residual addition
2. LayerNorm → Feed-forward network → Residual addition

Typical production models use 12-24 layers, but the 4-layer implementation still generates coherent text when trained on the [`names.txt`](https://github.com/karpathy/nn-zero-to-hero/blob/main/names.txt) dataset.

## Complete Implementation

The following self-contained module combines all components into a GPT-style decoder that you can train on the [`names.txt`](https://github.com/karpathy/nn-zero-to-hero/blob/main/names.txt) file provided in the repository. It mirrors the exact architecture described in `makemore_part7_gpt.ipynb`.

```python

# transformer.py  (stand-alone module)

import torch
import torch.nn as nn
import torch.nn.functional as F


class CharTokenizer:
    """Simple character-level tokenizer as built in Lecture 8."""
    def __init__(self, text):
        chars = sorted(set(text))
        self.stoi = {c: i + 1 for i, c in enumerate(chars)}   # 0 = start/stop token

        self.stoi['.'] = 0
        self.itos = {i: c for c, i in self.stoi.items()}

    def encode(self, s: str):
        return [self.stoi[c] for c in s]

    def decode(self, ids):
        return ''.join(self.itos[i] for i in ids)


class PositionalEmbedding(nn.Module):
    """Learned positional embeddings from makemore_part7_gpt.ipynb (lines 112-120)."""
    def __init__(self, vocab_size, d_model, max_len=128):
        super().__init__()
        self.token_emb = nn.Embedding(vocab_size, d_model)
        self.pos_emb   = nn.Embedding(max_len, d_model)

    def forward(self, idx):
        B, T = idx.shape
        token = self.token_emb(idx)                     # (B, T, d)

        positions = torch.arange(T, device=idx.device)  # (T,)

        pos = self.pos_emb(positions)                   # (T, d)

        return token + pos                              # (B, T, d)


class MultiHeadSelfAttention(nn.Module):
    """Multi-head attention split as shown in lines 167-175."""
    def __init__(self, d_model, n_head):
        super().__init__()
        assert d_model % n_head == 0
        self.n_head = n_head
        self.head_dim = d_model // n_head

        self.q_proj = nn.Linear(d_model, d_model, bias=False)
        self.k_proj = nn.Linear(d_model, d_model, bias=False)
        self.v_proj = nn.Linear(d_model, d_model, bias=False)
        self.out_proj = nn.Linear(d_model, d_model, bias=False)

    def forward(self, x):
        B, T, _ = x.shape

        # Project and split into heads: (B, n_head, T, head_dim)

        Q = self.q_proj(x).view(B, T, self.n_head, self.head_dim).transpose(1, 2)
        K = self.k_proj(x).view(B, T, self.n_head, self.head_dim).transpose(1, 2)
        V = self.v_proj(x).view(B, T, self.n_head, self.head_dim).transpose(1, 2)

        # Scaled dot-product attention (lines 139-159)

        attn_weights = (Q @ K.transpose(-2, -1)) / (self.head_dim ** 0.5)
        attn_weights = F.softmax(attn_weights, dim=-1)

        attn_output = attn_weights @ V  # (B, n_head, T, head_dim)

        attn_output = attn_output.transpose(1, 2).contiguous().view(B, T, -1)

        return self.out_proj(attn_output)


class FeedForward(nn.Module):
    """FFN block following makemore_part5_cnn1.ipynb (lines 70-78)."""
    def __init__(self, d_model, d_ff):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(d_model, d_ff),
            nn.GELU(),
            nn.Linear(d_ff, d_model)
        )

    def forward(self, x):
        return self.net(x)


class TransformerBlock(nn.Module):
    """Residual connections per lines 199-206."""
    def __init__(self, d_model, n_head, d_ff):
        super().__init__()
        self.ln1 = nn.LayerNorm(d_model)
        self.ln2 = nn.LayerNorm(d_model)
        self.attn = MultiHeadSelfAttention(d_model, n_head)
        self.ff   = FeedForward(d_model, d_ff)

    def forward(self, x):
        x = x + self.attn(self.ln1(x))  # Pre-norm residual

        x = x + self.ff(self.ln2(x))
        return x


class GPT(nn.Module):
    """Full GPT decoder stacking blocks as in lines 210-225."""
    def __init__(self, vocab_size, d_model=128, n_head=4, n_layer=4,
                 d_ff=4 * 128, max_len=128):
        super().__init__()
        self.emb = PositionalEmbedding(vocab_size, d_model, max_len)
        self.blocks = nn.ModuleList([
            TransformerBlock(d_model, n_head, d_ff) for _ in range(n_layer)
        ])
        self.ln_f = nn.LayerNorm(d_model)
        self.head = nn.Linear(d_model, vocab_size, bias=False)

    def forward(self, idx):
        x = self.emb(idx)
        for block in self.blocks:
            x = block(x)
        x = self.ln_f(x)
        return self.head(x)  # (B, T, vocab_size)


def train_gpt():
    """Training loop using autoregressive cross-entropy loss (lines 120-130)."""
    words = open('names.txt').read().splitlines()
    text = '\n'.join(words)
    tokenizer = CharTokenizer(text)
    vocab_size = len(tokenizer.stoi)

    # Build dataset of context -> next token pairs

    block_size = 8
    X, Y = [], []
    for w in words:
        ids = [0] + tokenizer.encode(w) + [0]  # start/stop token = 0

        for i in range(len(ids) - 1):
            X.append(ids[max(i - block_size + 1, 0):i + 1])
            Y.append(ids[i + 1])
    X = torch.tensor([x + [0] * (block_size - len(x)) for x in X], dtype=torch.long)
    Y = torch.tensor(Y, dtype=torch.long)

    model = GPT(vocab_size, d_model=128, n_head=4, n_layer=4, max_len=block_size)
    optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4)

    batch_size = 64
    for step in range(200_000):
        ix = torch.randint(0, X.size(0), (batch_size,))
        xb, yb = X[ix], Y[ix]

        logits = model(xb)
        loss = F.cross_entropy(logits.view(-1, vocab_size), yb)

        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        if step % 10_000 == 0:
            print(f"{step:7d}/200000 loss {loss.item():.4f}")

    # Sampling

    model.eval()
    with torch.no_grad():
        for _ in range(5):
            ctx = torch.zeros((1, block_size), dtype=torch.long)
            generated = []
            while True:
                logits = model(ctx)[:, -1, :]
                probs = F.softmax(logits, dim=-1)
                nxt = torch.multinomial(probs, 1).item()
                if nxt == 0:
                    break
                generated.append(nxt)
                ctx = torch.cat([ctx[:, 1:], torch.tensor([[nxt]])], dim=1)
            print(tokenizer.decode(generated))


# Uncomment to run:

# train_gpt()

```

## Training Objective and Optimization

The repository uses **autoregressive language modeling** as the training objective: the model learns to predict the next token in a sequence given all previous tokens. The loss function is standard **cross-entropy**, calculated between the predicted logits and the true next token index.

As implemented in `makemore_part5_cnn1.ipynb` (lines 120-130), the loss is computed as `F.cross_entropy(logits.view(-1, vocab_size), targets)`. The training loop uses **AdamW** optimizer with a learning rate of 3e-4, training for 200,000 steps on batches of 64 sequences. Despite the model's small size (4 layers, 128-dimensional embeddings), this configuration produces coherent character-level text generation after convergence on the [`names.txt`](https://github.com/karpathy/nn-zero-to-hero/blob/main/names.txt) dataset.

## Summary

- **Tokenization**: Character-level encoding using integer mappings from `makemore_part8_tokenizer.ipynb` provides the input representation.
- **Embeddings**: Learned token and positional embeddings summed together replace fixed sinusoidal encodings.
- **Attention**: Scaled dot-product attention (`softmax(Q·Kᵀ/√d)·V`) implemented manually in PyTorch tensors enables parallel sequence processing.
- **Multi-Head**: Splitting embeddings into 4 parallel heads allows the model to capture diverse relational patterns.
- **Architecture**: Stacks of 4 Transformer blocks, each containing pre-normalization, attention, feed-forward (GELU), and residual connections.
- **Training**: Cross-entropy loss on next-token prediction using AdamW optimizer trains the model to generate plausible text from scratch.

## Frequently Asked Questions

### What is the difference between a GPT-style decoder and the original Transformer encoder?

A GPT-style decoder, as implemented in nn-zero-to-hero, uses **causal (masked) self-attention** that prevents tokens from attending to future positions, making it autoregressive. The original Transformer encoder uses bidirectional attention (allowing full context visibility) and is designed for tasks like translation where the entire source sequence is available at once. The repository implements a decoder-only architecture suitable for language generation.

### Why does the implementation use LayerNorm instead of BatchNorm?

**LayerNorm** normalizes across the feature dimension for each token independently, making it invariant to sequence length and batch size. **BatchNorm** normalizes across the batch dimension, which becomes problematic for variable-length sequences and small batch sizes common in language modeling. The switch from BatchNorm to LayerNorm occurs in `makemore_part7_gpt.ipynb` (lines 185-193) specifically to stabilize the deep Transformer stack.

### How does the attention mechanism handle variable sequence lengths during training?

The implementation processes sequences in chunks of fixed `block_size` (context length), padding shorter sequences with the start token (0). The attention computation `Q @ K.transpose(-2, -1)` naturally handles the actual sequence length `T` at runtime, regardless of the maximum length used during initialization. The causal mask is implicitly handled by the autoregressive training setup where targets always follow inputs.

### Can this minimal implementation scale to full GPT-2 or GPT-3 sizes?

Yes, the architectural components remain identical; only the **hyperparameters** differ. To scale to GPT-2 size (1.5B parameters), you would increase `n_layer` to 48, `d_model` to 1600, and `n_head` to 25, while switching from character-level to **Byte Pair Encoding (BPE)** tokenization. The fundamental `MultiHeadSelfAttention` and `TransformerBlock` classes from the repository provide the correct structural foundation for these larger models.