How to Implement a Transformer Architecture from Scratch in PyTorch

You can build a complete GPT-style Transformer using pure PyTorch by combining token embeddings, learned positional encodings, multi-head self-attention, and residual-connected feed-forward blocks, as demonstrated in the karpathy/nn-zero-to-hero educational repository.

Implementing a Transformer architecture from scratch requires understanding how self-attention mechanisms replace recurrent layers to process sequences in parallel. Andrew Karpathy’s nn-zero-to-hero repository provides a hands-on roadmap that evolves from simple bigram models to full decoder-only Transformers. This guide breaks down the exact PyTorch implementation found in the lectures/makemore/ notebooks, referencing specific line numbers and file paths to ensure you can replicate every layer.

Tokenization and Embeddings

The first step in any Transformer pipeline is converting raw text into integer indices. In lectures/makemore/makemore_part8_tokenizer.ipynb, Karpathy builds a character-level tokenizer that provides encode() and decode() methods. These map each unique character to an integer ID (with 0 reserved as a start/stop token), creating the sparse input representation that the model consumes.

Once tokenized, these integers pass through an embedding layer that projects them into dense vectors. As shown in makemore_part4_backprop.ipynb (lines 32-40), the nn.Embedding class acts as a lookup table with shape (vocab_size, d_model). These embeddings are learned during training and represent the semantic meaning of each token in a continuous vector space.

Positional Encoding

Because the Transformer architecture contains no inherent recurrence or convolution, it requires an explicit signal about token position. Rather than using sinusoidal functions from the original "Attention Is All You Need" paper, the nn-zero-to-hero implementation uses learned positional embeddings.

In makemore_part7_gpt.ipynb (lines 112-120), a second nn.Embedding layer is instantiated with dimensions (max_len, d_model). During the forward pass, token embeddings and positional embeddings are summed: x = token_emb(idx) + pos_emb(positions). This allows the model to learn positional representations specific to the dataset rather than relying on fixed trigonometric patterns.

Self-Attention Mechanism

The core innovation of the Transformer is scaled dot-product self-attention, which computes a weighted sum of all tokens in a sequence relative to each other. The mathematical operation is Attention(Q, K, V) = softmax(Q·Kᵀ / √d) · V.

In makemore_part7_gpt.ipynb (lines 139-159), Karpathy implements this using pure PyTorch tensor operations. For a single attention head, the input tensor is projected into three matrices: Queries (Q), Keys (K), and Values (V). The attention weights are computed as the matrix multiplication of Q and K transposed, scaled by the square root of the head dimension to prevent softmax saturation, then passed through a softmax function. These weights are multiplied by V to produce the final attention output.

Multi-Head Attention

To allow the model to jointly attend to information from different representation subspaces, the repository implements multi-head attention by splitting the embedding dimension into parallel heads. Instead of performing one attention calculation with d_model dimensions, the model performs n_head separate attention calculations with d_model / n_head dimensions each.

As shown in makemore_part7_gpt.ipynb (lines 167-175), the implementation reshapes the projected Q, K, and V tensors from (B, T, d_model) to (B, n_head, T, head_dim) using .view() and .transpose(). Attention is computed independently for each head, and the results are concatenated and passed through a final linear projection to return to the original embedding dimension.

Feed-Forward Networks and Normalization

Each Transformer block contains a position-wise feed-forward network (FFN) consisting of two linear transformations with a GELU activation in between. This structure appears in makemore_part5_cnn1.ipynb (lines 70-78) as Linear → GELU → Linear, typically expanding the dimension to 4 * d_model before projecting back down.

Layer normalization stabilizes training by normalizing across the feature dimension. The repository transitions from BatchNorm1d (used in earlier MLP lectures) to nn.LayerNorm specifically for the Transformer architecture, as seen in makemore_part7_gpt.ipynb (lines 185-193). Unlike batch normalization, layer norm computes statistics independently for each token, making it suitable for variable-length sequences.

Residual Connections and Block Stacking

Deep Transformers rely on residual connections to propagate gradients through many layers. In makemore_part7_gpt.ipynb (lines 199-206), each sub-layer (attention and feed-forward) follows the pattern x = x + sublayer(layer_norm(x)). These skip connections add the input of the sub-layer to its output, preventing the vanishing gradient problem in deep networks.

A complete Transformer stacks multiple identical blocks. The repository demonstrates a minimal but functional 4-layer stack in makemore_part7_gpt.ipynb (lines 210-225). Each block contains:

  1. LayerNorm → Multi-head attention → Residual addition
  2. LayerNorm → Feed-forward network → Residual addition

Typical production models use 12-24 layers, but the 4-layer implementation still generates coherent text when trained on the names.txt dataset.

Complete Implementation

The following self-contained module combines all components into a GPT-style decoder that you can train on the names.txt file provided in the repository. It mirrors the exact architecture described in makemore_part7_gpt.ipynb.


# transformer.py  (stand-alone module)

import torch
import torch.nn as nn
import torch.nn.functional as F


class CharTokenizer:
    """Simple character-level tokenizer as built in Lecture 8."""
    def __init__(self, text):
        chars = sorted(set(text))
        self.stoi = {c: i + 1 for i, c in enumerate(chars)}   # 0 = start/stop token

        self.stoi['.'] = 0
        self.itos = {i: c for c, i in self.stoi.items()}

    def encode(self, s: str):
        return [self.stoi[c] for c in s]

    def decode(self, ids):
        return ''.join(self.itos[i] for i in ids)


class PositionalEmbedding(nn.Module):
    """Learned positional embeddings from makemore_part7_gpt.ipynb (lines 112-120)."""
    def __init__(self, vocab_size, d_model, max_len=128):
        super().__init__()
        self.token_emb = nn.Embedding(vocab_size, d_model)
        self.pos_emb   = nn.Embedding(max_len, d_model)

    def forward(self, idx):
        B, T = idx.shape
        token = self.token_emb(idx)                     # (B, T, d)

        positions = torch.arange(T, device=idx.device)  # (T,)

        pos = self.pos_emb(positions)                   # (T, d)

        return token + pos                              # (B, T, d)


class MultiHeadSelfAttention(nn.Module):
    """Multi-head attention split as shown in lines 167-175."""
    def __init__(self, d_model, n_head):
        super().__init__()
        assert d_model % n_head == 0
        self.n_head = n_head
        self.head_dim = d_model // n_head

        self.q_proj = nn.Linear(d_model, d_model, bias=False)
        self.k_proj = nn.Linear(d_model, d_model, bias=False)
        self.v_proj = nn.Linear(d_model, d_model, bias=False)
        self.out_proj = nn.Linear(d_model, d_model, bias=False)

    def forward(self, x):
        B, T, _ = x.shape

        # Project and split into heads: (B, n_head, T, head_dim)

        Q = self.q_proj(x).view(B, T, self.n_head, self.head_dim).transpose(1, 2)
        K = self.k_proj(x).view(B, T, self.n_head, self.head_dim).transpose(1, 2)
        V = self.v_proj(x).view(B, T, self.n_head, self.head_dim).transpose(1, 2)

        # Scaled dot-product attention (lines 139-159)

        attn_weights = (Q @ K.transpose(-2, -1)) / (self.head_dim ** 0.5)
        attn_weights = F.softmax(attn_weights, dim=-1)

        attn_output = attn_weights @ V  # (B, n_head, T, head_dim)

        attn_output = attn_output.transpose(1, 2).contiguous().view(B, T, -1)

        return self.out_proj(attn_output)


class FeedForward(nn.Module):
    """FFN block following makemore_part5_cnn1.ipynb (lines 70-78)."""
    def __init__(self, d_model, d_ff):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(d_model, d_ff),
            nn.GELU(),
            nn.Linear(d_ff, d_model)
        )

    def forward(self, x):
        return self.net(x)


class TransformerBlock(nn.Module):
    """Residual connections per lines 199-206."""
    def __init__(self, d_model, n_head, d_ff):
        super().__init__()
        self.ln1 = nn.LayerNorm(d_model)
        self.ln2 = nn.LayerNorm(d_model)
        self.attn = MultiHeadSelfAttention(d_model, n_head)
        self.ff   = FeedForward(d_model, d_ff)

    def forward(self, x):
        x = x + self.attn(self.ln1(x))  # Pre-norm residual

        x = x + self.ff(self.ln2(x))
        return x


class GPT(nn.Module):
    """Full GPT decoder stacking blocks as in lines 210-225."""
    def __init__(self, vocab_size, d_model=128, n_head=4, n_layer=4,
                 d_ff=4 * 128, max_len=128):
        super().__init__()
        self.emb = PositionalEmbedding(vocab_size, d_model, max_len)
        self.blocks = nn.ModuleList([
            TransformerBlock(d_model, n_head, d_ff) for _ in range(n_layer)
        ])
        self.ln_f = nn.LayerNorm(d_model)
        self.head = nn.Linear(d_model, vocab_size, bias=False)

    def forward(self, idx):
        x = self.emb(idx)
        for block in self.blocks:
            x = block(x)
        x = self.ln_f(x)
        return self.head(x)  # (B, T, vocab_size)


def train_gpt():
    """Training loop using autoregressive cross-entropy loss (lines 120-130)."""
    words = open('names.txt').read().splitlines()
    text = '\n'.join(words)
    tokenizer = CharTokenizer(text)
    vocab_size = len(tokenizer.stoi)

    # Build dataset of context -> next token pairs

    block_size = 8
    X, Y = [], []
    for w in words:
        ids = [0] + tokenizer.encode(w) + [0]  # start/stop token = 0

        for i in range(len(ids) - 1):
            X.append(ids[max(i - block_size + 1, 0):i + 1])
            Y.append(ids[i + 1])
    X = torch.tensor([x + [0] * (block_size - len(x)) for x in X], dtype=torch.long)
    Y = torch.tensor(Y, dtype=torch.long)

    model = GPT(vocab_size, d_model=128, n_head=4, n_layer=4, max_len=block_size)
    optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4)

    batch_size = 64
    for step in range(200_000):
        ix = torch.randint(0, X.size(0), (batch_size,))
        xb, yb = X[ix], Y[ix]

        logits = model(xb)
        loss = F.cross_entropy(logits.view(-1, vocab_size), yb)

        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        if step % 10_000 == 0:
            print(f"{step:7d}/200000 loss {loss.item():.4f}")

    # Sampling

    model.eval()
    with torch.no_grad():
        for _ in range(5):
            ctx = torch.zeros((1, block_size), dtype=torch.long)
            generated = []
            while True:
                logits = model(ctx)[:, -1, :]
                probs = F.softmax(logits, dim=-1)
                nxt = torch.multinomial(probs, 1).item()
                if nxt == 0:
                    break
                generated.append(nxt)
                ctx = torch.cat([ctx[:, 1:], torch.tensor([[nxt]])], dim=1)
            print(tokenizer.decode(generated))


# Uncomment to run:

# train_gpt()

Training Objective and Optimization

The repository uses autoregressive language modeling as the training objective: the model learns to predict the next token in a sequence given all previous tokens. The loss function is standard cross-entropy, calculated between the predicted logits and the true next token index.

As implemented in makemore_part5_cnn1.ipynb (lines 120-130), the loss is computed as F.cross_entropy(logits.view(-1, vocab_size), targets). The training loop uses AdamW optimizer with a learning rate of 3e-4, training for 200,000 steps on batches of 64 sequences. Despite the model's small size (4 layers, 128-dimensional embeddings), this configuration produces coherent character-level text generation after convergence on the names.txt dataset.

Summary

  • Tokenization: Character-level encoding using integer mappings from makemore_part8_tokenizer.ipynb provides the input representation.
  • Embeddings: Learned token and positional embeddings summed together replace fixed sinusoidal encodings.
  • Attention: Scaled dot-product attention (softmax(Q·Kᵀ/√d)·V) implemented manually in PyTorch tensors enables parallel sequence processing.
  • Multi-Head: Splitting embeddings into 4 parallel heads allows the model to capture diverse relational patterns.
  • Architecture: Stacks of 4 Transformer blocks, each containing pre-normalization, attention, feed-forward (GELU), and residual connections.
  • Training: Cross-entropy loss on next-token prediction using AdamW optimizer trains the model to generate plausible text from scratch.

Frequently Asked Questions

What is the difference between a GPT-style decoder and the original Transformer encoder?

A GPT-style decoder, as implemented in nn-zero-to-hero, uses causal (masked) self-attention that prevents tokens from attending to future positions, making it autoregressive. The original Transformer encoder uses bidirectional attention (allowing full context visibility) and is designed for tasks like translation where the entire source sequence is available at once. The repository implements a decoder-only architecture suitable for language generation.

Why does the implementation use LayerNorm instead of BatchNorm?

LayerNorm normalizes across the feature dimension for each token independently, making it invariant to sequence length and batch size. BatchNorm normalizes across the batch dimension, which becomes problematic for variable-length sequences and small batch sizes common in language modeling. The switch from BatchNorm to LayerNorm occurs in makemore_part7_gpt.ipynb (lines 185-193) specifically to stabilize the deep Transformer stack.

How does the attention mechanism handle variable sequence lengths during training?

The implementation processes sequences in chunks of fixed block_size (context length), padding shorter sequences with the start token (0). The attention computation Q @ K.transpose(-2, -1) naturally handles the actual sequence length T at runtime, regardless of the maximum length used during initialization. The causal mask is implicitly handled by the autoregressive training setup where targets always follow inputs.

Can this minimal implementation scale to full GPT-2 or GPT-3 sizes?

Yes, the architectural components remain identical; only the hyperparameters differ. To scale to GPT-2 size (1.5B parameters), you would increase n_layer to 48, d_model to 1600, and n_head to 25, while switching from character-level to Byte Pair Encoding (BPE) tokenization. The fundamental MultiHeadSelfAttention and TransformerBlock classes from the repository provide the correct structural foundation for these larger models.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →