How the Attention Mechanism and Self-Attention Work: A Deep Dive into nn-zero-to-hero

The attention mechanism enables neural networks to dynamically weigh the importance of different input tokens when producing an output, while self-attention specifically allows every token in a sequence to attend to every other token via learned Query, Key, and Value projections.

The attention mechanism is the foundational building block of modern Transformer architectures, including GPT-style language models. In the karpathy/nn-zero-to-hero repository, Andrej Karpathy introduces this concept as the critical evolution from simple bigram models to sophisticated sequence learners that can capture long-range dependencies. This article explains how self-attention computes pairwise token interactions using Query, Key, and Value matrices, based on the implementation patterns described in the repository's progression toward the Transformer architecture.

The Core Mechanism of Self-Attention

Self-attention operates by projecting each input token into three distinct vectors: Query, Key, and Value. Given an input sequence of token embeddings $\mathbf{X} = [\mathbf{x}_1, \dots, \mathbf{x}_n]$, the mechanism computes:

Projection Purpose Formula
Query $\mathbf{q}_i$ What the current token is looking for $\mathbf{q}_i = \mathbf{W}_Q \mathbf{x}_i$
Key $\mathbf{k}_j$ How each token describes itself $\mathbf{k}_j = \mathbf{W}_K \mathbf{x}_j$
Value $\mathbf{v}_j$ The information to be passed forward $\mathbf{v}_j = \mathbf{W}_V \mathbf{x}_j$

The attention weights are computed by measuring similarity between Queries and Keys using a scaled dot-product:

$$ \alpha_{ij} = \text{softmax}_j!\Big(\frac{\mathbf{q}_i \cdot \mathbf{k}_j}{\sqrt{d_k}}\Big) $$

Here $d_k$ represents the dimension of the key vectors. The scaling factor $\frac{1}{\sqrt{d_k}}$ prevents the dot products from growing too large in magnitude, which would push the softmax function into regions with extremely small gradients.

The final output representation for token $i$ is the weighted sum of all Values:

$$ \mathbf{z}i = \sum{j=1}^{n} \alpha_{ij},\mathbf{v}_j $$

When computed for all tokens simultaneously, this produces the self-attention matrix $\mathbf{Z} = \text{Attention}(\mathbf{Q}, \mathbf{K}, \mathbf{V})$.

Multi-Head Attention

A single attention head can only capture one type of relationship between tokens. Multi-head attention addresses this limitation by splitting the embedding dimension into $h$ parallel heads, each with its own projection matrices $\mathbf{W}_Q^h$, $\mathbf{W}_K^h$, and $\mathbf{W}_V^h$.

Each head independently computes attention over different subspaces of the embedding. The outputs from all heads are concatenated and linearly projected back to the original model dimension:

$$ \text{MultiHead}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{Concat}(\text{head}_1, \dots, \text{head}_h)\mathbf{W}_O $$

This parallelism allows the model to jointly attend to information from different representation subspaces at different positions, which is crucial for capturing complex linguistic patterns.

Why the Attention Mechanism Works

The attention mechanism in nn-zero-to-hero solves three fundamental limitations of earlier architectures like RNNs and early MLPs:

  • Dynamic weighting: Unlike fixed-window convolutions or sequential RNN states, attention learns to focus on relevant tokens regardless of their position in the sequence.
  • Parallel computation: All pairwise interactions are computed simultaneously via matrix multiplication, enabling efficient training on modern GPUs/TPUs.
  • Long-range dependencies: Direct connections between any two positions allow gradients to flow without vanishing across long sequences, a key advantage demonstrated in the repository's progression from makemore_part1_bigrams.ipynb to the full Transformer.

According to the repository's README.md, this mechanism forms the basis of the "modern Transformer language model" architecture introduced in the later lectures.

Implementation Examples

The following PyTorch implementations demonstrate the attention mechanism as practiced in the nn-zero-to-hero lecture series. These patterns reflect the transition from simple MLPs in makemore_part2_mlp.ipynb to the sophisticated attention layers required for GPT-style models.

Minimal Self-Attention Layer (Single Head)

This implementation shows the core QKV computation and scaled dot-product attention without any optimizations:

import torch
import torch.nn as nn
import math

class SelfAttention(nn.Module):
    def __init__(self, embed_dim):
        super().__init__()
        self.embed_dim = embed_dim
        self.query = nn.Linear(embed_dim, embed_dim)
        self.key   = nn.Linear(embed_dim, embed_dim)
        self.value = nn.Linear(embed_dim, embed_dim)

    def forward(self, x):
        # x: (batch, seq_len, embed_dim)

        Q = self.query(x)               # (B, T, D)

        K = self.key(x)                 # (B, T, D)

        V = self.value(x)               # (B, T, D)

        # scaled dot-product attention

        scores = torch.matmul(Q, K.transpose(-2, -1)) / math.sqrt(self.embed_dim)
        attn_weights = torch.softmax(scores, dim=-1)   # (B, T, T)

        out = torch.matmul(attn_weights, V)            # (B, T, D)

        return out, attn_weights

# Example usage

batch, seq_len, dim = 2, 5, 32
x = torch.randn(batch, seq_len, dim)
layer = SelfAttention(dim)
context, weights = layer(x)
print(context.shape)   # → torch.Size([2, 5, 32])

Multi-Head Self-Attention

This class implements the multi-head variant used in the original "Attention Is All You Need" paper, referenced in the repository's README.md:

class MultiHeadSelfAttention(nn.Module):
    def __init__(self, embed_dim, num_heads):
        super().__init__()
        assert embed_dim % num_heads == 0, "embed_dim must be divisible by num_heads"
        self.num_heads = num_heads
        self.head_dim = embed_dim // num_heads

        self.qkv = nn.Linear(embed_dim, embed_dim * 3)   # combined projection

        self.out_proj = nn.Linear(embed_dim, embed_dim)

    def forward(self, x):
        B, T, D = x.shape
        qkv = self.qkv(x)                     # (B, T, 3*D)

        qkv = qkv.reshape(B, T, 3, self.num_heads, self.head_dim)
        q, k, v = qkv.unbind(dim=2)           # each → (B, T, heads, head_dim)

        # transpose for batched matmul: (B, heads, T, head_dim)

        q = q.transpose(1, 2)
        k = k.transpose(1, 2)
        v = v.transpose(1, 2)

        scores = torch.matmul(q, k.transpose(-2, -1)) / math.sqrt(self.head_dim)
        attn = torch.softmax(scores, dim=-1)          # (B, heads, T, T)

        context = torch.matmul(attn, v)               # (B, heads, T, head_dim)

        # bring back to (B, T, D)

        context = context.transpose(1, 2).contiguous().view(B, T, D)
        return self.out_proj(context), attn

# Demo

mhsa = MultiHeadSelfAttention(embed_dim=64, num_heads=8)
out, attn = mhsa(torch.randn(1, 10, 64))
print(out.shape)   # → torch.Size([1, 10, 64])

TinyGPT Block

This example shows how self-attention integrates with feed-forward layers and residual connections to form a complete Transformer block, following the pattern used in the repository's GPT lecture:

class TinyGPTBlock(nn.Module):
    def __init__(self, embed_dim, num_heads, ff_hidden):
        super().__init__()
        self.ln1 = nn.LayerNorm(embed_dim)
        self.attn = MultiHeadSelfAttention(embed_dim, num_heads)
        self.ln2 = nn.LayerNorm(embed_dim)
        self.ff = nn.Sequential(
            nn.Linear(embed_dim, ff_hidden),
            nn.ReLU(),
            nn.Linear(ff_hidden, embed_dim)
        )

    def forward(self, x):
        # Self-attention sub-layer

        x = x + self.attn(self.ln1(x))[0]     # residual connection

        # Feed-forward sub-layer

        x = x + self.ff(self.ln2(x))           # residual connection

        return x

# Quick test

block = TinyGPTBlock(embed_dim=128, num_heads=4, ff_hidden=256)
tokens = torch.randn(2, 12, 128)   # (batch, seq_len, embed_dim)

out = block(tokens)
print(out.shape)   # → torch.Size([2, 12, 128])

Key Files in the Repository

The evolution of the attention mechanism across the nn-zero-to-hero curriculum is documented in several key files:

  • README.md: Introduces the progression from bigram models to the modern Transformer language model, explicitly citing "Attention Is All You Need" as the foundational paper for self-attention.
  • lectures/makemore/makemore_part1_bigrams.ipynb: Implements the simplest possible language model using only token frequencies, establishing the baseline that attention mechanisms later supersede.
  • lectures/makemore/makemore_part2_mlp.ipynb: Demonstrates multilayer perceptron architectures that process fixed-size contexts, highlighting the limitations that self-attention solves.
  • lectures/makemore/makemore_part3_bn.ipynb: Introduces batch normalization and deeper network architectures, concepts that combine with attention in the final Transformer implementation.

While the full GPT implementation resides primarily in the video lectures rather than committed notebooks, the README.md serves as the canonical reference for the repository's treatment of the attention mechanism.

Summary

  • Self-attention computes pairwise relationships between all tokens in a sequence using Query, Key, and Value projections, enabling dynamic context-aware representations.
  • The scaled dot-product attention formula prevents gradient vanishing by dividing by $\sqrt{d_k}$ before applying softmax.
  • Multi-head attention runs multiple attention operations in parallel to capture diverse types of relationships in different subspaces.
  • The nn-zero-to-hero repository demonstrates this evolution from simple bigram models in makemore_part1_bigrams.ipynb to sophisticated attention-based architectures capable of long-range dependency modeling.
  • When combined with residual connections and layer normalization, self-attention forms the core of Transformer blocks used in modern LLMs.

Frequently Asked Questions

What is the difference between attention and self-attention?

Self-attention is a specific type of attention where the Query, Key, and Value vectors all come from the same input sequence, allowing each token to attend to every other token in that sequence. General attention mechanisms can also involve cross-attention, where Queries come from one sequence (e.g., a decoder state) while Keys and Values come from another (e.g., an encoder output), as seen in original seq2seq models.

Why is the dot product scaled by the square root of dimension?

The scaling factor $\frac{1}{\sqrt{d_k}}$ prevents the dot products from becoming too large when dimensionality is high. Without this scaling, the softmax function would receive extremely large values, pushing the output distribution into a very sharp peak where gradients become vanishingly small, making training unstable. This is particularly important in the high-dimensional spaces used in nn-zero-to-hero's Transformer implementations.

How does multi-head attention improve over single-head attention?

Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions. A single attention head might focus only on local syntactic relationships, while another captures long-range semantic dependencies. By concatenating the outputs of multiple heads—each with its own learned projection matrices—the model achieves a richer representation than any single head could provide, as implemented in the MultiHeadSelfAttention class above.

Where does nn-zero-to-hero implement the full Transformer?

The repository's README.md references the "modern Transformer language model" as the culmination of the lecture series, specifically in the GPT-focused videos. While the foundational concepts appear across makemore_part2_mlp.ipynb (neural architectures) and makemore_part3_bn.ipynb (normalization techniques), the complete Transformer implementation with multi-head attention is primarily covered in the video lectures referenced in the README.md rather than standalone notebook files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →