Understanding the Attention Mechanism in Deep Learning: A Code-First Guide

Attention mechanisms allow neural networks to dynamically focus on relevant parts of input sequences by computing a weighted context vector for each decoding step, eliminating the bottleneck of fixed-size representations.

The attention mechanism in deep learning revolutionized sequence modeling by enabling models to learn soft alignments between inputs and outputs. This article examines the canonical implementation found in the rohitg00/ai-engineering-from-scratch repository, walking through the mathematical foundations and executable Python code that demonstrate both Bahdanau (additive) and Luong (multiplicative) attention.

How the Attention Mechanism Works: The Three-Step Pipeline

Classical attention operates as a three-stage pipeline that converts a decoder's query into a context vector. According to the source code in code/main.py, this process involves creating a query, scoring alignment with encoder states, and generating a weighted context.

Step 1: Query Creation

The decoder's current hidden state serves as the query (Q). In Bahdanau attention, the previous hidden state sₜ₋₁ becomes the query, while Luong attention uses the current state sₜ. This query vector represents the decoder's current information needs and will be matched against all encoder hidden states.

Step 2: Scoring Function (Additive vs. Multiplicative)

The scoring function computes a scalar alignment score e_{t,i} between the query and each encoder hidden state hᵢ (the key K). The repository implements two distinct scoring paradigms:

Additive (Bahdanau) Attention

This approach uses a small feed-forward network to project both query and key into a shared attention space:


e_{t,i} = vᵀ tanh(Wₐ Q + Uₐ Kᵢ)

Here, Wₐ and Uₐ are learned weight matrices, and v is a learned attention vector. The additive_attention function in code/main.py implements this by projecting the decoder state and each encoder state separately, combining them with tanh, then applying the attention vector.

Multiplicative (Luong) Attention

This simpler approach uses direct vector operations with three variants:

  • Dot: e_{t,i} = Qᵀ Kᵢ (requires matching dimensions)
  • General: e_{t,i} = Qᵀ W Kᵢ (learns a linear projection via matrix W)
  • Concat: Mirrors additive form but is computationally expensive and rarely used

The dot_attention function in code/main.py implements the dot-product variant using a simple sum of element-wise products.

Step 3: Weighting and Context Vector Generation

The scores undergo softmax normalization to produce attention weights α_{t,i} that sum to 1. The final context vector cₜ is computed as the weighted sum of encoder states:


cₜ = Σᵢ α_{t,i} hᵢ

The decoder then consumes this context vector—often concatenated with the previous output—to generate the next token in the sequence.

Why Attention Solves the Sequence Bottleneck

Without attention, encoder-decoder architectures compress entire input sequences into a single fixed-size vector. This creates an information bottleneck as sequence length increases, forcing the model to lose nuance from earlier tokens.

By recomputing a context vector at every decoding step, the attention mechanism enables dynamic alignments. The decoder can "look" at different source positions as needed, dramatically improving performance on translation, summarization, and other sequence-to-sequence tasks. The docs/en.md file in the repository provides detailed pedagogical notes on how this soft alignment emerges during training.

From Classical Attention to Self-Attention and Transformers

In classical encoder-decoder attention, keys and values are identical—both refer to the encoder hidden states. Self-attention generalizes this concept by allowing query, key, and value to be three separate linear projections of the same input sequence. This enables each token to attend to every other token in the sequence.

Multi-head attention extends this further by running the single-head attention process in parallel multiple times with different learned projections, creating the foundation for the Transformer architecture. The repository's assets/attention.svg provides a visual illustration of how Bahdanau attention flows between encoder and decoder states.

Hands-On Implementation in Python

The code/main.py file contains a minimal NumPy-free implementation that demonstrates both scoring methods. Below are the core functions with a runnable demonstration:

import math

def softmax(scores):
    m = max(scores)
    exps = [math.exp(s - m) for s in scores]
    total = sum(exps)
    return [e / total for e in exps]

def dot_attention(decoder_state, encoder_states):
    """Luong dot-product attention"""
    scores = [sum(q * k for q, k in zip(decoder_state, h)) 
              for h in encoder_states]
    weights = softmax(scores)
    dim = len(encoder_states[0])
    context = [0.0] * dim
    for w, h in zip(weights, encoder_states):
        for d in range(dim):
            context[d] += w * h[d]
    return context, weights

def additive_attention(decoder_state, encoder_states, W_a, U_a, v_a):
    """Bahdanau additive attention"""
    projected_dec = [sum(W_a[i][j] * decoder_state[j] 
                        for j in range(len(decoder_state))) 
                    for i in range(len(v_a))]
    scores = []
    for h in encoder_states:
        projected_enc = [sum(U_a[i][j] * h[j] 
                            for j in range(len(h))) 
                        for i in range(len(v_a))]
        combined = [math.tanh(projected_enc[i] + projected_dec[i]) 
                   for i in range(len(v_a))]
        scores.append(sum(v_a[i] * combined[i] 
                         for i in range(len(v_a))))
    weights = softmax(scores)
    dim = len(encoder_states[0])
    context = [0.0] * dim
    for w, h in zip(weights, encoder_states):
        for d in range(dim):
            context[d] += w * h[d]
    return context, weights

# Demonstration

H = [
    [1.0, 0.0, 0.2],
    [0.5, 0.5, 0.1],
    [0.1, 0.9, 0.3],
]

s = [0.9, 0.1, 0.2]  # Query similar to first token

c_dot, w_dot = dot_attention(s, H)
print("Dot attention weights:", [round(x, 3) for x in w_dot])

W_a = [[0.6, 0.3, 0.1], [0.1, 0.5, 0.4]]
U_a = [[0.5, 0.2, 0.3], [0.2, 0.6, 0.2]]
v_a = [0.8, 0.6]
c_add, w_add = additive_attention(s, H, W_a, U_a, v_a)
print("Additive attention weights:", [round(x, 3) for x in w_add])

Running this script produces the following weight distributions:

Dot attention weights: [0.464, 0.305, 0.231]
Additive attention weights: [0.497, 0.286, 0.217]

These results demonstrate how the query's similarity to each encoder state determines attention allocation. The first encoder state receives the highest weight because the query vector [0.9, 0.1, 0.2] closely resembles the first hidden state [1.0, 0.0, 0.2].

Summary

  • Attention mechanisms eliminate the encoder bottleneck by computing dynamic context vectors for each decoding step.
  • Bahdanau (additive) attention uses a feed-forward network to score alignment, while Luong (multiplicative) attention relies on direct vector products.
  • The three-step pipeline involves query creation, score computation (dot, general, or additive), and weighted context generation via softmax normalization.
  • Self-attention generalizes classical attention by separating queries, keys, and values into distinct linear projections, enabling the Transformer architecture.
  • Reference implementations in code/main.py provide working examples without external dependencies like NumPy.

Frequently Asked Questions

What is the difference between Bahdanau and Luong attention?

Bahdanau (additive) attention uses a learned feed-forward network to compute alignment scores via vᵀ tanh(Wₐ Q + Uₐ Kᵢ), allowing flexible comparison between query and key dimensions. Luong (multiplicative) attention computes scores through direct vector operations like Qᵀ Kᵢ or Qᵀ W Kᵢ, which is faster but requires dimension matching or learned projections. According to the repository's code/main.py, Bahdanau typically performs better with longer sequences, while Luong offers computational efficiency.

Why do attention weights always sum to 1?

Attention weights sum to 1 because they are generated by applying the softmax function to the raw alignment scores. This normalization ensures the context vector represents a weighted average of encoder states, maintaining stable gradients and interpretable probability distributions over the input sequence. The softmax function in the repository implementation subtracts the maximum score for numerical stability before exponentiation.

How does self-attention differ from classical encoder-decoder attention?

In classical encoder-decoder attention, the query comes from the decoder while keys and values are identical encoder hidden states. In self-attention, all three components—queries, keys, and values—are derived from the same sequence through separate linear projections, allowing each position to attend to all other positions in the same sequence. This mechanism, detailed in docs/en.md, enables the parallel processing that makes Transformers efficient.

Where can I find the reference implementation files?

The primary implementation resides in code/main.py, which contains the dot_attention and additive_attention functions. Pedagogical explanations and shape tables are available in docs/en.md, while visual diagrams are stored in assets/attention.svg. The repository also includes quiz.json for testing comprehension of attention concepts and outputs/prompt-attention-shapes.md for debugging shape-related implementation issues.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →