# Understanding the Attention Mechanism in Deep Learning: A Code-First Guide

> Master attention mechanisms in deep learning with this code-first guide. Learn how neural networks focus on input sequences to create powerful context vectors.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-19

---

**Attention mechanisms allow neural networks to dynamically focus on relevant parts of input sequences by computing a weighted context vector for each decoding step, eliminating the bottleneck of fixed-size representations.**

The **attention mechanism in deep learning** revolutionized sequence modeling by enabling models to learn soft alignments between inputs and outputs. This article examines the canonical implementation found in the `rohitg00/ai-engineering-from-scratch` repository, walking through the mathematical foundations and executable Python code that demonstrate both Bahdanau (additive) and Luong (multiplicative) attention.

## How the Attention Mechanism Works: The Three-Step Pipeline

Classical attention operates as a three-stage pipeline that converts a decoder's query into a context vector. According to the source code in [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py), this process involves creating a query, scoring alignment with encoder states, and generating a weighted context.

### Step 1: Query Creation

The decoder's current hidden state serves as the **query** (`Q`). In Bahdanau attention, the previous hidden state `sₜ₋₁` becomes the query, while Luong attention uses the current state `sₜ`. This query vector represents the decoder's current information needs and will be matched against all encoder hidden states.

### Step 2: Scoring Function (Additive vs. Multiplicative)

The scoring function computes a scalar alignment score `e_{t,i}` between the query and each encoder hidden state `hᵢ` (the **key** `K`). The repository implements two distinct scoring paradigms:

**Additive (Bahdanau) Attention**

This approach uses a small feed-forward network to project both query and key into a shared attention space:

```

e_{t,i} = vᵀ tanh(Wₐ Q + Uₐ Kᵢ)

```

Here, `Wₐ` and `Uₐ` are learned weight matrices, and `v` is a learned attention vector. The `additive_attention` function in [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py) implements this by projecting the decoder state and each encoder state separately, combining them with `tanh`, then applying the attention vector.

**Multiplicative (Luong) Attention**

This simpler approach uses direct vector operations with three variants:

- **Dot**: `e_{t,i} = Qᵀ Kᵢ` (requires matching dimensions)
- **General**: `e_{t,i} = Qᵀ W Kᵢ` (learns a linear projection via matrix `W`)
- **Concat**: Mirrors additive form but is computationally expensive and rarely used

The `dot_attention` function in [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py) implements the dot-product variant using a simple sum of element-wise products.

### Step 3: Weighting and Context Vector Generation

The scores undergo **softmax normalization** to produce attention weights `α_{t,i}` that sum to 1. The final **context vector** `cₜ` is computed as the weighted sum of encoder states:

```

cₜ = Σᵢ α_{t,i} hᵢ

```

The decoder then consumes this context vector—often concatenated with the previous output—to generate the next token in the sequence.

## Why Attention Solves the Sequence Bottleneck

Without attention, encoder-decoder architectures compress entire input sequences into a single fixed-size vector. This creates an **information bottleneck** as sequence length increases, forcing the model to lose nuance from earlier tokens.

By recomputing a context vector at every decoding step, the attention mechanism enables **dynamic alignments**. The decoder can "look" at different source positions as needed, dramatically improving performance on translation, summarization, and other sequence-to-sequence tasks. The [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md) file in the repository provides detailed pedagogical notes on how this soft alignment emerges during training.

## From Classical Attention to Self-Attention and Transformers

In classical encoder-decoder attention, **keys and values are identical**—both refer to the encoder hidden states. **Self-attention** generalizes this concept by allowing query, key, and value to be three separate linear projections of the same input sequence. This enables each token to attend to every other token in the sequence.

**Multi-head attention** extends this further by running the single-head attention process in parallel multiple times with different learned projections, creating the foundation for the Transformer architecture. The repository's `assets/attention.svg` provides a visual illustration of how Bahdanau attention flows between encoder and decoder states.

## Hands-On Implementation in Python

The [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py) file contains a minimal NumPy-free implementation that demonstrates both scoring methods. Below are the core functions with a runnable demonstration:

```python
import math

def softmax(scores):
    m = max(scores)
    exps = [math.exp(s - m) for s in scores]
    total = sum(exps)
    return [e / total for e in exps]

def dot_attention(decoder_state, encoder_states):
    """Luong dot-product attention"""
    scores = [sum(q * k for q, k in zip(decoder_state, h)) 
              for h in encoder_states]
    weights = softmax(scores)
    dim = len(encoder_states[0])
    context = [0.0] * dim
    for w, h in zip(weights, encoder_states):
        for d in range(dim):
            context[d] += w * h[d]
    return context, weights

def additive_attention(decoder_state, encoder_states, W_a, U_a, v_a):
    """Bahdanau additive attention"""
    projected_dec = [sum(W_a[i][j] * decoder_state[j] 
                        for j in range(len(decoder_state))) 
                    for i in range(len(v_a))]
    scores = []
    for h in encoder_states:
        projected_enc = [sum(U_a[i][j] * h[j] 
                            for j in range(len(h))) 
                        for i in range(len(v_a))]
        combined = [math.tanh(projected_enc[i] + projected_dec[i]) 
                   for i in range(len(v_a))]
        scores.append(sum(v_a[i] * combined[i] 
                         for i in range(len(v_a))))
    weights = softmax(scores)
    dim = len(encoder_states[0])
    context = [0.0] * dim
    for w, h in zip(weights, encoder_states):
        for d in range(dim):
            context[d] += w * h[d]
    return context, weights

# Demonstration

H = [
    [1.0, 0.0, 0.2],
    [0.5, 0.5, 0.1],
    [0.1, 0.9, 0.3],
]

s = [0.9, 0.1, 0.2]  # Query similar to first token

c_dot, w_dot = dot_attention(s, H)
print("Dot attention weights:", [round(x, 3) for x in w_dot])

W_a = [[0.6, 0.3, 0.1], [0.1, 0.5, 0.4]]
U_a = [[0.5, 0.2, 0.3], [0.2, 0.6, 0.2]]
v_a = [0.8, 0.6]
c_add, w_add = additive_attention(s, H, W_a, U_a, v_a)
print("Additive attention weights:", [round(x, 3) for x in w_add])

```

Running this script produces the following weight distributions:

```text
Dot attention weights: [0.464, 0.305, 0.231]
Additive attention weights: [0.497, 0.286, 0.217]

```

These results demonstrate how the query's similarity to each encoder state determines attention allocation. The first encoder state receives the highest weight because the query vector `[0.9, 0.1, 0.2]` closely resembles the first hidden state `[1.0, 0.0, 0.2]`.

## Summary

- **Attention mechanisms** eliminate the encoder bottleneck by computing dynamic context vectors for each decoding step.
- **Bahdanau (additive)** attention uses a feed-forward network to score alignment, while **Luong (multiplicative)** attention relies on direct vector products.
- The three-step pipeline involves query creation, score computation (dot, general, or additive), and weighted context generation via softmax normalization.
- **Self-attention** generalizes classical attention by separating queries, keys, and values into distinct linear projections, enabling the Transformer architecture.
- Reference implementations in [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py) provide working examples without external dependencies like NumPy.

## Frequently Asked Questions

### What is the difference between Bahdanau and Luong attention?

**Bahdanau (additive) attention** uses a learned feed-forward network to compute alignment scores via `vᵀ tanh(Wₐ Q + Uₐ Kᵢ)`, allowing flexible comparison between query and key dimensions. **Luong (multiplicative) attention** computes scores through direct vector operations like `Qᵀ Kᵢ` or `Qᵀ W Kᵢ`, which is faster but requires dimension matching or learned projections. According to the repository's [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py), Bahdanau typically performs better with longer sequences, while Luong offers computational efficiency.

### Why do attention weights always sum to 1?

Attention weights sum to 1 because they are generated by applying the **softmax function** to the raw alignment scores. This normalization ensures the context vector represents a weighted average of encoder states, maintaining stable gradients and interpretable probability distributions over the input sequence. The `softmax` function in the repository implementation subtracts the maximum score for numerical stability before exponentiation.

### How does self-attention differ from classical encoder-decoder attention?

In classical **encoder-decoder attention**, the query comes from the decoder while keys and values are identical encoder hidden states. In **self-attention**, all three components—queries, keys, and values—are derived from the same sequence through separate linear projections, allowing each position to attend to all other positions in the same sequence. This mechanism, detailed in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md), enables the parallel processing that makes Transformers efficient.

### Where can I find the reference implementation files?

The primary implementation resides in [`code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/code/main.py), which contains the `dot_attention` and `additive_attention` functions. Pedagogical explanations and shape tables are available in [`docs/en.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/docs/en.md), while visual diagrams are stored in `assets/attention.svg`. The repository also includes [`quiz.json`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/quiz.json) for testing comprehension of attention concepts and [`outputs/prompt-attention-shapes.md`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/outputs/prompt-attention-shapes.md) for debugging shape-related implementation issues.