# How to Implement Transformer Architecture from "Attention Is All You Need"

> Learn to implement the Transformer architecture from Attention Is All You Need. Build a stacked encoder-decoder model with multi-head self-attention and more.

- Repository: [Owain Lewis/awesome-artificial-intelligence](https://github.com/owainlewis/awesome-artificial-intelligence)
- Tags: tutorial
- Published: 2026-06-22

---

**Implement the Transformer by combining multi-head self-attention, position-wise feed-forward networks, residual connections with layer normalization, and sinusoidal positional encodings in a stacked encoder-decoder architecture.**

The Transformer architecture introduced in *Attention Is All You Need* revolutionized deep learning by replacing recurrence with fully attention-driven mechanisms. According to the owainlewis/awesome-artificial-intelligence repository, this paper appears in the "Landmark Papers" section of the [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md), serving as the primary reference for modern sequence-to-sequence models. To implement Transformer architecture from scratch, you must replicate five core mechanisms that enable parallel processing and long-range dependency modeling.

## Core Components of the Transformer Architecture

The original paper specifies exact hyperparameters: 8 attention heads, model dimension **d_model** of 512, feed-forward dimension **d_ff** of 2048, and dropout rate of 0.1. These components work together to process input sequences without recurrent connections.

### Multi-Head Self-Attention

**Multi-head self-attention** projects input sequences into queries, keys, and values across multiple parallel heads. Each head computes scaled dot-product attention using `softmax(QK^T / sqrt(d_k))`, allowing the model to attend to different representation subspaces simultaneously.

### Position-Wise Feed-Forward Networks

After attention, each position passes through identical **position-wise feed-forward networks** consisting of two linear transformations with ReLU activation. This shared MLP adds non-linearity while maintaining position independence.

### Add & Norm

**Residual connections** wrap both attention and feed-forward sub-layers, followed by **layer normalization** to stabilize gradients during training.

### Positional Encoding

Since attention mechanisms lack inherent sequence ordering, **sinusoidal positional encodings** inject position information into input embeddings using sine and cosine functions of varying frequencies.

## Implementing Multi-Head Self-Attention

The `MultiHeadAttention` class projects inputs into queries, keys, and values, then splits them into parallel heads. The implementation follows the scaled dot-product attention formula from the paper.

```python
import torch
import torch.nn as nn
import math

class MultiHeadAttention(nn.Module):
    def __init__(self, d_model, n_head, dropout=0.1):
        super().__init__()
        assert d_model % n_head == 0
        self.d_k = d_model // n_head
        self.n_head = n_head

        self.q_linear = nn.Linear(d_model, d_model)
        self.k_linear = nn.Linear(d_model, d_model)
        self.v_linear = nn.Linear(d_model, d_model)
        self.out_linear = nn.Linear(d_model, d_model)

        self.dropout = nn.Dropout(dropout)

    def forward(self, query, key, value, mask=None):
        batch_sz = query.size(0)

        # Linear projection + split into heads

        q = self.q_linear(query).view(batch_sz, -1, self.n_head, self.d_k).transpose(1, 2)
        k = self.k_linear(key).view(batch_sz, -1, self.n_head, self.d_k).transpose(1, 2)
        v = self.v_linear(value).view(batch_sz, -1, self.n_head, self.d_k).transpose(1, 2)

        # Scaled dot-product attention

        scores = torch.matmul(q, k.transpose(-2, -1)) / math.sqrt(self.d_k)
        if mask is not None:
            scores = scores.masked_fill(mask == 0, float('-inf'))
        attn = torch.softmax(scores, dim=-1)
        attn = self.dropout(attn)

        # Combine heads

        context = torch.matmul(attn, v).transpose(1, 2).contiguous()
        context = context.view(batch_sz, -1, self.n_head * self.d_k)
        output = self.out_linear(context)
        return output

```

## Position-Wise Feed-Forward Networks

The `PositionwiseFeedForward` class implements the two-layer MLP described in the paper. It expands the model dimension to 2048 (d_ff) before projecting back to 512 (d_model).

```python
class PositionwiseFeedForward(nn.Module):
    def __init__(self, d_model, d_ff, dropout=0.1):
        super().__init__()
        self.fc1 = nn.Linear(d_model, d_ff)
        self.fc2 = nn.Linear(d_ff, d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, x):
        return self.fc2(self.dropout(nn.functional.relu(self.fc1(x))))

```

## Positional Encoding for Sequence Order

The `PositionalEncoding` class generates fixed sinusoidal embeddings that align with the mathematical formulation in the original paper. These encodings are added to input token embeddings before the first attention layer.

```python
class PositionalEncoding(nn.Module):
    def __init__(self, d_model, max_len=5000):
        super().__init__()
        pe = torch.zeros(max_len, d_model)
        pos = torch.arange(0, max_len).unsqueeze(1).float()
        div_term = torch.exp(torch.arange(0, d_model, 2).float() *
                             -(math.log(10000.0) / d_model))
        pe[:, 0::2] = torch.sin(pos * div_term)
        pe[:, 1::2] = torch.cos(pos * div_term)
        self.register_buffer('pe', pe.unsqueeze(0))

    def forward(self, x):
        return x + self.pe[:, :x.size(1)]

```

## Building the Encoder and Decoder Stacks

The encoder and decoder each consist of N=6 identical layers. The `EncoderLayer` applies self-attention followed by feed-forward processing, while the `DecoderLayer` adds masked self-attention and encoder-decoder cross-attention.

### Encoder Layer Implementation

```python
class EncoderLayer(nn.Module):
    def __init__(self, d_model, n_head, d_ff, dropout=0.1):
        super().__init__()
        self.self_attn = MultiHeadAttention(d_model, n_head, dropout)
        self.ff = PositionwiseFeedForward(d_model, d_ff, dropout)
        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, src, src_mask=None):
        src2 = self.self_attn(src, src, src, src_mask)
        src = self.norm1(src + self.dropout(src2))
        src2 = self.ff(src)
        src = self.norm2(src + self.dropout(src2))
        return src

```

### Decoder Layer Implementation

The decoder includes an additional cross-attention mechanism that allows each position to attend over all encoder outputs.

```python
class DecoderLayer(nn.Module):
    def __init__(self, d_model, n_head, d_ff, dropout=0.1):
        super().__init__()
        self.self_attn = MultiHeadAttention(d_model, n_head, dropout)
        self.enc_attn = MultiHeadAttention(d_model, n_head, dropout)
        self.ff = PositionwiseFeedForward(d_model, d_ff, dropout)
        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)
        self.norm3 = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, tgt, memory, tgt_mask=None, memory_mask=None):
        tgt2 = self.self_attn(tgt, tgt, tgt, tgt_mask)
        tgt = self.norm1(tgt + self.dropout(tgt2))
        tgt2 = self.enc_attn(tgt, memory, memory, memory_mask)
        tgt = self.norm2(tgt + self.dropout(tgt2))
        tgt2 = self.ff(tgt)
        tgt = self.norm3(tgt + self.dropout(tgt2))
        return tgt

```

## Putting It All Together: The Complete Transformer

The `Transformer` class assembles the full architecture by stacking six encoder and decoder layers, adding embeddings scaled by `sqrt(d_model)`, and applying the final linear projection to vocabulary size.

```python
class Transformer(nn.Module):
    def __init__(self, vocab_sz, d_model=512, n_head=8,
                 num_enc_layers=6, num_dec_layers=6,
                 d_ff=2048, dropout=0.1, max_len=5000):
        super().__init__()
        self.embedding = nn.Embedding(vocab_sz, d_model)
        self.pos_encoder = PositionalEncoding(d_model, max_len)

        self.encoder = nn.ModuleList(
            [EncoderLayer(d_model, n_head, d_ff, dropout) for _ in range(num_enc_layers)]
        )
        self.decoder = nn.ModuleList(
            [DecoderLayer(d_model, n_head, d_ff, dropout) for _ in range(num_dec_layers)]
        )
        self.out_proj = nn.Linear(d_model, vocab_sz)

    def forward(self, src, tgt, src_mask=None, tgt_mask=None, memory_mask=None):
        src = self.embedding(src) * math.sqrt(self.embedding.embedding_dim)
        src = self.pos_encoder(src)

        for layer in self.encoder:
            src = layer(src, src_mask)

        tgt = self.embedding(tgt) * math.sqrt(self.embedding.embedding_dim)
        tgt = self.pos_encoder(tgt)

        for layer in self.decoder:
            tgt = layer(tgt, src, tgt_mask, memory_mask)

        return self.out_proj(tgt)

```

## Summary

- **Multi-head attention** splits queries, keys, and values into parallel heads to capture diverse contextual relationships.
- **Position-wise feed-forward networks** apply shared two-layer MLPs to each position independently.
- **Residual connections** and **layer normalization** stabilize training by wrapping sub-layers and normalizing activations.
- **Sinusoidal positional encodings** inject sequence order information into the model.
- The **encoder-decoder stack** processes source sequences through six identical encoder layers, then generates target sequences using six decoder layers with masked self-attention and cross-attention.

## Frequently Asked Questions

### What are the exact hyperparameters used in the original "Attention Is All You Need" paper?

The original implementation uses 8 attention heads, a model dimension (**d_model**) of 512, a feed-forward dimension (**d_ff**) of 2048, and a dropout rate of 0.1. Both the encoder and decoder contain 6 stacked layers. These specifications appear in the [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md) of the owainlewis/awesome-artificial-intelligence repository under the "Landmark Papers" section.

### How does the Transformer handle sequence order without recurrence?

The Transformer handles sequence order through **positional encodings** added to input embeddings. The implementation uses sinusoidal functions with varying frequencies to create unique position representations, allowing the model to learn relative positions while maintaining the parallel processing benefits of attention mechanisms.

### Why is the attention scaled by the square root of the key dimension?

The attention scores are divided by `sqrt(d_k)` to prevent the dot products from growing too large in magnitude, which would push the softmax function into regions with extremely small gradients. This **scaled dot-product attention** maintains stable gradient flow during training, particularly when dealing with high-dimensional key vectors.

### What is the difference between the encoder and decoder attention mechanisms?

The encoder uses standard **self-attention** where each position attends to all other positions in the input sequence. The decoder employs **masked self-attention** that prevents positions from attending to subsequent positions during training, plus an additional **encoder-decoder attention** layer that allows the decoder to attend over the full encoder output sequence.