How to Implement Transformer Architecture from "Attention Is All You Need"

Implement the Transformer by combining multi-head self-attention, position-wise feed-forward networks, residual connections with layer normalization, and sinusoidal positional encodings in a stacked encoder-decoder architecture.

The Transformer architecture introduced in Attention Is All You Need revolutionized deep learning by replacing recurrence with fully attention-driven mechanisms. According to the owainlewis/awesome-artificial-intelligence repository, this paper appears in the "Landmark Papers" section of the README.md, serving as the primary reference for modern sequence-to-sequence models. To implement Transformer architecture from scratch, you must replicate five core mechanisms that enable parallel processing and long-range dependency modeling.

Core Components of the Transformer Architecture

The original paper specifies exact hyperparameters: 8 attention heads, model dimension d_model of 512, feed-forward dimension d_ff of 2048, and dropout rate of 0.1. These components work together to process input sequences without recurrent connections.

Multi-Head Self-Attention

Multi-head self-attention projects input sequences into queries, keys, and values across multiple parallel heads. Each head computes scaled dot-product attention using softmax(QK^T / sqrt(d_k)), allowing the model to attend to different representation subspaces simultaneously.

Position-Wise Feed-Forward Networks

After attention, each position passes through identical position-wise feed-forward networks consisting of two linear transformations with ReLU activation. This shared MLP adds non-linearity while maintaining position independence.

Add & Norm

Residual connections wrap both attention and feed-forward sub-layers, followed by layer normalization to stabilize gradients during training.

Positional Encoding

Since attention mechanisms lack inherent sequence ordering, sinusoidal positional encodings inject position information into input embeddings using sine and cosine functions of varying frequencies.

Implementing Multi-Head Self-Attention

The MultiHeadAttention class projects inputs into queries, keys, and values, then splits them into parallel heads. The implementation follows the scaled dot-product attention formula from the paper.

import torch
import torch.nn as nn
import math

class MultiHeadAttention(nn.Module):
    def __init__(self, d_model, n_head, dropout=0.1):
        super().__init__()
        assert d_model % n_head == 0
        self.d_k = d_model // n_head
        self.n_head = n_head

        self.q_linear = nn.Linear(d_model, d_model)
        self.k_linear = nn.Linear(d_model, d_model)
        self.v_linear = nn.Linear(d_model, d_model)
        self.out_linear = nn.Linear(d_model, d_model)

        self.dropout = nn.Dropout(dropout)

    def forward(self, query, key, value, mask=None):
        batch_sz = query.size(0)

        # Linear projection + split into heads

        q = self.q_linear(query).view(batch_sz, -1, self.n_head, self.d_k).transpose(1, 2)
        k = self.k_linear(key).view(batch_sz, -1, self.n_head, self.d_k).transpose(1, 2)
        v = self.v_linear(value).view(batch_sz, -1, self.n_head, self.d_k).transpose(1, 2)

        # Scaled dot-product attention

        scores = torch.matmul(q, k.transpose(-2, -1)) / math.sqrt(self.d_k)
        if mask is not None:
            scores = scores.masked_fill(mask == 0, float('-inf'))
        attn = torch.softmax(scores, dim=-1)
        attn = self.dropout(attn)

        # Combine heads

        context = torch.matmul(attn, v).transpose(1, 2).contiguous()
        context = context.view(batch_sz, -1, self.n_head * self.d_k)
        output = self.out_linear(context)
        return output

Position-Wise Feed-Forward Networks

The PositionwiseFeedForward class implements the two-layer MLP described in the paper. It expands the model dimension to 2048 (d_ff) before projecting back to 512 (d_model).

class PositionwiseFeedForward(nn.Module):
    def __init__(self, d_model, d_ff, dropout=0.1):
        super().__init__()
        self.fc1 = nn.Linear(d_model, d_ff)
        self.fc2 = nn.Linear(d_ff, d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, x):
        return self.fc2(self.dropout(nn.functional.relu(self.fc1(x))))

Positional Encoding for Sequence Order

The PositionalEncoding class generates fixed sinusoidal embeddings that align with the mathematical formulation in the original paper. These encodings are added to input token embeddings before the first attention layer.

class PositionalEncoding(nn.Module):
    def __init__(self, d_model, max_len=5000):
        super().__init__()
        pe = torch.zeros(max_len, d_model)
        pos = torch.arange(0, max_len).unsqueeze(1).float()
        div_term = torch.exp(torch.arange(0, d_model, 2).float() *
                             -(math.log(10000.0) / d_model))
        pe[:, 0::2] = torch.sin(pos * div_term)
        pe[:, 1::2] = torch.cos(pos * div_term)
        self.register_buffer('pe', pe.unsqueeze(0))

    def forward(self, x):
        return x + self.pe[:, :x.size(1)]

Building the Encoder and Decoder Stacks

The encoder and decoder each consist of N=6 identical layers. The EncoderLayer applies self-attention followed by feed-forward processing, while the DecoderLayer adds masked self-attention and encoder-decoder cross-attention.

Encoder Layer Implementation

class EncoderLayer(nn.Module):
    def __init__(self, d_model, n_head, d_ff, dropout=0.1):
        super().__init__()
        self.self_attn = MultiHeadAttention(d_model, n_head, dropout)
        self.ff = PositionwiseFeedForward(d_model, d_ff, dropout)
        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, src, src_mask=None):
        src2 = self.self_attn(src, src, src, src_mask)
        src = self.norm1(src + self.dropout(src2))
        src2 = self.ff(src)
        src = self.norm2(src + self.dropout(src2))
        return src

Decoder Layer Implementation

The decoder includes an additional cross-attention mechanism that allows each position to attend over all encoder outputs.

class DecoderLayer(nn.Module):
    def __init__(self, d_model, n_head, d_ff, dropout=0.1):
        super().__init__()
        self.self_attn = MultiHeadAttention(d_model, n_head, dropout)
        self.enc_attn = MultiHeadAttention(d_model, n_head, dropout)
        self.ff = PositionwiseFeedForward(d_model, d_ff, dropout)
        self.norm1 = nn.LayerNorm(d_model)
        self.norm2 = nn.LayerNorm(d_model)
        self.norm3 = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, tgt, memory, tgt_mask=None, memory_mask=None):
        tgt2 = self.self_attn(tgt, tgt, tgt, tgt_mask)
        tgt = self.norm1(tgt + self.dropout(tgt2))
        tgt2 = self.enc_attn(tgt, memory, memory, memory_mask)
        tgt = self.norm2(tgt + self.dropout(tgt2))
        tgt2 = self.ff(tgt)
        tgt = self.norm3(tgt + self.dropout(tgt2))
        return tgt

Putting It All Together: The Complete Transformer

The Transformer class assembles the full architecture by stacking six encoder and decoder layers, adding embeddings scaled by sqrt(d_model), and applying the final linear projection to vocabulary size.

class Transformer(nn.Module):
    def __init__(self, vocab_sz, d_model=512, n_head=8,
                 num_enc_layers=6, num_dec_layers=6,
                 d_ff=2048, dropout=0.1, max_len=5000):
        super().__init__()
        self.embedding = nn.Embedding(vocab_sz, d_model)
        self.pos_encoder = PositionalEncoding(d_model, max_len)

        self.encoder = nn.ModuleList(
            [EncoderLayer(d_model, n_head, d_ff, dropout) for _ in range(num_enc_layers)]
        )
        self.decoder = nn.ModuleList(
            [DecoderLayer(d_model, n_head, d_ff, dropout) for _ in range(num_dec_layers)]
        )
        self.out_proj = nn.Linear(d_model, vocab_sz)

    def forward(self, src, tgt, src_mask=None, tgt_mask=None, memory_mask=None):
        src = self.embedding(src) * math.sqrt(self.embedding.embedding_dim)
        src = self.pos_encoder(src)

        for layer in self.encoder:
            src = layer(src, src_mask)

        tgt = self.embedding(tgt) * math.sqrt(self.embedding.embedding_dim)
        tgt = self.pos_encoder(tgt)

        for layer in self.decoder:
            tgt = layer(tgt, src, tgt_mask, memory_mask)

        return self.out_proj(tgt)

Summary

  • Multi-head attention splits queries, keys, and values into parallel heads to capture diverse contextual relationships.
  • Position-wise feed-forward networks apply shared two-layer MLPs to each position independently.
  • Residual connections and layer normalization stabilize training by wrapping sub-layers and normalizing activations.
  • Sinusoidal positional encodings inject sequence order information into the model.
  • The encoder-decoder stack processes source sequences through six identical encoder layers, then generates target sequences using six decoder layers with masked self-attention and cross-attention.

Frequently Asked Questions

What are the exact hyperparameters used in the original "Attention Is All You Need" paper?

The original implementation uses 8 attention heads, a model dimension (d_model) of 512, a feed-forward dimension (d_ff) of 2048, and a dropout rate of 0.1. Both the encoder and decoder contain 6 stacked layers. These specifications appear in the README.md of the owainlewis/awesome-artificial-intelligence repository under the "Landmark Papers" section.

How does the Transformer handle sequence order without recurrence?

The Transformer handles sequence order through positional encodings added to input embeddings. The implementation uses sinusoidal functions with varying frequencies to create unique position representations, allowing the model to learn relative positions while maintaining the parallel processing benefits of attention mechanisms.

Why is the attention scaled by the square root of the key dimension?

The attention scores are divided by sqrt(d_k) to prevent the dot products from growing too large in magnitude, which would push the softmax function into regions with extremely small gradients. This scaled dot-product attention maintains stable gradient flow during training, particularly when dealing with high-dimensional key vectors.

What is the difference between the encoder and decoder attention mechanisms?

The encoder uses standard self-attention where each position attends to all other positions in the input sequence. The decoder employs masked self-attention that prevents positions from attending to subsequent positions during training, plus an additional encoder-decoder attention layer that allows the decoder to attend over the full encoder output sequence.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →