Differences Between Transformer Architectures: Encoder, Decoder, and Encoder-Decoder Explained

Encoder-only models use bidirectional self-attention to build contextual representations, decoder-only models apply causal masking for autoregressive text generation, and encoder-decoder models combine both via cross-attention for sequence-to-sequence tasks.

Transformers have become the dominant architecture in modern AI, but implementations vary dramatically based on design goals. The rohitg00/ai-engineering-from-scratch repository provides hands-on implementations that illuminate the key differences between transformer architectures and their impact on training objectives and data flow.

Architectural Overview: Three Canonical Shapes

Transformers fundamentally differ in how they apply attention masks and process sequences. The following comparison summarizes the distinctions:

Feature Encoder-Only Decoder-Only Encoder-Decoder
Typical Use-Case Representation learning (e.g., BERT) Autoregressive generation (e.g., GPT) Sequence-to-sequence tasks (e.g., translation)
Attention Mask Bidirectional (no causal mask) Causal (lower-triangular) Encoder: bidirectional; Decoder: causal + cross-attention
Core Blocks encoder_block (self-attention + FFN) decoder_block with masked self-attention Stacks of encoder_block followed by decoder_block
Training Objective Masked language modeling (MLM) Next-token prediction Span corruption or denoising
Parameter Sharing Single embedding layer Token embedding tied to output head Separate encoder/decoder embeddings

Encoder-Only Architectures (BERT-Style)

Encoder-only models excel at understanding and representing text. In phases/07-transformers-deep-dive/06-bert-masked-language-modeling/code/main.py, the implementation demonstrates bidirectional attention, where each token attends to all other tokens in the sequence—both past and future.

This architecture trains via masked language modeling (MLM), randomly masking 15% of input tokens and requiring the model to predict the originals. Because there is no causal restriction, the final hidden states capture rich contextual embeddings suitable for classification or named entity recognition.


# From phases/07-transformers-deep-dive/06-bert-masked-language-modeling/code/main.py

tokens = [1, 5, 9, 2]            # example token ids

vocab_size = 1000
inp, labels = create_mlm_batch(tokens, vocab_size, mask_prob=0.15)
print("input ids:", inp)
print("labels:", labels)

The create_mlm_batch function prepares training data by selecting random tokens for prediction, allowing the multi_head_attention layers to attend bidirectionally without causal=True.

Decoder-Only Architectures (GPT-Style)

Decoder-only models are optimized for autoregressive generation. As implemented in phases/07-transformers-deep-dive/14-build-a-transformer-capstone/code/main.py, these models use causal self-attention where each position can only attend to previous positions via a lower-triangular mask created with torch.tril.

This constraint forces the model to predict the next token based solely on prior context, matching the requirements of open-ended text generation. The training objective is straightforward next-token prediction, and the architecture typically ties input token embeddings to the output language modeling head.


# Causal generation from phases/07-transformers-deep-dive/14-build-a-transformer-capstone/code/main.py

prompt = torch.tensor([[stoi["F"], stoi["i"], stoi["r"]]], device=device)
generated = model.generate(prompt, max_new_tokens=20, temperature=0.9, top_k=10)
print("generated text:", "".join(itos[i] for i in generated[0]))

The CausalSelfAttention class enforces the causal mask during the forward pass, ensuring no information leakage from future tokens during training or inference.

Encoder-Decoder Architectures (T5/BART-Style)

Encoder-decoder models handle sequence-to-sequence tasks by combining the strengths of both architectures. The encoder processes the source sequence with bidirectional attention, while the decoder generates the target sequence autoregressively using cross-attention to the encoder's final hidden states.

In phases/07-transformers-deep-dive/08-t5-bart-encoder-decoder/code/main.py, the implementation demonstrates span corruption, where random spans of the input are masked and the decoder must reconstruct the original text. The round_trip function coordinates this encode-decode cycle.


# Span corruption from phases/07-transformers-deep-dive/08-t5-bart-encoder-decoder/code/main.py

sentence = "the quick brown fox jumps over the lazy dog".split()
source, target = corrupt_spans(sentence, mask_rate=0.2)
reconstructed = round_trip(source, target)
print("reconstruction matches original:", reconstructed == sentence)

The cross-attention mechanism allows the decoder to query encoder outputs via the kv_source=enc_out parameter, as visible in the full transformer implementation.

The Full Transformer Implementation

The reference file phases/07-transformers-deep-dive/05-full-transformer/code/main.py demonstrates how both blocks interact in a complete system. The encoder_block processes input through self-attention and feed-forward networks, while the decoder_block first applies masked self-attention, then cross-attention to the encoder output before its own feed-forward layer.


# From phases/07-transformers-deep-dive/05-full-transformer/code/main.py

enc_out = src
for p in enc_params:
    enc_out = encoder_block(enc_out, p)

dec_out = tgt
for p in dec_params:
    dec_out = decoder_block(dec_out, enc_out, p)  # enc_out passed for cross-attention

This wiring illustrates how encoder-decoder models maintain separate pathways for encoding source context and generating target sequences, with the decoder block's additional cross-attention head serving as the critical link between them.

Summary

  • Encoder-only models employ bidirectional attention and MLM training to produce contextual embeddings, ideal for understanding tasks but incapable of autoregressive generation.
  • Decoder-only models use causal masking and next-token prediction to enable text generation, though they cannot leverage future context during training.
  • Encoder-decoder models combine bidirectional encoding with causal decoding and cross-attention, excelling at translation and summarization but requiring more compute.
  • Key implementation files in rohitg00/ai-engineering-from-scratch demonstrate these patterns through multi_head_attention configurations, corrupt_spans preprocessing, and explicit kv_source wiring in decoder_block functions.

Frequently Asked Questions

What is the fundamental difference between encoder and decoder transformer blocks?

Encoder blocks apply bidirectional self-attention where every token can see every other token, while decoder blocks use causal (masked) self-attention that restricts each token to view only previous positions. Additionally, decoder blocks in full transformer implementations include a cross-attention sublayer that queries encoder outputs, which encoder blocks lack entirely.

When should I use an encoder-decoder model instead of a decoder-only model?

Use encoder-decoder models when the task requires conditioning generation on a complete source sequence, such as machine translation or document summarization, where the encoder processes the full source text before the decoder begins generation. Decoder-only models suffice for unconditional generation or prompt-completion tasks where no separate source context needs encoding.

Why can't decoder-only models use bidirectional attention during training?

Decoder-only models must use causal masking (typically implemented with torch.tril) to prevent the model from attending to future tokens during training. If bidirectional attention were allowed, the model could cheat by looking at the target token it is supposed to predict, rendering the next-token prediction objective trivial and breaking the autoregressive property required for text generation.

How does cross-attention differ from self-attention in transformer architectures?

Self-attention computes queries, keys, and values from the same input sequence, allowing the model to relate different positions within that sequence. Cross-attention, as seen in phases/07-transformers-deep-dive/05-full-transformer/code/main.py, uses queries from the decoder but keys and values from the encoder output (kv_source=enc_out), enabling the decoder to retrieve relevant information from the encoded source sequence while generating each target token.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →