# Differences Between Transformer Architectures: Encoder, Decoder, and Encoder-Decoder Explained

> Understand the differences between Transformer encoder, decoder, and encoder-decoder architectures. Learn about their unique attention mechanisms and applications in NLP.

- Repository: [Rohit Ghumare/ai-engineering-from-scratch](https://github.com/rohitg00/ai-engineering-from-scratch)
- Tags: deep-dive
- Published: 2026-07-26

---

**Encoder-only models use bidirectional self-attention to build contextual representations, decoder-only models apply causal masking for autoregressive text generation, and encoder-decoder models combine both via cross-attention for sequence-to-sequence tasks.**

Transformers have become the dominant architecture in modern AI, but implementations vary dramatically based on design goals. The `rohitg00/ai-engineering-from-scratch` repository provides hands-on implementations that illuminate the key **differences between transformer architectures** and their impact on training objectives and data flow.

## Architectural Overview: Three Canonical Shapes

Transformers fundamentally differ in how they apply attention masks and process sequences. The following comparison summarizes the distinctions:

| Feature | Encoder-Only | Decoder-Only | Encoder-Decoder |
|---|---|---|---|
| **Typical Use-Case** | Representation learning (e.g., BERT) | Autoregressive generation (e.g., GPT) | Sequence-to-sequence tasks (e.g., translation) |
| **Attention Mask** | Bidirectional (no causal mask) | Causal (lower-triangular) | Encoder: bidirectional; Decoder: causal + cross-attention |
| **Core Blocks** | `encoder_block` (self-attention + FFN) | `decoder_block` with masked self-attention | Stacks of `encoder_block` followed by `decoder_block` |
| **Training Objective** | Masked language modeling (MLM) | Next-token prediction | Span corruption or denoising |
| **Parameter Sharing** | Single embedding layer | Token embedding tied to output head | Separate encoder/decoder embeddings |

## Encoder-Only Architectures (BERT-Style)

**Encoder-only models** excel at understanding and representing text. In [`phases/07-transformers-deep-dive/06-bert-masked-language-modeling/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/06-bert-masked-language-modeling/code/main.py), the implementation demonstrates **bidirectional attention**, where each token attends to all other tokens in the sequence—both past and future.

This architecture trains via **masked language modeling (MLM)**, randomly masking 15% of input tokens and requiring the model to predict the originals. Because there is no causal restriction, the final hidden states capture rich contextual embeddings suitable for classification or named entity recognition.

```python

# From phases/07-transformers-deep-dive/06-bert-masked-language-modeling/code/main.py

tokens = [1, 5, 9, 2]            # example token ids

vocab_size = 1000
inp, labels = create_mlm_batch(tokens, vocab_size, mask_prob=0.15)
print("input ids:", inp)
print("labels:", labels)

```

The `create_mlm_batch` function prepares training data by selecting random tokens for prediction, allowing the `multi_head_attention` layers to attend bidirectionally without `causal=True`.

## Decoder-Only Architectures (GPT-Style)

**Decoder-only models** are optimized for autoregressive generation. As implemented in [`phases/07-transformers-deep-dive/14-build-a-transformer-capstone/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/14-build-a-transformer-capstone/code/main.py), these models use **causal self-attention** where each position can only attend to previous positions via a lower-triangular mask created with `torch.tril`.

This constraint forces the model to predict the next token based solely on prior context, matching the requirements of open-ended text generation. The training objective is straightforward next-token prediction, and the architecture typically ties input token embeddings to the output language modeling head.

```python

# Causal generation from phases/07-transformers-deep-dive/14-build-a-transformer-capstone/code/main.py

prompt = torch.tensor([[stoi["F"], stoi["i"], stoi["r"]]], device=device)
generated = model.generate(prompt, max_new_tokens=20, temperature=0.9, top_k=10)
print("generated text:", "".join(itos[i] for i in generated[0]))

```

The `CausalSelfAttention` class enforces the causal mask during the forward pass, ensuring no information leakage from future tokens during training or inference.

## Encoder-Decoder Architectures (T5/BART-Style)

**Encoder-decoder models** handle sequence-to-sequence tasks by combining the strengths of both architectures. The encoder processes the source sequence with bidirectional attention, while the decoder generates the target sequence autoregressively using **cross-attention** to the encoder's final hidden states.

In [`phases/07-transformers-deep-dive/08-t5-bart-encoder-decoder/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/08-t5-bart-encoder-decoder/code/main.py), the implementation demonstrates **span corruption**, where random spans of the input are masked and the decoder must reconstruct the original text. The `round_trip` function coordinates this encode-decode cycle.

```python

# Span corruption from phases/07-transformers-deep-dive/08-t5-bart-encoder-decoder/code/main.py

sentence = "the quick brown fox jumps over the lazy dog".split()
source, target = corrupt_spans(sentence, mask_rate=0.2)
reconstructed = round_trip(source, target)
print("reconstruction matches original:", reconstructed == sentence)

```

The cross-attention mechanism allows the decoder to query encoder outputs via the `kv_source=enc_out` parameter, as visible in the full transformer implementation.

## The Full Transformer Implementation

The reference file [`phases/07-transformers-deep-dive/05-full-transformer/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/05-full-transformer/code/main.py) demonstrates how both blocks interact in a complete system. The `encoder_block` processes input through self-attention and feed-forward networks, while the `decoder_block` first applies masked self-attention, then **cross-attention** to the encoder output before its own feed-forward layer.

```python

# From phases/07-transformers-deep-dive/05-full-transformer/code/main.py

enc_out = src
for p in enc_params:
    enc_out = encoder_block(enc_out, p)

dec_out = tgt
for p in dec_params:
    dec_out = decoder_block(dec_out, enc_out, p)  # enc_out passed for cross-attention

```

This wiring illustrates how encoder-decoder models maintain separate pathways for encoding source context and generating target sequences, with the decoder block's additional cross-attention head serving as the critical link between them.

## Summary

- **Encoder-only** models employ bidirectional attention and MLM training to produce contextual embeddings, ideal for understanding tasks but incapable of autoregressive generation.
- **Decoder-only** models use causal masking and next-token prediction to enable text generation, though they cannot leverage future context during training.
- **Encoder-decoder** models combine bidirectional encoding with causal decoding and cross-attention, excelling at translation and summarization but requiring more compute.
- Key implementation files in `rohitg00/ai-engineering-from-scratch` demonstrate these patterns through `multi_head_attention` configurations, `corrupt_spans` preprocessing, and explicit `kv_source` wiring in `decoder_block` functions.

## Frequently Asked Questions

### What is the fundamental difference between encoder and decoder transformer blocks?

**Encoder blocks** apply bidirectional self-attention where every token can see every other token, while **decoder blocks** use causal (masked) self-attention that restricts each token to view only previous positions. Additionally, decoder blocks in full transformer implementations include a cross-attention sublayer that queries encoder outputs, which encoder blocks lack entirely.

### When should I use an encoder-decoder model instead of a decoder-only model?

Use **encoder-decoder models** when the task requires conditioning generation on a complete source sequence, such as machine translation or document summarization, where the encoder processes the full source text before the decoder begins generation. **Decoder-only models** suffice for unconditional generation or prompt-completion tasks where no separate source context needs encoding.

### Why can't decoder-only models use bidirectional attention during training?

Decoder-only models must use **causal masking** (typically implemented with `torch.tril`) to prevent the model from attending to future tokens during training. If bidirectional attention were allowed, the model could cheat by looking at the target token it is supposed to predict, rendering the next-token prediction objective trivial and breaking the autoregressive property required for text generation.

### How does cross-attention differ from self-attention in transformer architectures?

**Self-attention** computes queries, keys, and values from the same input sequence, allowing the model to relate different positions within that sequence. **Cross-attention**, as seen in [`phases/07-transformers-deep-dive/05-full-transformer/code/main.py`](https://github.com/rohitg00/ai-engineering-from-scratch/blob/main/phases/07-transformers-deep-dive/05-full-transformer/code/main.py), uses queries from the decoder but keys and values from the encoder output (`kv_source=enc_out`), enabling the decoder to retrieve relevant information from the encoded source sequence while generating each target token.