# How the MLP Expands and Compresses Embedding Dimensions in Transformers

> Understand how the MLP expands and compresses embedding dimensions in transformers. Learn its role in feature transformation and maintaining the residual pathway for LLM training.

- Repository: [Fareed Khan/train-llm-from-scratch](https://github.com/FareedKhan-dev/train-llm-from-scratch)
- Tags: deep-dive
- Published: 2026-05-31

---

**The MLP expands embeddings to 4× their original dimension via a hidden linear layer, applies ReLU activation, and compresses back to the original size through a projection layer, enabling rich feature transformation while maintaining the residual pathway.**

The feed-forward network inside each transformer block is responsible for the position-wise transformation of embeddings. In the `FareedKhan-dev/train-llm-from-scratch` repository, this mechanism is implemented in the `MLP` class located in [`src/models/mlp.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/mlp.py). Understanding how this component manipulates tensor dimensions is crucial for grasping why transformers can model complex functions without exploding parameter counts.

## The Expansion-Compression Pipeline

The MLP acts as a bottleneck that first expands the representation space to allow richer feature mixing, then compresses it back to preserve the residual connection. This process follows three distinct stages.

### Expanding to 4× Dimensions

Inside [`src/models/mlp.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/mlp.py), the expansion is handled by the first linear transformation. At line 24, the constructor creates:

```python
self.hidden = nn.Linear(n_embed, 4 * n_embed)

```

During the forward pass, this layer receives input of shape **(B, T, C)** where **C = n_embed** and projects it to **(B, T, 4×C)**. This four-fold expansion is the conventional scaling factor introduced in the original Transformer paper (Vaswani et al., 2017).

### Non-Linear Activation

Immediately after expansion, a ReLU activation function introduces non-linearity. As implemented on line 25:

```python
self.relu = nn.ReLU()

```

This activation is essential for the network's expressive power, allowing the model to learn complex, non-linear relationships within the expanded hidden space.

### Compressing Back to Embedding Size

The final stage projects the expanded representation back to the original embedding dimension. Line 26 defines the compression layer:

```python
self.proj = nn.Linear(4 * n_embed, n_embed)

```

This projection ensures the MLP output matches the input dimension, enabling the residual addition that follows. The output shape returns to **(B, T, C)**, identical to the input.

## Integration Within the Transformer Block

The MLP does not operate in isolation. In [`src/models/transformer_block.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer_block.py), each `Block` instantiates the MLP using the model's embedding dimension:

```python
self.mlp = MLP(n_embed)

```

According to lines 44-46, the block applies the MLP after layer normalization and a residual connection, following the pre-norm transformer architecture. The sequence inside each block is:

1. Layer normalization
2. Multi-head attention
3. Residual addition
4. Layer normalization
5. **MLP (expand → activate → compress)**
6. Residual addition

This placement ensures that the MLP refines the attention output without disrupting the gradient flow through the residual pathway.

## Code Examples

### Standalone MLP Usage

You can observe the dimension changes directly using the `MLP` class:

```python
import torch
from src.models.mlp import MLP

batch_size, seq_len, embed_dim = 2, 5, 16
x = torch.randn(batch_size, seq_len, embed_dim)

# Initialize MLP with expansion factor of 4

mlp = MLP(n_embed=embed_dim)
y = mlp(x)

print(f"Input shape:  {x.shape}")   # torch.Size([2, 5, 16])

print(f"Output shape: {y.shape}")   # torch.Size([2, 5, 16])

```

Despite the internal expansion to 64 dimensions (4 × 16), the output preserves the original embedding size, ready for the next layer or residual addition.

### MLP Inside a Complete Block

When integrated into a full transformer block, the MLP operates within the broader attention context:

```python
import torch
from src.models.transformer_block import Block

batch_size, seq_len, embed_dim = 2, 7, 32
num_heads = 4

x = torch.randn(batch_size, seq_len, embed_dim)
block = Block(n_head=num_heads, n_embed=embed_dim, context_length=seq_len)

# Forward pass includes attention + MLP + residuals

output = block(x)
print(f"Output shape: {output.shape}")   # torch.Size([2, 7, 32])

```

The block handles the expansion and compression transparently while maintaining the sequence length and embedding dimension throughout.

## Summary

- **Expansion layer**: `self.hidden` in [`src/models/mlp.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/mlp.py) (line 24) projects embeddings from `n_embed` to `4 * n_embed`.
- **Activation**: ReLU introduces non-linearity after expansion (line 25).
- **Compression layer**: `self.proj` (line 26) projects back to `n_embed`, enabling residual connections.
- **Context**: The MLP operates within each `Block` in [`src/models/transformer_block.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer_block.py), positioned after attention with layer normalization and residual pathways.
- **Purpose**: The 4× expansion allows richer feature mixing than the embedding dimension alone permits, following the standard transformer architecture.

## Frequently Asked Questions

### Why is the expansion factor exactly 4×?

The 4× expansion factor is the standard configuration established in the original "Attention Is All You Need" paper. This specific ratio provides enough capacity for rich feature transformation while keeping the parameter count manageable. Expanding to fewer dimensions would limit representational capacity, while larger expansions increase computation and memory costs without proportional gains in performance for most language modeling tasks.

### Where does the MLP reside within the transformer block?

According to [`src/models/transformer_block.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/transformer_block.py), the MLP is instantiated as `self.mlp = MLP(n_embed)` and appears after the second layer normalization in each block. Specifically, the forward pass applies layer normalization, feeds the result through the MLP, then adds the residual connection. This "pre-norm" positioning helps stabilize gradients during training.

### Does the MLP change the sequence length?

No. The MLP operates in a **position-wise** manner, meaning it applies the same expansion and compression logic to each token position independently. The sequence length (T) remains constant; only the embedding dimension (C) changes temporarily from C → 4C → C. The shape transformation is strictly **(B, T, C) → (B, T, 4C) → (B, T, C)**.

### Can I modify the expansion factor in this implementation?

Yes. While the standard implementation uses `4 * n_embed` in line 24 of [`src/models/mlp.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/src/models/mlp.py), you could pass a different multiplier as a constructor argument. However, doing so would require corresponding changes to the projection layer (line 26) and may affect model capacity and convergence characteristics. Any modification should maintain the matching input/output dimensions expected by the residual connections in [`transformer_block.py`](https://github.com/FareedKhan-dev/train-llm-from-scratch/blob/main/transformer_block.py).