How the MLP Expands and Compresses Embedding Dimensions in Transformers

The MLP expands embeddings to 4× their original dimension via a hidden linear layer, applies ReLU activation, and compresses back to the original size through a projection layer, enabling rich feature transformation while maintaining the residual pathway.

The feed-forward network inside each transformer block is responsible for the position-wise transformation of embeddings. In the FareedKhan-dev/train-llm-from-scratch repository, this mechanism is implemented in the MLP class located in src/models/mlp.py. Understanding how this component manipulates tensor dimensions is crucial for grasping why transformers can model complex functions without exploding parameter counts.

The Expansion-Compression Pipeline

The MLP acts as a bottleneck that first expands the representation space to allow richer feature mixing, then compresses it back to preserve the residual connection. This process follows three distinct stages.

Expanding to 4× Dimensions

Inside src/models/mlp.py, the expansion is handled by the first linear transformation. At line 24, the constructor creates:

self.hidden = nn.Linear(n_embed, 4 * n_embed)

During the forward pass, this layer receives input of shape (B, T, C) where C = n_embed and projects it to (B, T, 4×C). This four-fold expansion is the conventional scaling factor introduced in the original Transformer paper (Vaswani et al., 2017).

Non-Linear Activation

Immediately after expansion, a ReLU activation function introduces non-linearity. As implemented on line 25:

self.relu = nn.ReLU()

This activation is essential for the network's expressive power, allowing the model to learn complex, non-linear relationships within the expanded hidden space.

Compressing Back to Embedding Size

The final stage projects the expanded representation back to the original embedding dimension. Line 26 defines the compression layer:

self.proj = nn.Linear(4 * n_embed, n_embed)

This projection ensures the MLP output matches the input dimension, enabling the residual addition that follows. The output shape returns to (B, T, C), identical to the input.

Integration Within the Transformer Block

The MLP does not operate in isolation. In src/models/transformer_block.py, each Block instantiates the MLP using the model's embedding dimension:

self.mlp = MLP(n_embed)

According to lines 44-46, the block applies the MLP after layer normalization and a residual connection, following the pre-norm transformer architecture. The sequence inside each block is:

  1. Layer normalization
  2. Multi-head attention
  3. Residual addition
  4. Layer normalization
  5. MLP (expand → activate → compress)
  6. Residual addition

This placement ensures that the MLP refines the attention output without disrupting the gradient flow through the residual pathway.

Code Examples

Standalone MLP Usage

You can observe the dimension changes directly using the MLP class:

import torch
from src.models.mlp import MLP

batch_size, seq_len, embed_dim = 2, 5, 16
x = torch.randn(batch_size, seq_len, embed_dim)

# Initialize MLP with expansion factor of 4

mlp = MLP(n_embed=embed_dim)
y = mlp(x)

print(f"Input shape:  {x.shape}")   # torch.Size([2, 5, 16])

print(f"Output shape: {y.shape}")   # torch.Size([2, 5, 16])

Despite the internal expansion to 64 dimensions (4 × 16), the output preserves the original embedding size, ready for the next layer or residual addition.

MLP Inside a Complete Block

When integrated into a full transformer block, the MLP operates within the broader attention context:

import torch
from src.models.transformer_block import Block

batch_size, seq_len, embed_dim = 2, 7, 32
num_heads = 4

x = torch.randn(batch_size, seq_len, embed_dim)
block = Block(n_head=num_heads, n_embed=embed_dim, context_length=seq_len)

# Forward pass includes attention + MLP + residuals

output = block(x)
print(f"Output shape: {output.shape}")   # torch.Size([2, 7, 32])

The block handles the expansion and compression transparently while maintaining the sequence length and embedding dimension throughout.

Summary

  • Expansion layer: self.hidden in src/models/mlp.py (line 24) projects embeddings from n_embed to 4 * n_embed.
  • Activation: ReLU introduces non-linearity after expansion (line 25).
  • Compression layer: self.proj (line 26) projects back to n_embed, enabling residual connections.
  • Context: The MLP operates within each Block in src/models/transformer_block.py, positioned after attention with layer normalization and residual pathways.
  • Purpose: The 4× expansion allows richer feature mixing than the embedding dimension alone permits, following the standard transformer architecture.

Frequently Asked Questions

Why is the expansion factor exactly 4×?

The 4× expansion factor is the standard configuration established in the original "Attention Is All You Need" paper. This specific ratio provides enough capacity for rich feature transformation while keeping the parameter count manageable. Expanding to fewer dimensions would limit representational capacity, while larger expansions increase computation and memory costs without proportional gains in performance for most language modeling tasks.

Where does the MLP reside within the transformer block?

According to src/models/transformer_block.py, the MLP is instantiated as self.mlp = MLP(n_embed) and appears after the second layer normalization in each block. Specifically, the forward pass applies layer normalization, feeds the result through the MLP, then adds the residual connection. This "pre-norm" positioning helps stabilize gradients during training.

Does the MLP change the sequence length?

No. The MLP operates in a position-wise manner, meaning it applies the same expansion and compression logic to each token position independently. The sequence length (T) remains constant; only the embedding dimension (C) changes temporarily from C → 4C → C. The shape transformation is strictly (B, T, C) → (B, T, 4C) → (B, T, C).

Can I modify the expansion factor in this implementation?

Yes. While the standard implementation uses 4 * n_embed in line 24 of src/models/mlp.py, you could pass a different multiplier as a constructor argument. However, doing so would require corresponding changes to the projection layer (line 26) and may affect model capacity and convergence characteristics. Any modification should maintain the matching input/output dimensions expected by the residual connections in transformer_block.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →