Feedforward Transformer vs Standard Attention: Architectural Differences Explained
A Feedforward Transformer eliminates the multi-head self-attention mechanism entirely and processes tokens using only position-wise feed-forward networks, while a standard Transformer alternates between attention and FFN layers to capture inter-token dependencies at O(L²) complexity.
The labmlai/annotated_deep_learning_paper_implementations repository provides clean, annotated implementations of both architectures. Understanding the distinction between a Feedforward Transformer vs standard attention is crucial for selecting the right model based on your sequence length constraints and dependency modeling requirements.
How Standard Transformer Attention Works
In a standard Transformer (as implemented in labml_nn/transformers/mha.py), each layer contains a MultiHeadAttention module followed by a FeedForward network. The attention mechanism computes query-key dot products, enabling every token to attend to every other token in the sequence.
This inter-token communication captures long-range dependencies but incurs quadratic complexity of O(L² · d_model) with respect to sequence length L. The standard block alternates between attention-based mixing and position-wise FFN processing, as defined in the MultiHeadAttention class.
Understanding the Feedforward-Only Architecture
A Feedforward-Only Transformer (sometimes called an FFN-only model) removes the attention component completely. As seen in labml_nn/transformers/feed_forward.py, the architecture stacks FeedForward modules applied independently to each position, possibly with residual connections and layer normalization, but without any token-to-token interaction.
This design reduces computational complexity to linear O(L · d_model · d_ff), making it significantly faster for long sequences, though it sacrifices the ability to model relationships between different positions.
Key Differences Between Feedforward and Standard Transformers
When comparing Feedforward Transformer vs standard attention implementations in the LabML repository, three critical distinctions emerge:
-
Computational Complexity: Standard attention scales quadratically (O(L²)) with sequence length due to the attention matrix calculation in
labml_nn/transformers/mha.py, while Feedforward-only models scale linearly (O(L)) since they process each position independently. -
Token Interaction: The standard
MultiHeadAttentionmodule enables explicit token-to-token communication through attention scores. The Feedforward-only variant has no inter-token communication—each position is processed in isolation by theFeedForwardclass. -
Typical Applications: Standard Transformers dominate language modeling and translation tasks requiring context. Feedforward-only architectures appear in vision models like MLP-Mixer (see
labml_nn/transformers/mlp_mixer/experiment.py) and serve as experimental baselines to isolate attention's contribution.
Code Implementation Comparison
The LabML repository demonstrates both patterns using shared components from labml_nn/transformers/feed_forward.py.
Standard Transformer Block with Attention
This implementation combines MultiHeadAttention with the FFN sub-layer:
import torch
from labml_nn.transformers.mha import MultiHeadAttention
from labml_nn.transformers.feed_forward import FeedForward
from labml_nn.helpers import Residual, LayerNorm
class TransformerBlock(torch.nn.Module):
def __init__(self, d_model: int, heads: int, d_ff: int, dropout: float = 0.1):
super().__init__()
self.norm1 = LayerNorm(d_model)
self.attn = Residual(
MultiHeadAttention(heads=heads,
d_model=d_model,
dropout_prob=dropout))
self.norm2 = LayerNorm(d_model)
self.ffn = Residual(
FeedForward(d_model=d_model,
d_ff=d_ff,
dropout=dropout))
def forward(self, x, mask=None):
# Self-attention enables token interaction
x = self.norm1(x)
x = self.attn(x, x, x, mask=mask)
# Feed-forward processing
x = self.norm2(x)
x = self.ffn(x)
return x
Feedforward-Only Transformer Block
This variant omits the attention module entirely, using only the FeedForward class:
import torch
from labml_nn.transformers.feed_forward import FeedForward
from labml_nn.helpers import Residual, LayerNorm
class FFOnlyBlock(torch.nn.Module):
def __init__(self, d_model: int, d_ff: int, dropout: float = 0.1):
super().__init__()
self.norm = LayerNorm(d_model)
self.ffn = Residual(
FeedForward(d_model=d_model,
d_ff=d_ff,
dropout=dropout))
def forward(self, x):
x = self.norm(x)
return self.ffn(x)
Notice the absence of MultiHeadAttention and the simplified forward pass that processes each position independently.
Practical Implications for Model Selection
Speed and Memory: Removing the attention mechanism from labml_nn/transformers/mha.py significantly reduces GPU memory consumption and wall-clock time, particularly for long sequences where the O(L²) attention matrix becomes prohibitive.
Modeling Limitations: Without attention, the model cannot directly learn dependencies across distant tokens. Performance on tasks requiring contextual understanding—such as machine translation—typically degrades compared to standard Transformer architectures defined in labml_nn/transformers/models.py.
Research Applications: Feedforward-only variants serve as critical baselines in the labmlai/annotated_deep_learning_paper_implementations repository to isolate the specific contribution of attention mechanisms, or they combine with convolutional layers to create hybrid approaches.
Summary
- Standard Transformers alternate between
MultiHeadAttention(quadratic complexity) andFeedForwardlayers to capture inter-token dependencies. - Feedforward-Only Transformers eliminate attention entirely, using only the
FeedForwardmodule fromlabml_nn/transformers/feed_forward.pywith linear complexity. - The key trade-off is between computational efficiency (Feedforward-only) and the ability to model long-range context (Standard attention).
- LabML implements both variants, with attention-less models appearing in experiments like
labml_nn/transformers/mlp_mixer/experiment.py.
Frequently Asked Questions
Can a Feedforward Transformer replace a standard Transformer for language modeling?
No, removing attention typically causes significant performance degradation on language tasks because the model cannot capture word-to-word relationships across distances using only the position-wise FeedForward layers. However, it may suffice for character-level modeling or when combined with other mechanisms like convolution.
Why does the LabML repository implement both architectures?
The labmlai/annotated_deep_learning_paper_implementations repository provides both implementations to enable ablation studies—researchers can isolate the impact of attention by comparing identical architectures with and without the MultiHeadAttention module in labml_nn/transformers/mha.py.
What is the computational complexity difference between the two approaches?
Standard attention incurs O(L² · d_model) complexity due to the attention matrix multiplication in labml_nn/transformers/mha.py, while Feedforward-only models scale as O(L · d_model · d_ff), processing each of the L positions independently through the feed-forward network.
Where can I find a working example of a Feedforward-only model in the repository?
The MLP-Mixer implementation in labml_nn/transformers/mlp_mixer/experiment.py demonstrates a complete model using stacked FeedForward blocks without attention, serving as a practical reference for Feedforward-only architectures in computer vision tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →