# Feedforward Transformer vs Standard Attention: Architectural Differences Explained

> Discover the key differences between Feedforward Transformer and standard attention. Learn how Feedforward Transformers use FFNs solely, contrasting with attention's O(L²) complexity.

- Repository: [labml.ai/annotated_deep_learning_paper_implementations](https://github.com/labmlai/annotated_deep_learning_paper_implementations)
- Tags: deep-dive
- Published: 2026-03-04

---

**A Feedforward Transformer eliminates the multi-head self-attention mechanism entirely and processes tokens using only position-wise feed-forward networks, while a standard Transformer alternates between attention and FFN layers to capture inter-token dependencies at O(L²) complexity.**

The `labmlai/annotated_deep_learning_paper_implementations` repository provides clean, annotated implementations of both architectures. Understanding the distinction between a **Feedforward Transformer vs standard attention** is crucial for selecting the right model based on your sequence length constraints and dependency modeling requirements.

## How Standard Transformer Attention Works

In a standard Transformer (as implemented in [`labml_nn/transformers/mha.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mha.py)), each layer contains a **MultiHeadAttention** module followed by a **FeedForward** network. The attention mechanism computes query-key dot products, enabling every token to attend to every other token in the sequence.

This inter-token communication captures long-range dependencies but incurs **quadratic complexity** of *O(L² · d_model)* with respect to sequence length *L*. The standard block alternates between attention-based mixing and position-wise FFN processing, as defined in the `MultiHeadAttention` class.

## Understanding the Feedforward-Only Architecture

A **Feedforward-Only Transformer** (sometimes called an FFN-only model) removes the attention component completely. As seen in [`labml_nn/transformers/feed_forward.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/feed_forward.py), the architecture stacks **FeedForward** modules applied independently to each position, possibly with residual connections and layer normalization, but without any token-to-token interaction.

This design reduces computational complexity to **linear** *O(L · d_model · d_ff)*, making it significantly faster for long sequences, though it sacrifices the ability to model relationships between different positions.

## Key Differences Between Feedforward and Standard Transformers

When comparing **Feedforward Transformer vs standard attention** implementations in the LabML repository, three critical distinctions emerge:

- **Computational Complexity**: Standard attention scales quadratically (*O(L²)*) with sequence length due to the attention matrix calculation in [`labml_nn/transformers/mha.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mha.py), while Feedforward-only models scale linearly (*O(L)*) since they process each position independently.

- **Token Interaction**: The standard `MultiHeadAttention` module enables explicit token-to-token communication through attention scores. The Feedforward-only variant has **no inter-token communication**—each position is processed in isolation by the `FeedForward` class.

- **Typical Applications**: Standard Transformers dominate language modeling and translation tasks requiring context. Feedforward-only architectures appear in vision models like MLP-Mixer (see [`labml_nn/transformers/mlp_mixer/experiment.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mlp_mixer/experiment.py)) and serve as experimental baselines to isolate attention's contribution.

## Code Implementation Comparison

The LabML repository demonstrates both patterns using shared components from [`labml_nn/transformers/feed_forward.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/feed_forward.py).

### Standard Transformer Block with Attention

This implementation combines `MultiHeadAttention` with the FFN sub-layer:

```python
import torch
from labml_nn.transformers.mha import MultiHeadAttention
from labml_nn.transformers.feed_forward import FeedForward
from labml_nn.helpers import Residual, LayerNorm

class TransformerBlock(torch.nn.Module):
    def __init__(self, d_model: int, heads: int, d_ff: int, dropout: float = 0.1):
        super().__init__()
        self.norm1 = LayerNorm(d_model)
        self.attn = Residual(
            MultiHeadAttention(heads=heads,
                               d_model=d_model,
                               dropout_prob=dropout))
        self.norm2 = LayerNorm(d_model)
        self.ffn = Residual(
            FeedForward(d_model=d_model,
                        d_ff=d_ff,
                        dropout=dropout))

    def forward(self, x, mask=None):
        # Self-attention enables token interaction

        x = self.norm1(x)
        x = self.attn(x, x, x, mask=mask)

        # Feed-forward processing

        x = self.norm2(x)
        x = self.ffn(x)
        return x

```

### Feedforward-Only Transformer Block

This variant omits the attention module entirely, using only the `FeedForward` class:

```python
import torch
from labml_nn.transformers.feed_forward import FeedForward
from labml_nn.helpers import Residual, LayerNorm

class FFOnlyBlock(torch.nn.Module):
    def __init__(self, d_model: int, d_ff: int, dropout: float = 0.1):
        super().__init__()
        self.norm = LayerNorm(d_model)
        self.ffn = Residual(
            FeedForward(d_model=d_model,
                        d_ff=d_ff,
                        dropout=dropout))

    def forward(self, x):
        x = self.norm(x)
        return self.ffn(x)

```

Notice the absence of `MultiHeadAttention` and the simplified forward pass that processes each position independently.

## Practical Implications for Model Selection

**Speed and Memory**: Removing the attention mechanism from [`labml_nn/transformers/mha.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mha.py) significantly reduces GPU memory consumption and wall-clock time, particularly for long sequences where the *O(L²)* attention matrix becomes prohibitive.

**Modeling Limitations**: Without attention, the model cannot directly learn dependencies across distant tokens. Performance on tasks requiring contextual understanding—such as machine translation—typically degrades compared to standard Transformer architectures defined in [`labml_nn/transformers/models.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/models.py).

**Research Applications**: Feedforward-only variants serve as critical baselines in the `labmlai/annotated_deep_learning_paper_implementations` repository to isolate the specific contribution of attention mechanisms, or they combine with convolutional layers to create hybrid approaches.

## Summary

- **Standard Transformers** alternate between `MultiHeadAttention` (quadratic complexity) and `FeedForward` layers to capture inter-token dependencies.
- **Feedforward-Only Transformers** eliminate attention entirely, using only the `FeedForward` module from [`labml_nn/transformers/feed_forward.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/feed_forward.py) with linear complexity.
- The key trade-off is between computational efficiency (Feedforward-only) and the ability to model long-range context (Standard attention).
- LabML implements both variants, with attention-less models appearing in experiments like [`labml_nn/transformers/mlp_mixer/experiment.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mlp_mixer/experiment.py).

## Frequently Asked Questions

### Can a Feedforward Transformer replace a standard Transformer for language modeling?

No, removing attention typically causes significant performance degradation on language tasks because the model cannot capture word-to-word relationships across distances using only the position-wise `FeedForward` layers. However, it may suffice for character-level modeling or when combined with other mechanisms like convolution.

### Why does the LabML repository implement both architectures?

The `labmlai/annotated_deep_learning_paper_implementations` repository provides both implementations to enable ablation studies—researchers can isolate the impact of attention by comparing identical architectures with and without the `MultiHeadAttention` module in [`labml_nn/transformers/mha.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mha.py).

### What is the computational complexity difference between the two approaches?

Standard attention incurs *O(L² · d_model)* complexity due to the attention matrix multiplication in [`labml_nn/transformers/mha.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mha.py), while Feedforward-only models scale as *O(L · d_model · d_ff)*, processing each of the *L* positions independently through the feed-forward network.

### Where can I find a working example of a Feedforward-only model in the repository?

The MLP-Mixer implementation in [`labml_nn/transformers/mlp_mixer/experiment.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/mlp_mixer/experiment.py) demonstrates a complete model using stacked `FeedForward` blocks without attention, serving as a practical reference for Feedforward-only architectures in computer vision tasks.