# GAT vs GATv2: Key Differences in Graph Attention Networks Explained

> GATv2 enhances GAT with dynamic attention, learning query-dependent neighbor rankings impossible for original GAT. Explore key GAT vs GATv2 differences.

- Repository: [labml.ai/annotated_deep_learning_paper_implementations](https://github.com/labmlai/annotated_deep_learning_paper_implementations)
- Tags: deep-dive
- Published: 2026-03-04

---

**GATv2 replaces GAT's static attention mechanism with dynamic attention by applying separate linear transformations to source and target nodes, enabling the model to learn query-dependent neighbor rankings that the original GAT cannot express.**

The repository **labmlai/annotated_deep_learning_paper_implementations** provides clean, annotated PyTorch implementations of both Graph Attention Network architectures. Understanding the differences between GAT and GATv2 is essential for selecting the right graph neural network layer for your specific node classification or link prediction task.

## Attention Mechanism: Static vs Dynamic

The fundamental distinction lies in how each model computes attention scores between connected nodes.

**GAT (Static Attention)** computes edge scores using concatenated embeddings:

```python
e_ij = LeakyReLU(a^T [W h_i || W h_j]) = LeakyReLU(a_1^T W h_i + a_2^T W h_j)

```

In this formulation, the ranking of neighbors for any query node *i* depends only on the **key** term (`a_2^T W h_j`). Because the query contribution is constant across all neighbors, GAT produces a **static** ranking that cannot adapt based on the source node.

**GATv2 (Dynamic Attention)** computes edge scores using added embeddings:

```python
e_ij = a^T LeakyReLU(W_l h_i + W_r h_j)

```

By applying separate linear transforms `W_l` (query) and `W_r` (key) before the non-linearity, GATv2 enables **dynamic** attention where the importance of neighbor *j* depends on both nodes. This allows the model to express complex patterns like dictionary-lookup graphs where the same target node should have different importance for different source nodes.

## Implementation Differences in Source Code

The architectural differences are reflected directly in the layer implementations within `labmlai/annotated_deep_learning_paper_implementations`.

### Linear Projection Layers

In [`labml_nn/graphs/gat/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gat/__init__.py), the `GraphAttentionLayer` class uses a single shared projection:

```python
self.linear = nn.Linear(in_features,
                        self.n_hidden * n_heads,
                        bias=False)

```

In [`labml_nn/graphs/gatv2/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gatv2/__init__.py), the `GraphAttentionV2Layer` class defines distinct source and target projections:

```python
self.linear_l = nn.Linear(in_features,
                          self.n_hidden * n_heads,
                          bias=False)               # source (query)

if share_weights:
    self.linear_r = self.linear_l                # optional sharing

else:
    self.linear_r = nn.Linear(in_features,
                              self.n_hidden * n_heads,
                              bias=False)         # target (key)

```

The `share_weights` parameter in GATv2 provides flexibility to reduce parameters when computational constraints exist, while the default configuration learns separate transformations for maximum expressiveness.

### Attention Score Computation

**GAT** concatenates the transformed source and target embeddings, then applies a linear attention layer `self.attn` that accepts input of size `2 * self.n_hidden`:

```python

# Concatenated approach (GAT)

g = torch.cat([g_l, g_r], dim=-1)  # g_l and g_r are transformed node features

e = self.attn(g)                   # self.attn maps 2*n_hidden -> 1

```

**GATv2** adds the transformed embeddings, applies `LeakyReLU`, then uses a single-dimensional attention layer `self.attn` with input size `self.n_hidden`:

```python

# Addition approach (GATv2)

g = g_l + g_r                      # separate transforms W_l and W_r

e = self.attn(F.leaky_relu(g))     # self.attn maps n_hidden -> 1

```

This modification implements the paper's dynamic attention formula `a^T LeakyReLU(W_l h_i + W_r h_j)` exactly as specified in the GATv2 paper.

### Message Passing and Aggregation

Both implementations mask non-existent edges with `-inf` before applying softmax over the neighbor dimension, followed by dropout. However, the aggregation step differs subtly:

- **GAT** multiplies attention coefficients with the concatenated representation logic
- **GATv2** explicitly multiplies normalized attention coefficients with the **target** transformed vectors (`g_r`) using `torch.einsum`

Both use `torch.einsum` for the final head aggregation, supporting either concatenation (`is_concat=True`) or averaging (`is_concat=False`) of multi-head outputs.

## Training Configuration Comparison

The repository provides separate experiment scripts for benchmarking:

- **GAT training**: [`labml_nn/graphs/gat/experiment.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gat/experiment.py) defines `GATConfigs`
- **GATv2 training**: [`labml_nn/graphs/gatv2/experiment.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gatv2/experiment.py) defines `GATv2Configs`

Both scripts utilize the same `CoraDataset` class and optimizer configurations from [`labml_nn/optimizers/configs.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/optimizers/configs.py), ensuring fair comparison. You can switch between architectures by importing the respective config class:

```python
from labml_nn.graphs.gat.experiment import Configs as GATConfigs
from labml_nn.graphs.gatv2.experiment import Configs as GATv2Configs

# Instantiate GATv2 for dynamic attention experiments

conf = GATv2Configs()

```

## When to Use GATv2 Over GAT

Select **GATv2** when your task requires:

- **Query-dependent relationships**: Tasks where the same neighbor should have different importance depending on which node is asking (e.g., knowledge graphs, dictionary lookups)
- **Maximum expressiveness**: When parameter budget allows separate `W_l` and `W_r` matrices
- **Dynamic rankings**: Any graph structure where static attention assumptions fail

Select **GAT** when:

- **Simplicity suffices**: Standard node classification on homogeneous graphs like Cora or Citeseer
- **Parameter efficiency**: Single shared projection reduces model size
- **Baseline comparison**: Following original paper implementations exactly

## Summary

- **GAT** implements static attention with a single linear projection `W`, concatenating transformed features before scoring
- **GATv2** implements dynamic attention with optional separate projections `W_l` and `W_r`, adding transformed features inside the LeakyReLU non-linearity
- **Code location**: GAT layer is in [`labml_nn/graphs/gat/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gat/__init__.py) (class `GraphAttentionLayer`); GATv2 layer is in [`labml_nn/graphs/gatv2/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gatv2/__init__.py) (class `GraphAttentionV2Layer`)
- **Weight sharing**: GAT forces sharing; GATv2 offers `share_weights` flag for flexibility
- **Performance**: GATv2 handles synthetic tasks where static attention fails, while both perform comparably on standard benchmarks like Cora

## Frequently Asked Questions

### What is the main mathematical difference between GAT and GATv2?

GAT computes attention as `LeakyReLU(a^T [Wh_i || Wh_j])` using concatenation, which creates a static ranking because only the target node's features determine the relative scores. GATv2 computes attention as `a^T LeakyReLU(W_l h_i + W_r h_j)` using addition inside the non-linearity, enabling dynamic rankings where the query node's features influence which neighbors are prioritized.

### Does GATv2 require more parameters than GAT?

By default, yes. GATv2 learns separate weight matrices `W_l` and `W_r`, doubling the parameter count compared to GAT's single `W` matrix. However, GATv2 provides a `share_weights=True` option that reuses the same matrix for both transformations, matching GAT's parameter count while retaining the dynamic attention formulation.

### Can I use GATv2 on any graph dataset where GAT works?

Yes. GATv2 is a drop-in replacement for GAT in any PyTorch Geometric or native PyTorch graph pipeline. The `GraphAttentionV2Layer` class maintains the same interface as `GraphAttentionLayer`, accepting node features and adjacency information while outputting attended representations. Both implementations in `labmlai/annotated_deep_learning_paper_implementations` support identical training loops and evaluation protocols.

### Where are the official implementations located in the repository?

The GAT implementation resides in [`labml_nn/graphs/gat/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gat/__init__.py) defining `GraphAttentionLayer`, while GATv2 is implemented in [`labml_nn/graphs/gatv2/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gatv2/__init__.py) defining `GraphAttentionV2Layer`. Corresponding training experiments are available in [`labml_nn/graphs/gat/experiment.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gat/experiment.py) and [`labml_nn/graphs/gatv2/experiment.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/graphs/gatv2/experiment.py) respectively, both utilizing shared utilities from the graphs module.