GAT vs GATv2: Key Differences in Graph Attention Networks Explained
GATv2 replaces GAT's static attention mechanism with dynamic attention by applying separate linear transformations to source and target nodes, enabling the model to learn query-dependent neighbor rankings that the original GAT cannot express.
The repository labmlai/annotated_deep_learning_paper_implementations provides clean, annotated PyTorch implementations of both Graph Attention Network architectures. Understanding the differences between GAT and GATv2 is essential for selecting the right graph neural network layer for your specific node classification or link prediction task.
Attention Mechanism: Static vs Dynamic
The fundamental distinction lies in how each model computes attention scores between connected nodes.
GAT (Static Attention) computes edge scores using concatenated embeddings:
e_ij = LeakyReLU(a^T [W h_i || W h_j]) = LeakyReLU(a_1^T W h_i + a_2^T W h_j)
In this formulation, the ranking of neighbors for any query node i depends only on the key term (a_2^T W h_j). Because the query contribution is constant across all neighbors, GAT produces a static ranking that cannot adapt based on the source node.
GATv2 (Dynamic Attention) computes edge scores using added embeddings:
e_ij = a^T LeakyReLU(W_l h_i + W_r h_j)
By applying separate linear transforms W_l (query) and W_r (key) before the non-linearity, GATv2 enables dynamic attention where the importance of neighbor j depends on both nodes. This allows the model to express complex patterns like dictionary-lookup graphs where the same target node should have different importance for different source nodes.
Implementation Differences in Source Code
The architectural differences are reflected directly in the layer implementations within labmlai/annotated_deep_learning_paper_implementations.
Linear Projection Layers
In labml_nn/graphs/gat/__init__.py, the GraphAttentionLayer class uses a single shared projection:
self.linear = nn.Linear(in_features,
self.n_hidden * n_heads,
bias=False)
In labml_nn/graphs/gatv2/__init__.py, the GraphAttentionV2Layer class defines distinct source and target projections:
self.linear_l = nn.Linear(in_features,
self.n_hidden * n_heads,
bias=False) # source (query)
if share_weights:
self.linear_r = self.linear_l # optional sharing
else:
self.linear_r = nn.Linear(in_features,
self.n_hidden * n_heads,
bias=False) # target (key)
The share_weights parameter in GATv2 provides flexibility to reduce parameters when computational constraints exist, while the default configuration learns separate transformations for maximum expressiveness.
Attention Score Computation
GAT concatenates the transformed source and target embeddings, then applies a linear attention layer self.attn that accepts input of size 2 * self.n_hidden:
# Concatenated approach (GAT)
g = torch.cat([g_l, g_r], dim=-1) # g_l and g_r are transformed node features
e = self.attn(g) # self.attn maps 2*n_hidden -> 1
GATv2 adds the transformed embeddings, applies LeakyReLU, then uses a single-dimensional attention layer self.attn with input size self.n_hidden:
# Addition approach (GATv2)
g = g_l + g_r # separate transforms W_l and W_r
e = self.attn(F.leaky_relu(g)) # self.attn maps n_hidden -> 1
This modification implements the paper's dynamic attention formula a^T LeakyReLU(W_l h_i + W_r h_j) exactly as specified in the GATv2 paper.
Message Passing and Aggregation
Both implementations mask non-existent edges with -inf before applying softmax over the neighbor dimension, followed by dropout. However, the aggregation step differs subtly:
- GAT multiplies attention coefficients with the concatenated representation logic
- GATv2 explicitly multiplies normalized attention coefficients with the target transformed vectors (
g_r) usingtorch.einsum
Both use torch.einsum for the final head aggregation, supporting either concatenation (is_concat=True) or averaging (is_concat=False) of multi-head outputs.
Training Configuration Comparison
The repository provides separate experiment scripts for benchmarking:
- GAT training:
labml_nn/graphs/gat/experiment.pydefinesGATConfigs - GATv2 training:
labml_nn/graphs/gatv2/experiment.pydefinesGATv2Configs
Both scripts utilize the same CoraDataset class and optimizer configurations from labml_nn/optimizers/configs.py, ensuring fair comparison. You can switch between architectures by importing the respective config class:
from labml_nn.graphs.gat.experiment import Configs as GATConfigs
from labml_nn.graphs.gatv2.experiment import Configs as GATv2Configs
# Instantiate GATv2 for dynamic attention experiments
conf = GATv2Configs()
When to Use GATv2 Over GAT
Select GATv2 when your task requires:
- Query-dependent relationships: Tasks where the same neighbor should have different importance depending on which node is asking (e.g., knowledge graphs, dictionary lookups)
- Maximum expressiveness: When parameter budget allows separate
W_landW_rmatrices - Dynamic rankings: Any graph structure where static attention assumptions fail
Select GAT when:
- Simplicity suffices: Standard node classification on homogeneous graphs like Cora or Citeseer
- Parameter efficiency: Single shared projection reduces model size
- Baseline comparison: Following original paper implementations exactly
Summary
- GAT implements static attention with a single linear projection
W, concatenating transformed features before scoring - GATv2 implements dynamic attention with optional separate projections
W_landW_r, adding transformed features inside the LeakyReLU non-linearity - Code location: GAT layer is in
labml_nn/graphs/gat/__init__.py(classGraphAttentionLayer); GATv2 layer is inlabml_nn/graphs/gatv2/__init__.py(classGraphAttentionV2Layer) - Weight sharing: GAT forces sharing; GATv2 offers
share_weightsflag for flexibility - Performance: GATv2 handles synthetic tasks where static attention fails, while both perform comparably on standard benchmarks like Cora
Frequently Asked Questions
What is the main mathematical difference between GAT and GATv2?
GAT computes attention as LeakyReLU(a^T [Wh_i || Wh_j]) using concatenation, which creates a static ranking because only the target node's features determine the relative scores. GATv2 computes attention as a^T LeakyReLU(W_l h_i + W_r h_j) using addition inside the non-linearity, enabling dynamic rankings where the query node's features influence which neighbors are prioritized.
Does GATv2 require more parameters than GAT?
By default, yes. GATv2 learns separate weight matrices W_l and W_r, doubling the parameter count compared to GAT's single W matrix. However, GATv2 provides a share_weights=True option that reuses the same matrix for both transformations, matching GAT's parameter count while retaining the dynamic attention formulation.
Can I use GATv2 on any graph dataset where GAT works?
Yes. GATv2 is a drop-in replacement for GAT in any PyTorch Geometric or native PyTorch graph pipeline. The GraphAttentionV2Layer class maintains the same interface as GraphAttentionLayer, accepting node features and adjacency information while outputting attended representations. Both implementations in labmlai/annotated_deep_learning_paper_implementations support identical training loops and evaluation protocols.
Where are the official implementations located in the repository?
The GAT implementation resides in labml_nn/graphs/gat/__init__.py defining GraphAttentionLayer, while GATv2 is implemented in labml_nn/graphs/gatv2/__init__.py defining GraphAttentionV2Layer. Corresponding training experiments are available in labml_nn/graphs/gat/experiment.py and labml_nn/graphs/gatv2/experiment.py respectively, both utilizing shared utilities from the graphs module.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →