How the Attention Mechanism Works in Transformer XL with Relative Positional Embeddings

Transformer XL replaces absolute position embeddings with a four-term attention score that incorporates relative positional biases, enabling the model to reuse hidden states from previous segments via segment-level recurrence while maintaining temporal coherence across segment boundaries.

Transformer XL extends the standard Transformer architecture by introducing segment-level recurrence and relative positional embeddings to capture long-range dependencies beyond fixed-length contexts. In the labmlai/annotated_deep_learning_paper_implementations repository, the attention mechanism with relative positional embeddings is implemented in labml_nn/transformers/xl/__init__.py, where the four-component attention score enables seamless attention across segment boundaries.

Standard Self-Attention vs. Relative Positional Attention

Standard scaled dot-product attention computes scores as:

score = (q @ k.T) / sqrt(d)

In the classic Transformer, absolute position embeddings are added to token embeddings before the first layer. This approach fails in Transformer XL because the model reuses hidden states from previous segments (the memory). Absolute embeddings would incorrectly treat a token at position 5 in the current segment as position 5 in the entire sequence, breaking temporal continuity when attending across segment boundaries.

The Four-Component Attention Score in Transformer XL

The attention mechanism in Transformer XL decomposes the score between a query at position i and a key at position j into four distinct terms:

$$ \text{score}{i,j} = \underbrace{q_i^{\top} k_j}{\text{content-based}} + \underbrace{q_i^{\top} r_{i-j}}{\text{relative-positional}} + \underbrace{u^{\top} k_j}{\text{global content bias}} + \underbrace{v^{\top} r_{i-j}}_{\text{global positional bias}} $$

Content-Based Addressing

The term $q_i^{\top} k_j$ represents standard content-based attention, measuring how relevant the content of key j is to query i.

Relative Positional Bias

The term $q_i^{\top} r_{i-j}$ introduces relative positional embeddings $r_{i-j}$, which depend only on the distance between positions rather than absolute indices. This makes the attention mechanism translation-invariant, allowing the model to attend to memory from previous segments without positional confusion.

Global Content and Positional Biases

The terms $u^{\top} k_j$ and $v^{\top} r_{i-j}$ are learnable global bias vectors shared across all positions. These biases improve training stability by providing baseline attention scores independent of specific content or position.

Implementation in the Annotated Deep Learning Repository

The implementation resides in labml_nn/transformers/xl/__init__.py within the labmlai/annotated_deep_learning_paper_implementations repository.

TransformerXLLayer and RelativePositionalEmbedding

The TransformerXLLayer class (around line 45) constructs the attention block containing the relative positional logic. Inside this layer, RelativePositionalEmbedding (around line 70) creates the lookup table $r$ for each possible offset distance.

The _relative_attention_scores Method

The core computation occurs in _relative_attention_scores (around line 110), which sums the four attention terms:

  1. Content-based scores from query-key dot products
  2. Relative positional scores using $r_{i-j}$
  3. Global content bias $u$ applied to keys
  4. Global positional bias $v$ applied to relative positions

Memory Handling and Segment-Level Recurrence

The TransformerXL class (around line 107) manages segment-level recurrence by concatenating cached hidden states (mem) with the current segment before processing through each TransformerXLLayer. The forward method (around line 140) orchestrates this flow, computing queries, keys, and values, applying the relative attention scores, and producing the output logits alongside the updated memory cache.

Practical Code Example

Below is a minimal example demonstrating how to instantiate a Transformer-XL model with relative positional embeddings and execute a forward pass:

import torch
from labml_nn.transformers.xl import TransformerXL, TransformerXLLayer

# Hyperparameters matching the repository defaults

n_vocab = 10000
d_model = 512
n_layers = 6
n_heads = 8
d_head = d_model // n_heads
d_inner = 2048
dropout = 0.1
mem_len = 256

# Construct a single XL layer with relative attention

layer = TransformerXLLayer(
    n_heads=n_heads,
    d_model=d_model,
    d_head=d_head,
    d_inner=d_inner,
    dropout=dropout,
    dropout_attn=dropout,
    dropout_head=dropout,
    dropout_ff=dropout,
)

# Build the full Transformer-XL model

model = TransformerXL(layer, n_layers)

# Dummy token batch (batch_size=2, sequence_length=32)

tokens = torch.randint(0, n_vocab, (2, 32))

# Forward pass returns logits and updated memory

logits, new_mem = model(tokens, mem=None)

print(f"Logits shape: {logits.shape}")  # (2, 32, n_vocab)

print(f"Memory layers: {len(new_mem)}")  # n_layers

The TransformerXLLayer encapsulates the four-term relative attention mechanism, while TransformerXL manages the segment-level memory. Passing mem=None initializes the first segment; subsequent calls should pass new_mem to maintain context across segments.

Summary

  • Transformer XL extends the standard Transformer with segment-level recurrence and relative positional embeddings to model long-range dependencies.
  • The attention mechanism uses a four-component score: content-based attention, relative positional bias, global content bias, and global positional bias.
  • Relative positional embeddings $r_{i-j}$ depend only on the distance between query and key, making the model translation-invariant and compatible with cached memory from previous segments.
  • The implementation in labml_nn/transformers/xl/__init__.py provides TransformerXLLayer for the attention logic and TransformerXL for memory management, with the core computation in _relative_attention_scores.

Frequently Asked Questions

What makes relative positional embeddings different from absolute embeddings in Transformer XL?

Absolute embeddings encode specific positions (e.g., position 5, position 10) as unique vectors, which fails when reusing hidden states from previous segments because the same absolute index refers to different temporal positions across segments. Relative embeddings encode the distance between positions (e.g., distance of -3, +5), making the attention mechanism translation-invariant and allowing seamless attention across segment boundaries regardless of absolute position.

Why does Transformer XL use four terms in the attention score instead of two?

The four-term decomposition separates content-based interactions from positional biases while adding learnable global biases for stability. The content term ($q_i^{\top} k_j$) and relative position term ($q_i^{\top} r_{i-j}$) capture dynamic interactions, while the global content bias ($u^{\top} k_j$) and global positional bias ($v^{\top} r_{i-j}$) provide baseline scores that improve training convergence. This separation allows the model to learn distinct representations for "what" (content) and "where" (position) while maintaining stable gradients.

How does the memory mechanism interact with relative positional embeddings?

The memory mechanism caches hidden states from previous segments and concatenates them with the current segment's states before computing attention. Because relative positional embeddings depend only on the distance $i-j$ between query and key rather than absolute indices, the same embedding matrix can be used for both current-segment keys and memory keys. This allows the model to attend to arbitrarily long contexts (limited only by memory length) without positional ambiguity, as the relative distance correctly represents the temporal relationship between current queries and cached keys from previous segments.

Where can I find the complete implementation of relative attention in the labmlai repository?

The complete implementation resides in labml_nn/transformers/xl/__init__.py within the labmlai/annotated_deep_learning_paper_implementations repository. Key components include the TransformerXLLayer class (around line 45) which contains the attention logic, the RelativePositionalEmbedding lookup (around line 70), the learnable bias vectors u and v (around line 85), and the _relative_attention_scores method (around line 110) that computes the four-term attention equation. The TransformerXL class (around line 107) handles segment-level recurrence and memory management.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →