How Rotary Positional Embedding Is Implemented in PyTorch: A Deep Dive into labmlai/annotated_deep_learning_paper_implementations

The labmlai/annotated_deep_learning_paper_implementations repository implements Rotary Positional Embeddings (RoPE) as a pure-PyTorch module using cached sine/cosine matrices and a "negative-half" tensor trick to efficiently rotate query and key vectors by their position indices.

Rotary Positional Embedding (RoPE) encodes absolute positional information into transformer attention layers by rotating feature vectors in a two-dimensional subspace. According to the source code in labmlai/annotated_deep_learning_paper_implementations, the implementation provides drop-in modules that replace standard positional encodings while maintaining computational efficiency through smart caching mechanisms.

Core RoPE Module Architecture

The primary implementation resides in labml_nn/transformers/rope/__init__.py within the RotaryPositionalEmbeddings class (lines 30-71). This module implements the mathematical formulation from the RoFormer paper, applying rotation to paired dimensions of the input tensor.

Pair-wise Rotation Mechanism

RoPE splits the feature dimension d into d/2 pairs. Each pair (x^{(1)}_m, x^{(2)}_m) at position m rotates by an angle m·θ_i, where θ_i represents a frequency specific to that pair following the classic sinusoidal schedule. This rotation injects relative positional information directly into the dot-product attention calculation.

Cached Sine and Cosine Tensors

Computing trigonometric functions for every forward pass would create significant overhead. The implementation builds a cached tensor of shape (seq_len, 1, 1, d) the first time it encounters a sequence length exceeding the current cache size. Subsequent calls reuse these cached values:

def _build_cache(self, x: torch.Tensor):
    # Calculate theta values for each dimension pair

    theta = 1.0 / (self.base ** (torch.arange(0, self.d, 2).float() / self.d))
    
    # Position indices

    seq_idx = torch.arange(self.seq_len_cached, device=x.device).float()
    
    # Calculate the angles: m * theta_i

    idx_theta = torch.einsum('i,j->ij', seq_idx, theta)
    
    # Cache both sine and cosine

    self.cos_cached = idx_theta.cos()[None, None, :, :]
    self.sin_cached = idx_theta.sin()[None, None, :, :]

The Negative-Half Trick

To apply the rotation efficiently without explicit matrix multiplication, the code constructs a "neg-half" tensor. This operation concatenates the second half of features with a sign flip (-x_{d/2+1…d}) followed by the first half (x_{1…d/2}). This enables single element-wise multiplication with the cached sine tensor:

def _neg_half(self, x: torch.Tensor):
    # Split and flip with negation

    d_2 = self.d // 2
    return torch.cat([-x[:, :, :, d_2:], x[:, :, :, :d_2]], dim=-1)

Forward Pass Implementation

The forward method in RotaryPositionalEmbeddings orchestrates the rotation through five distinct steps:

  1. Cache initialization: self._build_cache(x) prepares or reuses sine/cosine matrices
  2. Tensor splitting: Separates rotatable dimensions from pass-through features
  3. Neg-half construction: Creates the rotated counterpart using _neg_half
  4. Rotation application: Combines original and neg-half tensors with cached values
  5. Concatenation: Reassembles rotated and unrotated features
def forward(self, x: torch.Tensor):
    # 1. Build or reuse cache

    self._build_cache(x)
    
    # 2. Split rotatable vs pass-through dimensions

    x_rope, x_pass = x[..., :self.d], x[..., self.d:]
    
    # 3. Create negative-half version

    neg_half_x = self._neg_half(x_rope)
    
    # 4. Apply rotation: x_rot = x*cos + neg_half*sin

    x_rope = (x_rope * self.cos_cached[:seq_len]) + \
             (neg_half_x * self.sin_cached[:seq_len])
    
    # 5. Concatenate results

    return torch.cat((x_rope, x_pass), dim=-1)

Integration with Multi-Head Attention

The RotaryPEMultiHeadAttention class demonstrates practical integration. It instantiates separate RoPE modules for queries and keys, applying them before the attention score calculation:

class RotaryPEMultiHeadAttention(MultiHeadAttention):
    def __init__(self, heads: int, d_model: int, dropout_prob: float = 0.0, 
                 rope_percentage: float = 0.5):
        super().__init__(heads, d_model, dropout_prob)
        self.query_rotary_pe = RotaryPositionalEmbeddings(
            d=int(rope_percentage * self.d_k)
        )
        self.key_rotary_pe = RotaryPositionalEmbeddings(
            d=int(rope_percentage * self.d_k)
        )
    
    def get_scores(self, query: torch.Tensor, key: torch.Tensor):
        # Apply RoPE to queries and keys

        query = self.query_rotary_pe(query)
        key = self.key_rotary_pe(key)
        
        # Calculate attention scores with position-aware vectors

        return torch.einsum('ibhd,jbhd->ijbh', query, key)

Reverse RoPE for Value-wise Encoding (RoPER)

The repository extends RoPE to value vectors through the ReverseRotaryPositionalEmbeddings class in labml_nn/transformers/rope/value_pe/__init__.py (lines 25-63). This variant rotates value tensors in the opposite direction using a flipped sine sign, enabling relative positional encoding in the attention output.

The reverse implementation inherits from the core class and modifies the rotation equation:

class ReverseRotaryPositionalEmbeddings(RotaryPositionalEmbeddings):
    def forward(self, x: torch.Tensor):
        # ... cache building logic ...

        
        # Reverse rotation: x_rot = x*cos - neg_half*sin

        x_rope = (x_rope * self.cos_cached[:seq_len]) - \
                 (neg_half_x * self.sin_cached[:seq_len])
        
        return torch.cat((x_rope, x_pass), dim=-1)

The RotaryValuePEMultiHeadAttention class coordinates both forward and reverse rotations: values enter through a standard RoPE module, participate in the weighted sum, then pass through ReverseRotaryPositionalEmbeddings to cancel the rotation while preserving relative positional information.

Practical Code Examples

Basic RoPE Module Usage

import torch
from labml_nn.transformers.rope import RotaryPositionalEmbeddings

# Input tensor: [seq_len, batch, heads, d]

x = torch.randn(5, 2, 4, 8)  # d=8 creates 4 rotation pairs

rope = RotaryPositionalEmbeddings(d=8)

# Apply positional embedding

x_rotated = rope(x)
print(x_rotated.shape)  # torch.Size([5, 2, 4, 8])

RoPE-Enhanced Multi-Head Attention

from labml_nn.transformers.rope import RotaryPEMultiHeadAttention

heads, d_model = 8, 256
attn = RotaryPEMultiHeadAttention(heads=heads, d_model=d_model)

# Dummy input tensors

seq_len, batch = 10, 3
query = torch.randn(seq_len, batch, d_model)
key = torch.randn(seq_len, batch, d_model)
value = torch.randn(seq_len, batch, d_model)

# Forward pass with rotary embeddings

output = attn(query=query, key=key, value=value)
print(output.shape)  # [seq_len, batch, d_model]

RoPER Implementation (Value-wise RoPE)

from labml_nn.transformers.rope.value_pe import RotaryValuePEMultiHeadAttention

attn = RotaryValuePEMultiHeadAttention(
    heads=8,
    d_model=256,
    rope_percentage=0.5,        # 50% of query/key dims rotated

    rope_value_percentage=0.5   # 50% of value dims rotated

)

# Same input tensors as above

output = attn(query=query, key=key, value=value)
print(output.shape)  # [seq_len, batch, d_model]

Summary

  • RotaryPositionalEmbeddings in labml_nn/transformers/rope/__init__.py provides the core implementation using cached sine/cosine matrices for computational efficiency.
  • The negative-half trick eliminates explicit rotation matrices by cleverly rearranging and negating tensor halves for element-wise multiplication.
  • Caching mechanism stores frequency tensors of shape (seq_len, 1, 1, d) to avoid recalculating trigonometric functions during training.
  • RotaryPEMultiHeadAttention demonstrates production integration, applying separate RoPE instances to queries and keys before attention score calculation.
  • ReverseRotaryPositionalEmbeddings enables the RoPER variant by reversing rotation direction, allowing positional encoding on value vectors.

Frequently Asked Questions

What makes RoPE different from standard positional encodings?

Standard positional encodings add sinusoidal vectors to input embeddings at the bottom of the transformer stack. RoPE instead rotates query and key vectors within the attention mechanism itself, encoding relative positions directly into the dot-product similarity calculation. This approach maintains translation invariance while providing better extrapolation to longer sequences than absolute encodings.

How does the caching mechanism improve performance?

The _build_cache method computes sine and cosine values only when encountering a sequence longer than previously seen. Since these values depend solely on position indices and fixed frequencies—not on input data—they can be reused across forward passes. This reduces the computational overhead from O(seq_len × d) trigonometric operations per layer to zero for typical training iterations after the first batch.

Can RoPE handle variable sequence lengths?

Yes, the implementation dynamically extends the cache when x.shape[0] exceeds self.seq_len_cached. The cache initialization checks if self.seq_len_cached < seq_len and rebuilds the frequency tensors only when necessary. This makes the module compatible with both training batches of varying length and inference autoregressive generation without manual cache management.

Where is reverse RoPE used in the codebase?

Reverse RoPE appears in labml_nn/transformers/rope/value_pe/__init__.py for the RoPER (Rotary Positional Embeddings for Values) variant. The RotaryValuePEMultiHeadAttention class applies standard RoPE to value vectors before the attention weighted sum, then uses ReverseRotaryPositionalEmbeddings to rotate them back. This encodes relative position information into the final attention output rather than the intermediate representations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →