# How Rotary Positional Embedding Is Implemented in PyTorch: A Deep Dive into labmlai/annotated_deep_learning_paper_implementations

> Learn how Rotary Positional Embedding is implemented in PyTorch using labmlai. Discover efficient rotation techniques with cached sine cosine matrices.

- Repository: [labml.ai/annotated_deep_learning_paper_implementations](https://github.com/labmlai/annotated_deep_learning_paper_implementations)
- Tags: deep-dive
- Published: 2026-03-04

---

**The labmlai/annotated_deep_learning_paper_implementations repository implements Rotary Positional Embeddings (RoPE) as a pure-PyTorch module using cached sine/cosine matrices and a "negative-half" tensor trick to efficiently rotate query and key vectors by their position indices.**

Rotary Positional Embedding (RoPE) encodes absolute positional information into transformer attention layers by rotating feature vectors in a two-dimensional subspace. According to the source code in `labmlai/annotated_deep_learning_paper_implementations`, the implementation provides drop-in modules that replace standard positional encodings while maintaining computational efficiency through smart caching mechanisms.

## Core RoPE Module Architecture

The primary implementation resides in [`labml_nn/transformers/rope/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/rope/__init__.py) within the `RotaryPositionalEmbeddings` class (lines 30-71). This module implements the mathematical formulation from the RoFormer paper, applying rotation to paired dimensions of the input tensor.

### Pair-wise Rotation Mechanism

RoPE splits the feature dimension `d` into `d/2` pairs. Each pair `(x^{(1)}_m, x^{(2)}_m)` at position `m` rotates by an angle `m·θ_i`, where `θ_i` represents a frequency specific to that pair following the classic sinusoidal schedule. This rotation injects relative positional information directly into the dot-product attention calculation.

### Cached Sine and Cosine Tensors

Computing trigonometric functions for every forward pass would create significant overhead. The implementation builds a **cached tensor** of shape `(seq_len, 1, 1, d)` the first time it encounters a sequence length exceeding the current cache size. Subsequent calls reuse these cached values:

```python
def _build_cache(self, x: torch.Tensor):
    # Calculate theta values for each dimension pair

    theta = 1.0 / (self.base ** (torch.arange(0, self.d, 2).float() / self.d))
    
    # Position indices

    seq_idx = torch.arange(self.seq_len_cached, device=x.device).float()
    
    # Calculate the angles: m * theta_i

    idx_theta = torch.einsum('i,j->ij', seq_idx, theta)
    
    # Cache both sine and cosine

    self.cos_cached = idx_theta.cos()[None, None, :, :]
    self.sin_cached = idx_theta.sin()[None, None, :, :]

```

### The Negative-Half Trick

To apply the rotation efficiently without explicit matrix multiplication, the code constructs a "neg-half" tensor. This operation concatenates the second half of features with a sign flip (`-x_{d/2+1…d}`) followed by the first half (`x_{1…d/2}`). This enables single element-wise multiplication with the cached sine tensor:

```python
def _neg_half(self, x: torch.Tensor):
    # Split and flip with negation

    d_2 = self.d // 2
    return torch.cat([-x[:, :, :, d_2:], x[:, :, :, :d_2]], dim=-1)

```

## Forward Pass Implementation

The `forward` method in `RotaryPositionalEmbeddings` orchestrates the rotation through five distinct steps:

1. **Cache initialization**: `self._build_cache(x)` prepares or reuses sine/cosine matrices
2. **Tensor splitting**: Separates rotatable dimensions from pass-through features
3. **Neg-half construction**: Creates the rotated counterpart using `_neg_half`
4. **Rotation application**: Combines original and neg-half tensors with cached values
5. **Concatenation**: Reassembles rotated and unrotated features

```python
def forward(self, x: torch.Tensor):
    # 1. Build or reuse cache

    self._build_cache(x)
    
    # 2. Split rotatable vs pass-through dimensions

    x_rope, x_pass = x[..., :self.d], x[..., self.d:]
    
    # 3. Create negative-half version

    neg_half_x = self._neg_half(x_rope)
    
    # 4. Apply rotation: x_rot = x*cos + neg_half*sin

    x_rope = (x_rope * self.cos_cached[:seq_len]) + \
             (neg_half_x * self.sin_cached[:seq_len])
    
    # 5. Concatenate results

    return torch.cat((x_rope, x_pass), dim=-1)

```

## Integration with Multi-Head Attention

The `RotaryPEMultiHeadAttention` class demonstrates practical integration. It instantiates separate RoPE modules for queries and keys, applying them before the attention score calculation:

```python
class RotaryPEMultiHeadAttention(MultiHeadAttention):
    def __init__(self, heads: int, d_model: int, dropout_prob: float = 0.0, 
                 rope_percentage: float = 0.5):
        super().__init__(heads, d_model, dropout_prob)
        self.query_rotary_pe = RotaryPositionalEmbeddings(
            d=int(rope_percentage * self.d_k)
        )
        self.key_rotary_pe = RotaryPositionalEmbeddings(
            d=int(rope_percentage * self.d_k)
        )
    
    def get_scores(self, query: torch.Tensor, key: torch.Tensor):
        # Apply RoPE to queries and keys

        query = self.query_rotary_pe(query)
        key = self.key_rotary_pe(key)
        
        # Calculate attention scores with position-aware vectors

        return torch.einsum('ibhd,jbhd->ijbh', query, key)

```

## Reverse RoPE for Value-wise Encoding (RoPER)

The repository extends RoPE to value vectors through the `ReverseRotaryPositionalEmbeddings` class in [`labml_nn/transformers/rope/value_pe/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/rope/value_pe/__init__.py) (lines 25-63). This variant rotates value tensors in the opposite direction using a flipped sine sign, enabling relative positional encoding in the attention output.

The reverse implementation inherits from the core class and modifies the rotation equation:

```python
class ReverseRotaryPositionalEmbeddings(RotaryPositionalEmbeddings):
    def forward(self, x: torch.Tensor):
        # ... cache building logic ...

        
        # Reverse rotation: x_rot = x*cos - neg_half*sin

        x_rope = (x_rope * self.cos_cached[:seq_len]) - \
                 (neg_half_x * self.sin_cached[:seq_len])
        
        return torch.cat((x_rope, x_pass), dim=-1)

```

The `RotaryValuePEMultiHeadAttention` class coordinates both forward and reverse rotations: values enter through a standard RoPE module, participate in the weighted sum, then pass through `ReverseRotaryPositionalEmbeddings` to cancel the rotation while preserving relative positional information.

## Practical Code Examples

### Basic RoPE Module Usage

```python
import torch
from labml_nn.transformers.rope import RotaryPositionalEmbeddings

# Input tensor: [seq_len, batch, heads, d]

x = torch.randn(5, 2, 4, 8)  # d=8 creates 4 rotation pairs

rope = RotaryPositionalEmbeddings(d=8)

# Apply positional embedding

x_rotated = rope(x)
print(x_rotated.shape)  # torch.Size([5, 2, 4, 8])

```

### RoPE-Enhanced Multi-Head Attention

```python
from labml_nn.transformers.rope import RotaryPEMultiHeadAttention

heads, d_model = 8, 256
attn = RotaryPEMultiHeadAttention(heads=heads, d_model=d_model)

# Dummy input tensors

seq_len, batch = 10, 3
query = torch.randn(seq_len, batch, d_model)
key = torch.randn(seq_len, batch, d_model)
value = torch.randn(seq_len, batch, d_model)

# Forward pass with rotary embeddings

output = attn(query=query, key=key, value=value)
print(output.shape)  # [seq_len, batch, d_model]

```

### RoPER Implementation (Value-wise RoPE)

```python
from labml_nn.transformers.rope.value_pe import RotaryValuePEMultiHeadAttention

attn = RotaryValuePEMultiHeadAttention(
    heads=8,
    d_model=256,
    rope_percentage=0.5,        # 50% of query/key dims rotated

    rope_value_percentage=0.5   # 50% of value dims rotated

)

# Same input tensors as above

output = attn(query=query, key=key, value=value)
print(output.shape)  # [seq_len, batch, d_model]

```

## Summary

- **RotaryPositionalEmbeddings** in [`labml_nn/transformers/rope/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/rope/__init__.py) provides the core implementation using cached sine/cosine matrices for computational efficiency.
- The **negative-half trick** eliminates explicit rotation matrices by cleverly rearranging and negating tensor halves for element-wise multiplication.
- **Caching mechanism** stores frequency tensors of shape `(seq_len, 1, 1, d)` to avoid recalculating trigonometric functions during training.
- **RotaryPEMultiHeadAttention** demonstrates production integration, applying separate RoPE instances to queries and keys before attention score calculation.
- **ReverseRotaryPositionalEmbeddings** enables the RoPER variant by reversing rotation direction, allowing positional encoding on value vectors.

## Frequently Asked Questions

### What makes RoPE different from standard positional encodings?

Standard positional encodings add sinusoidal vectors to input embeddings at the bottom of the transformer stack. RoPE instead rotates query and key vectors within the attention mechanism itself, encoding relative positions directly into the dot-product similarity calculation. This approach maintains translation invariance while providing better extrapolation to longer sequences than absolute encodings.

### How does the caching mechanism improve performance?

The `_build_cache` method computes sine and cosine values only when encountering a sequence longer than previously seen. Since these values depend solely on position indices and fixed frequencies—not on input data—they can be reused across forward passes. This reduces the computational overhead from `O(seq_len × d)` trigonometric operations per layer to zero for typical training iterations after the first batch.

### Can RoPE handle variable sequence lengths?

Yes, the implementation dynamically extends the cache when `x.shape[0]` exceeds `self.seq_len_cached`. The cache initialization checks `if self.seq_len_cached < seq_len` and rebuilds the frequency tensors only when necessary. This makes the module compatible with both training batches of varying length and inference autoregressive generation without manual cache management.

### Where is reverse RoPE used in the codebase?

Reverse RoPE appears in [`labml_nn/transformers/rope/value_pe/__init__.py`](https://github.com/labmlai/annotated_deep_learning_paper_implementations/blob/main/labml_nn/transformers/rope/value_pe/__init__.py) for the RoPER (Rotary Positional Embeddings for Values) variant. The `RotaryValuePEMultiHeadAttention` class applies standard RoPE to value vectors before the attention weighted sum, then uses `ReverseRotaryPositionalEmbeddings` to rotate them back. This encodes relative position information into the final attention output rather than the intermediate representations.