Implementing Transformer Positional Encodings (RoPE, ALiBi): From Sinusoidal to State-of-the-Art

RoPE and ALiBi are the dominant positional encoding schemes in modern LLMs, with RoPE rotating query/key vectors to encode relative position and ALiBi applying linear distance penalties to attention scores, both enabling efficient long-context extrapolation beyond traditional sinusoidal methods.

Transformers are permutation-invariant architectures that process tokens without inherent sequence order, making implementing transformer positional encodings critical for distinguishing "cat sat" from "sat cat". The ai-engineering-from-scratch curriculum teaches three production-ready schemes—absolute sinusoidal, Rotary Position Embedding (RoPE), and Attention with Linear Biases (ALiBi)—through pure Python implementations that clarify the mechanics behind Llama, Mistral, and BLOOM.

Why Transformers Need Positional Information

Self-attention mechanisms compute pairwise token interactions using dot-products, producing identical results regardless of token order. Without positional information, the attention matrix treats every position as interchangeable, rendering the model incapable of sequence understanding. Positional encodings inject order information either by adding vectors to input embeddings or by modifying the attention computation directly.

Absolute Sinusoidal Encodings: The Historical Baseline

The original Transformer architecture (Vaswani et al., 2017) introduced absolute sinusoidal encodings as a parameter-free method to represent position. A (max_len, d_model) matrix is filled with sinusoidal waves of varying frequencies using the formula sin(pos/10000^{2i/d}) and cos(pos/10000^{2i/d}) (see lines 30‑35 of phases/07-transformers-deep-dive/04-positional-encoding/docs/en.md). This matrix is added to token embeddings before the first attention layer.

When to use it: Sinusoidal encodings serve as a pedagogical baseline for tiny toy models and historic reproductions, though they extrapolate poorly beyond the training sequence length.

from phases.phase_07_transformers_deep_dive.positional_encoding.code.main import sinusoidal_pe

# Generate a sinusoidal table for max_len=8, d_model=8

pe = sinusoidal_pe(n=8, d=8)
print("First four positions, first four dimensions:")
for pos in range(4):
    print(f"pos={pos}:", pe[pos][:4])

RoPE (Rotary Position Embedding): The Modern Standard

RoPE has become the de-facto standard in Llama 2/3/4, Qwen 2/3, Mistral, DeepSeek-V3, and Kimi. Instead of adding position vectors to embeddings, RoPE rotates the query and key vectors of each head by a position-dependent angle before the dot-product. For each dimension pair (2i, 2i+1), a 2-D rotation matrix [[cos(pos·θ_i), -sin(pos·θ_i)], [sin(pos·θ_i), cos(pos·θ_i)]] is applied (see the derivation in lines 41‑48).

Because the same rotation applies to both Q and K vectors, their dot-product implicitly contains cos((pos_q - pos_k)·θ_i), making the attention score depend only on the relative distance between tokens. This property enables extrapolation to longer contexts when combined with base frequency scaling techniques like NTK-aware scaling, YaRN, or LongRoPE.

import random
from phases.phase_07_transformers_deep_dive.positional_encoding.code.main import apply_rope, dot

# Demonstrate that dot products depend only on relative distance

rng = random.Random(0)
d = 16
q = [rng.gauss(0, 1) for _ in range(d)]
k = [rng.gauss(0, 1) for _ in range(d)]

pairs = [(3, 5), (7, 9), (100, 102), (1024, 1026)]
for pq, pk in pairs:
    q_rot = apply_rope(q, pq)
    k_rot = apply_rope(k, pk)
    print(f"gap={pk-pq:2d}  dot={dot(q_rot, k_rot):.6f}")

# All rows with the same gap produce (almost) identical dot products.

ALiBi (Attention with Linear Biases): Zero-Cost Extrapolation

ALiBi provides robust length extrapolation without adding parameters or embeddings. For each head h, a slope m_h = 2^{-8·h/H} is multiplied by the absolute token distance and subtracted directly from the raw attention scores (see lines 58‑60). This biases nearby tokens upward and distant tokens downward, encouraging short-range focus while maintaining global attention capability.

When to use it: ALiBi powers models like BLOOM, MPT, and Baichuan that require extreme length extrapolation with zero training overhead, though it may slightly reduce short-range modeling power compared to learned embeddings.

from phases.phase_07_transformers_deep_dive.positional_encoding.code.main import alibi_bias

# Generate ALiBi bias matrix for 4 heads, seq_len=6

bias = alibi_bias(n_heads=4, seq_len=6, causal=False)
print("Bias for head 0 (closer tokens get smaller penalty):")
for row in bias[0]:
    print(row)

Selecting the Right Positional Encoding for Your Model

The curriculum provides a decision framework in phases/07-transformers-deep-dive/04-positional-encoding/outputs/skill-positional-encoding-picker.md to guide selection:

  • Sinusoidal: Choose for educational purposes, tiny models (< 100M parameters), or when reproducing 2017 Transformer results. Poor extrapolation beyond training length.

  • RoPE: Default for modern LLMs. Offers a single hyper-parameter (base frequency) that scales to millions of tokens via techniques like YaRN. Keeps model architecture unchanged while injecting relative position directly into attention.

  • ALiBi: Select when you need guaranteed extrapolation to sequences longer than training data without retraining. Ideal for document processing and retrieval-augmented generation with variable-length contexts.

Integration into the Transformer Stack

In the reference implementation (phases/07-transformers-deep-dive/04-positional-encoding/code/main.py), sinusoidal encodings are added to token embeddings before the first attention block. For RoPE or ALiBi, positional information is applied inside the attention block—either by rotating Q/K vectors or subtracting the bias matrix from attention scores—allowing you to omit the embedding sum entirely.

All implementations use pure Python/stdlib for clarity, making it straightforward to experiment with concepts before replacing them with high-performance kernels like Flash Attention or torch.nn.functional.rope.

Summary

  • RoPE rotates query and key vectors by position-dependent angles, making attention scores depend only on relative distance and enabling context window extension through base frequency scaling.

  • ALiBi subtracts a linear distance penalty from attention scores, providing zero-cost extrapolation without additional parameters or embeddings.

  • Sinusoidal encodings serve as a historical baseline but fail to generalize beyond training lengths.

  • The ai-engineering-from-scratch repository provides minimal, educational implementations in main.py that demonstrate the mathematical properties of each encoding through runnable demonstrations.

Frequently Asked Questions

What is the main advantage of RoPE over sinusoidal positional encodings?

RoPE encodes relative position rather than absolute position, meaning the model naturally generalizes to longer sequences and the attention score depends only on the distance between tokens, not their absolute locations. Additionally, RoPE requires no additional embedding parameters and allows context window extension through the base frequency hyper-parameter using techniques like NTK-aware scaling or YaRN.

When should I use ALiBi instead of RoPE?

Use ALiBi when you need guaranteed extrapolation to sequences significantly longer than training data without any retraining or architectural changes. ALiBi is ideal for production systems processing documents of variable and extreme lengths, though it may sacrifice some short-range modeling precision compared to RoPE's fine-grained rotational encoding.

How does RoPE handle longer sequences than seen during training?

RoPE extrapolates to longer sequences by adjusting the base frequency θ_i used in the rotation matrix. Techniques like NTK-aware interpolation, YaRN (Yet another RoPE extension method), and LongRoPE modify how position frequencies are computed or interpolated, allowing models trained on 4K tokens to handle 128K+ contexts without catastrophic attention degradation.

Can I combine sinusoidal encodings with RoPE or ALiBi?

No, these approaches are mutually exclusive in the attention block. Sinusoidal encodings are added to input embeddings before the first layer, while RoPE and ALiBi modify the attention computation internally—RoPE by rotating Q/K vectors and ALiBi by biasing attention scores. Using sinusoidal embeddings with RoPE or ALiBi would double-count positional information and distort the attention patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →