How ALiBi Removes Positional Embeddings While Preserving Attention Patterns
ALiBi eliminates traditional positional embedding layers by injecting deterministic, distance‑based linear biases directly into attention logits before the softmax, leveraging softmax translation invariance to maintain relative attention ordering without learned position vectors.
The Attention with Linear Biases (ALiBi) method redefines how transformer models encode sequence order. According to the labmlai/annotated_deep_learning_paper_implementations source code, this approach explicitly disables positional embeddings through a special 'no_pos' configuration while preserving the monotonic distance relationships that attention mechanisms rely on. Unlike learned absolute position embeddings that add vectors to token representations, ALiBi modifies the attention scoring function itself through head‑specific linear biases calibrated by query‑key distances.
Eliminating Learned Positional Embeddings
Configuration‑Level Removal
In the ALiBi transformer implementation, the embedding module is initialized with the string identifier 'no_pos'. This configuration—visible in labml_nn/transformers/alibi/experiment.py—signals the model to omit absolute positional vectors entirely, ensuring token representations remain purely semantic without positional interference.
Architecture Changes
Rather than augmenting input embeddings with position vectors, the architecture shifts positional information to the attention mechanism itself. The AlibiMultiHeadAttention class (defined in labml_nn/transformers/alibi/__init__.py) overrides the standard multi‑head attention forward pass to exclude positional dependencies from the initial embedding layer.
Injecting Distance‑Based Attention Biases
The Linear Bias Calculation
After computing scaled dot‑product attention scores (QK^T), ALiBi adds a predefined bias matrix before the softmax operation. The get_alibi_biases function generates these biases based on the distance between query and key positions, applying the formula bias = -m * |i - j| where m represents a head‑specific slope and i, j are token indices.
Head‑Specific Slopes
Each attention head receives a distinct slope parameter computed by get_slopes. This creates a family of linear bias functions—steeper slopes for some heads, shallower for others—allowing the model to learn diverse positional sensitivities across different representation subspaces. As implemented in the repository, these slopes follow a geometric progression calculated as m = 2^(-8/h) for head index h, ensuring each head specializes in different ranges of contextual dependencies.
Preserving Attention Patterns Through Softmax Invariance
Translation Invariance Properties
Softmax operations are translation invariant—adding a constant value to every element of a vector does not alter the resulting probability distribution. By adding the ALiBi bias matrix uniformly across attention logits (dependent only on distance, not token content), the relative ordering of scores remains unchanged while their absolute magnitudes shift according to query‑key distances.
Monotonic Distance Relationships
The linear bias grows negatively with the distance between query and key positions (larger distance = more negative bias). Because softmax preserves rank order under uniform shifts, faraway tokens receive proportionally lower attention weights while nearby tokens maintain their relative importance patterns. This deterministic approach replaces learned embeddings with a geometric prior grounded in sequence distance.
Practical Implementation in LabML
The following example demonstrates how to use the ALiBi attention module without explicit positional encoding:
import torch
from labml_nn.transformers.alibi import AlibiMultiHeadAttention, get_alibi_biases
# Configuration parameters
seq_len, batch_size, d_model = 8, 2, 64
n_heads = 4
# Dummy input tensor without positional encoding
x = torch.randn(seq_len, batch_size, d_model)
# Create causal mask (lower triangular boolean mask)
mask = torch.triu(torch.ones(seq_len, seq_len), diagonal=1).bool().logical_not().float()
# Initialize ALiBi multi-head attention
mha = AlibiMultiHeadAttention(heads=n_heads, d_model=d_model)
# Forward pass automatically applies distance biases
out = mha(query=x, key=x, value=x, mask=mask)
print(out.shape) # torch.Size([8, 2, 64])
The AlibiMultiHeadAttention class automatically computes and applies bias matrices during the forward pass via get_alibi_biases, eliminating the need for explicit positional encoding in the input preparation phase.
Summary
- ALiBi removes positional embeddings by configuring the transformer with
'no_pos'embeddings, completely omitting learned position vectors from the input layer as defined inlabml_nn/transformers/alibi/experiment.py. - Linear biases replace position IDs through the
get_alibi_biasesfunction, which computes distance‑dependent penalties added to attention logits before softmax calculation. - Attention patterns stay intact because softmax is translation invariant—adding uniform bias shifts preserves relative score rankings while encoding monotonic distance relationships.
- Head‑specific slopes enable diversity via
get_slopes, allowing different attention heads to specialize in short‑range or long‑range dependencies without increasing parameter count.
Frequently Asked Questions
What happens to tokens without positional embeddings?
Without traditional positional vectors, tokens rely solely on content‑based representations. The ALiBi mechanism restores positional awareness by penalizing attention scores between distant tokens through the get_alibi_biases calculation, ensuring the model understands sequence order through relative distance penalties rather than absolute position IDs.
Why doesn't adding a bias change the attention distribution?
Softmax normalization divides exponentiated scores by their sum, making the operation translation invariant. When ALiBi adds the same bias value to all attention logits for a given query (based solely on distance), the relative differences between candidate keys remain constant, preserving the original attention weight distribution while incorporating geometric distance information.
How are slopes assigned to different attention heads?
The get_slopes function geometrically partitions the range of attention biases across heads using the formula m = 2^(-8/h) where h indexes the head. This creates a set of linear functions with varying steepness, allowing some heads to focus on local context (steep slope = rapid distance penalty) while others attend to global patterns (shallow slope = gradual penalty) without additional learned parameters.
Can ALiBi handle sequences longer than training lengths?
Yes. ALiBi exhibits superior length generalization compared to learned positional embeddings because the linear bias function m * |i - j| extends naturally to arbitrary distances. The deterministic distance computation in get_alibi_biases requires no learned parameters for specific positions, enabling the model to process longer sequences than seen during training without interpolation errors or boundary effects.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →