How Rotary Embeddings Are Implemented in NanoChat: A Complete Guide
NanoChat implements rotary positional embeddings (RoPE) as a drop-in replacement for traditional absolute positional encodings, pre-computing rotation matrices in nanochat/gpt.py and applying them to query and key tensors inside each CausalSelfAttention layer.
The karpathy/nanochat repository leverages rotary embeddings to encode relative positional information without learned parameters. Unlike standard transformer implementations that add absolute position vectors to token embeddings, NanoChat’s approach rotates query and key representations within the attention mechanism itself, enabling better generalization to longer sequences.
Pre-computing the Rotary Buffers
The GPT class constructor initializes the rotary embedding caches during model instantiation. These tensors are computed once and reused across all layers to minimize overhead.
self.rotary_seq_len = config.sequence_len * 10 # over‑allocate (10×)
head_dim = config.n_embd // config.n_head
cos, sin = self._precompute_rotary_embeddings(self.rotary_seq_len, head_dim)
self.register_buffer("cos", cos, persistent=False)
self.register_buffer("sin", sin, persistent=False)
The _precompute_rotary_embeddings method generates cosine and sine tensors of shape (1, seq_len, 1, head_dim/2) using a frequency base of 100000, consistent with the original RoPE paper. These buffers are registered as non-persistent, meaning they are excluded from model checkpoints and can be recomputed on the fly during loading.
Applying RoPE to Queries and Keys
During the forward pass of each CausalSelfAttention layer, the model fetches the pre-computed buffers and applies them to both the query (q) and key (k) tensors before computing attention scores.
cos, sin = cos_sin # fetched from self.cos / self.sin
q, k = apply_rotary_emb(q, cos, sin), apply_rotary_emb(k, cos, sin)
The apply_rotary_emb function (defined at line 57 in nanochat/gpt.py) implements the rotation by splitting the last dimension of the input tensor into two halves and applying a 2D rotation matrix:
d = x.shape[3] // 2
x1, x2 = x[..., :d], x[..., d:] # split
y1 = x1 * cos + x2 * sin
y2 = x1 * (-sin) + x2 * cos
return torch.cat([y1, y2], 3)
This operation injects relative positional information directly into the dot-product attention computation without introducing any learnable parameters for absolute positions.
KV-Cache Safety and Dynamic Sequence Length
When using key-value caching during inference, NanoChat asserts that the current sequence length does not exceed the pre-allocated rotary buffer size:
assert T <= self.cos.size(1), \
f"Sequence length grew beyond the rotary embeddings cache: {T} > {self.cos.size(1)}"
The repository over-allocates the cache by 10× the configured sequence_len to accommodate longer sequences without recomputation, though the source notes that exceeding this buffer requires rebuilding the caches.
Practical Usage Examples
To inspect the rotary embedding mechanism in a live model:
import torch
from nanochat.gpt import GPT, GPTConfig
# Build a config and model
cfg = GPTConfig(sequence_len=2048, n_embd=768, n_head=6, n_kv_head=6, n_layer=12)
model = GPT(cfg)
# Initialise weights (creates the rotary buffers)
model.init_weights()
# Run a forward pass (rotary embeddings are applied internally)
x = torch.randint(0, cfg.vocab_size, (1, 16))
logits = model(x)
print(logits.shape) # torch.Size([1, 16, cfg.vocab_size])
For manual application to custom tensors:
# Access pre-computed buffers
cos, sin = model.cos, model.sin # shape (1, seq_len, 1, head_dim/2)
# Apply to a query tensor manually
q = torch.randn(1, 8, cfg.n_head, cfg.n_embd // cfg.n_head)
q_rot = model.apply_rotary_emb(q, cos[:, :8], sin[:, :8])
Summary
- RoPE replaces absolute embeddings: NanoChat eliminates traditional positional embeddings entirely, using rotary embeddings as the sole positional encoding mechanism.
- Pre-computation in
GPT.__init__: The_precompute_rotary_embeddingsmethod generates cosine and sine matrices shaped(1, seq_len, 1, head_dim/2)withbase=100000. - Non-persistent buffers: The
cosandsintensors are registered as buffers withpersistent=False, keeping checkpoints smaller. - Application in attention: The
apply_rotary_embfunction splits the final dimension of Q/K tensors and applies 2D rotation before the attention dot-product. - Safety checks: Inference with KV-cache triggers an assertion if
T > self.cos.size(1), preventing silent errors from sequence length overflow.
Frequently Asked Questions
Why does NanoChat use rotary embeddings instead of absolute positional embeddings?
Rotary embeddings provide relative positional information that extrapolates more naturally to sequence lengths beyond the training distribution. According to the source comments in nanochat/gpt.py, the design explicitly lists "rotary embeddings (and no positional embeddings)" as a core architectural decision, enabling smoother scaling to longer contexts without learned position parameters.
What is the exact shape of the pre-computed rotary buffers?
The cos and sin buffers have shape (1, seq_len, 1, head_dim/2), where head_dim equals config.n_embd // config.n_head. The dimensions are arranged to broadcast correctly over the batch and head dimensions during the rotation operation inside apply_rotary_emb.
How does apply_rotary_emb rotate the query and key tensors?
The function splits the last dimension of the input tensor into two equal halves (x1 and x2), then computes rotated outputs using y1 = x1 * cos + x2 * sin and y2 = x1 * (-sin) + x2 * cos, concatenating the results. This mathematical operation corresponds to multiplying each pair of channels by a rotation matrix parameterized by the position index.
What happens if the input sequence exceeds the rotary cache size?
During inference, the CausalSelfAttention layer asserts that the current sequence length T is less than or equal to self.cos.size(1). If this condition fails, the model raises an AssertionError indicating that the sequence length has grown beyond the rotary embeddings cache. The repository over-allocates by 10× the configured length to prevent this in most practical scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →