Precomputed Frequency Tensor for Rotary Embeddings in Llama: Purpose and Implementation

The precomputed frequency tensor (freqs_cis) in Llama stores complex rotation values for rotary positional embeddings (RoPE) to avoid expensive trigonometric calculations during every forward pass, enabling efficient position encoding across all attention layers.

The meta-llama/llama repository implements rotary positional embeddings (RoPE) to inject positional information directly into query and key vectors. Rather than computing rotation angles repeatedly during inference, the model generates a precomputed frequency tensor once during initialization in Transformer.__init__. This optimization eliminates redundant trigonometric operations and reduces per-token computation to a single broadcast-compatible multiplication.

Why Llama Precomputes Rotary Frequencies

RoPE uses complex exponentials e^{i·θ·pos} to rotate query and key vectors in the complex plane. Computing these values on-the-fly for every token position would require expensive trigonometric operations at every attention layer.

The Computational Cost of On-the-Fly RoPE

Calculating sine and cosine values for every position during each forward pass introduces significant computational overhead. Since the rotation pattern depends only on head dimension and maximum sequence length—not the actual input tokens—the values remain static across all inference steps.

Static Properties Enable Caching

The frequency values depend solely on model architecture hyperparameters: the head dimension (dim // n_heads) and the maximum sequence length. This immutability allows Llama to compute the tensor once and reuse it across all layers, dramatically reducing per-token compute. According to the source code in llama/model.py, the tensor is moved to the appropriate device alongside activations, eliminating host-to-device transfer costs during inference.

How the Frequency Tensor is Constructed

In llama/model.py, the function precompute_freqs_cis (lines 80-104) generates a complex-valued tensor containing cosine-sine pairs (cis) for all positions up to max_seq_len. The resulting freqs_cis tensor has shape (max_seq_len * 2, dim_head), where the factor of 2 accommodates the sin-cos pair for each dimension.

The implementation creates rotation frequencies based on the formula 1.0 / (theta ** (torch.arange(0, dim, 2)[: (dim // 2)].float() / dim)), then computes the complex values for all positions through outer product and exponentiation.

Using the Precomputed Tensor During Inference

During the forward pass, the model slices the cached tensor according to the current sequence position and applies it via complex multiplication to rotate the query and key representations.

Slicing for Variable Sequence Lengths

In Transformer.forward (lines 71-73), the code extracts the relevant window using freqs_cis[start_pos:start_pos+seqlen]. This dynamic slicing supports variable-length sequences and efficient key-value caching during autoregressive generation, ensuring only the necessary positional encodings are applied.

Broadcasting to Query and Key Vectors

The reshape_for_broadcast function (lines 106-129) aligns the sliced tensor dimensions with the query and key tensors for element-wise multiplication. The apply_rotary_emb function (lines 156-162) interprets the real-valued tensors as complex numbers, multiplies them by the broadcasted freqs_cis, and converts the results back to real representations. This operation occurs within Attention.forward (lines 80-82) for every layer.

Code Implementation in meta-llama/llama

The following examples demonstrate the frequency tensor lifecycle from initialization to application.

1. Generating the frequency tensor (internally done once)


# Inside Transformer.__init__

self.freqs_cis = precompute_freqs_cis(
    self.params.dim // self.params.n_heads,   # head-dim

    self.params.max_seq_len * 2               # double for sin/cos pair

)

2. Using the tensor during a forward pass


# Slice the tensor for the current window

freqs_cis = self.freqs_cis[start_pos : start_pos + seqlen]

# In Attention.forward (for each layer)

xq, xk = apply_rotary_emb(xq, xk, freqs_cis=freqs_cis)

3. Minimal standalone demonstration

import torch
from llama.model import precompute_freqs_cis, apply_rotary_emb

head_dim = 64               # example per-head dimension

max_len = 128
freqs = precompute_freqs_cis(head_dim, max_len)          # (128, 64) complex tensor

# Fake query/key tensors: (batch, seq, heads, head_dim)

xq = torch.randn(2, 10, 8, head_dim * 2)   # *2 because we store real+imag parts

xk = torch.randn(2, 10, 8, head_dim * 2)

# Apply RoPE

xq_rot, xk_rot = apply_rotary_emb(xq, xk, freqs[:10])  # use first 10 positions

print(xq_rot.shape, xk_rot.shape)   # → torch.Size([2, 10, 8, head_dim*2])

Summary

  • Purpose: The freqs_cis tensor encodes positional rotation patterns for RoPE, avoiding repeated trigonometric calculations during inference.
  • Construction: Created once in Transformer.__init__ by precompute_freqs_cis (lines 80-104) as a complex-valued tensor of shape (max_seq_len * 2, dim_head).
  • Optimization: Values depend only on static architecture parameters (head dimension, max sequence length), not input data, enabling safe caching across all layers.
  • Application: Sliced dynamically per position in Transformer.forward (lines 71-73) and applied via apply_rotary_emb (lines 156-162) using broadcast-compatible multiplication.
  • Performance: Reduces per-token computation from trigonometric operations to a single complex multiplication, significantly improving inference efficiency.

Frequently Asked Questions

What is the shape of the precomputed frequency tensor in Llama?

The tensor has shape (max_seq_len * 2, dim_head), where dim_head is the per-head dimension (dim // n_heads). The factor of 2 in the sequence dimension accommodates the cosine-sine pair required for the complex rotation representation.

Why does the frequency tensor have a factor of 2 in its sequence length dimension?

The dimension is doubled to store both sine and cosine components for each frequency. As implemented in llama/model.py lines 80-104, the precompute_freqs_cis function generates complex values where the real and imaginary parts correspond to cosine and sine respectively, requiring twice the storage for the interleaved representation.

How does the precomputed tensor improve inference performance?

By computing the trigonometric values once during model initialization, Llama eliminates the need for torch.sin() and torch.cos() calls during every forward pass. This reduces the per-token computational overhead to a single complex multiplication (xq_ * freqs_cis) in apply_rotary_emb, significantly speeding up autoregressive generation across all attention layers.

Is the frequency tensor trainable or fixed during training?

The frequency tensor is fixed and non-trainable. It is computed purely from mathematical constants (position indices and frequency bases) in precompute_freqs_cis and stored as a buffer, not a parameter. This immutability ensures consistent positional encoding across training and inference without adding trainable parameters to the model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →