Needle 2 Memory Architecture for Sliding-Window KV Sinks: A Technical Deep Dive

Needle 2 implements a bounded-memory sliding-window KV cache that caps memory usage at ~28 MiB by computing a dynamic token window from a KV budget, while preserving tool-related tokens as permanent "sinks" that never drop from cache.

This article examines how the Needle 2 inference engine from cactus-compute/needle solves the memory explosion problem of transformer KV caches. Rather than growing unbounded with conversation length, Needle 2's sliding-window KV sinks architecture maintains constant memory regardless of sequence length—critical for edge deployment.

How Needle 2 Calculates KV Memory Budget

The foundation of Needle 2's memory architecture is kv_budget_window() in needle/model/architecture.py (lines 1004-1012). This function determines the maximum tokens storable without exceeding KV_BUDGET_BYTES (≈ 11.5 MiB).

The budget computation accounts for:

  • KV-cache width: num_kv_heads * head_dim
  • Layer count: Number of transformer layers
  • Engram KV slots: Additional persistent memory slots
  • Group-wise packing: Processed in chunks of KV_GROUP = 32

The result is clamped to a minimum of 160 tokens and capped at config.max_seq_len. By default, this yields approximately 256 tokens for typical configurations—enabling the ~28 MiB total memory footprint claimed in the README.

Effective KV Window: Budget vs. User Override

Needle 2 allows explicit window control through effective_kv_window() (architecture.py line 1014).

Default behavior (kv_window=0): Uses budget-derived window.

User override (kv_window > 0): Takes min(budget_window, user_window).

from needle.model.architecture import TransformerConfig, effective_kv_window

# Budget-derived window (default)

cfg = TransformerConfig(vocab_size=8192, kv_window=0)
print("Effective KV window:", effective_kv_window(cfg))  # → ~256

# Force smaller window for aggressive memory limits

cfg_small = TransformerConfig(vocab_size=8192, kv_window=128)
print("Capped window:", effective_kv_window(cfg_small))  # → 128

Sliding-Window Mask Implementation

During forward passes, the hidden_cells function (architecture.py line 544) constructs the attention mask through a two-step intersection:

  1. Base causal mask: Standard autoregressive constraint
  2. Recent mask: Boolean filter keeping only last window positions

The recent mask uses positional arithmetic (line 545):

recent = ((pos[:, None] - pos[None, :]) < window)

This drops tokens outside the sliding window from attention, effectively pruning their KV entries from cache.

KV Sinks: Preserving Tool Tokens Beyond the Window

The critical innovation for tool-using agents is KV sinks—selectively pinned KV entries that survive the sliding window. This is implemented via mask OR-ing in hidden_cells (line 548):

keep = recent if sink is None else (recent | sink[:, None, None, :])

How sink masks work:

  • Tool tokens are identified by segment IDs in seg_ids
  • sink_mask = (seg_ids > 0) marks tool positions
  • The sink mask is broadcast and OR-ed with the recent mask
  • Result: Tool KV entries remain visible to all positions even after falling outside the 256-token window

The make_causal_packing_mask function (architecture.py line 1020) provides the public API for this masking logic:

import jax.numpy as jnp
from needle.model.architecture import make_causal_packing_mask

# seg_ids: 0 = conversation, 1 = tool segment

seg_ids = jnp.array([[0, 1, 1, 0, 0, 0, 0]])  # two tool tokens at positions 1-2

sink_mask = (seg_ids > 0)  # True where tools appear

# Build mask: causal + sliding window (256) + preserved sinks

mask = make_causal_packing_mask(
    seg_ids, 
    prefix=None, 
    window=256  # Tokens 0,3,4,5,6 subject to window; tokens 1,2 persist as sinks

)

Memory Architecture in Training vs. Inference

The sliding-window KV sinks design serves dual purposes:

Phase Window Behavior Sink Purpose
Training Enforces fixed memory per sample regardless of sequence length Tool demonstrations remain in context for imitation learning
Inference Cache never exceeds ~28 MiB Tool definitions and API schemas persist across long conversations

This bounded-memory guarantee enables deployment on memory-constrained devices where unbounded KV caches would exhaust RAM within thousands of tokens.

Key Source Files

Understanding Needle 2's memory architecture requires examining these files:

Summary

  • KV budget calculation: kv_budget_window() derives token capacity from fixed 11.5 MiB memory target
  • Effective window: effective_kv_window() merges budget constraints with optional user caps
  • Sliding-window mask: hidden_cells implements position-based token eviction via boolean masking
  • KV sinks: Tool tokens bypass eviction through OR-ed sink masks preserving cross-attention visibility
  • Constant memory: Architecture guarantees ~28 MiB regardless of conversation length

Frequently Asked Questions

How does Needle 2 prevent KV cache memory growth?

Needle 2 computes a fixed token window from KV_BUDGET_BYTES in kv_budget_window(), then enforces it through position masking in hidden_cells(). Tokens outside the window are excluded from attention, allowing their KV entries to be freed. This caps memory regardless of sequence length.

What happens when I set kv_window smaller than the budget-derived value?

effective_kv_window() takes the minimum of budget-derived and user-specified values. Setting kv_window=128 with a budget-derived 256 yields 128, reducing memory further at the cost of context retention.

Why are tool tokens called "sinks" and how do they persist?

"Sink" refers to permanent visibility in the attention sink (output). The mask logic ORs sink_mask with recent_mask, forcing attention to tool positions even when they're far outside the sliding window. This preserves tool functionality without expanding cache memory.

Can I disable sliding-window and use full attention?

No—Needle 2 is architecturally designed around bounded memory. The budget calculation runs unconditionally, and while kv_window can increase effective context, the underlying mask construction always applies window constraints. Full unbounded attention would violate the memory guarantees.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →