How to Implement Attention Sink for Managing Long Context in Video Generation with LongLive
Implement an attention sink in LongLive by setting the sink_size parameter when instantiating the Ulysses sequence-parallel model to permanently retain early video frames in the KV-cache while automatically rolling the oldest local tokens.
Long video generation with diffusion transformers creates a memory bottleneck because every new frame adds thousands of tokens (pixels × frames) to the Key-Value (KV) cache. The NVIDIA LongLive repository solves this through an attention sink mechanism that anchors initial frames to prevent context overflow while maintaining generative coherence. This guide explains exactly how to configure and implement attention sinks using the actual source code from NVlabs/LongLive.
What Is an Attention Sink and Why You Need It
Standard transformer attention stores every token's key and value tensors in a KV-cache that grows linearly with sequence length. For long videos, this quickly exhausts GPU memory.
The KV-Cache Bottleneck
In causal video generation, the model processes frames sequentially, appending each new frame's tokens to the cache. Without intervention, the cache size equals the total number of generated tokens, causing out-of-memory errors on long sequences. According to the LongLive source code in wan_5b/modules/causal_model_sp_ulysses.py, the solution involves anchoring a configurable number of early frames—the sink—that are never evicted when the cache rolls.
Sink Architecture Components
LongLive implements a hierarchical sink system with three configurable components:
sink_size: The legacy sink parameter defining how many initial frames remain permanently in the cacheglobal_sink_size: Frames that are globally anchored for the entire generation run, used specifically for multi-shot generation scenarios- Pinned region: A short block of frames immediately following the global sink that must stay together during cache operations
When the cache reaches capacity, the model rolls by evicting the oldest local tokens that appear after the effective sink boundary (global sink + pinned region). This bounds memory usage while preserving attention to the earliest context.
How Attention Sink Works Internally
The sink logic resides primarily in the UlyssesCausalWanSelfAttention class within wan_5b/modules/causal_model_sp_ulysses.py. A non-parallel fallback implementation exists in wan_5b/modules/causal_model.py as CausalWanSelfAttention.
Initialization Parameters
During model construction, the attention layer stores sink configuration as instance attributes:
# From wan_5b/modules/causal_model_sp_ulysses.py (lines 85-94)
self.sink_size = sink_size # Configurable legacy sink
self.global_sink_size = 0 # Defaults to 0; set later for multi-shot
These values determine how many tokens remain protected during cache rolling operations.
Effective Sink Computation
The _effective_sink(kv_cache, frame_seqlen) method calculates the exact token boundary that must be preserved:
def _effective_sink(self, kv_cache, frame_seqlen):
sink_tokens = self.sink_size * frame_seqlen
global_sink_tokens = self.global_sink_size * frame_seqlen
if has_pinned_region:
# Multi-shot case: pinned region sits right after global sink
effective_sink = global_sink_tokens + pinned_len
else:
effective_sink = max(global_sink_tokens, sink_tokens)
return effective_sink
This computation returns the number of leading tokens that cannot be overwritten during cache eviction, accounting for both the global anchor and any pinned multi-shot regions.
Cache Rolling Logic
The _update_cache_and_get_kv method implements the actual memory management. When new tokens would exceed cache capacity, the algorithm:
- Identifies the effective sink boundary from
_effective_sink - Evicts the oldest local tokens appearing after this boundary
- Updates the pinned start index if the pinned region resides within the sink
- Constructs final K/V tensors by concatenating:
[global sink] + [pinned region] + [local window]
This concatenation ensures the attention mechanism always sees both anchored historical context and recent local frames. The implementation spans lines 40-96 in wan_5b/modules/causal_model_sp_ulysses.py.
Configuring Attention Sink in Your Pipeline
You enable the attention sink through the diffusion wrapper constructor, which forwards the parameter to the underlying Ulysses model.
Pipeline Wrapper Configuration
In pipeline/causal_diffusion_inference_sp.py, instantiate SPWanDiffusionWrapper5B with your desired sink size:
class SPWanDiffusionWrapper5B(torch.nn.Module):
def __init__(self,
model_name="Wan2.2-TI2V-5B",
timestep_shift=5.0,
local_attn_size=-1,
sink_size=4, # Keep first 4 frames permanently
num_frame_per_block=1,
t_scale=1.0,
rope_method="linear",
original_seq_len=None):
# ...
self.model = UlyssesSPCausalWanModel(
# ...
sink_size=sink_size, # Passed to attention layers
# ...
)
Setting sink_size greater than zero activates the attention sink for that many initial frames.
Direct Model Instantiation
For custom implementations, instantiate the sequence-parallel model directly:
from wan_5b.modules.causal_model_sp_ulysses import UlyssesSPCausalWanModel
model = UlyssesSPCausalWanModel(
model_type="ti2v",
in_dim=48,
dim=3072,
ffn_dim=14336,
out_dim=48,
num_heads=24,
num_layers=30,
local_attn_size=-1, # Full attention within local window
sink_size=4, # Attention sink: first 4 frames persist
num_frame_per_block=1,
)
Runtime Management and Inspection
Once initialized, the sink operates automatically during inference, but you can inspect and partially modify its behavior at runtime.
Inspecting Effective Sink
Monitor how many tokens are currently protected by querying the effective sink during generation:
# kv_cache is returned from previous forward pass
effective_sink, _, _, _ = model.self_attn._effective_sink(kv_cache, frame_seqlen=64)
print(f"Effective sink tokens: {effective_sink}")
This debug output helps verify that your sink configuration matches your memory constraints and context requirements.
Modifying Global Sink
While the legacy sink_size cannot change after initialization (as it would require cache reallocation), you can update the global sink for multi-shot scenarios:
# Enable global anchoring for multi-shot generation
model.self_attn.global_sink_size = 2 # First 2 frames now globally anchored
This modification affects subsequent forward passes without requiring model reconstruction.
When to Use Attention Sinks
| Scenario | Recommended Configuration |
|---|---|
| Long-form video (>10 seconds) | sink_size between 2-8 frames depending on GPU memory budget |
| Multi-shot generation (reusing early frames across shots) | Set global_sink_size > 0 and configure pinned regions |
| Real-time streaming (low latency) | sink_size=0 to minimize memory footprint; rely on rolling window |
| High-resolution frames (memory-constrained) | Smaller sink_size (1-2 frames) to balance context and VRAM |
Summary
- Attention sinks prevent memory overflow in long video generation by permanently caching a configurable number of initial frames while rolling older local tokens.
- Configure the sink via
sink_sizeinSPWanDiffusionWrapper5Bor directly inUlyssesSPCausalWanModelfromwan_5b/modules/causal_model_sp_ulysses.py. - The
_effective_sinkmethod calculates protection boundaries, while_update_cache_and_get_kvhandles the actual eviction and concatenation logic. - For multi-shot generation, use
global_sink_sizeinstead of or alongside the standard sink. - The sink operates transparently—set it once at initialization and the KV-cache automatically respects the anchors across all inference steps.
Frequently Asked Questions
What is the difference between sink_size and global_sink_size in LongLive?
sink_size is the legacy parameter that permanently anchors the first N frames during the generation process, while global_sink_size is designed for multi-shot generation scenarios where specific frames must persist across entirely separate generation runs. When both are present, the effective sink becomes the maximum of the two values, or the sum if a pinned region follows the global sink.
How does the attention sink prevent out-of-memory errors during long video generation?
The attention sink bounds memory usage by ensuring that only tokens after the effective sink boundary are eligible for eviction when the cache rolls. According to the implementation in wan_5b/modules/causal_model_sp_ulysses.py, the _update_cache_and_get_kv method discards older local tokens while preserving the sink tokens and pinned regions, maintaining a constant memory footprint regardless of total video length.
Can I change the sink size after initializing the model?
No. You cannot modify sink_size after construction because the cache tensor layouts are allocated based on this initial value. However, you can update global_sink_size at runtime by setting model.self_attn.global_sink_size, which allows dynamic control over globally anchored frames without reallocating the underlying KV-cache structures.
Which files contain the core attention sink implementation?
The primary implementation resides in wan_5b/modules/causal_model_sp_ulysses.py within the UlyssesCausalWanSelfAttention class, specifically the _effective_sink and _update_cache_and_get_kv methods. A single-GPU fallback implementation exists in wan_5b/modules/causal_model.py as CausalWanSelfAttention. The parameter is exposed to users through pipeline/causal_diffusion_inference_sp.py in the SPWanDiffusionWrapper5B class.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →