How to Implement Attention Sink for Managing Long Context in Video Generation with LongLive

Implement an attention sink in LongLive by setting the sink_size parameter when instantiating the Ulysses sequence-parallel model to permanently retain early video frames in the KV-cache while automatically rolling the oldest local tokens.

Long video generation with diffusion transformers creates a memory bottleneck because every new frame adds thousands of tokens (pixels × frames) to the Key-Value (KV) cache. The NVIDIA LongLive repository solves this through an attention sink mechanism that anchors initial frames to prevent context overflow while maintaining generative coherence. This guide explains exactly how to configure and implement attention sinks using the actual source code from NVlabs/LongLive.

What Is an Attention Sink and Why You Need It

Standard transformer attention stores every token's key and value tensors in a KV-cache that grows linearly with sequence length. For long videos, this quickly exhausts GPU memory.

The KV-Cache Bottleneck

In causal video generation, the model processes frames sequentially, appending each new frame's tokens to the cache. Without intervention, the cache size equals the total number of generated tokens, causing out-of-memory errors on long sequences. According to the LongLive source code in wan_5b/modules/causal_model_sp_ulysses.py, the solution involves anchoring a configurable number of early frames—the sink—that are never evicted when the cache rolls.

Sink Architecture Components

LongLive implements a hierarchical sink system with three configurable components:

  • sink_size: The legacy sink parameter defining how many initial frames remain permanently in the cache
  • global_sink_size: Frames that are globally anchored for the entire generation run, used specifically for multi-shot generation scenarios
  • Pinned region: A short block of frames immediately following the global sink that must stay together during cache operations

When the cache reaches capacity, the model rolls by evicting the oldest local tokens that appear after the effective sink boundary (global sink + pinned region). This bounds memory usage while preserving attention to the earliest context.

How Attention Sink Works Internally

The sink logic resides primarily in the UlyssesCausalWanSelfAttention class within wan_5b/modules/causal_model_sp_ulysses.py. A non-parallel fallback implementation exists in wan_5b/modules/causal_model.py as CausalWanSelfAttention.

Initialization Parameters

During model construction, the attention layer stores sink configuration as instance attributes:


# From wan_5b/modules/causal_model_sp_ulysses.py (lines 85-94)

self.sink_size = sink_size               # Configurable legacy sink

self.global_sink_size = 0                # Defaults to 0; set later for multi-shot

These values determine how many tokens remain protected during cache rolling operations.

Effective Sink Computation

The _effective_sink(kv_cache, frame_seqlen) method calculates the exact token boundary that must be preserved:

def _effective_sink(self, kv_cache, frame_seqlen):
    sink_tokens = self.sink_size * frame_seqlen
    global_sink_tokens = self.global_sink_size * frame_seqlen
    
    if has_pinned_region:
        # Multi-shot case: pinned region sits right after global sink

        effective_sink = global_sink_tokens + pinned_len
    else:
        effective_sink = max(global_sink_tokens, sink_tokens)
    
    return effective_sink

This computation returns the number of leading tokens that cannot be overwritten during cache eviction, accounting for both the global anchor and any pinned multi-shot regions.

Cache Rolling Logic

The _update_cache_and_get_kv method implements the actual memory management. When new tokens would exceed cache capacity, the algorithm:

  1. Identifies the effective sink boundary from _effective_sink
  2. Evicts the oldest local tokens appearing after this boundary
  3. Updates the pinned start index if the pinned region resides within the sink
  4. Constructs final K/V tensors by concatenating: [global sink] + [pinned region] + [local window]

This concatenation ensures the attention mechanism always sees both anchored historical context and recent local frames. The implementation spans lines 40-96 in wan_5b/modules/causal_model_sp_ulysses.py.

Configuring Attention Sink in Your Pipeline

You enable the attention sink through the diffusion wrapper constructor, which forwards the parameter to the underlying Ulysses model.

Pipeline Wrapper Configuration

In pipeline/causal_diffusion_inference_sp.py, instantiate SPWanDiffusionWrapper5B with your desired sink size:

class SPWanDiffusionWrapper5B(torch.nn.Module):
    def __init__(self,
                 model_name="Wan2.2-TI2V-5B",
                 timestep_shift=5.0,
                 local_attn_size=-1,
                 sink_size=4,                     # Keep first 4 frames permanently

                 num_frame_per_block=1,
                 t_scale=1.0,
                 rope_method="linear",
                 original_seq_len=None):
        # ...

        self.model = UlyssesSPCausalWanModel(
            # ...

            sink_size=sink_size,                # Passed to attention layers

            # ...

        )

Setting sink_size greater than zero activates the attention sink for that many initial frames.

Direct Model Instantiation

For custom implementations, instantiate the sequence-parallel model directly:

from wan_5b.modules.causal_model_sp_ulysses import UlyssesSPCausalWanModel

model = UlyssesSPCausalWanModel(
    model_type="ti2v",
    in_dim=48,
    dim=3072,
    ffn_dim=14336,
    out_dim=48,
    num_heads=24,
    num_layers=30,
    local_attn_size=-1,          # Full attention within local window

    sink_size=4,                 # Attention sink: first 4 frames persist

    num_frame_per_block=1,
)

Runtime Management and Inspection

Once initialized, the sink operates automatically during inference, but you can inspect and partially modify its behavior at runtime.

Inspecting Effective Sink

Monitor how many tokens are currently protected by querying the effective sink during generation:


# kv_cache is returned from previous forward pass

effective_sink, _, _, _ = model.self_attn._effective_sink(kv_cache, frame_seqlen=64)
print(f"Effective sink tokens: {effective_sink}")

This debug output helps verify that your sink configuration matches your memory constraints and context requirements.

Modifying Global Sink

While the legacy sink_size cannot change after initialization (as it would require cache reallocation), you can update the global sink for multi-shot scenarios:


# Enable global anchoring for multi-shot generation

model.self_attn.global_sink_size = 2   # First 2 frames now globally anchored

This modification affects subsequent forward passes without requiring model reconstruction.

When to Use Attention Sinks

Scenario Recommended Configuration
Long-form video (>10 seconds) sink_size between 2-8 frames depending on GPU memory budget
Multi-shot generation (reusing early frames across shots) Set global_sink_size > 0 and configure pinned regions
Real-time streaming (low latency) sink_size=0 to minimize memory footprint; rely on rolling window
High-resolution frames (memory-constrained) Smaller sink_size (1-2 frames) to balance context and VRAM

Summary

  • Attention sinks prevent memory overflow in long video generation by permanently caching a configurable number of initial frames while rolling older local tokens.
  • Configure the sink via sink_size in SPWanDiffusionWrapper5B or directly in UlyssesSPCausalWanModel from wan_5b/modules/causal_model_sp_ulysses.py.
  • The _effective_sink method calculates protection boundaries, while _update_cache_and_get_kv handles the actual eviction and concatenation logic.
  • For multi-shot generation, use global_sink_size instead of or alongside the standard sink.
  • The sink operates transparently—set it once at initialization and the KV-cache automatically respects the anchors across all inference steps.

Frequently Asked Questions

What is the difference between sink_size and global_sink_size in LongLive?

sink_size is the legacy parameter that permanently anchors the first N frames during the generation process, while global_sink_size is designed for multi-shot generation scenarios where specific frames must persist across entirely separate generation runs. When both are present, the effective sink becomes the maximum of the two values, or the sum if a pinned region follows the global sink.

How does the attention sink prevent out-of-memory errors during long video generation?

The attention sink bounds memory usage by ensuring that only tokens after the effective sink boundary are eligible for eviction when the cache rolls. According to the implementation in wan_5b/modules/causal_model_sp_ulysses.py, the _update_cache_and_get_kv method discards older local tokens while preserving the sink tokens and pinned regions, maintaining a constant memory footprint regardless of total video length.

Can I change the sink size after initializing the model?

No. You cannot modify sink_size after construction because the cache tensor layouts are allocated based on this initial value. However, you can update global_sink_size at runtime by setting model.self_attn.global_sink_size, which allows dynamic control over globally anchored frames without reallocating the underlying KV-cache structures.

Which files contain the core attention sink implementation?

The primary implementation resides in wan_5b/modules/causal_model_sp_ulysses.py within the UlyssesCausalWanSelfAttention class, specifically the _effective_sink and _update_cache_and_get_kv methods. A single-GPU fallback implementation exists in wan_5b/modules/causal_model.py as CausalWanSelfAttention. The parameter is exposed to users through pipeline/causal_diffusion_inference_sp.py in the SPWanDiffusionWrapper5B class.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →