# How to Implement Attention Sink for Managing Long Context in Video Generation with LongLive

> Learn to implement attention sink in LongLive for long context video generation. Effortlessly manage early video frames in KV-cache with Ulysses model for better results.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: how-to-guide
- Published: 2026-05-24

---

**Implement an attention sink in LongLive by setting the `sink_size` parameter when instantiating the Ulysses sequence-parallel model to permanently retain early video frames in the KV-cache while automatically rolling the oldest local tokens.**

Long video generation with diffusion transformers creates a memory bottleneck because every new frame adds thousands of tokens (pixels × frames) to the Key-Value (KV) cache. The NVIDIA LongLive repository solves this through an **attention sink** mechanism that anchors initial frames to prevent context overflow while maintaining generative coherence. This guide explains exactly how to configure and implement attention sinks using the actual source code from `NVlabs/LongLive`.

## What Is an Attention Sink and Why You Need It

Standard transformer attention stores every token's key and value tensors in a KV-cache that grows linearly with sequence length. For long videos, this quickly exhausts GPU memory.

### The KV-Cache Bottleneck

In causal video generation, the model processes frames sequentially, appending each new frame's tokens to the cache. Without intervention, the cache size equals the total number of generated tokens, causing out-of-memory errors on long sequences. According to the LongLive source code in [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py), the solution involves anchoring a configurable number of early frames—the *sink*—that are never evicted when the cache rolls.

### Sink Architecture Components

LongLive implements a hierarchical sink system with three configurable components:

- **`sink_size`**: The legacy sink parameter defining how many initial frames remain permanently in the cache
- **`global_sink_size`**: Frames that are globally anchored for the entire generation run, used specifically for multi-shot generation scenarios
- **Pinned region**: A short block of frames immediately following the global sink that must stay together during cache operations

When the cache reaches capacity, the model **rolls** by evicting the oldest local tokens that appear after the effective sink boundary (global sink + pinned region). This bounds memory usage while preserving attention to the earliest context.

## How Attention Sink Works Internally

The sink logic resides primarily in the `UlyssesCausalWanSelfAttention` class within [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py). A non-parallel fallback implementation exists in [`wan_5b/modules/causal_model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model.py) as `CausalWanSelfAttention`.

### Initialization Parameters

During model construction, the attention layer stores sink configuration as instance attributes:

```python

# From wan_5b/modules/causal_model_sp_ulysses.py (lines 85-94)

self.sink_size = sink_size               # Configurable legacy sink

self.global_sink_size = 0                # Defaults to 0; set later for multi-shot

```

These values determine how many tokens remain protected during cache rolling operations.

### Effective Sink Computation

The `_effective_sink(kv_cache, frame_seqlen)` method calculates the exact token boundary that must be preserved:

```python
def _effective_sink(self, kv_cache, frame_seqlen):
    sink_tokens = self.sink_size * frame_seqlen
    global_sink_tokens = self.global_sink_size * frame_seqlen
    
    if has_pinned_region:
        # Multi-shot case: pinned region sits right after global sink

        effective_sink = global_sink_tokens + pinned_len
    else:
        effective_sink = max(global_sink_tokens, sink_tokens)
    
    return effective_sink

```

This computation returns the number of leading tokens that **cannot** be overwritten during cache eviction, accounting for both the global anchor and any pinned multi-shot regions.

### Cache Rolling Logic

The `_update_cache_and_get_kv` method implements the actual memory management. When new tokens would exceed cache capacity, the algorithm:

1. Identifies the effective sink boundary from `_effective_sink`
2. Evicts the oldest local tokens appearing after this boundary
3. Updates the pinned start index if the pinned region resides within the sink
4. Constructs final K/V tensors by concatenating: `[global sink] + [pinned region] + [local window]`

This concatenation ensures the attention mechanism always sees both anchored historical context and recent local frames. The implementation spans lines 40-96 in [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py).

## Configuring Attention Sink in Your Pipeline

You enable the attention sink through the diffusion wrapper constructor, which forwards the parameter to the underlying Ulysses model.

### Pipeline Wrapper Configuration

In [`pipeline/causal_diffusion_inference_sp.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference_sp.py), instantiate `SPWanDiffusionWrapper5B` with your desired sink size:

```python
class SPWanDiffusionWrapper5B(torch.nn.Module):
    def __init__(self,
                 model_name="Wan2.2-TI2V-5B",
                 timestep_shift=5.0,
                 local_attn_size=-1,
                 sink_size=4,                     # Keep first 4 frames permanently

                 num_frame_per_block=1,
                 t_scale=1.0,
                 rope_method="linear",
                 original_seq_len=None):
        # ...

        self.model = UlyssesSPCausalWanModel(
            # ...

            sink_size=sink_size,                # Passed to attention layers

            # ...

        )

```

Setting `sink_size` greater than zero activates the attention sink for that many initial frames.

### Direct Model Instantiation

For custom implementations, instantiate the sequence-parallel model directly:

```python
from wan_5b.modules.causal_model_sp_ulysses import UlyssesSPCausalWanModel

model = UlyssesSPCausalWanModel(
    model_type="ti2v",
    in_dim=48,
    dim=3072,
    ffn_dim=14336,
    out_dim=48,
    num_heads=24,
    num_layers=30,
    local_attn_size=-1,          # Full attention within local window

    sink_size=4,                 # Attention sink: first 4 frames persist

    num_frame_per_block=1,
)

```

## Runtime Management and Inspection

Once initialized, the sink operates automatically during inference, but you can inspect and partially modify its behavior at runtime.

### Inspecting Effective Sink

Monitor how many tokens are currently protected by querying the effective sink during generation:

```python

# kv_cache is returned from previous forward pass

effective_sink, _, _, _ = model.self_attn._effective_sink(kv_cache, frame_seqlen=64)
print(f"Effective sink tokens: {effective_sink}")

```

This debug output helps verify that your sink configuration matches your memory constraints and context requirements.

### Modifying Global Sink

While the legacy `sink_size` cannot change after initialization (as it would require cache reallocation), you can update the global sink for multi-shot scenarios:

```python

# Enable global anchoring for multi-shot generation

model.self_attn.global_sink_size = 2   # First 2 frames now globally anchored

```

This modification affects subsequent forward passes without requiring model reconstruction.

## When to Use Attention Sinks

| Scenario | Recommended Configuration |
|---|---|
| **Long-form video** (>10 seconds) | `sink_size` between 2-8 frames depending on GPU memory budget |
| **Multi-shot generation** (reusing early frames across shots) | Set `global_sink_size` > 0 and configure pinned regions |
| **Real-time streaming** (low latency) | `sink_size=0` to minimize memory footprint; rely on rolling window |
| **High-resolution frames** (memory-constrained) | Smaller `sink_size` (1-2 frames) to balance context and VRAM |

## Summary

- **Attention sinks** prevent memory overflow in long video generation by permanently caching a configurable number of initial frames while rolling older local tokens.
- Configure the sink via `sink_size` in `SPWanDiffusionWrapper5B` or directly in `UlyssesSPCausalWanModel` from [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py).
- The `_effective_sink` method calculates protection boundaries, while `_update_cache_and_get_kv` handles the actual eviction and concatenation logic.
- For multi-shot generation, use `global_sink_size` instead of or alongside the standard sink.
- The sink operates transparently—set it once at initialization and the KV-cache automatically respects the anchors across all inference steps.

## Frequently Asked Questions

### What is the difference between `sink_size` and `global_sink_size` in LongLive?

**`sink_size`** is the legacy parameter that permanently anchors the first N frames during the generation process, while **`global_sink_size`** is designed for multi-shot generation scenarios where specific frames must persist across entirely separate generation runs. When both are present, the effective sink becomes the maximum of the two values, or the sum if a pinned region follows the global sink.

### How does the attention sink prevent out-of-memory errors during long video generation?

The attention sink bounds memory usage by ensuring that only tokens after the effective sink boundary are eligible for eviction when the cache rolls. According to the implementation in [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py), the `_update_cache_and_get_kv` method discards older local tokens while preserving the sink tokens and pinned regions, maintaining a constant memory footprint regardless of total video length.

### Can I change the sink size after initializing the model?

**No.** You cannot modify `sink_size` after construction because the cache tensor layouts are allocated based on this initial value. However, you **can** update `global_sink_size` at runtime by setting `model.self_attn.global_sink_size`, which allows dynamic control over globally anchored frames without reallocating the underlying KV-cache structures.

### Which files contain the core attention sink implementation?

The primary implementation resides in [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py) within the `UlyssesCausalWanSelfAttention` class, specifically the `_effective_sink` and `_update_cache_and_get_kv` methods. A single-GPU fallback implementation exists in [`wan_5b/modules/causal_model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model.py) as `CausalWanSelfAttention`. The parameter is exposed to users through [`pipeline/causal_diffusion_inference_sp.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference_sp.py) in the `SPWanDiffusionWrapper5B` class.