# How to Configure Local Attention Size for Long Video Generation in LongLive

> Master local attention size for long video generation in LongLive. Learn to configure window_size for sliding-window attention control in NVlabs/LongLive.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: how-to-guide
- Published: 2026-05-24

---

**Set the `window_size` tuple `(left, right)` in the model configuration files or pass it directly to `WanSelfAttention` modules to control the sliding-window attention scope.**

LongLive implements a **sliding-window (local) attention** mechanism to limit the quadratic computational cost of full-sequence attention when processing very long video sequences. You configure the size of this attention window by adjusting the `window_size` parameter, which propagates through the transformer architecture from configuration files down to the FlashAttention implementation.

## How Local Attention Works in LongLive

LongLive’s local attention relies on FlashAttention’s native sliding-window masking capabilities. The window size constraints flow from model configurations through the self-attention layers to the underlying GPU kernels.

### FlashAttention Integration

In [`wan_5b/modules/attention.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/attention.py), the `flash_attention` routine accepts a `window_size` parameter that defines the attention scope. When this value is not `(-1, -1)`, the FlashAttention implementation applies a sliding-window mask so each token attends only to its neighbors.

```python
def flash_attention(..., window_size=(-1, -1), ...):
    ...
    # Inside the Flash-Attention call

    x = flash_attn.flash_attn_varlen_func(...,
                                          window_size=window_size, ...)

```

The tuple `(left, right)` specifies the number of tokens on the left and right that each position can attend to. The default value `(-1, -1)` disables the window mask, enabling global attention across the full sequence.

### WanSelfAttention Propagation

The `WanSelfAttention` module stores the `window_size` tuple and forwards it to the `flash_attention` function during the forward pass. This class is defined in [`wan_5b/modules/model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/model.py).

```python
class WanSelfAttention(nn.Module):
    def __init__(..., window_size=(-1, -1), ...):
        self.window_size = window_size
        ...
    
    def forward(...):
        ...
        x = flash_attention(..., window_size=self.window_size)

```

This propagation ensures that window size constraints apply consistently across all self-attention blocks in the network.

## Configuration Methods

You can configure the local attention window size at three different levels: through configuration files, during module instantiation, or by overriding configs at runtime.

### Method 1: Configuration Files

Every model entry point defines a `window_size` field in its configuration file. By default, this is set to `(-1, -1)` for global attention. Change this field to enable local attention.

For example, in [`wan_5b/configs/wan_i2v_A14B.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/configs/wan_i2v_A14B.py):

```python

# Default is global attention

i2v_A14B.window_size = (-1, -1)

# Change to symmetric local window of 128 tokens each side

i2v_A14B.window_size = (128, 128)

```

Similar configuration fields exist in:
- [`wan_5b/configs/wan_t2v_A14B.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/configs/wan_t2v_A14B.py) for text-to-video generation
- [`wan_5b/configs/wan_ti2v_5B.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/configs/wan_ti2v_5B.py) for the larger 5B parameter model

### Method 2: Direct Module Instantiation

For custom architectures, pass the `window_size` tuple directly when constructing `WanSelfAttention` instances:

```python
from wan_5b.modules.model import WanSelfAttention

# Asymmetric window: 64 tokens left, 32 tokens right

self_attn = WanSelfAttention(
    dim=1024, 
    num_heads=16,
    window_size=(64, 32), 
    qk_norm=True
)

# Use in forward pass

output = self_attn(x, seq_lens, grid_sizes, freqs)

```

### Method 3: Runtime Override

You can override window sizes without modifying source files by importing and altering the configuration object before model initialization:

```python
from wan_5b.configs.wan_i2v_A14B import i2v_A14B

# Override for this session only

i2v_A14B.window_size = (256, 256)

# Now initialize your model - it will use the modified window size

```

## Example Configurations for Long Videos

When generating long videos, select window sizes that balance memory constraints with temporal coherence:

**Symmetric 256-token window:**

```python

# config_256.py

from wan_5b.configs.base import Config

cfg = Config()
cfg.window_size = (128, 128)  # Total receptive field of 256 tokens

cfg.patch_size = (2, 2, 2)

```

**Asymmetric causal window:**

```python

# For video prediction tasks where future frames shouldn't be visible

cfg.window_size = (256, 0)  # Attend to 256 past tokens only

```

**Mixing global and local layers:**

```python

# First layer global for high-resolution spatial details

layer_0 = WanSelfAttention(dim=1024, window_size=(-1, -1), ...)

# Deep layers local for efficient temporal processing

layer_n = WanSelfAttention(dim=1024, window_size=(64, 64), ...)

```

## Summary

- LongLive controls attention scope via the **`window_size`** tuple `(left, right)` propagated through the model hierarchy.
- The `flash_attention` function in [`wan_5b/modules/attention.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/attention.py) implements the sliding-window mask using FlashAttention's native `window_size` parameter.
- **`WanSelfAttention`** modules in [`wan_5b/modules/model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/model.py) store and forward these constraints to the attention kernels.
- Configure windows through config files like [`wan_5b/configs/wan_i2v_A14B.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/configs/wan_i2v_A14B.py) or pass them directly to module constructors for custom architectures.

## Frequently Asked Questions

### What is the default attention mode in LongLive?

By default, LongLive uses **global attention** with `window_size` set to `(-1, -1)` in all configuration files. This allows every token to attend to every other token, maximizing quality for shorter sequences but requiring significant memory for long videos.

### How does the window size affect video generation quality?

Smaller windows reduce memory usage and computation time but limit the receptive field of each token. For very long videos (hundreds of frames), a moderate window size (e.g., 128–256 tokens) typically preserves quality while enabling feasible inference. You may need to tune this based on temporal coherence requirements in your specific video domain.

### Can I use different window sizes for different layers?

Yes. Since each `WanSelfAttention` module accepts its own `window_size` parameter, you can configure individual layers with different window sizes by modifying the model architecture initialization. For example, early layers might use larger windows for high-resolution spatial attention, while deeper temporal layers use smaller windows for efficiency.

### Where can I change the window size without modifying source files?

You can override the window size at runtime by importing the configuration class and modifying the attribute before model initialization, as shown in Method 3 above. This approach avoids editing the original configuration files in the repository while allowing quick experimentation with different attention spans.