How to Configure Local Attention Size for Long Video Generation in LongLive
Set the window_size tuple (left, right) in the model configuration files or pass it directly to WanSelfAttention modules to control the sliding-window attention scope.
LongLive implements a sliding-window (local) attention mechanism to limit the quadratic computational cost of full-sequence attention when processing very long video sequences. You configure the size of this attention window by adjusting the window_size parameter, which propagates through the transformer architecture from configuration files down to the FlashAttention implementation.
How Local Attention Works in LongLive
LongLive’s local attention relies on FlashAttention’s native sliding-window masking capabilities. The window size constraints flow from model configurations through the self-attention layers to the underlying GPU kernels.
FlashAttention Integration
In wan_5b/modules/attention.py, the flash_attention routine accepts a window_size parameter that defines the attention scope. When this value is not (-1, -1), the FlashAttention implementation applies a sliding-window mask so each token attends only to its neighbors.
def flash_attention(..., window_size=(-1, -1), ...):
...
# Inside the Flash-Attention call
x = flash_attn.flash_attn_varlen_func(...,
window_size=window_size, ...)
The tuple (left, right) specifies the number of tokens on the left and right that each position can attend to. The default value (-1, -1) disables the window mask, enabling global attention across the full sequence.
WanSelfAttention Propagation
The WanSelfAttention module stores the window_size tuple and forwards it to the flash_attention function during the forward pass. This class is defined in wan_5b/modules/model.py.
class WanSelfAttention(nn.Module):
def __init__(..., window_size=(-1, -1), ...):
self.window_size = window_size
...
def forward(...):
...
x = flash_attention(..., window_size=self.window_size)
This propagation ensures that window size constraints apply consistently across all self-attention blocks in the network.
Configuration Methods
You can configure the local attention window size at three different levels: through configuration files, during module instantiation, or by overriding configs at runtime.
Method 1: Configuration Files
Every model entry point defines a window_size field in its configuration file. By default, this is set to (-1, -1) for global attention. Change this field to enable local attention.
For example, in wan_5b/configs/wan_i2v_A14B.py:
# Default is global attention
i2v_A14B.window_size = (-1, -1)
# Change to symmetric local window of 128 tokens each side
i2v_A14B.window_size = (128, 128)
Similar configuration fields exist in:
wan_5b/configs/wan_t2v_A14B.pyfor text-to-video generationwan_5b/configs/wan_ti2v_5B.pyfor the larger 5B parameter model
Method 2: Direct Module Instantiation
For custom architectures, pass the window_size tuple directly when constructing WanSelfAttention instances:
from wan_5b.modules.model import WanSelfAttention
# Asymmetric window: 64 tokens left, 32 tokens right
self_attn = WanSelfAttention(
dim=1024,
num_heads=16,
window_size=(64, 32),
qk_norm=True
)
# Use in forward pass
output = self_attn(x, seq_lens, grid_sizes, freqs)
Method 3: Runtime Override
You can override window sizes without modifying source files by importing and altering the configuration object before model initialization:
from wan_5b.configs.wan_i2v_A14B import i2v_A14B
# Override for this session only
i2v_A14B.window_size = (256, 256)
# Now initialize your model - it will use the modified window size
Example Configurations for Long Videos
When generating long videos, select window sizes that balance memory constraints with temporal coherence:
Symmetric 256-token window:
# config_256.py
from wan_5b.configs.base import Config
cfg = Config()
cfg.window_size = (128, 128) # Total receptive field of 256 tokens
cfg.patch_size = (2, 2, 2)
Asymmetric causal window:
# For video prediction tasks where future frames shouldn't be visible
cfg.window_size = (256, 0) # Attend to 256 past tokens only
Mixing global and local layers:
# First layer global for high-resolution spatial details
layer_0 = WanSelfAttention(dim=1024, window_size=(-1, -1), ...)
# Deep layers local for efficient temporal processing
layer_n = WanSelfAttention(dim=1024, window_size=(64, 64), ...)
Summary
- LongLive controls attention scope via the
window_sizetuple(left, right)propagated through the model hierarchy. - The
flash_attentionfunction inwan_5b/modules/attention.pyimplements the sliding-window mask using FlashAttention's nativewindow_sizeparameter. WanSelfAttentionmodules inwan_5b/modules/model.pystore and forward these constraints to the attention kernels.- Configure windows through config files like
wan_5b/configs/wan_i2v_A14B.pyor pass them directly to module constructors for custom architectures.
Frequently Asked Questions
What is the default attention mode in LongLive?
By default, LongLive uses global attention with window_size set to (-1, -1) in all configuration files. This allows every token to attend to every other token, maximizing quality for shorter sequences but requiring significant memory for long videos.
How does the window size affect video generation quality?
Smaller windows reduce memory usage and computation time but limit the receptive field of each token. For very long videos (hundreds of frames), a moderate window size (e.g., 128–256 tokens) typically preserves quality while enabling feasible inference. You may need to tune this based on temporal coherence requirements in your specific video domain.
Can I use different window sizes for different layers?
Yes. Since each WanSelfAttention module accepts its own window_size parameter, you can configure individual layers with different window sizes by modifying the model architecture initialization. For example, early layers might use larger windows for high-resolution spatial attention, while deeper temporal layers use smaller windows for efficiency.
Where can I change the window size without modifying source files?
You can override the window size at runtime by importing the configuration class and modifying the attribute before model initialization, as shown in Method 3 above. This approach avoids editing the original configuration files in the repository while allowing quick experimentation with different attention spans.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →