How to Implement KV Cache Recaching for Video Streaming in LongLive

To implement KV cache recaching for video streaming in LongLive, configure shot_clean_recache: true and multi_shot_sink: true in your inference settings, allowing the CausalDiffusionInferencePipeline to automatically detect scene cuts and reset local cache entries via _zero_kv_data while preserving essential global sink tokens.

LongLive, developed by NVIDIA Labs (NVlabs/LongLive), is a causal transformer-based video generation system that maintains a key-value (KV) cache to avoid recomputing attention for previously processed frames. When generating long-form video content or handling dynamic scene transitions, stale cache entries can contaminate new footage unless properly managed through strategic recaching. This guide explains how to leverage LongLive's built-in recaching mechanisms to optimize streaming inference performance.

Understanding KV Cache Architecture in LongLive

The CausalDiffusionInferencePipeline

The CausalDiffusionInferencePipeline in pipeline/causal_diffusion_inference.py orchestrates the entire inference lifecycle, including cache initialization, chunk-wise updates, and recaching triggers. During initialization (lines 192-199), the pipeline creates empty KV cache tensors that persist across video chunks:

if self.kv_cache_pos is None:
    self._initialize_kv_cache(...)

The pipeline handles per-chunk cache updates through self._apply_cache_updates and manages the cache lifecycle during streaming scenarios.

Global Sinks and Local Cache Regions

LongLive distinguishes between global sink tokens (permanent attention anchors that survive across chunks) and local cache entries (rolling window data subject to eviction). This separation allows the system to preserve critical temporal context while discarding obsolete frame information.

The cache structure tracks:

  • global_sink_tokens: Fixed attention anchors maintained throughout generation
  • local_end_index: Dynamic boundary for rolling cache entries
  • pinned_start and pinned_len: Metadata for multi-shot sink behavior

Configuring KV Cache Recaching

To activate recaching for video streaming scenarios, modify your YAML configuration or command-line arguments:

inference:
  shot_clean_recache: true    # Zero KV cache on scene cuts

  multi_shot_sink: true      # Enable adaptive pinned sinks

  sink_size: 64              # Number of tokens reserved as global sink

The shot_clean_recache flag is read in the pipeline's __init__ (line 73) and evaluated during the chunk processing loop (lines 520-523).

Core Recaching Methods

Resetting Cache with _zero_kv_data

When the pipeline detects a scene cut using the detect_scene_cut helper, the _zero_kv_data method (lines 882-889 in pipeline/causal_diffusion_inference.py) clears the local portion of the cache while preserving global sinks:

if is_scene_cut and self.shot_clean_recache:
    print("[inference] Scene cut at chunk ..., zeroing KV before recache")
    self._zero_kv_data(self.kv_cache_pos, current_start_tokens)

This method specifically resets the local_end_index to prevent old KV entries from contaminating attention calculations for the new scene, while keeping global_sink_tokens intact for temporal coherence.

Pinning Chunks with _pin_current_chunk

For multi-shot video streams, _pin_current_chunk (lines 665-678) marks the current chunk as pinned so its tokens become the new sink after the next rollover:

def _pin_current_chunk(self, kv_cache, pinned_start, pinned_len):
    for block_cache in kv_cache:
        block_cache['pinned_start'] = pinned_start
        block_cache['pinned_len'] = pinned_len

This enables adaptive sink behavior where the model adjusts its attention anchors based on recent content rather than fixed initial tokens.

Low-Level Cache Operations

The actual rolling, eviction, and de-quantization logic resides in:

These modules handle the tensor operations that physically shift cache contents and manage dequantize_kv_cache operations when quantization is enabled.

Practical Implementation Example

Complete setup for streaming video generation with automatic recaching:

from pipeline.causal_diffusion_inference import CausalDiffusionInferencePipeline

# Configure with recaching enabled

config = {
    'shot_clean_recache': True,
    'multi_shot_sink': True,
    'sink_size': 64
}

pipeline = CausalDiffusionInferencePipeline(
    args=config,
    device="cuda",
    generator=None,
    text_encoder=None,
    vae=None,
)

# Run streaming inference; recaching occurs automatically on scene cuts

video = pipeline.inference(
    noise=noise_tensor,
    text_prompts=["A sunrise over mountains"],
    start_frame_index=0,
)

When recaching triggers successfully, the console outputs:


[inference] Scene cut at chunk 7, zeroing KV before recache

Summary

  • Enable recaching by setting shot_clean_recache: true and multi_shot_sink: true in your configuration to handle scene transitions
  • Preserve context through global sink tokens while clearing local cache entries on scene cuts to prevent attention contamination
  • Use _zero_kv_data to safely reset cache regions without losing critical attention anchors required for temporal consistency
  • Leverage _pin_current_chunk to maintain continuity across multi-shot video streams by adapting sink regions to recent content
  • Monitor logs for scene cut detection messages confirming successful recaching operations at chunk boundaries

Frequently Asked Questions

What is KV cache recaching in video generation?

KV cache recaching is the process of selectively clearing or resetting key-value attention cache entries during long-form video generation to prevent stale frame information from influencing newly generated content. In LongLive, this occurs automatically at scene boundaries when shot_clean_recache is enabled, allowing the causal transformer to start fresh attention patterns for new visual segments while preserving essential temporal anchors.

When should I enable shot_clean_recache?

Enable shot_clean_recache when processing video streams with distinct scene transitions, cuts, or lighting changes that require the model to forget previous visual contexts. This prevents attention contamination between unrelated segments while maintaining the global sink tokens necessary for coherent motion. Disable it for single continuous shots where temporal consistency across the entire sequence is desired.

How does multi_shot_sink differ from standard recaching?

Standard recaching clears local cache entries while preserving fixed global sinks defined at model initialization. The multi_shot_sink feature additionally invokes _pin_current_chunk to mark the most recent chunk as a dynamic sink, allowing the model to adapt its attention anchors based on recent content. This creates a rolling memory effect where the sink region updates to reflect the current shot's characteristics rather than static initial tokens.

Which files control the KV cache rolling mechanism?

The high-level orchestration occurs in pipeline/causal_diffusion_inference.py through methods like _zero_kv_data (lines 882-889) and _pin_current_chunk (lines 665-678). The low-level tensor operations, eviction logic, and sequence-parallel handling are implemented in wan_5b/modules/causal_model.py and wan_5b/modules/causal_model_sp_ulysses.py, specifically within _apply_cache_updates and the _effective_sink property used during attention computation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →