# How to Implement KV Cache Recaching for Video Streaming in LongLive

> Implement KV cache recaching for video streaming in LongLive by setting shot_clean_recache and multi_shot_sink to true. Expert guide to optimize inference with automatic scene cut detection.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: how-to-guide
- Published: 2026-05-24

---

**To implement KV cache recaching for video streaming in LongLive, configure `shot_clean_recache: true` and `multi_shot_sink: true` in your inference settings, allowing the `CausalDiffusionInferencePipeline` to automatically detect scene cuts and reset local cache entries via `_zero_kv_data` while preserving essential global sink tokens.**

LongLive, developed by NVIDIA Labs (NVlabs/LongLive), is a causal transformer-based video generation system that maintains a **key-value (KV) cache** to avoid recomputing attention for previously processed frames. When generating long-form video content or handling dynamic scene transitions, stale cache entries can contaminate new footage unless properly managed through strategic recaching. This guide explains how to leverage LongLive's built-in recaching mechanisms to optimize streaming inference performance.

## Understanding KV Cache Architecture in LongLive

### The CausalDiffusionInferencePipeline

The `CausalDiffusionInferencePipeline` in [`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py) orchestrates the entire inference lifecycle, including cache initialization, chunk-wise updates, and recaching triggers. During initialization (lines 192-199), the pipeline creates empty KV cache tensors that persist across video chunks:

```python
if self.kv_cache_pos is None:
    self._initialize_kv_cache(...)

```

The pipeline handles per-chunk cache updates through `self._apply_cache_updates` and manages the cache lifecycle during streaming scenarios.

### Global Sinks and Local Cache Regions

LongLive distinguishes between **global sink tokens** (permanent attention anchors that survive across chunks) and **local cache entries** (rolling window data subject to eviction). This separation allows the system to preserve critical temporal context while discarding obsolete frame information.

The cache structure tracks:
- `global_sink_tokens`: Fixed attention anchors maintained throughout generation
- `local_end_index`: Dynamic boundary for rolling cache entries
- `pinned_start` and `pinned_len`: Metadata for multi-shot sink behavior

## Configuring KV Cache Recaching

To activate recaching for video streaming scenarios, modify your YAML configuration or command-line arguments:

```yaml
inference:
  shot_clean_recache: true    # Zero KV cache on scene cuts

  multi_shot_sink: true      # Enable adaptive pinned sinks

  sink_size: 64              # Number of tokens reserved as global sink

```

The `shot_clean_recache` flag is read in the pipeline's `__init__` (line 73) and evaluated during the chunk processing loop (lines 520-523).

## Core Recaching Methods

### Resetting Cache with _zero_kv_data

When the pipeline detects a scene cut using the `detect_scene_cut` helper, the `_zero_kv_data` method (lines 882-889 in [`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py)) clears the local portion of the cache while preserving global sinks:

```python
if is_scene_cut and self.shot_clean_recache:
    print("[inference] Scene cut at chunk ..., zeroing KV before recache")
    self._zero_kv_data(self.kv_cache_pos, current_start_tokens)

```

This method specifically resets the `local_end_index` to prevent old KV entries from contaminating attention calculations for the new scene, while keeping `global_sink_tokens` intact for temporal coherence.

### Pinning Chunks with _pin_current_chunk

For multi-shot video streams, `_pin_current_chunk` (lines 665-678) marks the current chunk as pinned so its tokens become the new sink after the next rollover:

```python
def _pin_current_chunk(self, kv_cache, pinned_start, pinned_len):
    for block_cache in kv_cache:
        block_cache['pinned_start'] = pinned_start
        block_cache['pinned_len'] = pinned_len

```

This enables **adaptive sink behavior** where the model adjusts its attention anchors based on recent content rather than fixed initial tokens.

### Low-Level Cache Operations

The actual rolling, eviction, and de-quantization logic resides in:
- [`wan_5b/modules/causal_model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model.py): Core KV cache updates and cache rolling
- [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py): Sequence-parallel variant implementing `_effective_sink` for distributed attention

These modules handle the tensor operations that physically shift cache contents and manage `dequantize_kv_cache` operations when quantization is enabled.

## Practical Implementation Example

Complete setup for streaming video generation with automatic recaching:

```python
from pipeline.causal_diffusion_inference import CausalDiffusionInferencePipeline

# Configure with recaching enabled

config = {
    'shot_clean_recache': True,
    'multi_shot_sink': True,
    'sink_size': 64
}

pipeline = CausalDiffusionInferencePipeline(
    args=config,
    device="cuda",
    generator=None,
    text_encoder=None,
    vae=None,
)

# Run streaming inference; recaching occurs automatically on scene cuts

video = pipeline.inference(
    noise=noise_tensor,
    text_prompts=["A sunrise over mountains"],
    start_frame_index=0,
)

```

When recaching triggers successfully, the console outputs:

```

[inference] Scene cut at chunk 7, zeroing KV before recache

```

## Summary

- **Enable recaching** by setting `shot_clean_recache: true` and `multi_shot_sink: true` in your configuration to handle scene transitions
- **Preserve context** through global sink tokens while clearing local cache entries on scene cuts to prevent attention contamination
- **Use `_zero_kv_data`** to safely reset cache regions without losing critical attention anchors required for temporal consistency
- **Leverage `_pin_current_chunk`** to maintain continuity across multi-shot video streams by adapting sink regions to recent content
- **Monitor logs** for scene cut detection messages confirming successful recaching operations at chunk boundaries

## Frequently Asked Questions

### What is KV cache recaching in video generation?

KV cache recaching is the process of selectively clearing or resetting key-value attention cache entries during long-form video generation to prevent stale frame information from influencing newly generated content. In LongLive, this occurs automatically at scene boundaries when `shot_clean_recache` is enabled, allowing the causal transformer to start fresh attention patterns for new visual segments while preserving essential temporal anchors.

### When should I enable shot_clean_recache?

Enable `shot_clean_recache` when processing video streams with distinct scene transitions, cuts, or lighting changes that require the model to forget previous visual contexts. This prevents attention contamination between unrelated segments while maintaining the global sink tokens necessary for coherent motion. Disable it for single continuous shots where temporal consistency across the entire sequence is desired.

### How does multi_shot_sink differ from standard recaching?

Standard recaching clears local cache entries while preserving fixed global sinks defined at model initialization. The `multi_shot_sink` feature additionally invokes `_pin_current_chunk` to mark the most recent chunk as a dynamic sink, allowing the model to adapt its attention anchors based on recent content. This creates a rolling memory effect where the sink region updates to reflect the current shot's characteristics rather than static initial tokens.

### Which files control the KV cache rolling mechanism?

The high-level orchestration occurs in [`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py) through methods like `_zero_kv_data` (lines 882-889) and `_pin_current_chunk` (lines 665-678). The low-level tensor operations, eviction logic, and sequence-parallel handling are implemented in [`wan_5b/modules/causal_model.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model.py) and [`wan_5b/modules/causal_model_sp_ulysses.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/modules/causal_model_sp_ulysses.py), specifically within `_apply_cache_updates` and the `_effective_sink` property used during attention computation.