# How to Configure the Streaming VAE Pipeline for Long Video Generation in LongLive

> Learn to configure the streaming VAE pipeline for efficient, long video generation in LongLive. Reduce memory usage with incremental decoding. Get started now.

- Repository: [NVIDIA Research Projects/LongLive](https://github.com/NVlabs/LongLive)
- Tags: how-to-guide
- Published: 2026-05-24

---

**The streaming VAE pipeline is an incremental decoding mode that processes video in temporal chunks during diffusion to reduce memory usage, enabled by setting `inference.streaming_vae=true` and optionally configuring `vae_device` for multi-GPU or `async_vae` for CUDA stream parallelization.**

The LongLive repository by NVlabs implements a **streaming VAE pipeline** that enables generation of arbitrarily long videos without requiring tens of gigabytes of GPU memory. Unlike standard diffusion pipelines that accumulate the entire latent tensor before decoding, LongLive processes and decodes temporal blocks concurrently with denoising. This guide explains the architecture and the specific configuration flags required to activate this mode.

## What Is the Streaming VAE Pipeline?

The **streaming VAE pipeline** is an inference mode in `CausalDiffusionInferencePipeline` (located in [`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py)) that integrates decoding directly into the diffusion loop. Instead of waiting for the complete latent representation—which can exceed available VRAM for long sequences—the pipeline splits generation into blocks of `num_frame_per_block` frames, decodes each block immediately after denoising, and appends the result to the output video stream.

This approach reduces peak GPU memory consumption from tens of gigabytes to a level determined by chunk size. When using asynchronous execution, the pipeline also overlaps VAE decode with subsequent diffusion steps, significantly improving throughput for long-form content.

## How the Streaming VAE Architecture Works

The implementation selects between three execution strategies based on configuration flags parsed at lines 83–86 of the inference pipeline.

### Chunk-Wise Processing Logic

At the core of the system is a temporal chunking mechanism in the `inference` method (line 52 of [`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py)):

```python
for chunk_index, current_num_frames in enumerate(all_num_frames):

```

For each iteration, the code runs the spatial denoising loop, stores the latent block in an output buffer, and immediately triggers decode via the `self.vae.model.cached_decode` method before proceeding to the next temporal segment. This prevents accumulation of full-sequence latents in GPU memory.

### Memory Optimization Strategies

The helper function `place_vae_for_streaming` in [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py) (lines 50–72) implements three distinct placement modes controlled by the conditional block at lines 100–106:

- **`streaming`**: Decode occurs synchronously on the same device as the diffusion model.
- **`streaming-pipeline`**: The VAE module and its `mean`/`std` buffers transfer to a separate device specified by `vae_device`, isolating decode memory from diffusion compute using a background Python thread (lines 71–84).
- **`streaming-async`**: Decode executes on a dedicated CUDA stream within the same GPU context (lines 60–68), allowing overlap between VAE execution and the next diffusion steps without requiring a second physical GPU.

When `async_vae` is enabled, the code creates a CUDA stream and records synchronization events to manage the parallel execution path safely.

## Configuration Guide

Enable and tune the streaming VAE pipeline by setting these keys in your inference configuration (see [`wan_5b/configs/wan_ti2v_5B.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/configs/wan_ti2v_5B.py) for a complete template):

| Config Key | Description | Recommended Value |
|------------|-------------|-------------------|
| `inference.streaming_vae` | Master flag enabling chunk-wise decode mode. | `true` |
| `inference.async_vae` | Enables CUDA stream-based overlap; requires `streaming_vae`. | `true` (or `false` for thread-based path) |
| `inference.vae_device` | Target GPU for VAE decode (e.g., `"cuda:1"`). Must differ from diffusion device for offloading. | `"cuda:1"` |
| `inference.num_frame_per_block` | Temporal chunk size; smaller values reduce memory but increase kernel launch overhead. | `8`–`16` for 4K video |
| `inference.return_latents` | Debug flag to skip VAE decoding and return raw latents. | `false` |

## Implementation Example

The following example demonstrates how to configure and run the streaming pipeline using LongLive's utility functions:

```python
import torch
from utils.config import load_config_from_yaml
from pipeline.causal_diffusion_inference import CausalDiffusionInferencePipeline
from utils.inference_utils import place_vae_for_streaming, prepare_single_prompt_inputs, save_video

# Load configuration and enable streaming

cfg = load_config_from_yaml("wan_5b/configs/wan_ti2v_5B.yaml")
cfg.inference.streaming_vae = True          # Enable streaming mode

cfg.inference.async_vae = True              # Use async CUDA stream

cfg.inference.vae_device = "cuda:1"         # Offload VAE to second GPU

cfg.inference.num_frame_per_block = 16      # Decode 16 frames per chunk

# Initialize pipeline on primary GPU (cuda:0)

device = torch.device("cuda:0")
pipeline = CausalDiffusionInferencePipeline(args=cfg, device=device)

# Relocate VAE to dedicated device (constructor also calls this internally)

place_vae_for_streaming(pipeline, cfg)

# Prepare inputs

noise, prompts = prepare_single_prompt_inputs(
    cfg,
    prompt="A timelapse of a city from sunrise to sunset",
    device=device,
)

# Run inference with live decoding

video = pipeline.inference(
    noise=noise,
    text_prompts=prompts,
    return_latents=False,  # Return decoded video tensor [B, T, C, H, W]

)

# Export result

save_video(video, "output/streaming_timelapse.mp4", fps=24)

```

In this configuration, the diffusion model runs on `cuda:0` while the VAE executes on `cuda:1`, eliminating memory contention. The `async_vae` flag ensures the CUDA stream on `cuda:1` operates concurrently with the next diffusion steps on `cuda:0`.

## Key Source Files

| File | Purpose |
|------|---------|
| [`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py) | Contains the `CausalDiffusionInferencePipeline` class with the chunk-wise loop (line 52), async decode logic (lines 60–68), and thread-based pipeline fallback (lines 71–84). |
| [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py) | Implements `place_vae_for_streaming` (lines 50–72) for device relocation and helper utilities for input preparation and video export. |
| [`wan_5b/configs/wan_ti2v_5B.py`](https://github.com/NVlabs/LongLive/blob/main/wan_5b/configs/wan_ti2v_5B.py) | Reference configuration showing inference flag declarations including `streaming_vae`, `async_vae`, and `vae_device`. |
| [`utils/wan_5b_wrapper.py`](https://github.com/NVlabs/LongLive/blob/main/utils/wan_5b_wrapper.py) | VAE construction logic (`build_vae_5b`) used when initializing models for multi-device deployment. |

## Summary

- The **streaming VAE pipeline** enables arbitrary-length video generation by decoding temporal chunks during the diffusion process rather than waiting for full sequence completion.
- Activate the feature by setting `inference.streaming_vae: true` in your configuration.
- Offload memory by setting `vae_device` to a secondary GPU, or enable `async_vae` for same-GPU parallel execution via CUDA streams.
- Tune `num_frame_per_block` (typically 8–16 frames) to balance memory usage against per-chunk computational overhead.
- The core logic resides in `CausalDiffusionInferencePipeline` with device management handled by `place_vae_for_streaming` in [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py).

## Frequently Asked Questions

### What is the difference between `streaming`, `streaming-pipeline`, and `streaming-async` modes?

According to the conditional block at lines 100–106 in [`pipeline/causal_diffusion_inference.py`](https://github.com/NVlabs/LongLive/blob/main/pipeline/causal_diffusion_inference.py), the three modes dictate where the VAE executes. The `streaming` mode decodes synchronously on the same device as the diffusion model. The `streaming-pipeline` mode moves the VAE to a separate device specified by `vae_device` and manages decode in a background Python thread. The `streaming-async` mode keeps the VAE on the same device but executes decode on a dedicated CUDA stream to overlap with the next diffusion steps.

### Can I run the streaming VAE pipeline on a single GPU?

Yes. If `vae_device` is unset or matches the diffusion device, and `async_vae` is disabled, the pipeline defaults to synchronous on-device decoding. To maximize throughput on a single GPU, enable `async_vae: true`, which uses separate CUDA streams (lines 60–68) to decode the current chunk while the diffusion model computes the next latent block.

### How does `num_frame_per_block` affect performance and quality?

The `num_frame_per_block` parameter controls the temporal granularity of streaming chunks. Values of 1 frame minimize memory usage but increase kernel launch overhead and may reduce temporal consistency across chunk boundaries. Values of 8–16 frames provide a practical balance for 4K video generation, maintaining quality while keeping peak memory well below the requirements of full-sequence decoding.

### Where is the VAE model moved when `vae_device` is specified?

The relocation is handled by `place_vae_for_streaming` in [`utils/inference_utils.py`](https://github.com/NVlabs/LongLive/blob/main/utils/inference_utils.py) (lines 50–72). This function transfers the VAE module along with its running statistics buffers (`mean` and `std`) to the specified `torch.device`, ensuring that decode operations execute entirely on the target GPU without paging memory back to the diffusion device.