How to Configure the Streaming VAE Pipeline for Long Video Generation in LongLive

The streaming VAE pipeline is an incremental decoding mode that processes video in temporal chunks during diffusion to reduce memory usage, enabled by setting inference.streaming_vae=true and optionally configuring vae_device for multi-GPU or async_vae for CUDA stream parallelization.

The LongLive repository by NVlabs implements a streaming VAE pipeline that enables generation of arbitrarily long videos without requiring tens of gigabytes of GPU memory. Unlike standard diffusion pipelines that accumulate the entire latent tensor before decoding, LongLive processes and decodes temporal blocks concurrently with denoising. This guide explains the architecture and the specific configuration flags required to activate this mode.

What Is the Streaming VAE Pipeline?

The streaming VAE pipeline is an inference mode in CausalDiffusionInferencePipeline (located in pipeline/causal_diffusion_inference.py) that integrates decoding directly into the diffusion loop. Instead of waiting for the complete latent representation—which can exceed available VRAM for long sequences—the pipeline splits generation into blocks of num_frame_per_block frames, decodes each block immediately after denoising, and appends the result to the output video stream.

This approach reduces peak GPU memory consumption from tens of gigabytes to a level determined by chunk size. When using asynchronous execution, the pipeline also overlaps VAE decode with subsequent diffusion steps, significantly improving throughput for long-form content.

How the Streaming VAE Architecture Works

The implementation selects between three execution strategies based on configuration flags parsed at lines 83–86 of the inference pipeline.

Chunk-Wise Processing Logic

At the core of the system is a temporal chunking mechanism in the inference method (line 52 of pipeline/causal_diffusion_inference.py):

for chunk_index, current_num_frames in enumerate(all_num_frames):

For each iteration, the code runs the spatial denoising loop, stores the latent block in an output buffer, and immediately triggers decode via the self.vae.model.cached_decode method before proceeding to the next temporal segment. This prevents accumulation of full-sequence latents in GPU memory.

Memory Optimization Strategies

The helper function place_vae_for_streaming in utils/inference_utils.py (lines 50–72) implements three distinct placement modes controlled by the conditional block at lines 100–106:

  • streaming: Decode occurs synchronously on the same device as the diffusion model.
  • streaming-pipeline: The VAE module and its mean/std buffers transfer to a separate device specified by vae_device, isolating decode memory from diffusion compute using a background Python thread (lines 71–84).
  • streaming-async: Decode executes on a dedicated CUDA stream within the same GPU context (lines 60–68), allowing overlap between VAE execution and the next diffusion steps without requiring a second physical GPU.

When async_vae is enabled, the code creates a CUDA stream and records synchronization events to manage the parallel execution path safely.

Configuration Guide

Enable and tune the streaming VAE pipeline by setting these keys in your inference configuration (see wan_5b/configs/wan_ti2v_5B.py for a complete template):

Config Key Description Recommended Value
inference.streaming_vae Master flag enabling chunk-wise decode mode. true
inference.async_vae Enables CUDA stream-based overlap; requires streaming_vae. true (or false for thread-based path)
inference.vae_device Target GPU for VAE decode (e.g., "cuda:1"). Must differ from diffusion device for offloading. "cuda:1"
inference.num_frame_per_block Temporal chunk size; smaller values reduce memory but increase kernel launch overhead. 8–16 for 4K video
inference.return_latents Debug flag to skip VAE decoding and return raw latents. false

Implementation Example

The following example demonstrates how to configure and run the streaming pipeline using LongLive's utility functions:

import torch
from utils.config import load_config_from_yaml
from pipeline.causal_diffusion_inference import CausalDiffusionInferencePipeline
from utils.inference_utils import place_vae_for_streaming, prepare_single_prompt_inputs, save_video

# Load configuration and enable streaming

cfg = load_config_from_yaml("wan_5b/configs/wan_ti2v_5B.yaml")
cfg.inference.streaming_vae = True          # Enable streaming mode

cfg.inference.async_vae = True              # Use async CUDA stream

cfg.inference.vae_device = "cuda:1"         # Offload VAE to second GPU

cfg.inference.num_frame_per_block = 16      # Decode 16 frames per chunk

# Initialize pipeline on primary GPU (cuda:0)

device = torch.device("cuda:0")
pipeline = CausalDiffusionInferencePipeline(args=cfg, device=device)

# Relocate VAE to dedicated device (constructor also calls this internally)

place_vae_for_streaming(pipeline, cfg)

# Prepare inputs

noise, prompts = prepare_single_prompt_inputs(
    cfg,
    prompt="A timelapse of a city from sunrise to sunset",
    device=device,
)

# Run inference with live decoding

video = pipeline.inference(
    noise=noise,
    text_prompts=prompts,
    return_latents=False,  # Return decoded video tensor [B, T, C, H, W]

)

# Export result

save_video(video, "output/streaming_timelapse.mp4", fps=24)

In this configuration, the diffusion model runs on cuda:0 while the VAE executes on cuda:1, eliminating memory contention. The async_vae flag ensures the CUDA stream on cuda:1 operates concurrently with the next diffusion steps on cuda:0.

Key Source Files

File Purpose
pipeline/causal_diffusion_inference.py Contains the CausalDiffusionInferencePipeline class with the chunk-wise loop (line 52), async decode logic (lines 60–68), and thread-based pipeline fallback (lines 71–84).
utils/inference_utils.py Implements place_vae_for_streaming (lines 50–72) for device relocation and helper utilities for input preparation and video export.
wan_5b/configs/wan_ti2v_5B.py Reference configuration showing inference flag declarations including streaming_vae, async_vae, and vae_device.
utils/wan_5b_wrapper.py VAE construction logic (build_vae_5b) used when initializing models for multi-device deployment.

Summary

  • The streaming VAE pipeline enables arbitrary-length video generation by decoding temporal chunks during the diffusion process rather than waiting for full sequence completion.
  • Activate the feature by setting inference.streaming_vae: true in your configuration.
  • Offload memory by setting vae_device to a secondary GPU, or enable async_vae for same-GPU parallel execution via CUDA streams.
  • Tune num_frame_per_block (typically 8–16 frames) to balance memory usage against per-chunk computational overhead.
  • The core logic resides in CausalDiffusionInferencePipeline with device management handled by place_vae_for_streaming in utils/inference_utils.py.

Frequently Asked Questions

What is the difference between streaming, streaming-pipeline, and streaming-async modes?

According to the conditional block at lines 100–106 in pipeline/causal_diffusion_inference.py, the three modes dictate where the VAE executes. The streaming mode decodes synchronously on the same device as the diffusion model. The streaming-pipeline mode moves the VAE to a separate device specified by vae_device and manages decode in a background Python thread. The streaming-async mode keeps the VAE on the same device but executes decode on a dedicated CUDA stream to overlap with the next diffusion steps.

Can I run the streaming VAE pipeline on a single GPU?

Yes. If vae_device is unset or matches the diffusion device, and async_vae is disabled, the pipeline defaults to synchronous on-device decoding. To maximize throughput on a single GPU, enable async_vae: true, which uses separate CUDA streams (lines 60–68) to decode the current chunk while the diffusion model computes the next latent block.

How does num_frame_per_block affect performance and quality?

The num_frame_per_block parameter controls the temporal granularity of streaming chunks. Values of 1 frame minimize memory usage but increase kernel launch overhead and may reduce temporal consistency across chunk boundaries. Values of 8–16 frames provide a practical balance for 4K video generation, maintaining quality while keeping peak memory well below the requirements of full-sequence decoding.

Where is the VAE model moved when vae_device is specified?

The relocation is handled by place_vae_for_streaming in utils/inference_utils.py (lines 50–72). This function transfers the VAE module along with its running statistics buffers (mean and std) to the specified torch.device, ensuring that decode operations execute entirely on the target GPU without paging memory back to the diffusion device.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →