How Async VAE Decoding Improves Inference Throughput in LongLive

Enabling async VAE decoding in LongLive runs the VAE decoder on a separate CUDA stream that overlaps with the diffusion model's computation, reducing total inference time from the sum of both stages to roughly their maximum and delivering 1.5×–2× higher throughput for long-video generation.

LongLive (NVlabs/LongLive) is an open-source framework for long-video generation that optimizes inference pipelines through streaming architectures. When generating extended video sequences, the async VAE decoding feature eliminates the bottleneck caused by sequential latent decoding, allowing the diffusion model and VAE decoder to execute concurrently on the same GPU.

How Async VAE Decoding Works

The implementation in pipeline/causal_diffusion_inference.py (lines 1645–1697) leverages PyTorch CUDA streams to parallelize the diffusion and decoding stages. The mechanism revolves around four core steps that ensure correct synchronization while maximizing GPU utilization.

Detecting Streaming Mode

The pipeline first evaluates configuration flags to determine whether asynchronous decoding should activate. In causal_diffusion_inference.py (lines 1645–1649), the code checks if streaming VAE is enabled and whether the VAE resides on a separate device:

streaming_decode = self.streaming_vae and not return_latents
pipeline_vae = streaming_decode and self.vae_device is not None
async_vae = streaming_decode and self.async_vae and not pipeline_vae

When async_vae evaluates to True, the system prepares to decode latent chunks on an auxiliary stream rather than blocking the main diffusion loop.

Allocating a Dedicated CUDA Stream

Upon detecting the async condition, the pipeline allocates a non-default CUDA stream exclusively for VAE operations. This occurs at lines 1650–1652:

if async_vae:
    vae_stream = torch.cuda.Stream(device=noise.device)
    prev_vae_done = None

Creating a separate stream allows the VAE decoder to execute kernels independently from the diffusion model's default stream.

Overlapping Diffusion and Decode

During each iteration of the diffusion loop, the pipeline records completion events and launches VAE decodes on the auxiliary stream. This overlap logic (lines 1655–1670) ensures that while the VAE processes the previous chunk, the diffusion model already computes the next latent representation:

if async_vae:
    diffusion_done = torch.cuda.Event()
    diffusion_done.record()
    if prev_vae_done is not None:
        prev_vae_done.synchronize()
    with torch.cuda.stream(vae_stream):
        vae_stream.wait_event(diffusion_done)
        chunk_bcthw = latents.permute(0, 2, 1, 3, 4).contiguous()
        decoded_chunk = self.vae.model.cached_decode(
            chunk_bcthw, vae_scale).float().clamp_(-1, 1)
        video_chunks.append(decoded_chunk)
    prev_vae_done = torch.cuda.Event()
    prev_vae_done.record(vae_stream)

The torch.cuda.Event markers enforce happens-before relationships without stalling the GPU. The diffusion loop proceeds immediately after recording diffusion_done, while the VAE stream waits only for that specific event before beginning its decode.

Final Synchronization

After the diffusion loop completes all chunks, the pipeline synchronizes the VAE stream to ensure all pending decodes finish before returning the final video. This final barrier appears at lines 1695–1697:

if async_vae:
    vae_stream.synchronize()

Configuration and Implementation

YAML Configuration

Enable async VAE decoding through the configuration file by setting the streaming and async flags:


# configs/inference.yaml

inference:
  streaming_vae: true          # Stream chunks as they are generated

  async_vae: true              # Decode each chunk asynchronously

  vae_device: null             # Keep VAE on the same GPU as diffusion

  sampling_steps: 50

Python API Usage

Instantiate the pipeline with async VAE enabled via arguments:

from inference import CausalDiffusionInferencePipeline
from utils.config import load_config

cfg = load_config("configs/inference.yaml")
pipeline = CausalDiffusionInferencePipeline(
    args=cfg,
    device="cuda",
)

noise = torch.randn(1, 64, 4, 256, 256, device="cuda")
video = pipeline.inference(
    noise=noise,
    text_prompts=["A sunrise over a mountain range"],
    streaming_vae=True,
    async_vae=True,
)

Manual Stream Implementation

For custom pipelines, implement the same concurrency pattern manually:

import torch

vae_stream = torch.cuda.Stream()
event = torch.cuda.Event()

latent = diffusion_step(prev_latent)
event.record()

with torch.cuda.stream(vae_stream):
    vae_stream.wait_event(event)
    decoded = vae.decode(latent)
    decoded_cpu = decoded.cpu()

Performance Impact

Async VAE decoding transforms the inference timeline from sequential execution to parallel execution. Because the VAE operates concurrently with the diffusion loop, the overall wall-clock time approaches the maximum of the two stage durations rather than their sum.

In practice, this architecture yields 1.5×–2× higher throughput for long-video generation, particularly when the VAE model becomes the bottleneck. This commonly occurs on GPUs where the decoder's compute bandwidth limits overall performance. The async approach also preserves deterministic latent ordering by copying each chunk's tensor before decoding, ensuring the diffusion side never stalls waiting for the decoder to complete.

LongLive also supports a pipeline_vae mode that moves decoding to a separate worker thread and device, but async VAE provides a lightweight alternative that avoids thread-management overhead while still achieving significant overlap.

Summary

  • Async VAE decoding launches the decoder on a separate CUDA stream to overlap with diffusion computation.
  • The implementation in causal_diffusion_inference.py (lines 1645–1697) uses torch.cuda.Event and torch.cuda.Stream to synchronize work without blocking.
  • Configuration requires streaming_vae=True and async_vae=True with vae_device set to null.
  • Performance gains of 1.5×–2× are typical when the VAE decoder represents the throughput bottleneck.
  • This method offers lower overhead than the multi-threaded pipeline_vae alternative while maintaining full parallelism on a single GPU.

Frequently Asked Questions

What is async VAE decoding?

Async VAE decoding is a GPU optimization technique where the VAE decoder runs on a separate CUDA stream from the diffusion model. In LongLive, this allows the system to decode latent representations into RGB pixels while simultaneously computing the next diffusion chunk, rather than waiting for each stage to complete sequentially.

When should I use async VAE instead of pipeline VAE?

Use async VAE when running both the diffusion model and VAE decoder on the same GPU and you want minimal overhead. Choose pipeline VAE (enabled via vae_device parameter) only when you have a second GPU available to dedicate entirely to decoding, as it uses a multi-threaded queue that adds communication overhead but can utilize separate hardware.

How much throughput improvement does async VAE provide?

According to the LongLive source implementation, enabling async VAE decoding typically improves inference throughput by 1.5× to 2× for long-video generation. The gains are most pronounced when the VAE decoder is the bottleneck, which depends on the specific GPU's compute characteristics and the video resolution.

Does async VAE decoding affect video quality?

No, async VAE decoding does not affect output quality. The implementation ensures deterministic processing by copying each latent chunk before decoding begins. The use of CUDA events guarantees that each chunk is decoded only after its corresponding diffusion computation completes, preserving the exact same pixel values that would result from synchronous decoding.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →