How Async VAE Decoding Improves Inference Throughput in LongLive
Enabling async VAE decoding in LongLive runs the VAE decoder on a separate CUDA stream that overlaps with the diffusion model's computation, reducing total inference time from the sum of both stages to roughly their maximum and delivering 1.5×–2× higher throughput for long-video generation.
LongLive (NVlabs/LongLive) is an open-source framework for long-video generation that optimizes inference pipelines through streaming architectures. When generating extended video sequences, the async VAE decoding feature eliminates the bottleneck caused by sequential latent decoding, allowing the diffusion model and VAE decoder to execute concurrently on the same GPU.
How Async VAE Decoding Works
The implementation in pipeline/causal_diffusion_inference.py (lines 1645–1697) leverages PyTorch CUDA streams to parallelize the diffusion and decoding stages. The mechanism revolves around four core steps that ensure correct synchronization while maximizing GPU utilization.
Detecting Streaming Mode
The pipeline first evaluates configuration flags to determine whether asynchronous decoding should activate. In causal_diffusion_inference.py (lines 1645–1649), the code checks if streaming VAE is enabled and whether the VAE resides on a separate device:
streaming_decode = self.streaming_vae and not return_latents
pipeline_vae = streaming_decode and self.vae_device is not None
async_vae = streaming_decode and self.async_vae and not pipeline_vae
When async_vae evaluates to True, the system prepares to decode latent chunks on an auxiliary stream rather than blocking the main diffusion loop.
Allocating a Dedicated CUDA Stream
Upon detecting the async condition, the pipeline allocates a non-default CUDA stream exclusively for VAE operations. This occurs at lines 1650–1652:
if async_vae:
vae_stream = torch.cuda.Stream(device=noise.device)
prev_vae_done = None
Creating a separate stream allows the VAE decoder to execute kernels independently from the diffusion model's default stream.
Overlapping Diffusion and Decode
During each iteration of the diffusion loop, the pipeline records completion events and launches VAE decodes on the auxiliary stream. This overlap logic (lines 1655–1670) ensures that while the VAE processes the previous chunk, the diffusion model already computes the next latent representation:
if async_vae:
diffusion_done = torch.cuda.Event()
diffusion_done.record()
if prev_vae_done is not None:
prev_vae_done.synchronize()
with torch.cuda.stream(vae_stream):
vae_stream.wait_event(diffusion_done)
chunk_bcthw = latents.permute(0, 2, 1, 3, 4).contiguous()
decoded_chunk = self.vae.model.cached_decode(
chunk_bcthw, vae_scale).float().clamp_(-1, 1)
video_chunks.append(decoded_chunk)
prev_vae_done = torch.cuda.Event()
prev_vae_done.record(vae_stream)
The torch.cuda.Event markers enforce happens-before relationships without stalling the GPU. The diffusion loop proceeds immediately after recording diffusion_done, while the VAE stream waits only for that specific event before beginning its decode.
Final Synchronization
After the diffusion loop completes all chunks, the pipeline synchronizes the VAE stream to ensure all pending decodes finish before returning the final video. This final barrier appears at lines 1695–1697:
if async_vae:
vae_stream.synchronize()
Configuration and Implementation
YAML Configuration
Enable async VAE decoding through the configuration file by setting the streaming and async flags:
# configs/inference.yaml
inference:
streaming_vae: true # Stream chunks as they are generated
async_vae: true # Decode each chunk asynchronously
vae_device: null # Keep VAE on the same GPU as diffusion
sampling_steps: 50
Python API Usage
Instantiate the pipeline with async VAE enabled via arguments:
from inference import CausalDiffusionInferencePipeline
from utils.config import load_config
cfg = load_config("configs/inference.yaml")
pipeline = CausalDiffusionInferencePipeline(
args=cfg,
device="cuda",
)
noise = torch.randn(1, 64, 4, 256, 256, device="cuda")
video = pipeline.inference(
noise=noise,
text_prompts=["A sunrise over a mountain range"],
streaming_vae=True,
async_vae=True,
)
Manual Stream Implementation
For custom pipelines, implement the same concurrency pattern manually:
import torch
vae_stream = torch.cuda.Stream()
event = torch.cuda.Event()
latent = diffusion_step(prev_latent)
event.record()
with torch.cuda.stream(vae_stream):
vae_stream.wait_event(event)
decoded = vae.decode(latent)
decoded_cpu = decoded.cpu()
Performance Impact
Async VAE decoding transforms the inference timeline from sequential execution to parallel execution. Because the VAE operates concurrently with the diffusion loop, the overall wall-clock time approaches the maximum of the two stage durations rather than their sum.
In practice, this architecture yields 1.5×–2× higher throughput for long-video generation, particularly when the VAE model becomes the bottleneck. This commonly occurs on GPUs where the decoder's compute bandwidth limits overall performance. The async approach also preserves deterministic latent ordering by copying each chunk's tensor before decoding, ensuring the diffusion side never stalls waiting for the decoder to complete.
LongLive also supports a pipeline_vae mode that moves decoding to a separate worker thread and device, but async VAE provides a lightweight alternative that avoids thread-management overhead while still achieving significant overlap.
Summary
- Async VAE decoding launches the decoder on a separate CUDA stream to overlap with diffusion computation.
- The implementation in
causal_diffusion_inference.py(lines 1645–1697) usestorch.cuda.Eventandtorch.cuda.Streamto synchronize work without blocking. - Configuration requires
streaming_vae=Trueandasync_vae=Truewithvae_deviceset tonull. - Performance gains of 1.5×–2× are typical when the VAE decoder represents the throughput bottleneck.
- This method offers lower overhead than the multi-threaded
pipeline_vaealternative while maintaining full parallelism on a single GPU.
Frequently Asked Questions
What is async VAE decoding?
Async VAE decoding is a GPU optimization technique where the VAE decoder runs on a separate CUDA stream from the diffusion model. In LongLive, this allows the system to decode latent representations into RGB pixels while simultaneously computing the next diffusion chunk, rather than waiting for each stage to complete sequentially.
When should I use async VAE instead of pipeline VAE?
Use async VAE when running both the diffusion model and VAE decoder on the same GPU and you want minimal overhead. Choose pipeline VAE (enabled via vae_device parameter) only when you have a second GPU available to dedicate entirely to decoding, as it uses a multi-threaded queue that adds communication overhead but can utilize separate hardware.
How much throughput improvement does async VAE provide?
According to the LongLive source implementation, enabling async VAE decoding typically improves inference throughput by 1.5× to 2× for long-video generation. The gains are most pronounced when the VAE decoder is the bottleneck, which depends on the specific GPU's compute characteristics and the video resolution.
Does async VAE decoding affect video quality?
No, async VAE decoding does not affect output quality. The implementation ensures deterministic processing by copying each latent chunk before decoding begins. The use of CUDA events guarantees that each chunk is decoded only after its corresponding diffusion computation completes, preserving the exact same pixel values that would result from synchronous decoding.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →