Achieving 27FPS for Minute-Length Video Generation with LongSANA: A Complete Technical Guide
LongSANA achieves approximately 27 frames-per-second for minute-length videos by combining chunk-causal inference with KV-cache reuse, streaming training, and the optimized LongLiveFlowEuler sampler.
LongSANA extends the SANA diffusion model from NVIDIA Labs to enable high-resolution video generation at real-time speeds. This guide explains the architectural innovations in the NVlabs/Sana repository that make 27FPS video generation possible on a single GPU, including specific implementation details from the source code.
Architectural Foundations for Constant-Time Generation
LongSANA maintains high throughput across thousands of frames by rearchitecting both the training and inference pipelines to work with chunks rather than full sequences.
Streaming Training via LongSANATrainer
The model learns to generate extended sequences through overlapping chunks using the LongSANATrainer class defined in diffusion/longsana/trainer/longsana_trainer.py. During training, videos are split into manageable blocks so the model learns continuation without requiring the entire sequence in memory. The trainer initializes a StreamingSANATrainingModel wrapper that handles chunk boundaries and gradient accumulation across the streaming buffer.
Chunk-Causal Inference with KV-Cache Accumulation
At inference time, the LongLiveFlowEuler class in diffusion/longsana/sampler.py generates video in discrete blocks (base_model_frames) while caching attention keys and values. The private method _accumulate_kv_cache merges tensors across a configurable window defined by num_cached_blocks, allowing the model to reference previous chunks without recomputing attention over the entire history. This design ensures that per-frame cost remains constant after the first few blocks, rather than growing quadratically with sequence length.
Flow-Match Scheduler Optimization
LongSANA replaces traditional diffusion schedulers with FlowMatchScheduler, instantiated in the model wrapper with shift=self.flow_shift, sigma_min=0.0, and extra_one_step=True. This lightweight scheduler uses a single noise scale per step, dramatically reducing the computational overhead of the sampling loop compared to conventional multi-step diffusion processes.
Implementation Guide: Configuring for 27FPS Performance
To achieve the benchmark 27FPS on an A100-class GPU when generating 480p videos, you must configure the sampling algorithm, chunk size, and cache settings precisely.
Command-Line Configuration
Use the inference_sana_video.py script with the self-forcing checkpoint and specific caching parameters:
python -m scripts.inference_sana_video \
--config configs/sana_video_config/longsana/480ms/self_forcing.yaml \
--model_path hf://Efficient-Large-Model/LongSANA_2B_480p_self_forcing/checkpoints/LongSANA_2B_480p_self_forcing.pt \
--sampling_algo longlive_flow_euler \
--base_model_frames 40 \
--num_cached_blocks 2 \
--fps 27 \
--bs 1 \
--num_frames 800 \
--prompt "A serene sunrise over a misty lake, gentle waves, soft pastel colors." \
--output ./generated_video
Critical parameters explained:
--sampling_algo longlive_flow_euler: Activates theLongLiveFlowEulersampler, the only implementation that supports streaming block generation and KV-cache persistence.--base_model_frames 40: Sets each generation chunk to 40 frames (approximately 1.6 seconds at 25FPS). This matches the model's stride pattern and balances memory usage with computational overhead.--num_cached_blocks 2: Retains the previous two chunks in GPU memory, enabling attention reuse while preventing memory exhaustion on long sequences.--fps 27: Targets the output frame rate; when combined with the above optimizations, the generation pipeline achieves this throughput.--num_frames 800: Generates roughly 30 seconds of video; increase to 1500 frames for a full minute.
Programmatic API Implementation
For custom integrations, instantiate the pipeline components directly using the classes from the LongSANA module:
from diffusion.longsana.sana_video_pipeline import LongSANAVideoInference
from diffusion.longsana.sampler import LongLiveFlowEuler
import torch
# Configure inference parameters
cfg = LongSANAVideoInference(
model_path="hf://Efficient-Large-Model/LongSANA_2B_480p_self_forcing/checkpoints/LongSANA_2B_480p_self_forcing.pt",
prompt="A quiet forest path in early autumn, golden leaves rustling.",
num_frames=800,
sampling_algo="longlive_flow_euler",
base_model_frames=40,
num_cached_blocks=2,
fps=27,
)
# Initialize model and sampler
model = cfg.build_model()
model.eval()
sampler = LongLiveFlowEuler(
model_fn=model,
condition=cfg.get_condition_embeddings(),
model_kwargs={"mask": None},
flow_shift=cfg.flow_shift,
base_chunk_frames=cfg.base_model_frames,
num_cached_blocks=cfg.num_cached_blocks,
)
# Generate video with automatic KV-cache management
latents = torch.randn(1, cfg.latent_dim, cfg.num_frames // cfg.stride[0],
cfg.latent_h, cfg.latent_w, device="cuda")
videos = sampler.sample(latents) # Returns [B, C, T, H, W]
decoded = cfg.decode(videos) # VAE decoding
decoded.save("output_27fps.mp4", fps=cfg.fps)
The LongSANAVideoInference class handles configuration validation and provides the build_model, get_condition_embeddings, and decode helper methods used internally by the inference script.
Verifying 27FPS Throughput
To confirm your configuration achieves the target performance, profile the generation loop with explicit CUDA synchronization:
import time
import torch
from diffusion.longsana.sampler import LongLiveFlowEuler
# ... model and sampler initialization ...
latents = torch.randn(1, 4, 800 // 4, 30, 30, device="cuda")
torch.cuda.synchronize()
start = time.time()
output = sampler.sample(latents)
torch.cuda.synchronize()
elapsed = time.time() - start
fps = 800 / elapsed
print(f"Generated 800 frames in {elapsed:.2f}s → {fps:.1f} FPS")
On an NVIDIA A100 (40GB), this configuration typically yields approximately 27.3 FPS, generating an 800-frame video in under 30 seconds.
Summary
- LongSANA achieves 27FPS through chunk-causal inference that reuses KV-caches across generation blocks, avoiding the quadratic cost of full attention over long sequences.
- The
LongLiveFlowEulersampler indiffusion/longsana/sampler.pyimplements the critical_accumulate_kv_cachemethod that persists attention states across chunk boundaries. - Configuration requirements include setting
base_model_framesto 40,num_cached_blocksto 2, and using theLongSANA_2B_480p_self_forcing.ptcheckpoint with theself_forcing.yamlconfiguration. - Memory efficiency comes from the
FlowMatchSchedulerand aggressive cache management inLongSANATrainer, allowing minute-length 480p video generation on a single GPU.
Frequently Asked Questions
What hardware is required to achieve 27FPS with LongSANA?
An NVIDIA A100 GPU with 40GB of VRAM is the reference hardware for achieving 27FPS at 480p resolution. The implementation relies on torch.cuda.empty_cache() calls within scripts/inference_sana_video.py to manage memory between chunks, but consumer GPUs with less VRAM will need to reduce base_model_frames or num_cached_blocks, which lowers the effective frame rate.
Why does LongSANA use chunk-causal inference instead of full autoregression?
Full autoregressive attention requires computing relationships between every pair of frames, resulting in O(N²) complexity that becomes prohibitively expensive for minute-long videos. LongSANA's chunk-causal approach in diffusion/longsana/sampler.py processes blocks independently while caching keys and values from previous blocks, reducing the per-step cost to O(M) where M is the block size, enabling the constant-time generation necessary for 27FPS throughput.
How does the KV-cache limit affect maximum video length?
The num_cached_blocks parameter controls how many previous chunks remain in GPU memory for attention reuse. Setting this to 2 (as recommended) means the model can reference the past 80 frames (2 chunks × 40 frames) while generating the current block. While this limits direct attention to recent history, the streaming training in diffusion/longsana/trainer/longsana_trainer.py ensures the model learns to maintain coherent long-range motion without requiring access to the entire sequence history.
What distinguishes LongLiveFlowEuler from standard Euler samplers?
LongLiveFlowEuler extends standard Euler sampling by implementing block-wise generation with stateful caching. Unlike conventional samplers that process entire sequences or independent frames, it manages the denoising_step_list across chunks and accumulates KV-caches via _accumulate_kv_cache. This stateful design, combined with the FlowMatchScheduler using extra_one_step=True, eliminates redundant computation between blocks, whereas standard Euler implementations would need to recompute attention for the full growing sequence at each step.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →