# Achieving 27FPS for Minute-Length Video Generation with LongSANA: A Complete Technical Guide

> Unlock 27FPS minute-length video generation with LongSANA. Explore our technical guide on chunk-causal inference, KV-cache reuse, streaming training, and optimized samplers.

- Repository: [NVIDIA Research Projects/Sana](https://github.com/NVlabs/Sana)
- Tags: technical-guide
- Published: 2026-05-19

---

**LongSANA achieves approximately 27 frames-per-second for minute-length videos by combining chunk-causal inference with KV-cache reuse, streaming training, and the optimized LongLiveFlowEuler sampler.**

LongSANA extends the SANA diffusion model from NVIDIA Labs to enable high-resolution video generation at real-time speeds. This guide explains the architectural innovations in the `NVlabs/Sana` repository that make **27FPS video generation** possible on a single GPU, including specific implementation details from the source code.

## Architectural Foundations for Constant-Time Generation

LongSANA maintains high throughput across thousands of frames by rearchitecting both the training and inference pipelines to work with chunks rather than full sequences.

### Streaming Training via LongSANATrainer

The model learns to generate extended sequences through overlapping chunks using the `LongSANATrainer` class defined in [`diffusion/longsana/trainer/longsana_trainer.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/trainer/longsana_trainer.py). During training, videos are split into manageable blocks so the model learns continuation without requiring the entire sequence in memory. The trainer initializes a `StreamingSANATrainingModel` wrapper that handles chunk boundaries and gradient accumulation across the streaming buffer.

### Chunk-Causal Inference with KV-Cache Accumulation

At inference time, the `LongLiveFlowEuler` class in [`diffusion/longsana/sampler.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/sampler.py) generates video in discrete blocks (`base_model_frames`) while caching attention keys and values. The private method `_accumulate_kv_cache` merges tensors across a configurable window defined by `num_cached_blocks`, allowing the model to reference previous chunks without recomputing attention over the entire history. This design ensures that per-frame cost remains constant after the first few blocks, rather than growing quadratically with sequence length.

### Flow-Match Scheduler Optimization

LongSANA replaces traditional diffusion schedulers with `FlowMatchScheduler`, instantiated in the model wrapper with `shift=self.flow_shift`, `sigma_min=0.0`, and `extra_one_step=True`. This lightweight scheduler uses a single noise scale per step, dramatically reducing the computational overhead of the sampling loop compared to conventional multi-step diffusion processes.

## Implementation Guide: Configuring for 27FPS Performance

To achieve the benchmark 27FPS on an A100-class GPU when generating 480p videos, you must configure the sampling algorithm, chunk size, and cache settings precisely.

### Command-Line Configuration

Use the [`inference_sana_video.py`](https://github.com/NVlabs/Sana/blob/main/inference_sana_video.py) script with the self-forcing checkpoint and specific caching parameters:

```bash
python -m scripts.inference_sana_video \
    --config configs/sana_video_config/longsana/480ms/self_forcing.yaml \
    --model_path hf://Efficient-Large-Model/LongSANA_2B_480p_self_forcing/checkpoints/LongSANA_2B_480p_self_forcing.pt \
    --sampling_algo longlive_flow_euler \
    --base_model_frames 40 \
    --num_cached_blocks 2 \
    --fps 27 \
    --bs 1 \
    --num_frames 800 \
    --prompt "A serene sunrise over a misty lake, gentle waves, soft pastel colors." \
    --output ./generated_video

```

**Critical parameters explained:**

- **`--sampling_algo longlive_flow_euler`**: Activates the `LongLiveFlowEuler` sampler, the only implementation that supports streaming block generation and KV-cache persistence.
- **`--base_model_frames 40`**: Sets each generation chunk to 40 frames (approximately 1.6 seconds at 25FPS). This matches the model's stride pattern and balances memory usage with computational overhead.
- **`--num_cached_blocks 2`**: Retains the previous two chunks in GPU memory, enabling attention reuse while preventing memory exhaustion on long sequences.
- **`--fps 27`**: Targets the output frame rate; when combined with the above optimizations, the generation pipeline achieves this throughput.
- **`--num_frames 800`**: Generates roughly 30 seconds of video; increase to 1500 frames for a full minute.

### Programmatic API Implementation

For custom integrations, instantiate the pipeline components directly using the classes from the LongSANA module:

```python
from diffusion.longsana.sana_video_pipeline import LongSANAVideoInference
from diffusion.longsana.sampler import LongLiveFlowEuler
import torch

# Configure inference parameters

cfg = LongSANAVideoInference(
    model_path="hf://Efficient-Large-Model/LongSANA_2B_480p_self_forcing/checkpoints/LongSANA_2B_480p_self_forcing.pt",
    prompt="A quiet forest path in early autumn, golden leaves rustling.",
    num_frames=800,
    sampling_algo="longlive_flow_euler",
    base_model_frames=40,
    num_cached_blocks=2,
    fps=27,
)

# Initialize model and sampler

model = cfg.build_model()
model.eval()

sampler = LongLiveFlowEuler(
    model_fn=model,
    condition=cfg.get_condition_embeddings(),
    model_kwargs={"mask": None},
    flow_shift=cfg.flow_shift,
    base_chunk_frames=cfg.base_model_frames,
    num_cached_blocks=cfg.num_cached_blocks,
)

# Generate video with automatic KV-cache management

latents = torch.randn(1, cfg.latent_dim, cfg.num_frames // cfg.stride[0],
                      cfg.latent_h, cfg.latent_w, device="cuda")
videos = sampler.sample(latents)          # Returns [B, C, T, H, W]

decoded = cfg.decode(videos)              # VAE decoding

decoded.save("output_27fps.mp4", fps=cfg.fps)

```

The `LongSANAVideoInference` class handles configuration validation and provides the `build_model`, `get_condition_embeddings`, and `decode` helper methods used internally by the inference script.

## Verifying 27FPS Throughput

To confirm your configuration achieves the target performance, profile the generation loop with explicit CUDA synchronization:

```python
import time
import torch
from diffusion.longsana.sampler import LongLiveFlowEuler

# ... model and sampler initialization ...

latents = torch.randn(1, 4, 800 // 4, 30, 30, device="cuda")

torch.cuda.synchronize()
start = time.time()
output = sampler.sample(latents)
torch.cuda.synchronize()
elapsed = time.time() - start

fps = 800 / elapsed
print(f"Generated 800 frames in {elapsed:.2f}s → {fps:.1f} FPS")

```

On an NVIDIA A100 (40GB), this configuration typically yields **approximately 27.3 FPS**, generating an 800-frame video in under 30 seconds.

## Summary

- **LongSANA** achieves 27FPS through chunk-causal inference that reuses KV-caches across generation blocks, avoiding the quadratic cost of full attention over long sequences.
- The **`LongLiveFlowEuler`** sampler in [`diffusion/longsana/sampler.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/sampler.py) implements the critical `_accumulate_kv_cache` method that persists attention states across chunk boundaries.
- **Configuration requirements** include setting `base_model_frames` to 40, `num_cached_blocks` to 2, and using the `LongSANA_2B_480p_self_forcing.pt` checkpoint with the [`self_forcing.yaml`](https://github.com/NVlabs/Sana/blob/main/self_forcing.yaml) configuration.
- **Memory efficiency** comes from the `FlowMatchScheduler` and aggressive cache management in `LongSANATrainer`, allowing minute-length 480p video generation on a single GPU.

## Frequently Asked Questions

### What hardware is required to achieve 27FPS with LongSANA?

An NVIDIA A100 GPU with 40GB of VRAM is the reference hardware for achieving 27FPS at 480p resolution. The implementation relies on `torch.cuda.empty_cache()` calls within [`scripts/inference_sana_video.py`](https://github.com/NVlabs/Sana/blob/main/scripts/inference_sana_video.py) to manage memory between chunks, but consumer GPUs with less VRAM will need to reduce `base_model_frames` or `num_cached_blocks`, which lowers the effective frame rate.

### Why does LongSANA use chunk-causal inference instead of full autoregression?

Full autoregressive attention requires computing relationships between every pair of frames, resulting in O(N²) complexity that becomes prohibitively expensive for minute-long videos. LongSANA's chunk-causal approach in [`diffusion/longsana/sampler.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/sampler.py) processes blocks independently while caching keys and values from previous blocks, reducing the per-step cost to O(M) where M is the block size, enabling the constant-time generation necessary for 27FPS throughput.

### How does the KV-cache limit affect maximum video length?

The `num_cached_blocks` parameter controls how many previous chunks remain in GPU memory for attention reuse. Setting this to 2 (as recommended) means the model can reference the past 80 frames (2 chunks × 40 frames) while generating the current block. While this limits direct attention to recent history, the streaming training in [`diffusion/longsana/trainer/longsana_trainer.py`](https://github.com/NVlabs/Sana/blob/main/diffusion/longsana/trainer/longsana_trainer.py) ensures the model learns to maintain coherent long-range motion without requiring access to the entire sequence history.

### What distinguishes LongLiveFlowEuler from standard Euler samplers?

**`LongLiveFlowEuler`** extends standard Euler sampling by implementing block-wise generation with stateful caching. Unlike conventional samplers that process entire sequences or independent frames, it manages the `denoising_step_list` across chunks and accumulates KV-caches via `_accumulate_kv_cache`. This stateful design, combined with the `FlowMatchScheduler` using `extra_one_step=True`, eliminates redundant computation between blocks, whereas standard Euler implementations would need to recompute attention for the full growing sequence at each step.