# LongLive | NVIDIA Research Projects | Knowledge Base | Instagit

LongLive 2.0: Infra - Long Video Gen

GitHub Stars: 1.9k

Repository: https://github.com/NVlabs/LongLive

---

## Articles

### [How to Integrate TriAttention KV Cache Compression with LongLive: Complete Implementation Guide](/NVlabs/LongLive/how-to-integrate-triattention-kv-cache-compression-with-longlive)

Learn to integrate TriAttention KV cache compression with LongLive. Reduce GPU memory by 50% with this guide and boost your model's efficiency without sacrificing quality.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Configure Local Attention Size for Long Video Generation in LongLive](/NVlabs/LongLive/configure-local-attention-size-long-video-generation-longlive)

Master local attention size for long video generation in LongLive. Learn to configure window_size for sliding-window attention control in NVlabs/LongLive.

- Tags: how-to-guide
- Published: 2026-05-24

### [Best Practices for NVFP4 Checkpoint Loading and Weight Materialization in LongLive](/NVlabs/LongLive/best-practices-nvfp4-checkpoint-loading-weight-materialization-longlive)

Optimize NVFP4 checkpoint loading and weight materialization in LongLive. Safely load checkpoints, unwrap generators, clean FSDP prefixes, and drop master weights to save GPU memory.

- Tags: best-practices
- Published: 2026-05-24

### [How to Use the WanDiffusionWrapper with Custom Model Configurations in LongLive](/NVlabs/LongLive/how-to-use-wandiffusionwrapper-with-custom-model-configurations-longlive)

Learn to customize the WanDiffusionWrapper in LongLive using model_kwargs for checkpoint loading, causal mode, sampling, and VAEs. Tailor your diffusion pipeline without code changes.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Debug Memory Issues with the ErrorBuffer Utility in LongLive](/NVlabs/LongLive/how-to-debug-memory-issues-with-error-buffer-utility-longlive)

Debug memory issues in LongLive using the ErrorBuffer utility. Monitor stats, verify shard size, and ensure buffer warmup for efficient memory management.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Implement Few-Step DMD Distillation for Faster Inference in LongLive](/NVlabs/LongLive/how-to-implement-few-step-dmd-distillation-for-faster-inference-longlive)

Learn how to implement few-step DMD distillation in LongLive for faster inference. Reduce num train timestep and use the score distillation trainer to compress the diffusion schedule.

- Tags: how-to-guide
- Published: 2026-05-24

### [Causal Diffusion Pipeline Architecture in LongLive 2.0: A Technical Deep Dive](/NVlabs/LongLive/causal-diffusion-pipeline-architecture-longlive-2-0)

Explore the causal diffusion pipeline architecture in LongLive 2.0. Learn how it generates videos autoregressively using causal masking and KV-cache for efficient temporal attention. Deep dive into NVlabs/LongLive.

- Tags: deep-dive
- Published: 2026-05-24

### [How to Configure the Timestep Shift Parameter for Flow Matching Schedulers in LongLive](/NVlabs/LongLive/configure-timestep-shift-parameter-flow-matching-schedulers-longlive)

Configure the timestep shift parameter in LongLive Flow Matching Schedulers via YAML or WanDiffusionWrapper to control noise distribution. Learn how to optimize your model training.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Optimize Memory Usage During AR Training with Sequence Parallelism in LongLive](/NVlabs/LongLive/optimize-memory-usage-ar-training-sequence-parallelism-longlive)

Optimize AR training memory with LongLive sequence parallelism. Shard temporal data across GPUs to slash activation memory from 40GB+ to 8-16GB per device.

- Tags: performance
- Published: 2026-05-24

### [Score Distillation vs Diffusion Training in LongLive: Technical Differences Explained](/NVlabs/LongLive/score-distillation-vs-diffusion-training-modes-longlive)

Explore score distillation vs diffusion training in NVlabs LongLive. Understand key technical differences: KL-gradient matching versus direct MSE flow prediction for your generative models.

- Tags: deep-dive
- Published: 2026-05-24

### [How to Implement KV Cache Recaching for Video Streaming in LongLive](/NVlabs/LongLive/how-to-implement-kv-cache-recaching-for-video-streaming-longlive)

Implement KV cache recaching for video streaming in LongLive by setting shot_clean_recache and multi_shot_sink to true. Expert guide to optimize inference with automatic scene cut detection.

- Tags: how-to-guide
- Published: 2026-05-24

### [How RoPE Position Embedding Offset Works for Multi-Shot Sequences in LongLive](/NVlabs/LongLive/rope-position-embedding-offset-multi-shot-sequences-longlive)

Discover how LongLive's RoPE position embedding offset revolutionizes multi-shot video generation by maintaining temporal context without KV cache resets.

- Tags: deep-dive
- Published: 2026-05-24

### [How to Configure Multi-Shot Video Generation with Scene Cut Prefixes in LongLive](/NVlabs/LongLive/configure-multi-shot-video-generation-scene-cut-prefixes-longlive)

Unlock multi-shot video generation in LongLive by configuring scene cut prefixes. Learn how to pin KV cache across scenes for seamless video creation. Get started now!

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Switch Between TransformerEngine and FourOverSix NVFP4 Backends in LongLive](/NVlabs/LongLive/switch-between-transformerengine-and-fourover-six-nvfp4-backends-longlive)

Easily switch between TransformerEngine and FourOverSix NVFP4 backends in LongLive via configuration or command line. Optimize your model performance now.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Use LoRA Adaptation with LongLive 2.0 Generator Models: A Complete Guide](/NVlabs/LongLive/how-to-use-lora-adaptation-with-longlive-2-0-generator-models)

Master LoRA adaptation with LongLive 2.0 generator models. Configure adapters, wrap transformers, and let DistillationTrainer manage low-rank weights for efficient training.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Implement Attention Sink for Managing Long Context in Video Generation with LongLive](/NVlabs/LongLive/implement-attention-sink-for-long-context-video-generation-longlive)

Learn to implement attention sink in LongLive for long context video generation. Effortlessly manage early video frames in KV-cache with Ulysses model for better results.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Configure the Streaming VAE Pipeline for Long Video Generation in LongLive](/NVlabs/LongLive/streaming-vae-pipeline-configuration-long-video-generation)

Learn to configure the streaming VAE pipeline for efficient, long video generation in LongLive. Reduce memory usage with incremental decoding. Get started now.

- Tags: how-to-guide
- Published: 2026-05-24

### [How Async VAE Decoding Improves Inference Throughput in LongLive](/NVlabs/LongLive/how-async-vae-decoding-improves-inference-throughput-longlive)

Discover how async VAE decoding in LongLive boosts inference throughput by 1.5x-2x for long video generation by overlapping computations and reducing wait times.

- Tags: performance
- Published: 2026-05-24

### [How to Use KV Cache Quantization with LongLive's LongLiveQuantizationConfig](/NVlabs/LongLive/how-to-use-kv-cache-quantization-with-longlivequantizationconfig)

Learn to use KV cache quantization with LongLiveQuantizationConfig in NVlabs/LongLive. Compress key-value tensors to FP4 with ease for faster inference.

- Tags: how-to-guide
- Published: 2026-05-24

### [FlowUniPCMultistepScheduler vs FlowDPMSolverMultistepScheduler in LongLive: Technical Differences Explained](/NVlabs/LongLive/difference-between-flowunipcscheduler-and-flowdpm-solvermultistep-scheduler-longlive)

Compare FlowUniPC and FlowDPMSolver schedulers in NVlabs LongLive. Understand their technical differences in ODE solving for video generation and choose the best for your needs.

- Tags: deep-dive
- Published: 2026-05-24

### [How to Configure Sequence Parallel Training with Balanced Workload Distribution in LongLive](/NVlabs/LongLive/how-to-configure-sequence-parallel-training-with-balanced-workload-distribution)

Learn to configure LongLive sequence parallel training for balanced workload distribution. Optimize your video data and model with NVlabs/LongLive for efficient deep learning.

- Tags: how-to-guide
- Published: 2026-05-24

### [How to Implement Multi-Shot Attention Sink for Generating Long Videos with LongLive](/NVlabs/LongLive/how-to-implement-multi-shot-attention-sink-for-generating-long-videos)

Learn how to implement multi-shot attention sink for long video generation with LongLive. Pin KV-cache regions to maintain attention across long sequences.

- Tags: how-to-guide
- Published: 2026-05-24

### [How NVFP4 Quantization Works in LongLive 2.0 for Transformer Inference](/NVlabs/LongLive/how-does-nvfp4-quantization-work-in-longlive-2-0-for-transformer-inference)

Explore NVFP4 quantization in LongLive 2.0 for transformer inference. Discover how this 4-bit floating-point technique offers 8x memory savings with minimal quality loss.

- Tags: deep-dive
- Published: 2026-05-24

