LongLive
LongLive 2.0: Infra - Long Video Gen
Learn to integrate TriAttention KV cache compression with LongLive. Reduce GPU memory by 50% with this guide and boost your model's efficiency without sacrificing quality.
How to Configure Local Attention Size for Long Video Generation in LongLiveMaster local attention size for long video generation in LongLive. Learn to configure window_size for sliding-window attention control in NVlabs/LongLive.
Best Practices for NVFP4 Checkpoint Loading and Weight Materialization in LongLiveOptimize NVFP4 checkpoint loading and weight materialization in LongLive. Safely load checkpoints, unwrap generators, clean FSDP prefixes, and drop master weights to save GPU memory.
How to Use the WanDiffusionWrapper with Custom Model Configurations in LongLiveLearn to customize the WanDiffusionWrapper in LongLive using model_kwargs for checkpoint loading, causal mode, sampling, and VAEs. Tailor your diffusion pipeline without code changes.
How to Debug Memory Issues with the ErrorBuffer Utility in LongLiveDebug memory issues in LongLive using the ErrorBuffer utility. Monitor stats, verify shard size, and ensure buffer warmup for efficient memory management.
How to Implement Few-Step DMD Distillation for Faster Inference in LongLiveLearn how to implement few-step DMD distillation in LongLive for faster inference. Reduce num train timestep and use the score distillation trainer to compress the diffusion schedule.
Causal Diffusion Pipeline Architecture in LongLive 2.0: A Technical Deep DiveExplore the causal diffusion pipeline architecture in LongLive 2.0. Learn how it generates videos autoregressively using causal masking and KV-cache for efficient temporal attention. Deep dive into NVlabs/LongLive.
How to Configure the Timestep Shift Parameter for Flow Matching Schedulers in LongLiveConfigure the timestep shift parameter in LongLive Flow Matching Schedulers via YAML or WanDiffusionWrapper to control noise distribution. Learn how to optimize your model training.
How to Optimize Memory Usage During AR Training with Sequence Parallelism in LongLiveOptimize AR training memory with LongLive sequence parallelism. Shard temporal data across GPUs to slash activation memory from 40GB+ to 8-16GB per device.
Score Distillation vs Diffusion Training in LongLive: Technical Differences ExplainedExplore score distillation vs diffusion training in NVlabs LongLive. Understand key technical differences: KL-gradient matching versus direct MSE flow prediction for your generative models.
How to Implement KV Cache Recaching for Video Streaming in LongLiveImplement KV cache recaching for video streaming in LongLive by setting shot_clean_recache and multi_shot_sink to true. Expert guide to optimize inference with automatic scene cut detection.
How RoPE Position Embedding Offset Works for Multi-Shot Sequences in LongLiveDiscover how LongLive's RoPE position embedding offset revolutionizes multi-shot video generation by maintaining temporal context without KV cache resets.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →