How to Set Up Model Offloading to CPU or Disk for VRAM Constraints in LTX-2
LTX-2 provides three built-in offload modes—cpu, disk, and none—controlled via a single --offload flag that routes model weights through CPU RAM or disk storage instead of loading everything into GPU VRAM.
The LTX-2 video generation model from Lightricks requires approximately 28 GB of VRAM for full GPU inference, which exceeds the capacity of most consumer hardware. To address this, the codebase includes a sophisticated weight streaming system that offloads model parameters to CPU memory or even local disk, trading inference speed for dramatically reduced GPU memory requirements. This guide walks through the OffloadMode enum, CLI configuration, and underlying streaming architecture.
Understanding the Three Offload Modes
The OffloadMode enum in packages/ltx-pipelines/src/ltx_pipelines/utils/types.py (lines 29-45) defines three strategies:
| Mode | Weight Location | Memory Profile | Best For |
|---|---|---|---|
| NONE | Fully resident on GPU | ~28 GB VRAM | Maximum inference speed on high-end hardware (A100, H100) |
| CPU | Pinned in system RAM, streamed to small GPU buffer | ~36 GB RAM + ~5 GB VRAM | Machines with abundant RAM but limited VRAM |
| DISK | Stored on SSD/HDD, streamed through minimal CPU buffer | ~5 GB RAM + ~5 GB VRAM | Minimal RAM environments where disk bandwidth is acceptable |
The CPU mode caches weights in RAM after the first inference pass, making subsequent passes faster. The DISK mode minimizes RAM usage entirely by reading layers directly from storage—useful on systems with fast NVMe drives but limited memory.
Activating Offloading via Command Line
The offloading strategy is activated through the --offload argument parsed in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 643-652). This flag accepts lowercase strings that map directly to the enum values.
CPU Offloading Example
python -m ltx_pipelines.ti2vid_one_stage \
--input video.mp4 \
--output result.mp4 \
--offload cpu
This configuration keeps ~36 GB of model weights in system RAM while using only ~5 GB of VRAM for active computation.
Disk Offloading Example
python -m ltx_pipelines.ti2vid_one_stage \
--input video.mp4 \
--output result.mp4 \
--offload disk
Ideal for 16 GB RAM laptops or cloud instances with large attached storage. Each forward pass reads required layers from disk, so NVMe SSDs are strongly recommended.
Disabling Offloading (Default)
python -m ltx_pipelines.ti2vid_one_stage \
--input video.mp4 \
--output result.mp4 \
--offload none
This yields the fastest inference but requires the full ~28 GB of GPU memory.
How the Streaming Builder Implements Offloading
Once the CLI flag is parsed, the selected mode propagates to packages/ltx-core/src/ltx_core/block_streaming/builder.py (lines 68-80). The StreamingModelBuilder class allocates CPU slots—memory buffers that hold weights staged for GPU transfer.
For CPU mode, the builder allocates standard slots based on layer count. For DISK mode, it uses the special DISK_CPU_SLOTS constant to minimize memory footprint. The builder coordinates with StreamingExecutor to prefetch upcoming layers while the GPU computes on current ones, maintaining pipeline efficiency.
Key implementation detail from the source: the builder distinguishes between "active" weights on GPU and "staged" weights in CPU/disk buffers, swapping them through non-blocking transfers when possible.
Additional VRAM Optimization: Validation-Time Optimizer Offloading
During training, LTX-2 can further reduce VRAM pressure during validation phases. The packages/ltx-trainer/src/ltx_trainer/trainer.py implementation (lines 825-856) supports offloading optimizer state to CPU when running validation loops.
Enable this in your training configuration:
# config.yaml
acceleration:
offload_optimizer_during_validation: true
This temporarily moves optimizer buffers (often 2× model size for Adam/AdamW) to CPU RAM, freeing substantial GPU memory for larger validation batch sizes or longer sequences.
Performance Considerations and Hardware Matching
Choose your offloading strategy based on your hardware profile:
- >28 GB VRAM (A100 40GB/80GB, H100): Use
--offload nonefor maximum throughput - >36 GB system RAM, <28 GB VRAM (RTX 4090 24GB, A10 24GB): Use
--offload cpuwith minimal speed penalty after warm-up - Limited RAM, fast NVMe (cloud instances, laptops): Use
--offload diskwith expected 2-5× slowdown depending on storage speed - Network storage or slow HDD: Expect severe degradation; consider CPU mode with swap space instead
The streaming system in ltx_core/block_streaming is optimized for sequential layer access patterns common in transformer inference, making partial offloading more efficient than naive implementations.
Summary
- LTX-2 model offloading requires only the
--offload cpu|disk|noneflag on any pipeline script - Three modes trade speed for memory:
NONE(28 GB VRAM),CPU(36 GB RAM + 5 GB VRAM),DISK(5 GB RAM + 5 GB VRAM + storage) - Core files:
types.pydefines modes,args.pyparses CLI flags,builder.pyexecutes streaming strategy - Training optimization:
offload_optimizer_during_validationmoves optimizer state to CPU during validation - Recommended pairing: CPU offloading for RAM-rich systems, disk offloading for storage-rich, RAM-constrained environments
Frequently Asked Questions
What is the minimum hardware to run LTX-2 with offloading?
The disk offloading mode requires approximately 5 GB of VRAM and 5 GB of RAM, making it feasible on consumer GPUs like the RTX 3060 12GB or RTX 4060 8GB (with gradient checkpointing). A fast NVMe SSD is strongly recommended to minimize the I/O bottleneck inherent in disk streaming.
How much slower is CPU offloading compared to full GPU inference?
First inference passes with --offload cpu incur roughly 20-40% overhead due to initial weight staging. Subsequent passes achieve near-native speed because weights remain cached in RAM. Disk offloading typically sustains 2-5× slowdown depending on storage bandwidth and layer access patterns.
Can I combine model offloading with other memory optimizations?
Yes. The LTX-2 trainer supports gradient checkpointing, mixed precision training, and optimizer state offloading as complementary techniques. Configure these through the acceleration section of your training config while using --offload cpu or --offload disk for the model weights themselves.
Does offloading work for both inference and training?
The --offload flag applies to inference pipelines (ti2vid_one_stage.py, train.py evaluation modes). For training, the streaming builder is reused but optimizer state requires separate handling via offload_optimizer_during_validation. Full training with disk-offloaded weights is supported but significantly slower due to frequent weight updates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →