How to Configure Model Offloading (CPU vs Disk) for Memory-Constrained Systems in LTX-2
LTX-2 provides three offloading strategies—none, cpu, and disk—controlled via the OffloadMode enum and --offload CLI flag, allowing massive diffusion models to run on GPUs with limited VRAM by streaming weights from system RAM or NVMe storage.
LTX-2 is Lightricks' open-source video diffusion framework, and model offloading (CPU vs disk) for memory-constrained systems is a core capability that makes large transformer-based models accessible on consumer hardware. The offloading system is built around a clean abstraction that lets you trade latency for memory capacity without modifying model code.
Understanding the OffloadMode Enum
The foundation of LTX-2's offloading system is the OffloadMode enumeration defined in packages/ltx-pipelines/src/ltx_pipelines/utils/types.py at line 129. This enum provides three distinct memory tiers:
class OffloadMode(Enum):
NONE = auto() # All weights remain on GPU (default)
CPU = auto() # Weights stored in system RAM, streamed to GPU per layer
DISK = auto() # Weights memory-mapped from NVMe/SSD, minimal RAM footprint
Each mode represents a different point on the memory-performance spectrum, and the consistent naming across the codebase makes switching strategies trivial.
Command-Line Configuration
All LTX-2 pipeline entry points expose the --offload flag through the argument parser in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py at line 643. This unified interface ensures consistent behavior whether you're running text-to-image-to-video, image-to-video, or video-to-video pipelines.
CPU Offloading
Use CPU offloading when GPU VRAM is insufficient but your host machine has ample system RAM (typically 2× the model size):
python -m ltx_pipelines.ti2vid_one_stage \
--model_path /path/to/large_model \
--offload cpu \
--output_dir ./outputs
Under this mode, weight tensors are copied once to CPU memory. During inference, each transformer layer's weights are lazily moved to GPU, executed, and immediately released. This pattern eliminates VRAM pressure at the cost of PCIe bandwidth and higher per-step latency.
Disk Offloading
Use disk offloading when both GPU VRAM and system RAM are constrained, but you have fast NVMe storage available:
python -m ltx_pipelines.ti2vid_two_stages \
--model_path /path/to/huge_model \
--offload disk \
--output_dir ./outputs
Disk offloading leverages memory-mapped weight files. The StreamingModelBuilder (detailed below) creates lazy weight objects that read only the necessary tensor slices from disk on demand. Because inactive layers never reside in RAM, you can run models that would otherwise exceed even host memory capacity—ideal for 24GB+ parameter models on laptops or edge devices.
Python API Configuration
For programmatic use, pass OffloadMode directly to pipeline constructors:
from ltx_pipelines.utils.types import OffloadMode
from ltx_pipelines.ti2vid_one_stage import TI2VIDOneStage
# CPU offloading for moderate memory constraints
pipeline_cpu = TI2VIDOneStage(
model_path="models/large",
offload_mode=OffloadMode.CPU,
)
# Disk offloading for severe memory constraints
pipeline_disk = TI2VIDOneStage(
model_path="models/huge",
offload_mode=OffloadMode.DISK,
)
pipeline_cpu.run()
This programmatic approach integrates cleanly with configuration management systems and hyperparameter sweeps.
How Streaming Blocks Implement Offloading
The actual weight streaming logic lives in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py. At line 308, the pipeline builder inspects offload_mode and conditionally swaps standard transformer blocks for streaming variants:
- When
offload_mode == OffloadMode.NONE: Standard in-memory blocks are used - When
offload_mode != OffloadMode.NONE:StreamingModelBuilderconstructs blocks with_cpuor_diskweight backends
Lines 391–394 define the slot bookkeeping system that controls how many weight chunks remain resident simultaneously. For disk offloading specifically, the constant DISK_CPU_SLOTS governs this prefetch window, balancing I/O parallelism against memory consumption.
This architecture means offloading is transparent to the model forward pass—the same diffusion code executes regardless of where weights reside.
Training: Optimizer State Offloading
LTX-2 extends the offloading concept to training workflows through optimizer state offloading. Unlike inference weight offloading, this feature moves the optimizer's moment buffers (momentum, variance estimates in Adam/AdamW) off GPU specifically during validation phases.
Control this behavior in ltx_trainer/config.py at line 400:
from ltx_trainer.config import TrainerConfig
cfg = TrainerConfig(
acceleration={
"offload_optimizer_during_validation": True,
},
# ... other training parameters
)
The implementation in packages/ltx-trainer/src/ltx_trainer/trainer.py (lines 825–839) provides a context manager that:
- Detects when entering validation
- Moves optimizer state tensors to CPU via
optimizer.state_dict()manipulation - Runs the validation loop with reduced GPU memory pressure
- Restores optimizer state to GPU before resuming training
This is particularly valuable for training large models with AdamW, where optimizer states can consume 2× the model's parameter memory.
Choosing the Right Offloading Strategy
| Scenario | Recommended Mode | Expected Trade-off |
|---|---|---|
| GPU VRAM sufficient for full model | none |
Maximum speed, no overhead |
| GPU insufficient, RAM ≥ 2× model size | cpu |
Moderate latency increase, full throughput |
| Both GPU and RAM insufficient, fast NVMe | disk |
Higher latency, I/O-bound, minimal RAM |
| Training with large optimizer states | offload_optimizer_during_validation=True |
Slower validation, stable training batch sizes |
Performance Considerations
- CPU offloading: Bandwidth-bound by PCIe speed. Modern PCIe 4.0 x16 links can sustain ~32 GB/s, making this practical for 7B–13B parameter models at modest frame rates.
- Disk offloading: Latency-bound by random read IOPS. NVMe SSDs with 500K+ IOPS and high queue depths are essential; SATA SSDs or HDDs will create severe bottlenecks.
- Slot tuning: Advanced users can modify
cpu_slots_countorDISK_CPU_SLOTSinblocks.pyto increase parallelism at the cost of memory (default configurations target 16GB–24GB consumer GPUs).
Summary
- LTX-2 model offloading is controlled by the
OffloadModeenum with valuesNONE,CPU, andDISK - Use the
--offloadCLI flag oroffload_modeparameter in Python to select strategies - CPU offloading streams weights from system RAM; disk offloading uses memory-mapped files for minimal RAM footprint
- The
StreamingModelBuilderinblocks.pytransparently handles weight placement without model code changes - Training workflows support additional optimizer state offloading during validation via
offload_optimizer_during_validation
Frequently Asked Questions
What is the minimum RAM requirement for CPU offloading in LTX-2?
You need system RAM approximately equal to the model's checkpoint size. For a 24GB checkpoint, plan for 24GB RAM plus overhead for activations and operating system. Disk offloading relaxes this to roughly 2–4GB regardless of model size, bounded only by the DISK_CPU_SLOTS prefetch window.
Does disk offloading work with network-attached storage?
Technically yes, but performance will likely be unacceptable. Disk offloading in LTX-2 uses memory-mapped file I/O with random access patterns. Network storage introduces latency that multiplies across thousands of layer accesses per diffusion step. Local NVMe is strongly recommended.
Can I use optimizer offloading without weight offloading?
Yes. These features operate independently. You can train with full GPU-resident weights (offload_mode=OffloadMode.NONE) while still enabling offload_optimizer_during_validation to fit larger validation batch sizes. Conversely, you can offload weights for inference without any optimizer involved.
Where does LTX-2 store temporary files for disk offloading?
The framework memory-maps directly from the model checkpoint path specified via --model_path. No additional temporary copies are created, so ensure your checkpoint location has both read bandwidth and sufficient capacity.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →