How to Configure Model Offloading (CPU/Disk) to Minimize VRAM Usage in LTX-2
LTX-2 supports three offload modes—NONE, CPU, and DISK—that stream model weights from GPU, CPU RAM, or local disk to reduce VRAM consumption as low as ~5 GB.
The LTX-2 video generation framework from Lightricks provides flexible model offloading strategies to run large transformer models on hardware with limited GPU memory. By configuring where weights are stored and how they flow to the GPU, you can trade inference speed for dramatically reduced VRAM requirements.
Understanding the OffloadMode Enumeration
The core abstraction for offloading lives in packages/ltx-pipelines/src/ltx_pipelines/utils/types.py. The OffloadMode enum defines three distinct strategies:
NONE— All model weights remain on the GPU. This yields the fastest generation but consumes maximum VRAM.CPU— Weights are pinned in CPU RAM and streamed layer-by-layer to the GPU during forward passes. Requires approximately 36 GB system RAM and reduces VRAM to ~5 GB.DISK— Weights are read from local disk on-demand through a small CPU cache (~2 GB RAM). This is the most memory-efficient option when both GPU and CPU memory are constrained.
Source: OffloadMode definition (lines 129–144)
Selecting an Offload Strategy via CLI or Configuration
Command-Line Interface
The pipelines expose a --offload argument defined in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py. Valid values are none, cpu, or disk.
python -m ltx_pipelines.ti2vid_two_stages \
--config my_config.yaml \
--offload cpu
The flag sets offload_mode, which propagates through stage builders to the streaming transformer constructor in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py (lines 370–404). When offloading is enabled, the builder configures cpu_slots_count for disk caching and validates that incompatible components—such as text_encoder_builder—are not used.
Source: --offload argument definition (lines 643–658)
YAML Configuration
You can also specify offloading directly in your pipeline configuration:
pipeline:
offload_mode: cpu # or: disk, none
The CLI flag takes precedence when both are present.
Optimizer State Offloading During Validation
Even when model weights stay on GPU, optimizer state can dominate VRAM during validation phases. LTX-2 provides a separate toggle for this scenario.
In packages/ltx-trainer/src/ltx_trainer/config.py (line 400), the boolean field acceleration.offload_optimizer_during_validation controls this behavior:
acceleration:
offload_optimizer_during_validation: true # default: false
When enabled, the trainer's _offloaded_optimizer_state context manager—implemented in packages/ltx-trainer/src/ltx_trainer/trainer.py (lines 825–856)—moves optimizer tensors to CPU during validation and restores them afterward. This frees GPU memory for the forward pass without interrupting training state.
Source: offload_optimizer_during_validation field
Source: _offloaded_optimizer_state context manager (lines 825–856)
Complete Minimal-VRAM Configuration
For the most aggressive memory savings suitable for consumer GPUs:
# minimal_vram_config.yaml
acceleration:
offload_optimizer_during_validation: true
pipeline:
offload_mode: disk # stream from disk; use 'cpu' if disk I/O is too slow
Run with:
python -m ltx_pipelines.ti2vid_two_stages \
--config minimal_vram_config.yaml \
--offload disk
Performance expectations:
DISKmode: Slowest inference, minimal RAM requirements (~2 GB CPU cache)CPUmode: Moderate speed, requires ~36 GB system RAMNONEmode: Fastest, requires full model in VRAM
Summary
OffloadModeintypes.pycontrols where weights live: GPU, CPU RAM, or disk--offload cpu|diskCLI flag orpipeline.offload_modein YAML selects the strategycpu_slots_countconfigures the disk cache size when usingDISKmodeoffload_optimizer_during_validationreduces VRAM pressure specifically during validation loops- Builders in
blocks.pyvalidate incompatible configurations when offloading is active
Frequently Asked Questions
What is the minimum VRAM required to run LTX-2 with offloading?
With CPU or DISK offloading enabled, VRAM requirements drop to approximately 5 GB. The DISK mode further reduces CPU RAM needs to roughly 2 GB through its on-disk weight streaming with small CPU cache. These figures assume no additional memory-heavy components like certain text encoders are active.
Does model offloading affect video generation quality?
No. Offloading only changes where weights are stored and how they reach the GPU—it does not alter model architecture, precision, or inference computations. Output quality remains identical across NONE, CPU, and DISK modes. The trade-off is purely between memory efficiency and inference speed.
Can I use model offloading during training, or only inference?
The OffloadMode system primarily targets inference pipelines in ltx-pipelines. For training, the separate offload_optimizer_during_validation toggle in ltx-trainer addresses optimizer state, but full weight offloading during training would require additional implementation not present in the current source.
Why does the builder reject certain configurations when offloading is enabled?
The stage builders in blocks.py validate that components incompatible with streaming—such as text_encoder_builder—are not used when offload_mode != OffloadMode.NONE. This prevents runtime errors from components that expect full GPU-resident weights. Check the validation logic at lines 370–404 for specific restrictions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →