CPU and Disk Offloading for Large Model Inference in LTX-2
LTX-2 implements transparent CPU and disk offloading through the OffloadMode enum, enabling inference of diffusion models that exceed GPU memory by automatically streaming weights from system RAM or memory-mapped disk storage.
The LTX-2 video generation framework from Lightricks is engineered for production-scale diffusion models whose parameters often surpass the capacity of consumer and even professional GPUs. According to the LTX-2 source code, the framework solves this through a unified offloading architecture that treats memory hierarchy—GPU VRAM, system RAM, and persistent storage—as a single addressable space for model weights.
How Offloading Works in LTX-2
LTX-2 provides three operating modes defined in ltx_pipelines/utils/types.py:
| Mode | Behavior | Use Case |
|---|---|---|
NONE |
All weights remain in GPU memory | Models that fit entirely in VRAM |
CPU |
Weights stored in system RAM, copied to GPU on demand | Large models that fit in RAM but exceed VRAM |
DISK |
Weights memory-mapped from disk, streamed as needed | Extremely large models exceeding total RAM |
The OffloadMode enum is the central abstraction. Pipeline builders receive this value and propagate it to the streaming transformer, which configures weight-loading callbacks accordingly.
Core Implementation Files
OffloadMode Definition
The enumeration resides in [ltx_pipelines/utils/types.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py):
from enum import Enum
class OffloadMode(Enum):
NONE = "none"
CPU = "cpu"
DISK = "disk"
Command-Line Interface
Users select offloading via the --offload flag, implemented in [ltx_pipelines/utils/args.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py):
python -m ltx_pipelines.ti2vid_two_stage --offload cpu --model-dir /path/to/model
Streaming Transformer Integration
The core offloading logic lives in [ltx_core/block_streaming/builder.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/block_streaming/builder.py). Here, the selected OffloadMode configures callbacks that:
- Fetch weights from their offloaded location when a layer executes
- Release GPU copies after computation, allowing eviction
Pipeline scripts like [ti2vid_two_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stage.py) demonstrate propagation:
# Simplified excerpt showing offload_mode forwarding
streaming_transformer = build_streaming_transformer(
config=model_config,
offload_mode=args.offload_mode, # OffloadMode.CPU, etc.
)
Running Inference with CPU Offloading
CPU offloading keeps weights in system RAM—accessible faster than disk but outside GPU memory pressure.
python -m ltx_pipelines.ti2vid_two_stage \
--offload cpu \
--model-dir /models/ltx-video-large \
--output-dir /results/t2v_output \
--prompt "A drone shot over a coastal city at sunset"
What happens internally:
- Weights load initially into pinned system memory
- Each transformer block copies its weights to GPU during forward pass
- Post-execution, GPU memory is freed immediately
This incurs PCIe transfer overhead but prevents out-of-memory errors for models 2-4× larger than available VRAM.
Running Inference with Disk Offloading
Disk offloading targets models exceeding total system RAM, using memory-mapped files for transparent paging.
python -m ltx_pipelines.retake \
--offload disk \
--model-dir /models/ltx-video-xl \
--output-dir /results/retake \
--input-video /input/source.mp4
Implementation details from ltx_core/block_streaming/builder.py:
- Weights serialize to a cache directory (typically
./.cache/ltx_offload/) - The transformer
mmaps files on demand - Kernel page cache optimizations reduce redundant disk reads
Trade-off: Higher latency per layer due to storage I/O, but unbounded model size support.
Optimizer State Offloading During Training
For training scenarios, LTX-2 adds CPU offloading specifically for optimizer states—distinct from weight offloading. Controlled via [ltx_trainer/config.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py):
# training_config.yaml
acceleration:
offload_optimizer_during_validation: true
The [ltx_trainer/trainer.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) implements this through the _offloaded_optimizer_state context manager:
from ltx_trainer import Trainer, Config
cfg = Config.from_yaml("training_config.yaml")
trainer = Trainer(cfg)
trainer.train() # Automatically offloads optimizer to CPU during validation
This temporarily moves Adam moments and other optimizer tensors to CPU RAM, freeing VRAM for full-batch validation without terminating training state.
Inspecting Offload Configuration at Runtime
Debug or confirm your setup programmatically:
from ltx_pipelines.utils.types import OffloadMode
# Access via pipeline builder
print(f"Active offload mode: {pipeline_builder.offload_mode}")
# → Active offload mode: OffloadMode.CPU
# Check if optimizer offloading is enabled
print(f"Optimizer offload: {trainer.config.acceleration.offload_optimizer_during_validation}")
# → Optimizer offload: True
Performance Considerations
| Factor | CPU Offload | Disk Offload |
|---|---|---|
| Typical bandwidth | 16-64 GB/s (PCIe 4.0/5.0) | 0.5-7 GB/s (SSD) |
| Latency per layer | ~1-5 ms | ~10-100 ms |
| Maximum model size | System RAM limit | Storage capacity |
| Best for | Interactive generation, moderate scaling | Batch processing, extreme scale |
Recommendation from the LTX-2 architecture: Start with CPU mode; downgrade to DISK only when RAM is exhausted. The streaming transformer minimizes transfers through activation checkpointing and layerwise execution.
Summary
- LTX-2 offloading centers on the
OffloadModeenum (NONE,CPU,DISK) defined inltx_pipelines/utils/types.py - CPU mode streams weights from system RAM via
ltx_core/block_streaming/builder.pycallbacks - Disk mode memory-maps weights from storage for arbitrarily large models
- Optimizer offloading separately handles training states through
ltx_trainer/trainer.py - Configuration accepts both CLI flags (
--offload) and YAML config files - Pipeline scripts like
ti2vid_two_stage.pyandretake.pydemonstrate end-to-end usage
Frequently Asked Questions
What hardware requirements exist for disk offloading in LTX-2?
Disk offloading requires only sufficient storage space for the model weights—typically 10-50 GB for LTX-2 variants—and an SSD strongly recommended to prevent I/O bottlenecks. No specialized hardware is needed; the memory-mapping operates through standard Linux/Windows file systems.
Can I combine CPU and disk offloading?
The LTX-2 implementation treats these as mutually exclusive modes in OffloadMode. However, the underlying streaming transformer architecture could theoretically layer them—weights in CPU with overflow to disk—but this is not exposed in the current pipeline API. Use DISK mode as the fallback when CPU RAM is insufficient.
How does optimizer offloading differ from model weight offloading?
Model weight offloading (CPU/DISK modes) applies during inference and training forward passes, moving network parameters through the memory hierarchy. Optimizer offloading (controlled by offload_optimizer_during_validation) specifically targets first and second moment tensors used by optimizers like AdamW, activating only during validation phases to maximize transient VRAM availability.
Where does LTX-2 cache disk-offloaded weights?
The framework uses a default cache location—typically .cache/ltx_offload/ relative to the working directory—managed through the streaming transformer builder in ltx_core/block_streaming/builder.py. This directory persists across runs, enabling faster warm-start for repeated inference on the same model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →