How to Configure Model Offloading (CPU vs Disk) for Memory-Constrained Systems in LTX-2

LTX-2 provides three offloading strategies—none, cpu, and disk—controlled via the OffloadMode enum and --offload CLI flag, allowing massive diffusion models to run on GPUs with limited VRAM by streaming weights from system RAM or NVMe storage.

LTX-2 is Lightricks' open-source video diffusion framework, and model offloading (CPU vs disk) for memory-constrained systems is a core capability that makes large transformer-based models accessible on consumer hardware. The offloading system is built around a clean abstraction that lets you trade latency for memory capacity without modifying model code.

Understanding the OffloadMode Enum

The foundation of LTX-2's offloading system is the OffloadMode enumeration defined in packages/ltx-pipelines/src/ltx_pipelines/utils/types.py at line 129. This enum provides three distinct memory tiers:

class OffloadMode(Enum):
    NONE = auto()   # All weights remain on GPU (default)

    CPU = auto()    # Weights stored in system RAM, streamed to GPU per layer

    DISK = auto()   # Weights memory-mapped from NVMe/SSD, minimal RAM footprint

Each mode represents a different point on the memory-performance spectrum, and the consistent naming across the codebase makes switching strategies trivial.

Command-Line Configuration

All LTX-2 pipeline entry points expose the --offload flag through the argument parser in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py at line 643. This unified interface ensures consistent behavior whether you're running text-to-image-to-video, image-to-video, or video-to-video pipelines.

CPU Offloading

Use CPU offloading when GPU VRAM is insufficient but your host machine has ample system RAM (typically 2× the model size):

python -m ltx_pipelines.ti2vid_one_stage \
    --model_path /path/to/large_model \
    --offload cpu \
    --output_dir ./outputs

Under this mode, weight tensors are copied once to CPU memory. During inference, each transformer layer's weights are lazily moved to GPU, executed, and immediately released. This pattern eliminates VRAM pressure at the cost of PCIe bandwidth and higher per-step latency.

Disk Offloading

Use disk offloading when both GPU VRAM and system RAM are constrained, but you have fast NVMe storage available:

python -m ltx_pipelines.ti2vid_two_stages \
    --model_path /path/to/huge_model \
    --offload disk \
    --output_dir ./outputs

Disk offloading leverages memory-mapped weight files. The StreamingModelBuilder (detailed below) creates lazy weight objects that read only the necessary tensor slices from disk on demand. Because inactive layers never reside in RAM, you can run models that would otherwise exceed even host memory capacity—ideal for 24GB+ parameter models on laptops or edge devices.

Python API Configuration

For programmatic use, pass OffloadMode directly to pipeline constructors:

from ltx_pipelines.utils.types import OffloadMode
from ltx_pipelines.ti2vid_one_stage import TI2VIDOneStage

# CPU offloading for moderate memory constraints

pipeline_cpu = TI2VIDOneStage(
    model_path="models/large",
    offload_mode=OffloadMode.CPU,
)

# Disk offloading for severe memory constraints

pipeline_disk = TI2VIDOneStage(
    model_path="models/huge",
    offload_mode=OffloadMode.DISK,
)

pipeline_cpu.run()

This programmatic approach integrates cleanly with configuration management systems and hyperparameter sweeps.

How Streaming Blocks Implement Offloading

The actual weight streaming logic lives in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py. At line 308, the pipeline builder inspects offload_mode and conditionally swaps standard transformer blocks for streaming variants:

  • When offload_mode == OffloadMode.NONE: Standard in-memory blocks are used
  • When offload_mode != OffloadMode.NONE: StreamingModelBuilder constructs blocks with _cpu or _disk weight backends

Lines 391–394 define the slot bookkeeping system that controls how many weight chunks remain resident simultaneously. For disk offloading specifically, the constant DISK_CPU_SLOTS governs this prefetch window, balancing I/O parallelism against memory consumption.

This architecture means offloading is transparent to the model forward pass—the same diffusion code executes regardless of where weights reside.

Training: Optimizer State Offloading

LTX-2 extends the offloading concept to training workflows through optimizer state offloading. Unlike inference weight offloading, this feature moves the optimizer's moment buffers (momentum, variance estimates in Adam/AdamW) off GPU specifically during validation phases.

Control this behavior in ltx_trainer/config.py at line 400:

from ltx_trainer.config import TrainerConfig

cfg = TrainerConfig(
    acceleration={
        "offload_optimizer_during_validation": True,
    },
    # ... other training parameters

)

The implementation in packages/ltx-trainer/src/ltx_trainer/trainer.py (lines 825–839) provides a context manager that:

  1. Detects when entering validation
  2. Moves optimizer state tensors to CPU via optimizer.state_dict() manipulation
  3. Runs the validation loop with reduced GPU memory pressure
  4. Restores optimizer state to GPU before resuming training

This is particularly valuable for training large models with AdamW, where optimizer states can consume 2× the model's parameter memory.

Choosing the Right Offloading Strategy

Scenario Recommended Mode Expected Trade-off
GPU VRAM sufficient for full model none Maximum speed, no overhead
GPU insufficient, RAM ≥ 2× model size cpu Moderate latency increase, full throughput
Both GPU and RAM insufficient, fast NVMe disk Higher latency, I/O-bound, minimal RAM
Training with large optimizer states offload_optimizer_during_validation=True Slower validation, stable training batch sizes

Performance Considerations

  • CPU offloading: Bandwidth-bound by PCIe speed. Modern PCIe 4.0 x16 links can sustain ~32 GB/s, making this practical for 7B–13B parameter models at modest frame rates.
  • Disk offloading: Latency-bound by random read IOPS. NVMe SSDs with 500K+ IOPS and high queue depths are essential; SATA SSDs or HDDs will create severe bottlenecks.
  • Slot tuning: Advanced users can modify cpu_slots_count or DISK_CPU_SLOTS in blocks.py to increase parallelism at the cost of memory (default configurations target 16GB–24GB consumer GPUs).

Summary

  • LTX-2 model offloading is controlled by the OffloadMode enum with values NONE, CPU, and DISK
  • Use the --offload CLI flag or offload_mode parameter in Python to select strategies
  • CPU offloading streams weights from system RAM; disk offloading uses memory-mapped files for minimal RAM footprint
  • The StreamingModelBuilder in blocks.py transparently handles weight placement without model code changes
  • Training workflows support additional optimizer state offloading during validation via offload_optimizer_during_validation

Frequently Asked Questions

What is the minimum RAM requirement for CPU offloading in LTX-2?

You need system RAM approximately equal to the model's checkpoint size. For a 24GB checkpoint, plan for 24GB RAM plus overhead for activations and operating system. Disk offloading relaxes this to roughly 2–4GB regardless of model size, bounded only by the DISK_CPU_SLOTS prefetch window.

Does disk offloading work with network-attached storage?

Technically yes, but performance will likely be unacceptable. Disk offloading in LTX-2 uses memory-mapped file I/O with random access patterns. Network storage introduces latency that multiplies across thousands of layer accesses per diffusion step. Local NVMe is strongly recommended.

Can I use optimizer offloading without weight offloading?

Yes. These features operate independently. You can train with full GPU-resident weights (offload_mode=OffloadMode.NONE) while still enabling offload_optimizer_during_validation to fit larger validation batch sizes. Conversely, you can offload weights for inference without any optimizer involved.

Where does LTX-2 store temporary files for disk offloading?

The framework memory-maps directly from the model checkpoint path specified via --model_path. No additional temporary copies are created, so ensure your checkpoint location has both read bandwidth and sufficient capacity.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →