CPU and Disk Offloading for Large Model Inference in LTX-2

LTX-2 implements transparent CPU and disk offloading through the OffloadMode enum, enabling inference of diffusion models that exceed GPU memory by automatically streaming weights from system RAM or memory-mapped disk storage.

The LTX-2 video generation framework from Lightricks is engineered for production-scale diffusion models whose parameters often surpass the capacity of consumer and even professional GPUs. According to the LTX-2 source code, the framework solves this through a unified offloading architecture that treats memory hierarchy—GPU VRAM, system RAM, and persistent storage—as a single addressable space for model weights.


How Offloading Works in LTX-2

LTX-2 provides three operating modes defined in ltx_pipelines/utils/types.py:

Mode Behavior Use Case
NONE All weights remain in GPU memory Models that fit entirely in VRAM
CPU Weights stored in system RAM, copied to GPU on demand Large models that fit in RAM but exceed VRAM
DISK Weights memory-mapped from disk, streamed as needed Extremely large models exceeding total RAM

The OffloadMode enum is the central abstraction. Pipeline builders receive this value and propagate it to the streaming transformer, which configures weight-loading callbacks accordingly.


Core Implementation Files

OffloadMode Definition

The enumeration resides in [ltx_pipelines/utils/types.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py):

from enum import Enum

class OffloadMode(Enum):
    NONE = "none"
    CPU = "cpu"
    DISK = "disk"

Command-Line Interface

Users select offloading via the --offload flag, implemented in [ltx_pipelines/utils/args.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py):

python -m ltx_pipelines.ti2vid_two_stage --offload cpu --model-dir /path/to/model

Streaming Transformer Integration

The core offloading logic lives in [ltx_core/block_streaming/builder.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/block_streaming/builder.py). Here, the selected OffloadMode configures callbacks that:

  1. Fetch weights from their offloaded location when a layer executes
  2. Release GPU copies after computation, allowing eviction

Pipeline scripts like [ti2vid_two_stage.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stage.py) demonstrate propagation:


# Simplified excerpt showing offload_mode forwarding

streaming_transformer = build_streaming_transformer(
    config=model_config,
    offload_mode=args.offload_mode,  # OffloadMode.CPU, etc.

)

Running Inference with CPU Offloading

CPU offloading keeps weights in system RAM—accessible faster than disk but outside GPU memory pressure.

python -m ltx_pipelines.ti2vid_two_stage \
    --offload cpu \
    --model-dir /models/ltx-video-large \
    --output-dir /results/t2v_output \
    --prompt "A drone shot over a coastal city at sunset"

What happens internally:

  • Weights load initially into pinned system memory
  • Each transformer block copies its weights to GPU during forward pass
  • Post-execution, GPU memory is freed immediately

This incurs PCIe transfer overhead but prevents out-of-memory errors for models 2-4× larger than available VRAM.


Running Inference with Disk Offloading

Disk offloading targets models exceeding total system RAM, using memory-mapped files for transparent paging.

python -m ltx_pipelines.retake \
    --offload disk \
    --model-dir /models/ltx-video-xl \
    --output-dir /results/retake \
    --input-video /input/source.mp4

Implementation details from ltx_core/block_streaming/builder.py:

  • Weights serialize to a cache directory (typically ./.cache/ltx_offload/)
  • The transformer mmaps files on demand
  • Kernel page cache optimizations reduce redundant disk reads

Trade-off: Higher latency per layer due to storage I/O, but unbounded model size support.


Optimizer State Offloading During Training

For training scenarios, LTX-2 adds CPU offloading specifically for optimizer states—distinct from weight offloading. Controlled via [ltx_trainer/config.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py):


# training_config.yaml

acceleration:
  offload_optimizer_during_validation: true

The [ltx_trainer/trainer.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) implements this through the _offloaded_optimizer_state context manager:

from ltx_trainer import Trainer, Config

cfg = Config.from_yaml("training_config.yaml")
trainer = Trainer(cfg)

trainer.train()  # Automatically offloads optimizer to CPU during validation

This temporarily moves Adam moments and other optimizer tensors to CPU RAM, freeing VRAM for full-batch validation without terminating training state.


Inspecting Offload Configuration at Runtime

Debug or confirm your setup programmatically:

from ltx_pipelines.utils.types import OffloadMode

# Access via pipeline builder

print(f"Active offload mode: {pipeline_builder.offload_mode}")

# → Active offload mode: OffloadMode.CPU

# Check if optimizer offloading is enabled

print(f"Optimizer offload: {trainer.config.acceleration.offload_optimizer_during_validation}")

# → Optimizer offload: True

Performance Considerations

Factor CPU Offload Disk Offload
Typical bandwidth 16-64 GB/s (PCIe 4.0/5.0) 0.5-7 GB/s (SSD)
Latency per layer ~1-5 ms ~10-100 ms
Maximum model size System RAM limit Storage capacity
Best for Interactive generation, moderate scaling Batch processing, extreme scale

Recommendation from the LTX-2 architecture: Start with CPU mode; downgrade to DISK only when RAM is exhausted. The streaming transformer minimizes transfers through activation checkpointing and layerwise execution.


Summary


Frequently Asked Questions

What hardware requirements exist for disk offloading in LTX-2?

Disk offloading requires only sufficient storage space for the model weights—typically 10-50 GB for LTX-2 variants—and an SSD strongly recommended to prevent I/O bottlenecks. No specialized hardware is needed; the memory-mapping operates through standard Linux/Windows file systems.

Can I combine CPU and disk offloading?

The LTX-2 implementation treats these as mutually exclusive modes in OffloadMode. However, the underlying streaming transformer architecture could theoretically layer them—weights in CPU with overflow to disk—but this is not exposed in the current pipeline API. Use DISK mode as the fallback when CPU RAM is insufficient.

How does optimizer offloading differ from model weight offloading?

Model weight offloading (CPU/DISK modes) applies during inference and training forward passes, moving network parameters through the memory hierarchy. Optimizer offloading (controlled by offload_optimizer_during_validation) specifically targets first and second moment tensors used by optimizers like AdamW, activating only during validation phases to maximize transient VRAM availability.

Where does LTX-2 cache disk-offloaded weights?

The framework uses a default cache location—typically .cache/ltx_offload/ relative to the working directory—managed through the streaming transformer builder in ltx_core/block_streaming/builder.py. This directory persists across runs, enabling faster warm-start for repeated inference on the same model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →