# CPU and Disk Offloading for Large Model Inference in LTX-2

> Discover how LTX-2 enables large model inference by transparently offloading weights to CPU and disk. Stream weights from RAM or memory-mapped disk to overcome GPU memory limits.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: performance
- Published: 2026-08-15

---

**LTX-2 implements transparent CPU and disk offloading through the `OffloadMode` enum, enabling inference of diffusion models that exceed GPU memory by automatically streaming weights from system RAM or memory-mapped disk storage.**

The LTX-2 video generation framework from Lightricks is engineered for production-scale diffusion models whose parameters often surpass the capacity of consumer and even professional GPUs. According to the LTX-2 source code, the framework solves this through a unified offloading architecture that treats memory hierarchy—GPU VRAM, system RAM, and persistent storage—as a single addressable space for model weights.

---

## How Offloading Works in LTX-2

LTX-2 provides **three operating modes** defined in [`ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/types.py):

| Mode | Behavior | Use Case |
|------|----------|----------|
| `NONE` | All weights remain in GPU memory | Models that fit entirely in VRAM |
| `CPU` | Weights stored in system RAM, copied to GPU on demand | Large models that fit in RAM but exceed VRAM |
| `DISK` | Weights memory-mapped from disk, streamed as needed | Extremely large models exceeding total RAM |

The `OffloadMode` enum is the central abstraction. Pipeline builders receive this value and propagate it to the streaming transformer, which configures weight-loading callbacks accordingly.

---

## Core Implementation Files

### OffloadMode Definition

The enumeration resides in [[`ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/types.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py):

```python
from enum import Enum

class OffloadMode(Enum):
    NONE = "none"
    CPU = "cpu"
    DISK = "disk"

```

### Command-Line Interface

Users select offloading via the `--offload` flag, implemented in [[`ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/args.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py):

```bash
python -m ltx_pipelines.ti2vid_two_stage --offload cpu --model-dir /path/to/model

```

### Streaming Transformer Integration

The core offloading logic lives in [[`ltx_core/block_streaming/builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/block_streaming/builder.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/block_streaming/builder.py). Here, the selected `OffloadMode` configures callbacks that:

1. **Fetch** weights from their offloaded location when a layer executes
2. **Release** GPU copies after computation, allowing eviction

Pipeline scripts like [[`ti2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stage.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stage.py) demonstrate propagation:

```python

# Simplified excerpt showing offload_mode forwarding

streaming_transformer = build_streaming_transformer(
    config=model_config,
    offload_mode=args.offload_mode,  # OffloadMode.CPU, etc.

)

```

---

## Running Inference with CPU Offloading

CPU offloading keeps weights in system RAM—accessible faster than disk but outside GPU memory pressure.

```bash
python -m ltx_pipelines.ti2vid_two_stage \
    --offload cpu \
    --model-dir /models/ltx-video-large \
    --output-dir /results/t2v_output \
    --prompt "A drone shot over a coastal city at sunset"

```

**What happens internally:**

- Weights load initially into pinned system memory
- Each transformer block copies its weights to GPU during forward pass
- Post-execution, GPU memory is freed immediately

This incurs **PCIe transfer overhead** but prevents out-of-memory errors for models 2-4× larger than available VRAM.

---

## Running Inference with Disk Offloading

Disk offloading targets models exceeding total system RAM, using memory-mapped files for transparent paging.

```bash
python -m ltx_pipelines.retake \
    --offload disk \
    --model-dir /models/ltx-video-xl \
    --output-dir /results/retake \
    --input-video /input/source.mp4

```

**Implementation details from [`ltx_core/block_streaming/builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/block_streaming/builder.py):**

- Weights serialize to a cache directory (typically `./.cache/ltx_offload/`)
- The transformer `mmap`s files on demand
- Kernel page cache optimizations reduce redundant disk reads

**Trade-off:** Higher latency per layer due to storage I/O, but unbounded model size support.

---

## Optimizer State Offloading During Training

For training scenarios, LTX-2 adds CPU offloading specifically for optimizer states—distinct from weight offloading. Controlled via [[`ltx_trainer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/config.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py):

```yaml

# training_config.yaml

acceleration:
  offload_optimizer_during_validation: true

```

The [[`ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/trainer.py)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) implements this through the `_offloaded_optimizer_state` context manager:

```python
from ltx_trainer import Trainer, Config

cfg = Config.from_yaml("training_config.yaml")
trainer = Trainer(cfg)

trainer.train()  # Automatically offloads optimizer to CPU during validation

```

This temporarily moves **Adam moments and other optimizer tensors** to CPU RAM, freeing VRAM for full-batch validation without terminating training state.

---

## Inspecting Offload Configuration at Runtime

Debug or confirm your setup programmatically:

```python
from ltx_pipelines.utils.types import OffloadMode

# Access via pipeline builder

print(f"Active offload mode: {pipeline_builder.offload_mode}")

# → Active offload mode: OffloadMode.CPU

# Check if optimizer offloading is enabled

print(f"Optimizer offload: {trainer.config.acceleration.offload_optimizer_during_validation}")

# → Optimizer offload: True

```

---

## Performance Considerations

| Factor | CPU Offload | Disk Offload |
|--------|-------------|--------------|
| **Typical bandwidth** | 16-64 GB/s (PCIe 4.0/5.0) | 0.5-7 GB/s (SSD) |
| **Latency per layer** | ~1-5 ms | ~10-100 ms |
| **Maximum model size** | System RAM limit | Storage capacity |
| **Best for** | Interactive generation, moderate scaling | Batch processing, extreme scale |

**Recommendation from the LTX-2 architecture:** Start with `CPU` mode; downgrade to `DISK` only when RAM is exhausted. The streaming transformer minimizes transfers through activation checkpointing and layerwise execution.

---

## Summary

- **LTX-2 offloading** centers on the `OffloadMode` enum (`NONE`, `CPU`, `DISK`) defined in [`ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/types.py)
- **CPU mode** streams weights from system RAM via [`ltx_core/block_streaming/builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/block_streaming/builder.py) callbacks
- **Disk mode** memory-maps weights from storage for arbitrarily large models
- **Optimizer offloading** separately handles training states through [`ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/trainer.py)
- **Configuration** accepts both CLI flags (`--offload`) and YAML config files
- **Pipeline scripts** like [`ti2vid_two_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/ti2vid_two_stage.py) and [`retake.py`](https://github.com/Lightricks/LTX-2/blob/main/retake.py) demonstrate end-to-end usage

---

## Frequently Asked Questions

### What hardware requirements exist for disk offloading in LTX-2?

Disk offloading requires only sufficient storage space for the model weights—typically 10-50 GB for LTX-2 variants—and an SSD strongly recommended to prevent I/O bottlenecks. No specialized hardware is needed; the memory-mapping operates through standard Linux/Windows file systems.

### Can I combine CPU and disk offloading?

The LTX-2 implementation treats these as mutually exclusive modes in `OffloadMode`. However, the underlying streaming transformer architecture could theoretically layer them—weights in CPU with overflow to disk—but this is not exposed in the current pipeline API. Use `DISK` mode as the fallback when CPU RAM is insufficient.

### How does optimizer offloading differ from model weight offloading?

**Model weight offloading** (`CPU`/`DISK` modes) applies during inference and training forward passes, moving network parameters through the memory hierarchy. **Optimizer offloading** (controlled by `offload_optimizer_during_validation`) specifically targets first and second moment tensors used by optimizers like AdamW, activating only during validation phases to maximize transient VRAM availability.

### Where does LTX-2 cache disk-offloaded weights?

The framework uses a default cache location—typically `.cache/ltx_offload/` relative to the working directory—managed through the streaming transformer builder in [`ltx_core/block_streaming/builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/block_streaming/builder.py). This directory persists across runs, enabling faster warm-start for repeated inference on the same model.