# How to Configure Model Offloading (CPU/Disk) to Minimize VRAM Usage in LTX-2

> Minimize VRAM usage in LTX-2 with model offloading. Configure CPU or Disk modes to stream weights and reduce consumption to ~5 GB.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-18

---

**LTX-2 supports three offload modes—`NONE`, `CPU`, and `DISK`—that stream model weights from GPU, CPU RAM, or local disk to reduce VRAM consumption as low as ~5 GB.**

The LTX-2 video generation framework from Lightricks provides flexible **model offloading** strategies to run large transformer models on hardware with limited GPU memory. By configuring where weights are stored and how they flow to the GPU, you can trade inference speed for dramatically reduced VRAM requirements.

## Understanding the OffloadMode Enumeration

The core abstraction for offloading lives in [`packages/ltx-pipelines/src/ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py). The `OffloadMode` enum defines three distinct strategies:

- **`NONE`** — All model weights remain on the GPU. This yields the fastest generation but consumes maximum VRAM.
- **`CPU`** — Weights are pinned in **CPU RAM** and streamed layer-by-layer to the GPU during forward passes. Requires approximately 36 GB system RAM and reduces VRAM to ~5 GB.
- **`DISK`** — Weights are read from **local disk** on-demand through a small CPU cache (~2 GB RAM). This is the most memory-efficient option when both GPU and CPU memory are constrained.

Source: [`OffloadMode` definition (lines 129–144)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py#L129-L144)

## Selecting an Offload Strategy via CLI or Configuration

### Command-Line Interface

The pipelines expose a `--offload` argument defined in [`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py). Valid values are `none`, `cpu`, or `disk`.

```bash
python -m ltx_pipelines.ti2vid_two_stages \
    --config my_config.yaml \
    --offload cpu

```

The flag sets `offload_mode`, which propagates through stage builders to the streaming transformer constructor in [`packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) (lines 370–404). When offloading is enabled, the builder configures `cpu_slots_count` for disk caching and validates that incompatible components—such as `text_encoder_builder`—are not used.

Source: [`--offload` argument definition (lines 643–658)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py#L643-L658)

### YAML Configuration

You can also specify offloading directly in your pipeline configuration:

```yaml
pipeline:
  offload_mode: cpu  # or: disk, none

```

The CLI flag takes precedence when both are present.

## Optimizer State Offloading During Validation

Even when model weights stay on GPU, **optimizer state** can dominate VRAM during validation phases. LTX-2 provides a separate toggle for this scenario.

In [`packages/ltx-trainer/src/ltx_trainer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py) (line 400), the boolean field `acceleration.offload_optimizer_during_validation` controls this behavior:

```yaml
acceleration:
  offload_optimizer_during_validation: true  # default: false

```

When enabled, the trainer's `_offloaded_optimizer_state` context manager—implemented in [`packages/ltx-trainer/src/ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) (lines 825–856)—moves optimizer tensors to CPU during validation and restores them afterward. This frees GPU memory for the forward pass without interrupting training state.

Source: [`offload_optimizer_during_validation` field](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py#L400)  
Source: [`_offloaded_optimizer_state` context manager (lines 825–856)](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py#L825-L856)

## Complete Minimal-VRAM Configuration

For the most aggressive memory savings suitable for consumer GPUs:

```yaml

# minimal_vram_config.yaml

acceleration:
  offload_optimizer_during_validation: true

pipeline:
  offload_mode: disk  # stream from disk; use 'cpu' if disk I/O is too slow

```

Run with:

```bash
python -m ltx_pipelines.ti2vid_two_stages \
    --config minimal_vram_config.yaml \
    --offload disk

```

**Performance expectations:**
- `DISK` mode: Slowest inference, minimal RAM requirements (~2 GB CPU cache)
- `CPU` mode: Moderate speed, requires ~36 GB system RAM
- `NONE` mode: Fastest, requires full model in VRAM

## Summary

- **`OffloadMode`** in [`types.py`](https://github.com/Lightricks/LTX-2/blob/main/types.py) controls where weights live: GPU, CPU RAM, or disk
- **`--offload cpu|disk`** CLI flag or `pipeline.offload_mode` in YAML selects the strategy
- **`cpu_slots_count`** configures the disk cache size when using `DISK` mode
- **`offload_optimizer_during_validation`** reduces VRAM pressure specifically during validation loops
- Builders in [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py) validate incompatible configurations when offloading is active

## Frequently Asked Questions

### What is the minimum VRAM required to run LTX-2 with offloading?

With `CPU` or `DISK` offloading enabled, VRAM requirements drop to approximately **5 GB**. The `DISK` mode further reduces CPU RAM needs to roughly 2 GB through its on-disk weight streaming with small CPU cache. These figures assume no additional memory-heavy components like certain text encoders are active.

### Does model offloading affect video generation quality?

No. Offloading only changes **where** weights are stored and how they reach the GPU—it does not alter model architecture, precision, or inference computations. Output quality remains identical across `NONE`, `CPU`, and `DISK` modes. The trade-off is purely between memory efficiency and inference speed.

### Can I use model offloading during training, or only inference?

The `OffloadMode` system primarily targets **inference pipelines** in `ltx-pipelines`. For **training**, the separate `offload_optimizer_during_validation` toggle in `ltx-trainer` addresses optimizer state, but full weight offloading during training would require additional implementation not present in the current source.

### Why does the builder reject certain configurations when offloading is enabled?

The stage builders in [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py) validate that components incompatible with streaming—such as `text_encoder_builder`—are not used when `offload_mode != OffloadMode.NONE`. This prevents runtime errors from components that expect full GPU-resident weights. Check the validation logic at lines 370–404 for specific restrictions.