# How to Run LTX-2 Inference with Limited VRAM Using Block Streaming: A Complete Guide

> Run LTX-2 inference with limited VRAM using block streaming. Efficiently process large models on GPUs with only 4GB VRAM by streaming transformer blocks between CPU and GPU.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-06-20

---

**Block streaming enables LTX-2 inference on GPUs with as little as 4 GB of VRAM by transferring transformer blocks between CPU and GPU on-the-fly, keeping only the active computation in GPU memory.**

The LTX-2 models from Lightricks require substantial memory—full-precision checkpoints can demand tens of gigabytes of GPU memory. The `ltx-core` package solves this through a sophisticated block-streaming system that treats VRAM as a cache while streaming weights from disk. This approach allows you to run inference on consumer hardware without compromising model quality.

## Understanding the Block Streaming Architecture

The block-streaming pipeline consists of three core components that work together to minimize memory footprint.

### StreamingModelBuilder

Located in [`ltx_core/block_streaming/builder.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/block_streaming/builder.py), the `StreamingModelBuilder` class parses the checkpoint and initializes the streaming infrastructure. It creates a disk-weight source, manages a GPU buffer pool, and constructs a `WeightsProvider`. The builder ultimately returns a `BlockStreamingWrapper` instance that transparently replaces your original model.

### BlockStreamingWrapper

The `BlockStreamingWrapper` in [`ltx_core/block_streaming/wrapper.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/block_streaming/wrapper.py) is a thin `nn.Module` that intercepts forward passes. It registers forward-pre and forward-post hooks on each transformer block. The pre-hook pulls weights from the provider and injects them into the block, while the post-hook records a CUDA event and releases the GPU buffer for reuse.

### WeightsProvider

Found in [`ltx_core/block_streaming/provider.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/block_streaming/provider.py), the `WeightsProvider` manages a pinned-CPU buffer pool and handles asynchronous host-to-device (H2D) copies. It optionally fuses LoRA adapters on the copy stream and guarantees that a block's GPU buffer persists until computation completes.

## Step-by-Step Implementation

To enable block streaming, you configure the `StreamingModelBuilder` with memory-specific parameters and wrap your existing pipeline.

### Basic Configuration for Minimal VRAM

For the smallest possible VRAM footprint, set `cpu_slots_count=1` and `prefetch_depth=0`. This keeps only one block in CPU memory and disables asynchronous prefetching.

```python
import torch
from ltx_core.block_streaming import StreamingModelBuilder
from ltx_pipelines.t2a_one_stage import T2AOneStagePipeline

device = torch.device("cuda:0")

builder = StreamingModelBuilder(
    checkpoint_path="path/to/model.safetensors",
    target_device=device,
    dtype=torch.float16,
    cpu_slots_count=1,
    prefetch_depth=0,
)

model = builder.build()
pipeline = T2AOneStagePipeline(model=model)

result = pipeline(
    prompt="A photorealistic portrait of a cat",
    seed=123,
    num_inference_steps=30,
    height=512,
    width=512,
)

model.teardown()

```

The `dtype=torch.float16` parameter reduces VRAM usage by approximately 50% compared to full precision. With these settings, the system allocates only a single CPU-pinned buffer, keeping the rest of the model weights on disk until needed.

### Optimizing with Prefetch

If you have additional system RAM (8 GB or more), increase `cpu_slots_count` to 2 and `prefetch_depth` to 1. This allows the `WeightsProvider` to preload the next block while the current block processes on the GPU, hiding transfer latency without increasing VRAM consumption.

```python
builder = StreamingModelBuilder(
    checkpoint_path="model.safetensors",
    target_device=device,
    dtype=torch.float16,
    cpu_slots_count=2,
    prefetch_depth=1,
)

model = builder.build()

# ... use with any pipeline

```

### Adding LoRA Adapters

The streaming system automatically fuses LoRA adapters during the copy operation. Pass a list of LoRA paths to the builder, and the `WeightsProvider` merges deltas in-place via `_fuse_block_loras` without consuming additional GPU memory.

```python
builder = StreamingModelBuilder(
    checkpoint_path="model.safetensors",
    target_device=device,
    dtype=torch.float16,
    cpu_slots_count=2,
    prefetch_depth=1,
    lora_paths=["path/to/lora_a.safetensors", "path/to/lora_b.safetensors"],
)

model = builder.build()

```

## Complete Pipeline Examples

The block-streaming wrapper is compatible with any LTX-2 pipeline that accepts a `model` argument. For text-to-video generation using the `Ti2VidOneStagePipeline`:

```python
from ltx_pipelines.ti2vid_one_stage import Ti2VidOneStagePipeline

builder = StreamingModelBuilder(
    checkpoint_path="ltx2_video.safetensors",
    target_device=torch.device("cuda:0"),
    dtype=torch.float16,
    cpu_slots_count=2,
    prefetch_depth=1,
)

model = builder.build()
pipeline = Ti2VidOneStagePipeline(model=model)

video = pipeline(
    prompt="A sunset over a mountain lake",
    seed=42,
    num_inference_steps=40,
    height=720,
    width=1280,
    fps=24,
)

model.teardown()

```

## Verifying VRAM Usage

To confirm that only one transformer block resides in GPU memory at any time, monitor peak VRAM allocation:

```python
torch.cuda.reset_peak_memory_stats()
output = pipeline(...)
print("Peak VRAM (MiB):", torch.cuda.max_memory_allocated() / 2**20)

```

With FP16 precision and the minimal configuration above, typical peak usage remains around 2 GB regardless of the original checkpoint size.

## Summary

- **Block streaming** moves transformer blocks between CPU and GPU on-demand, enabling LTX-2 inference on limited VRAM.
- The **`StreamingModelBuilder`** creates a disk-weight source and buffer pool, while **`BlockStreamingWrapper`** manages block lifecycle through forward hooks.
- Control memory usage with **`cpu_slots_count`** (CPU-pinned buffers) and **`prefetch_depth`** (async prefetching).
- The wrapper is drop-in compatible with all LTX-2 pipelines, including **`T2AOneStagePipeline`** and **`Ti2VidOneStagePipeline`**.
- LoRA adapters are fused automatically during the copy stream without additional memory overhead.
- Always call **`model.teardown()`** after inference to release hooks and provider resources.

## Frequently Asked Questions

### What is the minimum VRAM required for LTX-2 block streaming?

With `cpu_slots_count=1`, `prefetch_depth=0`, and FP16 precision, you can run LTX-2 inference on GPUs with approximately 4 GB of VRAM. The system keeps only the actively executing transformer block in GPU memory while streaming others from disk.

### Does block streaming affect inference quality?

No. Block streaming is a memory management technique that does not modify model weights or computation graphs. The inference results are identical to running the full model in VRAM, as the same checkpoint weights are loaded—just transferred dynamically rather than statically.

### How does the prefetch_depth parameter work?

The `prefetch_depth` parameter in `StreamingModelBuilder` controls how many future blocks the `WeightsProvider` loads asynchronously into CPU-pinned buffers while the GPU processes the current block. A value of 1 hides H2D transfer latency by preloading the next block, while 0 minimizes memory usage by disabling prefetching entirely.

### Can I use multiple LoRA adapters with block streaming?

Yes. Pass multiple LoRA paths to the `lora_paths` parameter when constructing the `StreamingModelBuilder`. The `WeightsProvider` fuses all specified adapters into each block's weights during the async copy operation, applying them in the order specified without requiring additional GPU memory.