# Configuring Block Streaming for Memory‑Constrained GPUs in LTX‑2

> Optimize LTX-2 diffusion models for memory-constrained GPUs by configuring block streaming. Reduce VRAM needs by streaming transformer blocks from disk. Learn how to set gpu_slots_count and cpu_slots_count.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-15

---

**Set `gpu_slots_count=1` and `cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS` to run LTX‑2 diffusion models on GPUs with 4‑6 GB VRAM by streaming transformer blocks from disk instead of holding them in GPU memory.**

LTX‑2 implements a **block‑streaming** architecture that loads transformer weights on‑the‑fly, keeping only a handful of blocks resident on the GPU at any moment. This makes large video diffusion models accessible on consumer‑grade hardware. This guide explains how to configure the streaming parameters in `ltx_core.block_streaming` to match your GPU's memory constraints.

## How Block Streaming Works in LTX‑2

The streaming system revolves around five core components defined in `packages/ltx-core/src/ltx_core/block_streaming/`:

| Component | Responsibility | Key Source File |
|-----------|---------------|---------------|
| **`StreamingModelBuilder`** | Parses checkpoints, configures memory tiers, creates GPU buffer pools, returns `BlockStreamingWrapper` | [`builder.py`](https://github.com/Lightricks/LTX-2/blob/main/builder.py) |
| **`BlockStreamingWrapper`** | Wraps the meta‑model, registers forward hooks `_pre_hook`/`_post_hook` to fetch and release weights | [`wrapper.py`](https://github.com/Lightricks/LTX-2/blob/main/wrapper.py) |
| **`WeightsProvider`** | Manages GPU buffer pool, handles H2D copies, fuses LoRA adapters during transfer | [`provider.py`](https://github.com/Lightricks/LTX-2/blob/main/provider.py) |
| **`BufferPool`** | Pre‑allocates fixed `uint8` GPU buffers (`slot_nbytes` each), enforces event‑driven reuse | [`pool.py`](https://github.com/Lightricks/LTX-2/blob/main/pool.py) |
| **`utils`** | Buffer carving, layout derivation, 16‑byte aligned memory allocation | [`utils.py`](https://github.com/Lightricks/LTX-2/blob/main/utils.py) |

### Forward Pass Flow

1. **Pre‑hook activation**: `BlockStreamingWrapper._pre_hook` calls `provider.get(idx)` to obtain the block's GPU weights
2. **Cache management**: On a miss, `WeightsProvider` evicts the oldest GPU slot via `_copy_to_gpu`, copies bytes from pinned CPU or disk, fuses LoRA deltas with `_fuse_block_loras`, and returns tensor views via `carve_buffer`
3. **Post‑hook cleanup**: `_post_hook` calls `provider.mark_block_done` to record a `StreamEvent`, preventing slot reuse until compute completes

VRAM usage equals approximately `gpu_slots_count × slot_nbytes`. With `gpu_slots_count=2` by default, footprint remains minimal; setting it to `1` shrinks it further.

## Memory Configuration Parameters

The `StreamingModelBuilder.build()` method exposes two critical knobs for **configuring block streaming for memory‑constrained GPUs in LTX‑2**:

| Parameter | Effect | Low‑VRAM Recommendation |
|-----------|--------|------------------------|
| **`gpu_slots_count`** | Number of GPU buffer slots (one block per slot) | `1` (absolute minimum) or `2` |
| **`cpu_slots_count`** | Pinned CPU slots before spilling to disk | `DISK_CPU_SLOTS` (`2`) for disk streaming; `None` or `len(blocks)` for RAM‑only |

When `cpu_slots_count < number_of_blocks`, the builder instantiates `DiskWeightSource` to stream from the safetensors file. This trades sequential read latency for RAM savings.

### The `DISK_CPU_SLOTS` Constant

```python

# From builder.py — the sentinel value that enables disk-backed streaming

DISK_CPU_SLOTS: int = 2

```

Setting `cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS` keeps only two blocks in pinned host memory, reading the remainder on demand from disk.

## Complete Configuration Examples

### Minimal VRAM: Single GPU Slot + Disk Streaming

```python
import torch
from torch import device
from ltx_core.block_streaming import StreamingModelBuilder
from ltx_core.model.model_protocol import MyModelConfigurator  # LTX-2 configurator

from ltx_core.devices import synchronize_device

builder = StreamingModelBuilder(
    model_class_configurator=MyModelConfigurator,
    model_path="ltx2_checkpoint.safetensors",
    blocks_attr="transformer.blocks",      # Path to nn.ModuleList in model

    blocks_prefix="transformer.blocks",    # State-dict prefix for block weights

)

# Configure for 4-6 GB VRAM GPU

streaming_model = builder.build(
    device=device("cuda:0"),
    dtype=torch.bfloat16,
    gpu_slots_count=1,                                      # One block on GPU

    cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS,   # Two blocks pinned, rest from disk

)

output = streaming_model(input_tensor.to(device("cuda:0")))
synchronize_device()

```

**What this achieves:**
- `BufferPool` allocates exactly **one** GPU slot (`gpu_slots_count=1`)
- `DiskWeightSource` provides blocks with minimal pinned‑RAM overhead
- Non‑block weights (embeddings, final layers) still load via `_load_non_block_weights`

### RAM‑Only Streaming (More RAM, Faster Access)

```python
builder = StreamingModelBuilder(
    model_class_configurator=MyModelConfigurator,
    model_path="ltx2_checkpoint.safetensors",
    blocks_attr="transformer.blocks",
    blocks_prefix="transformer.blocks",
    cpu_slots_count=None,           # All blocks pinned in host RAM

)

model = builder.build(
    device=torch.device("cuda"),
    dtype=torch.bfloat16,
    gpu_slots_count=1,              # Minimal GPU footprint

)

```

### Streaming with LoRA Adapter Fusion

```python
from ltx_core.block_streaming.lora_types import LoraPathStrengthAndSDOps, SDOps

builder = builder.with_loras((
    LoraPathStrengthAndSDOps(
        path="style_lora.safetensors",
        strength=0.8,
        sd_ops=SDOps("lora")
    ),
))

model = builder.build(
    device=torch.device("cuda"),
    dtype=torch.bfloat16,
    gpu_slots_count=1,
    cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS,
)

```

LoRA weights merge during the H2D copy inside `WeightsProvider._fuse_block_loras`—no additional GPU memory required for adapter tensors.

## Key Implementation Files

Understanding these source files helps debug and extend the streaming behavior:

- **[`builder.py`](https://github.com/Lightricks/LTX-2/blob/main/builder.py)**: Immutable builder pattern; validates `gpu_slots_count` and `cpu_slots_count`, constructs `DiskWeightSource` or `PinnedCpuWeightSource` accordingly
- **[`provider.py`](https://github.com/Lightricks/LTX-2/blob/main/provider.py)**: `get()` method implements LRU eviction; `_fuse_block_loras()` applies deltas
- **[`pool.py`](https://github.com/Lightricks/LTX-2/blob/main/pool.py)**: `acquire()`/`release()` with `StreamEvent` synchronization
- **[`wrapper.py`](https://github.com/Lightricks/LTX-2/blob/main/wrapper.py)**: Hook registration in `__init__`, automatic cleanup on `forward` completion
- **[`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py)** (in `ltx_pipelines`): Pipeline integration helper that auto‑selects streaming builders

## Summary

- **Block streaming** enables LTX‑2 inference on 4‑6 GB GPUs by limiting GPU‑resident weights to `gpu_slots_count` blocks
- Set `gpu_slots_count=1` for absolute minimum VRAM usage
- Use `cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS` to stream from disk when host RAM is also constrained
- The `WeightsProvider` fuses LoRA adapters during H2D copy, maintaining memory efficiency
- Non‑block components always load fully to GPU; verify they fit within your remaining budget

## Frequently Asked Questions

### What is the minimum GPU VRAM for LTX‑2 with block streaming?

With `gpu_slots_count=1` and `cpu_slots_count=DISK_CPU_SLOTS`, LTX‑2 runs on GPUs with approximately **4‑6 GB VRAM**. The exact requirement depends on activation size and non‑block weights; test with your specific checkpoint.

### Does disk streaming hurt inference speed?

Disk streaming adds **sequential read latency** proportional to block size. With NVMe storage and `cpu_slots_count=DISK_CPU_SLOTS`, overhead is typically 10‑20% for standard LTX‑2 models. CPU‑only streaming (`cpu_slots_count=None`) eliminates this at higher RAM cost.

### How do I verify which blocks are currently cached on GPU?

The `WeightsProvider` holds this state internally; expose it by inspecting `provider._cached_blocks` (LRU cache dictionary) or add logging hooks to `BlockStreamingWrapper._pre_hook` in [`wrapper.py`](https://github.com/Lightricks/LTX-2/blob/main/wrapper.py).

### Can I stream with multiple LoRA adapters?

Yes. Pass multiple `LoraPathStrengthAndSDOps` tuples to `builder.with_loras()`. All adapters fuse during the H2D transfer; only the fused weights occupy GPU memory.