Configuring Block Streaming for Memory‑Constrained GPUs in LTX‑2

Set gpu_slots_count=1 and cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS to run LTX‑2 diffusion models on GPUs with 4‑6 GB VRAM by streaming transformer blocks from disk instead of holding them in GPU memory.

LTX‑2 implements a block‑streaming architecture that loads transformer weights on‑the‑fly, keeping only a handful of blocks resident on the GPU at any moment. This makes large video diffusion models accessible on consumer‑grade hardware. This guide explains how to configure the streaming parameters in ltx_core.block_streaming to match your GPU's memory constraints.

How Block Streaming Works in LTX‑2

The streaming system revolves around five core components defined in packages/ltx-core/src/ltx_core/block_streaming/:

Component Responsibility Key Source File
StreamingModelBuilder Parses checkpoints, configures memory tiers, creates GPU buffer pools, returns BlockStreamingWrapper builder.py
BlockStreamingWrapper Wraps the meta‑model, registers forward hooks _pre_hook/_post_hook to fetch and release weights wrapper.py
WeightsProvider Manages GPU buffer pool, handles H2D copies, fuses LoRA adapters during transfer provider.py
BufferPool Pre‑allocates fixed uint8 GPU buffers (slot_nbytes each), enforces event‑driven reuse pool.py
utils Buffer carving, layout derivation, 16‑byte aligned memory allocation utils.py

Forward Pass Flow

  1. Pre‑hook activation: BlockStreamingWrapper._pre_hook calls provider.get(idx) to obtain the block's GPU weights
  2. Cache management: On a miss, WeightsProvider evicts the oldest GPU slot via _copy_to_gpu, copies bytes from pinned CPU or disk, fuses LoRA deltas with _fuse_block_loras, and returns tensor views via carve_buffer
  3. Post‑hook cleanup: _post_hook calls provider.mark_block_done to record a StreamEvent, preventing slot reuse until compute completes

VRAM usage equals approximately gpu_slots_count × slot_nbytes. With gpu_slots_count=2 by default, footprint remains minimal; setting it to 1 shrinks it further.

Memory Configuration Parameters

The StreamingModelBuilder.build() method exposes two critical knobs for configuring block streaming for memory‑constrained GPUs in LTX‑2:

Parameter Effect Low‑VRAM Recommendation
gpu_slots_count Number of GPU buffer slots (one block per slot) 1 (absolute minimum) or 2
cpu_slots_count Pinned CPU slots before spilling to disk DISK_CPU_SLOTS (2) for disk streaming; None or len(blocks) for RAM‑only

When cpu_slots_count < number_of_blocks, the builder instantiates DiskWeightSource to stream from the safetensors file. This trades sequential read latency for RAM savings.

The DISK_CPU_SLOTS Constant


# From builder.py — the sentinel value that enables disk-backed streaming

DISK_CPU_SLOTS: int = 2

Setting cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS keeps only two blocks in pinned host memory, reading the remainder on demand from disk.

Complete Configuration Examples

Minimal VRAM: Single GPU Slot + Disk Streaming

import torch
from torch import device
from ltx_core.block_streaming import StreamingModelBuilder
from ltx_core.model.model_protocol import MyModelConfigurator  # LTX-2 configurator

from ltx_core.devices import synchronize_device

builder = StreamingModelBuilder(
    model_class_configurator=MyModelConfigurator,
    model_path="ltx2_checkpoint.safetensors",
    blocks_attr="transformer.blocks",      # Path to nn.ModuleList in model

    blocks_prefix="transformer.blocks",    # State-dict prefix for block weights

)

# Configure for 4-6 GB VRAM GPU

streaming_model = builder.build(
    device=device("cuda:0"),
    dtype=torch.bfloat16,
    gpu_slots_count=1,                                      # One block on GPU

    cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS,   # Two blocks pinned, rest from disk

)

output = streaming_model(input_tensor.to(device("cuda:0")))
synchronize_device()

What this achieves:

  • BufferPool allocates exactly one GPU slot (gpu_slots_count=1)
  • DiskWeightSource provides blocks with minimal pinned‑RAM overhead
  • Non‑block weights (embeddings, final layers) still load via _load_non_block_weights

RAM‑Only Streaming (More RAM, Faster Access)

builder = StreamingModelBuilder(
    model_class_configurator=MyModelConfigurator,
    model_path="ltx2_checkpoint.safetensors",
    blocks_attr="transformer.blocks",
    blocks_prefix="transformer.blocks",
    cpu_slots_count=None,           # All blocks pinned in host RAM

)

model = builder.build(
    device=torch.device("cuda"),
    dtype=torch.bfloat16,
    gpu_slots_count=1,              # Minimal GPU footprint

)

Streaming with LoRA Adapter Fusion

from ltx_core.block_streaming.lora_types import LoraPathStrengthAndSDOps, SDOps

builder = builder.with_loras((
    LoraPathStrengthAndSDOps(
        path="style_lora.safetensors",
        strength=0.8,
        sd_ops=SDOps("lora")
    ),
))

model = builder.build(
    device=torch.device("cuda"),
    dtype=torch.bfloat16,
    gpu_slots_count=1,
    cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS,
)

LoRA weights merge during the H2D copy inside WeightsProvider._fuse_block_loras—no additional GPU memory required for adapter tensors.

Key Implementation Files

Understanding these source files helps debug and extend the streaming behavior:

  • builder.py: Immutable builder pattern; validates gpu_slots_count and cpu_slots_count, constructs DiskWeightSource or PinnedCpuWeightSource accordingly
  • provider.py: get() method implements LRU eviction; _fuse_block_loras() applies deltas
  • pool.py: acquire()/release() with StreamEvent synchronization
  • wrapper.py: Hook registration in __init__, automatic cleanup on forward completion
  • blocks.py (in ltx_pipelines): Pipeline integration helper that auto‑selects streaming builders

Summary

  • Block streaming enables LTX‑2 inference on 4‑6 GB GPUs by limiting GPU‑resident weights to gpu_slots_count blocks
  • Set gpu_slots_count=1 for absolute minimum VRAM usage
  • Use cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS to stream from disk when host RAM is also constrained
  • The WeightsProvider fuses LoRA adapters during H2D copy, maintaining memory efficiency
  • Non‑block components always load fully to GPU; verify they fit within your remaining budget

Frequently Asked Questions

What is the minimum GPU VRAM for LTX‑2 with block streaming?

With gpu_slots_count=1 and cpu_slots_count=DISK_CPU_SLOTS, LTX‑2 runs on GPUs with approximately 4‑6 GB VRAM. The exact requirement depends on activation size and non‑block weights; test with your specific checkpoint.

Does disk streaming hurt inference speed?

Disk streaming adds sequential read latency proportional to block size. With NVMe storage and cpu_slots_count=DISK_CPU_SLOTS, overhead is typically 10‑20% for standard LTX‑2 models. CPU‑only streaming (cpu_slots_count=None) eliminates this at higher RAM cost.

How do I verify which blocks are currently cached on GPU?

The WeightsProvider holds this state internally; expose it by inspecting provider._cached_blocks (LRU cache dictionary) or add logging hooks to BlockStreamingWrapper._pre_hook in wrapper.py.

Can I stream with multiple LoRA adapters?

Yes. Pass multiple LoraPathStrengthAndSDOps tuples to builder.with_loras(). All adapters fuse during the H2D transfer; only the fused weights occupy GPU memory.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →