Configuring Block Streaming for Memory‑Constrained GPUs in LTX‑2
Set gpu_slots_count=1 and cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS to run LTX‑2 diffusion models on GPUs with 4‑6 GB VRAM by streaming transformer blocks from disk instead of holding them in GPU memory.
LTX‑2 implements a block‑streaming architecture that loads transformer weights on‑the‑fly, keeping only a handful of blocks resident on the GPU at any moment. This makes large video diffusion models accessible on consumer‑grade hardware. This guide explains how to configure the streaming parameters in ltx_core.block_streaming to match your GPU's memory constraints.
How Block Streaming Works in LTX‑2
The streaming system revolves around five core components defined in packages/ltx-core/src/ltx_core/block_streaming/:
| Component | Responsibility | Key Source File |
|---|---|---|
StreamingModelBuilder |
Parses checkpoints, configures memory tiers, creates GPU buffer pools, returns BlockStreamingWrapper |
builder.py |
BlockStreamingWrapper |
Wraps the meta‑model, registers forward hooks _pre_hook/_post_hook to fetch and release weights |
wrapper.py |
WeightsProvider |
Manages GPU buffer pool, handles H2D copies, fuses LoRA adapters during transfer | provider.py |
BufferPool |
Pre‑allocates fixed uint8 GPU buffers (slot_nbytes each), enforces event‑driven reuse |
pool.py |
utils |
Buffer carving, layout derivation, 16‑byte aligned memory allocation | utils.py |
Forward Pass Flow
- Pre‑hook activation:
BlockStreamingWrapper._pre_hookcallsprovider.get(idx)to obtain the block's GPU weights - Cache management: On a miss,
WeightsProviderevicts the oldest GPU slot via_copy_to_gpu, copies bytes from pinned CPU or disk, fuses LoRA deltas with_fuse_block_loras, and returns tensor views viacarve_buffer - Post‑hook cleanup:
_post_hookcallsprovider.mark_block_doneto record aStreamEvent, preventing slot reuse until compute completes
VRAM usage equals approximately gpu_slots_count × slot_nbytes. With gpu_slots_count=2 by default, footprint remains minimal; setting it to 1 shrinks it further.
Memory Configuration Parameters
The StreamingModelBuilder.build() method exposes two critical knobs for configuring block streaming for memory‑constrained GPUs in LTX‑2:
| Parameter | Effect | Low‑VRAM Recommendation |
|---|---|---|
gpu_slots_count |
Number of GPU buffer slots (one block per slot) | 1 (absolute minimum) or 2 |
cpu_slots_count |
Pinned CPU slots before spilling to disk | DISK_CPU_SLOTS (2) for disk streaming; None or len(blocks) for RAM‑only |
When cpu_slots_count < number_of_blocks, the builder instantiates DiskWeightSource to stream from the safetensors file. This trades sequential read latency for RAM savings.
The DISK_CPU_SLOTS Constant
# From builder.py — the sentinel value that enables disk-backed streaming
DISK_CPU_SLOTS: int = 2
Setting cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS keeps only two blocks in pinned host memory, reading the remainder on demand from disk.
Complete Configuration Examples
Minimal VRAM: Single GPU Slot + Disk Streaming
import torch
from torch import device
from ltx_core.block_streaming import StreamingModelBuilder
from ltx_core.model.model_protocol import MyModelConfigurator # LTX-2 configurator
from ltx_core.devices import synchronize_device
builder = StreamingModelBuilder(
model_class_configurator=MyModelConfigurator,
model_path="ltx2_checkpoint.safetensors",
blocks_attr="transformer.blocks", # Path to nn.ModuleList in model
blocks_prefix="transformer.blocks", # State-dict prefix for block weights
)
# Configure for 4-6 GB VRAM GPU
streaming_model = builder.build(
device=device("cuda:0"),
dtype=torch.bfloat16,
gpu_slots_count=1, # One block on GPU
cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS, # Two blocks pinned, rest from disk
)
output = streaming_model(input_tensor.to(device("cuda:0")))
synchronize_device()
What this achieves:
BufferPoolallocates exactly one GPU slot (gpu_slots_count=1)DiskWeightSourceprovides blocks with minimal pinned‑RAM overhead- Non‑block weights (embeddings, final layers) still load via
_load_non_block_weights
RAM‑Only Streaming (More RAM, Faster Access)
builder = StreamingModelBuilder(
model_class_configurator=MyModelConfigurator,
model_path="ltx2_checkpoint.safetensors",
blocks_attr="transformer.blocks",
blocks_prefix="transformer.blocks",
cpu_slots_count=None, # All blocks pinned in host RAM
)
model = builder.build(
device=torch.device("cuda"),
dtype=torch.bfloat16,
gpu_slots_count=1, # Minimal GPU footprint
)
Streaming with LoRA Adapter Fusion
from ltx_core.block_streaming.lora_types import LoraPathStrengthAndSDOps, SDOps
builder = builder.with_loras((
LoraPathStrengthAndSDOps(
path="style_lora.safetensors",
strength=0.8,
sd_ops=SDOps("lora")
),
))
model = builder.build(
device=torch.device("cuda"),
dtype=torch.bfloat16,
gpu_slots_count=1,
cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTS,
)
LoRA weights merge during the H2D copy inside WeightsProvider._fuse_block_loras—no additional GPU memory required for adapter tensors.
Key Implementation Files
Understanding these source files helps debug and extend the streaming behavior:
builder.py: Immutable builder pattern; validatesgpu_slots_countandcpu_slots_count, constructsDiskWeightSourceorPinnedCpuWeightSourceaccordinglyprovider.py:get()method implements LRU eviction;_fuse_block_loras()applies deltaspool.py:acquire()/release()withStreamEventsynchronizationwrapper.py: Hook registration in__init__, automatic cleanup onforwardcompletionblocks.py(inltx_pipelines): Pipeline integration helper that auto‑selects streaming builders
Summary
- Block streaming enables LTX‑2 inference on 4‑6 GB GPUs by limiting GPU‑resident weights to
gpu_slots_countblocks - Set
gpu_slots_count=1for absolute minimum VRAM usage - Use
cpu_slots_count=StreamingModelBuilder.DISK_CPU_SLOTSto stream from disk when host RAM is also constrained - The
WeightsProviderfuses LoRA adapters during H2D copy, maintaining memory efficiency - Non‑block components always load fully to GPU; verify they fit within your remaining budget
Frequently Asked Questions
What is the minimum GPU VRAM for LTX‑2 with block streaming?
With gpu_slots_count=1 and cpu_slots_count=DISK_CPU_SLOTS, LTX‑2 runs on GPUs with approximately 4‑6 GB VRAM. The exact requirement depends on activation size and non‑block weights; test with your specific checkpoint.
Does disk streaming hurt inference speed?
Disk streaming adds sequential read latency proportional to block size. With NVMe storage and cpu_slots_count=DISK_CPU_SLOTS, overhead is typically 10‑20% for standard LTX‑2 models. CPU‑only streaming (cpu_slots_count=None) eliminates this at higher RAM cost.
How do I verify which blocks are currently cached on GPU?
The WeightsProvider holds this state internally; expose it by inspecting provider._cached_blocks (LRU cache dictionary) or add logging hooks to BlockStreamingWrapper._pre_hook in wrapper.py.
Can I stream with multiple LoRA adapters?
Yes. Pass multiple LoraPathStrengthAndSDOps tuples to builder.with_loras(). All adapters fuse during the H2D transfer; only the fused weights occupy GPU memory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →