How to Run LTX-2 Inference with Limited VRAM Using Block Streaming: A Complete Guide
Block streaming enables LTX-2 inference on GPUs with as little as 4 GB of VRAM by transferring transformer blocks between CPU and GPU on-the-fly, keeping only the active computation in GPU memory.
The LTX-2 models from Lightricks require substantial memory—full-precision checkpoints can demand tens of gigabytes of GPU memory. The ltx-core package solves this through a sophisticated block-streaming system that treats VRAM as a cache while streaming weights from disk. This approach allows you to run inference on consumer hardware without compromising model quality.
Understanding the Block Streaming Architecture
The block-streaming pipeline consists of three core components that work together to minimize memory footprint.
StreamingModelBuilder
Located in ltx_core/block_streaming/builder.py, the StreamingModelBuilder class parses the checkpoint and initializes the streaming infrastructure. It creates a disk-weight source, manages a GPU buffer pool, and constructs a WeightsProvider. The builder ultimately returns a BlockStreamingWrapper instance that transparently replaces your original model.
BlockStreamingWrapper
The BlockStreamingWrapper in ltx_core/block_streaming/wrapper.py is a thin nn.Module that intercepts forward passes. It registers forward-pre and forward-post hooks on each transformer block. The pre-hook pulls weights from the provider and injects them into the block, while the post-hook records a CUDA event and releases the GPU buffer for reuse.
WeightsProvider
Found in ltx_core/block_streaming/provider.py, the WeightsProvider manages a pinned-CPU buffer pool and handles asynchronous host-to-device (H2D) copies. It optionally fuses LoRA adapters on the copy stream and guarantees that a block's GPU buffer persists until computation completes.
Step-by-Step Implementation
To enable block streaming, you configure the StreamingModelBuilder with memory-specific parameters and wrap your existing pipeline.
Basic Configuration for Minimal VRAM
For the smallest possible VRAM footprint, set cpu_slots_count=1 and prefetch_depth=0. This keeps only one block in CPU memory and disables asynchronous prefetching.
import torch
from ltx_core.block_streaming import StreamingModelBuilder
from ltx_pipelines.t2a_one_stage import T2AOneStagePipeline
device = torch.device("cuda:0")
builder = StreamingModelBuilder(
checkpoint_path="path/to/model.safetensors",
target_device=device,
dtype=torch.float16,
cpu_slots_count=1,
prefetch_depth=0,
)
model = builder.build()
pipeline = T2AOneStagePipeline(model=model)
result = pipeline(
prompt="A photorealistic portrait of a cat",
seed=123,
num_inference_steps=30,
height=512,
width=512,
)
model.teardown()
The dtype=torch.float16 parameter reduces VRAM usage by approximately 50% compared to full precision. With these settings, the system allocates only a single CPU-pinned buffer, keeping the rest of the model weights on disk until needed.
Optimizing with Prefetch
If you have additional system RAM (8 GB or more), increase cpu_slots_count to 2 and prefetch_depth to 1. This allows the WeightsProvider to preload the next block while the current block processes on the GPU, hiding transfer latency without increasing VRAM consumption.
builder = StreamingModelBuilder(
checkpoint_path="model.safetensors",
target_device=device,
dtype=torch.float16,
cpu_slots_count=2,
prefetch_depth=1,
)
model = builder.build()
# ... use with any pipeline
Adding LoRA Adapters
The streaming system automatically fuses LoRA adapters during the copy operation. Pass a list of LoRA paths to the builder, and the WeightsProvider merges deltas in-place via _fuse_block_loras without consuming additional GPU memory.
builder = StreamingModelBuilder(
checkpoint_path="model.safetensors",
target_device=device,
dtype=torch.float16,
cpu_slots_count=2,
prefetch_depth=1,
lora_paths=["path/to/lora_a.safetensors", "path/to/lora_b.safetensors"],
)
model = builder.build()
Complete Pipeline Examples
The block-streaming wrapper is compatible with any LTX-2 pipeline that accepts a model argument. For text-to-video generation using the Ti2VidOneStagePipeline:
from ltx_pipelines.ti2vid_one_stage import Ti2VidOneStagePipeline
builder = StreamingModelBuilder(
checkpoint_path="ltx2_video.safetensors",
target_device=torch.device("cuda:0"),
dtype=torch.float16,
cpu_slots_count=2,
prefetch_depth=1,
)
model = builder.build()
pipeline = Ti2VidOneStagePipeline(model=model)
video = pipeline(
prompt="A sunset over a mountain lake",
seed=42,
num_inference_steps=40,
height=720,
width=1280,
fps=24,
)
model.teardown()
Verifying VRAM Usage
To confirm that only one transformer block resides in GPU memory at any time, monitor peak VRAM allocation:
torch.cuda.reset_peak_memory_stats()
output = pipeline(...)
print("Peak VRAM (MiB):", torch.cuda.max_memory_allocated() / 2**20)
With FP16 precision and the minimal configuration above, typical peak usage remains around 2 GB regardless of the original checkpoint size.
Summary
- Block streaming moves transformer blocks between CPU and GPU on-demand, enabling LTX-2 inference on limited VRAM.
- The
StreamingModelBuildercreates a disk-weight source and buffer pool, whileBlockStreamingWrappermanages block lifecycle through forward hooks. - Control memory usage with
cpu_slots_count(CPU-pinned buffers) andprefetch_depth(async prefetching). - The wrapper is drop-in compatible with all LTX-2 pipelines, including
T2AOneStagePipelineandTi2VidOneStagePipeline. - LoRA adapters are fused automatically during the copy stream without additional memory overhead.
- Always call
model.teardown()after inference to release hooks and provider resources.
Frequently Asked Questions
What is the minimum VRAM required for LTX-2 block streaming?
With cpu_slots_count=1, prefetch_depth=0, and FP16 precision, you can run LTX-2 inference on GPUs with approximately 4 GB of VRAM. The system keeps only the actively executing transformer block in GPU memory while streaming others from disk.
Does block streaming affect inference quality?
No. Block streaming is a memory management technique that does not modify model weights or computation graphs. The inference results are identical to running the full model in VRAM, as the same checkpoint weights are loaded—just transferred dynamically rather than statically.
How does the prefetch_depth parameter work?
The prefetch_depth parameter in StreamingModelBuilder controls how many future blocks the WeightsProvider loads asynchronously into CPU-pinned buffers while the GPU processes the current block. A value of 1 hides H2D transfer latency by preloading the next block, while 0 minimizes memory usage by disabling prefetching entirely.
Can I use multiple LoRA adapters with block streaming?
Yes. Pass multiple LoRA paths to the lora_paths parameter when constructing the StreamingModelBuilder. The WeightsProvider fuses all specified adapters into each block's weights during the async copy operation, applying them in the order specified without requiring additional GPU memory.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →