Debugging Out-of-Memory Errors in LTX-2: GPU Profiling and Gradient-Estimation Sampling
Enable verbose memory logging, wrap heavy operations with free_gpu_memory_context, and switch to gradient_estimating_euler_denoising_loop to cut diffusion steps and peak VRAM usage.
Debugging out-of-memory errors in LTX-2 requires a systematic approach that combines runtime profiling, deterministic memory cleanup, and algorithmic optimizations. LTX-2 (Live-Texture eXchange) is a multi-modal diffusion pipeline capable of simultaneous video and audio generation, but its large model sizes—often exceeding 10 GB VRAM—and high-resolution frame processing create substantial memory pressure. The repository provides three complementary mechanisms to track, reduce, and resolve GPU memory issues.
GPU-Memory Utilities in ltx_trainer/gpu_utils.py
The foundation of memory debugging in LTX-2 lies in two utilities: free_gpu_memory and free_gpu_memory_context.
Immediate Cleanup with free_gpu_memory
free_gpu_memory calls torch.cuda.empty_cache() and Python's garbage collector. When invoked with log=True, it reports both allocated and reserved memory:
from ltx_trainer.gpu_utils import free_gpu_memory, get_gpu_memory_gb
free_gpu_memory(log=True) # Immediate clean-up + log
current = get_gpu_memory_gb(torch.device('cuda')) # GB used right now
Context-Manager for Deterministic Reclamation
The free_gpu_memory_context wrapper guarantees memory is freed before and/or after any code block:
from ltx_trainer.gpu_utils import free_gpu_memory_context
with free_gpu_memory_context(before=True, after=True, log=True):
# Heavy operation that may temporarily blow memory
result = run_inference(model, batch)
These utilities appear throughout the Trainer class to maintain predictable VRAM footprints when switching between video and audio modalities.
Gradient-Estimation Sampling for Reduced Memory Pressure
The standard Euler denoising loop (euler_denoising_loop) evaluates the model at every diffusion step. The gradient_estimating_euler_denoising_loop in ltx_pipelines/utils/samplers.py adds velocity-correction to accelerate convergence:
def gradient_estimating_euler_denoising_loop(
sigmas, video_state, audio_state,
stepper, transformer, denoiser,
ge_gamma: float = 2.0,
) -> tuple[LatentState | None, LatentState | None]:
"""Run diffusion with gradient-estimation sampling (paper: https://openreview.net/pdf?id=o2ND9v0CeK)."""
...
Key parameters and behavior:
ge_gamma(default 2.0) controls velocity-correction aggressiveness- Velocity reuse via
previous_video_velocityandprevious_audio_velocityamortizes computation cost - Fewer diffusion steps (e.g., 50 → 35) directly reduces intermediate activation storage and peak VRAM
Activate this sampler by replacing the default loop in your pipeline configuration:
from ltx_pipelines.utils.samplers import gradient_estimating_euler_denoising_loop
pipeline_cfg = {
"denoising_loop": gradient_estimating_euler_denoising_loop,
"ge_gamma": 2.5, # tweak for faster convergence
}
pipeline = build_pipeline(pipeline_cfg)
Trainer-Level Memory Profiling
The Trainer class in ltx_trainer/trainer.py implements systematic memory monitoring at strategic points.
Post-Initialization Snapshot
# After model initialization
vram_usage_gb = torch.cuda.memory_allocated() / 1024**3
logger.debug(f"GPU memory usage after models preparation: {vram_usage_gb:.2f} GB")
Periodic Training Logs
if step % cfg.logging.log_memory_every == 0:
current_mem = get_gpu_memory_gb(device)
logger.info(f"[Step {step}] Current GPU memory: {current_mem:.2f} GB")
Set log_memory_every to a low value (e.g., 10 steps) to capture fine-grained memory timelines.
Practical Debugging Workflow
Follow this systematic approach to identify and resolve out-of-memory errors in LTX-2:
- Enable verbose logging – Set
log_memory_every=10in your training configuration - Instrument suspect blocks – Wrap video loading, tiling operations, or denoising calls with
free_gpu_memory_context(before=True, after=True, log=True) - Switch sampling algorithms – Replace
euler_denoising_loopwithgradient_estimating_euler_denoising_loop - Analyze peak memory – Locate the highest "Current GPU memory" log entry to identify the problematic stage
- Apply additional optimizations – Enable
cfg.optimization.enable_gradient_checkpointing=Trueor spatial tiling (use_tiling=True)
Example: Profiling Video Loading
from ltx_trainer.gpu_utils import free_gpu_memory_context, get_gpu_memory_gb
import logging
logger = logging.getLogger(__name__)
def load_chunk(path):
with free_gpu_memory_context(before=True, after=True, log=True):
chunk = read_video_file(path) # heavy I/O
mem = get_gpu_memory_gb(torch.device('cuda'))
logger.info(f"Memory after loading {path}: {mem:.2f} GB")
return chunk
Example: Trainer Memory Profiling Integration
from ltx_trainer.gpu_utils import get_gpu_memory_gb
class Trainer:
...
def _log_memory(self, step):
mem = get_gpu_memory_gb(self.device)
self.logger.info(f"[Step {step}] GPU memory: {mem:.2f} GB")
def train(self):
for step in range(self.total_steps):
# … training logic …
if step % self.cfg.logging.log_memory_every == 0:
self._log_memory(step)
Key Source Files for Memory Debugging
| File | Location | Purpose |
|---|---|---|
gpu_utils.py |
packages/ltx-trainer/src/ltx_trainer/gpu_utils.py |
Core utilities for memory cleanup and GB-level queries |
samplers.py |
packages/ltx-pipelines/src/ltx_pipelines/utils/samplers.py |
Gradient-estimation Euler implementation |
trainer.py |
packages/ltx-trainer/src/ltx_trainer/trainer.py |
Training orchestration with built-in memory logging |
config.py |
packages/ltx-trainer/src/ltx_trainer/config.py |
Optimization flags: gradient checkpointing, accumulation steps |
process_videos.py |
packages/ltx-trainer/scripts/process_videos.py |
Video processing with spatial tiling support |
Summary
free_gpu_memoryandfree_gpu_memory_contextprovide deterministic CUDA cache cleanup and precise memory measurementgradient_estimating_euler_denoising_loopreduces diffusion steps through velocity correction, lowering peak VRAM requirements- Trainer-level profiling with configurable
log_memory_everyintervals creates actionable memory timelines - Iterative workflow combining instrumentation, algorithm swapping, and configuration tuning eliminates most OOM crashes without quality degradation
Frequently Asked Questions
What is the default ge_gamma value in gradient-estimation sampling?
The default ge_gamma is 2.0, as defined in the gradient_estimating_euler_denoising_loop signature. Increasing this value (e.g., to 2.5) applies more aggressive velocity correction, potentially reducing steps further at the cost of sample stability.
Where does LTX-2 log peak GPU memory during training?
According to the Trainer implementation in ltx_trainer/trainer.py, peak memory is logged immediately after model preparation and periodically during training when step % cfg.logging.log_memory_every == 0. Both use torch.cuda.memory_allocated() divided by 1024**3 to report gigabytes.
Can I use free_gpu_memory_context outside the trainer?
Yes. The context manager is defined in ltx_trainer/gpu_utils.py and has no trainer dependencies. It works in any Python block where you need deterministic CUDA cache clearing—common use cases include dataset loading, inference pipelines, and multi-modal switching operations.
How does gradient-estimation sampling reduce memory usage?
The velocity-correction mechanism in gradient_estimating_euler_denoising_loop reuses previous velocity vectors to reach the target distribution faster. This convergence speedup allows reducing total diffusion steps (e.g., from 50 to 35), which proportionally decreases the number of forward passes and intermediate activations stored in VRAM.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →