# Debugging Out-of-Memory Errors in LTX-2: GPU Profiling and Gradient-Estimation Sampling

> Debug LTX-2 out-of-memory errors with GPU profiling and gradient-estimation sampling. Reduce VRAM usage and diffusion steps for smoother performance.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-15

---

**Enable verbose memory logging, wrap heavy operations with `free_gpu_memory_context`, and switch to `gradient_estimating_euler_denoising_loop` to cut diffusion steps and peak VRAM usage.**

Debugging out-of-memory errors in LTX-2 requires a systematic approach that combines runtime profiling, deterministic memory cleanup, and algorithmic optimizations. LTX-2 (Live-Texture eXchange) is a multi-modal diffusion pipeline capable of simultaneous video and audio generation, but its large model sizes—often exceeding 10 GB VRAM—and high-resolution frame processing create substantial memory pressure. The repository provides three complementary mechanisms to track, reduce, and resolve GPU memory issues.

## GPU-Memory Utilities in [`ltx_trainer/gpu_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/gpu_utils.py)

The foundation of memory debugging in LTX-2 lies in two utilities: `free_gpu_memory` and `free_gpu_memory_context`.

### Immediate Cleanup with `free_gpu_memory`

**`free_gpu_memory`** calls `torch.cuda.empty_cache()` and Python's garbage collector. When invoked with `log=True`, it reports both allocated and reserved memory:

```python
from ltx_trainer.gpu_utils import free_gpu_memory, get_gpu_memory_gb

free_gpu_memory(log=True)                     # Immediate clean-up + log

current = get_gpu_memory_gb(torch.device('cuda'))  # GB used right now

```

### Context-Manager for Deterministic Reclamation

The **`free_gpu_memory_context`** wrapper guarantees memory is freed before and/or after any code block:

```python
from ltx_trainer.gpu_utils import free_gpu_memory_context

with free_gpu_memory_context(before=True, after=True, log=True):
    # Heavy operation that may temporarily blow memory

    result = run_inference(model, batch)

```

These utilities appear throughout the `Trainer` class to maintain predictable VRAM footprints when switching between video and audio modalities.

## Gradient-Estimation Sampling for Reduced Memory Pressure

The standard Euler denoising loop (`euler_denoising_loop`) evaluates the model at every diffusion step. The **`gradient_estimating_euler_denoising_loop`** in [`ltx_pipelines/utils/samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/samplers.py) adds velocity-correction to accelerate convergence:

```python
def gradient_estimating_euler_denoising_loop(
    sigmas, video_state, audio_state,
    stepper, transformer, denoiser,
    ge_gamma: float = 2.0,
) -> tuple[LatentState | None, LatentState | None]:
    """Run diffusion with gradient-estimation sampling (paper: https://openreview.net/pdf?id=o2ND9v0CeK)."""
    ...

```

**Key parameters and behavior:**

- **`ge_gamma`** (default 2.0) controls velocity-correction aggressiveness
- **Velocity reuse** via `previous_video_velocity` and `previous_audio_velocity` amortizes computation cost
- **Fewer diffusion steps** (e.g., 50 → 35) directly reduces intermediate activation storage and peak VRAM

Activate this sampler by replacing the default loop in your pipeline configuration:

```python
from ltx_pipelines.utils.samplers import gradient_estimating_euler_denoising_loop

pipeline_cfg = {
    "denoising_loop": gradient_estimating_euler_denoising_loop,
    "ge_gamma": 2.5,          # tweak for faster convergence

}
pipeline = build_pipeline(pipeline_cfg)

```

## Trainer-Level Memory Profiling

The `Trainer` class in [`ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/trainer.py) implements systematic memory monitoring at strategic points.

### Post-Initialization Snapshot

```python

# After model initialization

vram_usage_gb = torch.cuda.memory_allocated() / 1024**3
logger.debug(f"GPU memory usage after models preparation: {vram_usage_gb:.2f} GB")

```

### Periodic Training Logs

```python
if step % cfg.logging.log_memory_every == 0:
    current_mem = get_gpu_memory_gb(device)
    logger.info(f"[Step {step}] Current GPU memory: {current_mem:.2f} GB")

```

Set `log_memory_every` to a low value (e.g., 10 steps) to capture fine-grained memory timelines.

## Practical Debugging Workflow

Follow this systematic approach to identify and resolve out-of-memory errors in LTX-2:

1. **Enable verbose logging** – Set `log_memory_every=10` in your training configuration
2. **Instrument suspect blocks** – Wrap video loading, tiling operations, or denoising calls with `free_gpu_memory_context(before=True, after=True, log=True)`
3. **Switch sampling algorithms** – Replace `euler_denoising_loop` with `gradient_estimating_euler_denoising_loop`
4. **Analyze peak memory** – Locate the highest "Current GPU memory" log entry to identify the problematic stage
5. **Apply additional optimizations** – Enable `cfg.optimization.enable_gradient_checkpointing=True` or spatial tiling (`use_tiling=True`)

### Example: Profiling Video Loading

```python
from ltx_trainer.gpu_utils import free_gpu_memory_context, get_gpu_memory_gb
import logging

logger = logging.getLogger(__name__)

def load_chunk(path):
    with free_gpu_memory_context(before=True, after=True, log=True):
        chunk = read_video_file(path)          # heavy I/O

        mem = get_gpu_memory_gb(torch.device('cuda'))
        logger.info(f"Memory after loading {path}: {mem:.2f} GB")
    return chunk

```

### Example: Trainer Memory Profiling Integration

```python
from ltx_trainer.gpu_utils import get_gpu_memory_gb

class Trainer:
    ...
    def _log_memory(self, step):
        mem = get_gpu_memory_gb(self.device)
        self.logger.info(f"[Step {step}] GPU memory: {mem:.2f} GB")
    
    def train(self):
        for step in range(self.total_steps):
            # … training logic …

            if step % self.cfg.logging.log_memory_every == 0:
                self._log_memory(step)

```

## Key Source Files for Memory Debugging

| File | Location | Purpose |
|------|----------|---------|
| [`gpu_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/gpu_utils.py) | [`packages/ltx-trainer/src/ltx_trainer/gpu_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/gpu_utils.py) | Core utilities for memory cleanup and GB-level queries |
| [`samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/samplers.py) | [`packages/ltx-pipelines/src/ltx_pipelines/utils/samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/samplers.py) | Gradient-estimation Euler implementation |
| [`trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/trainer.py) | [`packages/ltx-trainer/src/ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/trainer.py) | Training orchestration with built-in memory logging |
| [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py) | [`packages/ltx-trainer/src/ltx_trainer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py) | Optimization flags: gradient checkpointing, accumulation steps |
| [`process_videos.py`](https://github.com/Lightricks/LTX-2/blob/main/process_videos.py) | [`packages/ltx-trainer/scripts/process_videos.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/scripts/process_videos.py) | Video processing with spatial tiling support |

## Summary

- **`free_gpu_memory` and `free_gpu_memory_context`** provide deterministic CUDA cache cleanup and precise memory measurement
- **`gradient_estimating_euler_denoising_loop`** reduces diffusion steps through velocity correction, lowering peak VRAM requirements
- **Trainer-level profiling** with configurable `log_memory_every` intervals creates actionable memory timelines
- **Iterative workflow** combining instrumentation, algorithm swapping, and configuration tuning eliminates most OOM crashes without quality degradation

## Frequently Asked Questions

### What is the default `ge_gamma` value in gradient-estimation sampling?

The default `ge_gamma` is **2.0**, as defined in the `gradient_estimating_euler_denoising_loop` signature. Increasing this value (e.g., to 2.5) applies more aggressive velocity correction, potentially reducing steps further at the cost of sample stability.

### Where does LTX-2 log peak GPU memory during training?

According to the `Trainer` implementation in [`ltx_trainer/trainer.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/trainer.py), peak memory is logged immediately after model preparation and periodically during training when `step % cfg.logging.log_memory_every == 0`. Both use `torch.cuda.memory_allocated()` divided by `1024**3` to report gigabytes.

### Can I use `free_gpu_memory_context` outside the trainer?

Yes. The context manager is defined in [`ltx_trainer/gpu_utils.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/gpu_utils.py) and has no trainer dependencies. It works in any Python block where you need deterministic CUDA cache clearing—common use cases include dataset loading, inference pipelines, and multi-modal switching operations.

### How does gradient-estimation sampling reduce memory usage?

The velocity-correction mechanism in `gradient_estimating_euler_denoising_loop` reuses previous velocity vectors to reach the target distribution faster. This convergence speedup allows reducing total diffusion steps (e.g., from 50 to 35), which proportionally decreases the number of forward passes and intermediate activations stored in VRAM.