# How to Reduce Inference Steps with Gradient Estimation in LTX-2 for Faster Video Generation

> Learn how LTX-2 reduces video generation inference steps by 50% using gradient estimation and velocity correction without sacrificing quality. Optimize your diffusion models today.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: performance
- Published: 2026-08-14

---

**LTX-2 reduces diffusion inference steps through gradient-estimation sampling, which augments the standard Euler denoising loop with a velocity-correction term that carries information from consecutive gradient estimates, typically enabling 50% fewer steps without quality loss.**

The LTX-2 audio-video diffusion model from Lightricks implements an advanced inference optimization that directly addresses one of the biggest bottlenecks in generative AI: the number of expensive denoising steps required. By swapping the default `euler_denoising_loop` for `gradient_estimating_euler_denoising_loop`, you can significantly accelerate generation while preserving perceptual fidelity. This guide explains the mechanics, implementation, and tuning of gradient-estimation sampling in the LTX-2 codebase.

## What Is Gradient-Estimation Sampling?

Gradient-estimation sampling is a technique introduced in *"Gradient-Estimation Sampling for Diffusion Models"* (OpenReview) that exploits the smoothness of diffusion trajectories to predict and correct for discretization error. Rather than treating each denoising step in isolation, it maintains a **velocity buffer** that remembers how the latent state changed between consecutive steps.

In [`packages/ltx-pipelines/src/ltx_pipelines/utils/samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/samplers.py), the `gradient_estimating_euler_denoising_loop` function implements this by tracking `previous_video_velocity` and `previous_audio_velocity` across the diffusion schedule. When a previous velocity exists, the algorithm computes a **delta correction** scaled by the `ge_gamma` hyperparameter and applies it to the current denoised estimate.

## How Gradient-Estimation Sampling Works: Step-by-Step

The gradient-estimation loop follows a precise six-stage process that differentiates it from standard Euler sampling:

### 1. Initialize Velocity Buffers

Two state-level buffers store velocity history across modalities:

```python
previous_audio_velocity = None
previous_video_velocity = None

```

These are declared in [`samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/samplers.py) at lines 106-108 and remain `None` until the second denoising step when sufficient history exists.

### 2. Run the Denoiser

At each step, the denoiser returns denoised latents for video and audio:

```python
video_result, audio_result = denoiser(
    transformer=transformer,
    video_state=video_state,
    audio_state=audio_state,
    sigma=sigma,
)

```

This occurs at lines 120-122 in [`samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/samplers.py), using the same `denoiser` callable employed in training.

### 3. Compute Current Velocity

The `to_velocity` utility converts noisy samples and sigma values into velocity vectors:

```python
current_velocity = to_velocity(noisy_sample, sigma, denoised_sample)

```

Implemented at lines 111-113, this represents the instantaneous direction of latent change.

### 4. Apply Gradient-Estimation Correction

When previous velocity exists, the correction activates:

```python
delta_v = current_velocity - previous_velocity
total_velocity = ge_gamma * delta_v + previous_velocity
denoised_sample = to_denoised(noisy_sample, total_velocity, sigma)

```

This is the core innovation at lines 114-117: the **extrapolated velocity** (`total_velocity`) combines current measurement with a momentum-like term from previous steps. The default `ge_gamma=2.0` weights this extrapolation aggressively.

### 5. Advance the Latent State

The corrected denoised latent feeds into the ODE stepper:

```python
video_state = replace(video_state, latent=stepper.step(
    sigma_from=sigma,
    sigma_to=next_sigma,
    denoised_x=denoised_sample,
    current_x=noisy_sample,
))

```

Lines 141-143 in [`samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/samplers.py) show this state update using the `dataclasses.replace` pattern for immutable state management.

### 6. Iterate Over Shortened Schedule

The loop header at lines 119-121 processes `sigmas[:-1]`, but critically, **this schedule can be truncated** because the velocity correction compensates for larger step sizes:

```python
for step_idx, _ in enumerate(tqdm(sigmas[:-1])):
    # ... denoising logic

```

Typical configurations reduce 50-step schedules to 25-30 steps without visible degradation.

## Implementation: Enabling Gradient-Estimation Sampling

To activate gradient-estimation sampling in your LTX-2 pipeline, swap the sampler function and adjust the schedule length:

### Complete Working Example

```python
import torch
from ltx_pipelines.utils.samplers import gradient_estimating_euler_denoising_loop
from ltx_pipelines.utils.types import LatentState
from ltx_core.components.diffusion_steps import EulerAncestralDiffusionStep
from ltx_core.model.transformer import X0Model

# Custom denoiser implementing the Denoiser protocol

from my_project.denoiser import my_denoiser

# 1. Prepare shortened diffusion schedule

# Standard: 50 steps → Gradient-estimation: 25-30 steps

sigmas = torch.logspace(start=0, end=-6, steps=25, base=10.0)

# 2. Initialize video latent state

video_state = LatentState(
    latent=torch.randn(1, 3, 64, 64, device="cuda"),
    denoise_mask=torch.ones(1, 1, 64, 64, device="cuda"),
    clean_latent=None,
)

# Audio optional — set to None if not generating audio

audio_state = None

# 3. Select ODE stepper

stepper = EulerAncestralDiffusionStep()

# 4. Load transformer backbone

transformer = X0Model.from_pretrained("ltx-base")

# 5. Execute gradient-estimation sampling loop

final_video, final_audio = gradient_estimating_euler_denoising_loop(
    sigmas=sigmas,
    video_state=video_state,
    audio_state=audio_state,
    stepper=stepper,
    transformer=transformer,
    denoiser=my_denoiser,
    ge_gamma=2.0,  # Default; increase for more aggressive speedups

)

```

### Key Implementation Details

- **Schedule truncation**: The `steps=25` in `logspace` directly cuts computation in half compared to standard Euler.
- **No model retraining**: The `denoiser` callable remains identical to training; only the sampling algorithm changes.
- **Modality flexibility**: Pass `audio_state=None` for video-only generation without code modification.

## Tuning the `ge_gamma` Hyperparameter

The `ge_gamma` parameter (default 2.0) controls the strength of gradient-estimation correction:

| `ge_gamma` Value | Behavior | Use Case |
|------------------|----------|----------|
| **1.0** | Minimal correction; conservative but stable | Maximum quality priority, uncertain schedules |
| **2.0** | Balanced extrapolation (default) | Recommended starting point |
| **3.0-4.0** | Aggressive correction; larger effective steps | Speed-critical applications, tolerant quality review |

Modify `ge_gamma` in the loop call:

```python
final_video, _ = gradient_estimating_euler_denoising_loop(
    # ... other arguments

    ge_gamma=3.0,  # More aggressive speedup

)

```

Higher values enable shorter schedules but may introduce temporal flickering or audio-video sync artifacts in extreme cases.

## When to Use Gradient-Estimation Sampling

**Inference-only scenarios** — This technique applies exclusively to the sampling phase. Training continues with standard schedules and `euler_denoising_loop`.

**Latency-critical deployments** — Real-time or interactive applications benefit most. Typical speedups include:
- 50 → 25 steps: **2× faster inference**
- Proportional GPU memory reduction from shorter sequence lengths

**Consumer hardware constraints** — Shorter sequences reduce peak VRAM usage, enabling LTX-2 on GPUs that would otherwise OOM with full step counts.

**Avoid when**: You require bit-exact reproducibility with standard Euler outputs, or when working with highly irregular diffusion schedules where velocity assumptions break down.

## Source Code Reference Map

Understanding the file structure helps with debugging and customization:

| File | Purpose | Key Symbols |
|------|---------|-------------|
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/samplers.py) | Core sampling loops | `gradient_estimating_euler_denoising_loop`, `euler_denoising_loop` |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/types.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/types.py) | Protocol definitions | `LatentState`, `Denoiser`, `DiffusionStepProtocol` |
| [`packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) | Post-processing utilities | `post_process_latent`, `to_velocity`, `to_denoised` |
| [`packages/ltx-core/src/ltx_core/components/diffusion_steps.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/diffusion_steps.py) | ODE step implementations | `EulerAncestralDiffusionStep`, `EulerDiffusionStep` |
| [`packages/ltx-core/src/ltx_core/model/transformer.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/model/transformer.py) | Neural network backbone | `X0Model.forward` (called by denoiser) |

The velocity conversion utilities in [`helpers.py`](https://github.com/Lightricks/LTX-2/blob/main/helpers.py) are particularly important—these implement the mathematical transformation between noise-space and velocity-space representations that underlie the gradient-estimation correction.

## Common Issues and Debugging

**Velocity buffer not populating**: Ensure the loop runs at least 2 steps; correction only activates on step ≥ 2 when `previous_velocity` is non-`None`.

**Artifacts with high `ge_gamma`**: Reduce step count more gradually or lower `ge_gamma` to 1.5. The optimal trade-off is content-dependent.

**Audio-video desynchronization**: Verify both modalities have velocity buffers initialized. The `LatentState` for audio must not be `None` if audio generation is intended.

## Summary

- **Gradient-estimation sampling** in LTX-2 reduces inference steps by tracking velocity history and extrapolating corrections between consecutive denoising steps.

- **Implementation** requires only swapping `gradient_estimating_euler_denoising_loop` for the standard Euler loop and optionally shortening the sigma schedule.

- **`ge_gamma=2.0`** provides a reliable default; increase for speed, decrease for stability.

- **No model changes** are needed—the same trained weights and denoiser work with both sampling methods.

- **Source locations**: [`samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/samplers.py) contains the loop (lines 106-143), [`types.py`](https://github.com/Lightricks/LTX-2/blob/main/types.py) defines state containers, and [`diffusion_steps.py`](https://github.com/Lightricks/LTX-2/blob/main/diffusion_steps.py) provides ODE integrators.

## Frequently Asked Questions

### What is the minimum number of steps with gradient-estimation sampling?

Practical deployments typically use 25-30 steps, roughly half the 50 steps required for standard Euler sampling in LTX-2. Below 20 steps, quality degradation becomes noticeable regardless of correction strength. The theoretical limit depends on content complexity and `ge_gamma` tuning.

### Does gradient-estimation sampling affect training or require fine-tuning?

No. The technique is **inference-only** and operates purely through modified sampling dynamics. The same model weights trained with standard Euler sampling work without any adjustment. Training code in `ltx-core` continues to use the original diffusion schedule and `euler_denoising_loop`.

### How does `ge_gamma` relate to other diffusion samplers like DPM++?

`ge_gamma` serves a similar purpose to the predictor-corrector steps in DPM++ or the momentum terms in Heun sampling—it compensates for discretization error. However, gradient-estimation specifically exploits the temporal smoothness of velocity estimates rather than higher-order derivatives of the score function. The two approaches could theoretically combine but are not implemented together in LTX-2.

### Can I use gradient-estimation sampling with other ODE steppers besides Euler?

The current implementation in [`samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/samplers.py) is bound to `EulerAncestralDiffusionStep` and its variants through the `stepper.step` interface. Extending to other `DiffusionStepProtocol` implementations (such as implicit Adams methods) would require modifying the velocity correction logic to match the stepper's mathematical formulation.