# LTX-2 Pipeline Architecture: How DiffusionStage, VideoDecoder, AudioDecoder, and PromptEncoder Work Together

> Explore the LTX-2 pipeline architecture. Learn how DiffusionStage, VideoDecoder, AudioDecoder, and PromptEncoder work together for efficient memory usage and stateless reuse. Discover the build-use-free lifecycle.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: architecture
- Published: 2026-08-18

---

**LTX-2's generation pipeline is built from four self-contained building blocks—DiffusionStage, VideoDecoder, AudioDecoder, and PromptEncoder—that follow a strict "build-use-free" lifecycle to maximize memory efficiency and enable stateless reuse across calls.**

The LTX-2 video generation system, developed by Lightricks, implements a modular architecture where each pipeline component owns its model lifecycle independently. Rather than maintaining persistent GPU-resident models, every block reconstructs its underlying neural network on demand and releases weights immediately after use. This design eliminates central model registries and makes the pipeline naturally compatible with multi-GPU streaming, quantization, and LoRA-based fine-tuning.

## The Four Core Building Blocks of LTX-2

All four pipeline blocks are defined in [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) and follow identical construction patterns using specialized builder classes. Each block manages either a transformer, a VAE decoder, or a text encoder from the underlying `ltx_core` package.

### DiffusionStage: The Transformer Lifecycle Manager

**DiffusionStage** owns the transformer—the core diffusion model that performs iterative denoising. In [`blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/blocks.py) (lines 53–120), this class encapsulates checkpoint loading, optional LoRA injection, quantization, compilation, and the complete on-call lifecycle.

Construction begins with `DiffusionStage.from_checkpoint()`, which instantiates a builder configured from checkpoint paths and hyperparameters. The class supports several optimization modes:

- **Streaming/off-load mode** – activated when `_is_streaming=True`, wraps the model in `_streaming_model` context manager
- **Quantization** – applied via `_chain_quantization()` for reduced precision inference
- **Compilation** – enabled through `_apply_compile_ops()` for kernel fusion

During the forward pass, `DiffusionStage.__call__()` builds the transformer, executes the denoising loop across sigma timesteps, and automatically frees GPU weights via `gpu_model` or `_streaming_model` context managers. This ensures the transformer exists only during active generation.

```python
from ltx_pipelines.utils.blocks import DiffusionStage

# Build from checkpoint with optional optimizations

stage = DiffusionStage.from_checkpoint(
    checkpoint_path="/path/to/transformer.pt",
    dtype=torch.bfloat16,
    device=torch.device("cuda"),
).with_loras(lora_weights)           # optional LoRA injection

  .with_attention("flash-attn")      # optional attention backend

# Execute diffusion (model built, used, and freed automatically)

video_state, audio_state = stage(
    denoiser=my_denoiser,
    sigmas=torch.linspace(1.0, 0.01, 30),
    noiser=my_noiser,
    width=512,
    height=768,
    frames=30,
    video={"conditionings": [], "noise_scale": 1.0},
    audio={"conditionings": [], "noise_scale": 1.0},
)

```

### VideoDecoder: Latent-to-Pixel Conversion

**VideoDecoder** (defined at `blocks.py#L1048`) reconstructs video frames from diffusion-produced latents. It dynamically selects between convolutional and diffusion VAE decoders based on checkpoint metadata, specifically the `is_diffusion_video_vae` field.

The decoder configuration flows through `VideoDecoderConfigurator`, which instantiates the appropriate decoder class from [`ltx_core/model/video_vae/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/video_vae.py). Key capabilities include:

- **Decoder selection** – Conv VAE for standard checkpoints, Diffusion VAE when `is_diffusion_video_vae=True`
- **DSL acceleration** – native attention kernels when `na_dsl_available` and compatible hardware detected
- **Memory-efficient decoding** – enabled via `memory_efficient=True`, adds `CHANNELS_LAST_3D_WEIGHTS` and `MEMORY_EFFICIENT_DECODE` flags

The `__call__` method returns an iterator that yields decoded frame tensors. Critical to the architecture: the iterator holds the built decoder model, and exhaustion triggers automatic GPU memory release. This streaming pattern supports arbitrarily long videos without proportional memory growth.

```python
from ltx_pipelines.utils.blocks import VideoDecoder

video_decoder = VideoDecoder(
    checkpoint_path="/path/to/video_vae.pt",
    dtype=torch.bfloat16,
    device=torch.device("cuda"),
    memory_efficient=True,    # enable tiling for large resolutions

)

# Iterator pattern: decoder built on first __next__, freed on StopIteration

for frame_tensor in video_decoder(video_latent):
    # Process frame (H, W, C) tensor

    save_frame(frame_tensor)

```

### AudioDecoder: Spectrogram Reconstruction

**AudioDecoder** (class definition at `ltx_core/model/audio_vae/audio_vae.py#L77`) mirrors the video decoder structure for audio latents. It reconstructs time-domain waveforms or spectrograms from diffusion-generated audio representations.

Configuration occurs through `AudioDecoderConfigurator` in [`model_configurator.py`](https://github.com/Lightricks/LTX-2/blob/main/model_configurator.py). The architecture implements:

- **Per-channel statistics** – `PerChannelStatistics` handles de-normalization of latent channels
- **Causal convolution support** – `CausalityAxis.WIDTH` enables streaming audio generation
- **Attention blocks** – optional transformer layers in the decoder path for high-fidelity reconstruction

The forward pass executes a fixed pipeline: `conv_in → mid_block → upsampling → final_conv → shape_adjustment`. The `AudioPatchifier` class manages latent patching and unpatching operations.

```python
from ltx_pipelines.utils.blocks import AudioDecoder

audio_decoder = AudioDecoder(
    checkpoint_path="/path/to/audio_vae.pt",
    dtype=torch.bfloat16,
    device=torch.device("cuda"),
)

# Direct tensor return (no iterator for audio)

audio_waveform = audio_decoder(audio_latent)  # shape: (batch, channels, samples)

```

### PromptEncoder: Text-to-Embedding Pipeline

**PromptEncoder** (`blocks.py#L1029`) handles the complete text conditioning path. It loads a Gemma-based text encoder, optionally runs prompt enhancement, and feeds hidden states through an embeddings processor.

The class manages three distinct model builders:

- **`_text_encoder_builder`** – the primary Gemma encoder (`ltx_core/text_encoders/gemma/`)
- **`_enhancer_text_encoder_builder`** – optional prompt enhancement model
- **`_streaming_text_encoder_builder`** – off-load variant for memory-constrained scenarios

The `__call__` method orchestrates: (1) optional enhancement of the first prompt via the enhancer model, (2) encoding of all prompts through the main Gemma encoder, (3) processing of hidden states through `EmbeddingsProcessorConfigurator` to produce final conditioning embeddings.

```python
from ltx_pipelines.utils.blocks import PromptEncoder

prompt_encoder = PromptEncoder(
    model_paths=paths,           # ModelPaths instance with checkpoint locations

    dtype=torch.bfloat16,
    device=torch.device("cuda"),
)

embeddings = prompt_encoder(
    prompts=["A cinematic drone shot of ocean waves"],
    enhance_first_prompt=True,   # run through enhancement model

)

# Returns processed embeddings ready for diffusion injection

```

## How LTX-2 Pipeline Blocks Orchestrate Generation

The four blocks connect in a strict data-flow sequence that maps user prompts to final media output:

1. **PromptEncoder** converts raw text into embedding tensors compatible with the transformer's cross-attention layers
2. **DiffusionStage** consumes embeddings, builds the transformer, runs iterative denoising, and emits separate video and audio latent states
3. **VideoDecoder** and **AudioDecoder** independently reconstruct their respective modalities from latent representations

This architecture enforces **stateless operation**: no block retains model weights between calls. Each invocation rebuilds its neural network from checkpoint, executes computation, and frees parameters to the meta device. The pattern eliminates explicit memory management and enables:

- **Interleaved media generation** – video and audio decoding proceed in parallel without weight duplication
- **Dynamic checkpoint swapping** – different LoRAs, VAE architectures, or precision modes per generation
- **Streaming inference** – model shards load and unload across pipeline stages for limited GPU memory

## Source File Reference Map

| Block | Primary Source | Model Implementation |
|-------|---------------|----------------------|
| DiffusionStage | [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) (lines 53–120) | `ltx_core/model/transformer/` |
| VideoDecoder | [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) (lines 1048–1130) | [`ltx_core/model/video_vae/video_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/video_vae/video_vae.py) |
| AudioDecoder | [`ltx_core/model/audio_vae/audio_vae.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/model/audio_vae/audio_vae.py) (line 77+) | `ltx_core/model/audio_vae/` |
| PromptEncoder | [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) (lines 1029–1076) | `ltx_core/text_encoders/gemma/` |

Builder configurators reside in [`model_configurator.py`](https://github.com/Lightricks/LTX-2/blob/main/model_configurator.py) files for each modality and in [`gemma_assets.py`](https://github.com/Lightricks/LTX-2/blob/main/gemma_assets.py) for text encoder assets.

## Complete Pipeline Assembly Example

The following runnable script demonstrates end-to-end block coordination:

```python
import torch
from ltx_pipelines.utils.blocks import (
    DiffusionStage, VideoDecoder, AudioDecoder, PromptEncoder
)
from ltx_pipelines.utils.model_paths import ModelPaths

# Unified checkpoint management

paths = ModelPaths.from_monolith(
    monolith_path="/path/to/ltx2_monolith.pt"
)

# 1. Encode conditioning

prompt_encoder = PromptEncoder(
    model_paths=paths,
    dtype=torch.bfloat16,
    device=torch.device("cuda"),
)
embeddings = prompt_encoder(
    prompts=["A robot dancing in a neon-lit warehouse"],
    enhance_first_prompt=True,
)

# 2. Run diffusion

stage = DiffusionStage.from_checkpoint(
    checkpoint_path=paths.transformer(),
    dtype=torch.bfloat16,
    device=torch.device("cuda"),
)
video_state, audio_state = stage(
    denoiser=lambda *a, **kw: a,
    sigmas=torch.linspace(1.0, 0.01, 40),
    noiser=lambda *a, **kw: a,
    width=640,
    height=480,
    frames=45,
    video={"conditionings": [embeddings], "noise_scale": 1.0},
    audio={"conditionings": [], "noise_scale": 1.0},
)

# 3. Decode outputs

video_decoder = VideoDecoder(
    checkpoint_path=paths.video_vae(),
    dtype=torch.bfloat16,
    device=torch.device("cuda"),
    memory_efficient=True,
)

audio_decoder = AudioDecoder(
    checkpoint_path=paths.audio_vae(),
    dtype=torch.bfloat16,
    device=torch.device("cuda"),
)

# Collect frames and audio (models auto-freed after use)

frames = list(video_decoder(video_state.latent))
audio = audio_decoder(audio_state.latent)

```

## Summary

- **DiffusionStage**, **VideoDecoder**, **AudioDecoder**, and **PromptEncoder** form the complete LTX-2 pipeline architecture, each defined in [`ltx_pipelines/utils/blocks.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_pipelines/utils/blocks.py) with model implementations in `ltx_core`
- All blocks follow identical **builder-construction** patterns and **on-call lifecycles** via `gpu_model`/`_streaming_model` context managers
- **Stateless design** eliminates persistent GPU memory usage—models rebuild and release on every generation
- **Modular configuration** enables per-call customization of checkpoints, LoRAs, quantization, and attention backends without pipeline restructuring
- File paths and class names from the source analysis: `DiffusionStage.from_checkpoint()`, `VideoDecoderConfigurator`, `AudioDecoder` at `audio_vae.py#L77`, `PromptEncoder` with `EmbeddingsProcessorConfigurator`

## Frequently Asked Questions

### What makes LTX-2 pipeline blocks "stateless"?

Each block reconstructs its underlying neural network from checkpoint on every call and releases GPU weights immediately after use via context managers like `gpu_model` and `_streaming_model`. No model parameters persist between pipeline invocations, which eliminates memory leaks and enables dynamic checkpoint swapping without process restart.

### How does DiffusionStage support LoRA and quantization?

The `DiffusionStage` builder pattern exposes chainable configuration methods: `.with_loras()` injects low-rank adaptation weights, `.with_attention()` selects attention implementations (flash-attention, DSL kernels), and `_chain_quantization()` applies precision reduction. These modifiers update the builder state before `__call__()` constructs the final optimized transformer.

### Why does VideoDecoder return an iterator instead of a tensor?

The iterator pattern enables **streaming video generation** where the decoder model exists only during active frame production. Each `__next__()` call decodes a temporal chunk, yields frames, and can offload completed segments. Iterator exhaustion triggers automatic model destruction, supporting arbitrarily long videos without proportional VRAM growth—critical for memory-efficient decoding with `MEMORY_EFFICIENT_DECODE` enabled.

### Can AudioDecoder and VideoDecoder run with different precisions or devices?

Yes. Each block is fully independent with its own builder configuration. You can instantiate `VideoDecoder` with `torch.bfloat16` on `cuda:0` and `AudioDecoder` with `torch.float16` on `cuda:1`, or quantize one decoder while keeping the other in full precision. The stateless architecture ensures these configurations don't interfere across pipeline stages.