# Monolith vs Split Checkpoint Layouts in LTX-2: A Complete Guide

> Understand monolith vs split checkpoint layouts in LTX-2. Discover how to choose the best option for your project, balancing ease of use with modular flexibility.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-20

---

**The monolith layout packages all LTX-2 components (transformer, video VAE, audio VAE) into a single `.safetensors` file, while the split layout separates them into individual checkpoint files for modular flexibility.**

LTX-2 supports two distinct ways of organizing model weights and auxiliary components. Understanding the difference between monolith and split checkpoint layouts helps you optimize storage, simplify CLI usage, and enable advanced workflows like VAE swapping. This guide breaks down both approaches using the actual implementation in `Lightricks/LTX-2`.

---

## Monolith (Unified) Checkpoint Layout

The **monolith layout** combines the transformer, video VAE, and optional audio VAE into one "fat" checkpoint file.

### File Structure

A typical monolith checkpoint follows this naming convention:

- `ltx-2.x-checkpoint.safetensors` (e.g., `ltx-2.3-22b-dev.safetensors`)

The Gemma text encoder remains in a separate folder by design, but all other model parts live inside the single file.

### Code Implementation

In [`packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py), the `ModelPaths.from_monolith()` static method handles this layout:

```python
from ltx_pipelines.utils.model_paths import ModelPaths

paths = ModelPaths.from_monolith(
    ltx_model_path="/models/checkpoints/ltx-2.3-22b-dev.safetensors",
    gemma_root_path="/models/gemma3",
)

```

The method automatically discovers the internal VAE from checkpoint metadata—no separate path required.

### YAML Configuration

```yaml
model:
  model_path: "/models/checkpoints/ltx-2.3-22b-dev.safetensors"
  gemma_root_path: "/models/gemma3"

```

### When to Use Monolith

- **Default choice** for most LTX-2, LTX-2.3, and LTX-2.5 releases
- When simplicity outweighs storage constraints
- When you want minimal CLI flags (`--model-path` only)

---

## Split Checkpoint Layout

The **split layout** decomposes the model into separate files, enabling component-level flexibility.

### File Structure

| Component | Typical Filename |
|-----------|----------------|
| Transformer | `ltx-2-transformer.safetensors` |
| Video VAE | `ltx-2-video_vae.safetensors` |
| Audio VAE (optional) | `ltx-2-audio_vae.safetensors` |
| Text encoder | Separate Gemma folder |

### Code Implementation

Use `ModelPaths` constructor directly or explicit split methods:

```python
from ltx_pipelines.utils.model_paths import ModelPaths

# Explicit construction

paths = ModelPaths(
    transformer_path="/models/checkpoints/ltx-2-transformer.safetensors",
    video_vae_path="/models/checkpoints/ltx-2-video_vae.safetensors",
    gemma_root_path="/models/gemma4-ltx-v1",
)

```

### YAML Configuration

```yaml
model:
  model_path: "/models/checkpoints/ltx-2-transformer.safetensors"
  video_vae_path: "/models/checkpoints/ltx-2-video_vae.safetensors"
  gemma_root_path: "/models/gemma4-ltx-v1"

```

### When to Use Split

- **VAE experimentation**: Swap the default VAE for a distilled DiffVAE
- **Storage constraints**: Download only components you need
- **Version mixing**: Upgrade individual parts without re-downloading everything

---

## Key Differences at a Glance

| Aspect | Monolith | Split |
|--------|----------|-------|
| **Files** | Single `.safetensors` | Multiple separate files |
| **CLI flags** | `--model-path` only | `--model-path` + `--video-vae-path` (+ optional `--audio-vae-path`) |
| **Code entry** | `ModelPaths.from_monolith()` | `ModelPaths()` constructor or `from_split()` |
| **Metadata** | Single `model_type` and `vcd` fields | Per-component metadata with cross-validation |
| **Flexibility** | Fixed bundle | Modular replacement |

---

## Metadata and Validation

Both layouts rely on checkpoint metadata, but validation differs:

- **Monolith**: The `model_type` field describes architecture; `vcd` specifies the decoder type (conv vs diffusion). The pipeline extracts VAE settings automatically.

- **Split**: Each component carries independent metadata. The pipeline validates consistency— for example, checking that `is_diffusion_video_vae` matches between transformer and supplied VAE.

According to [`packages/ltx-pipelines/docs/installation.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/docs/installation.md), this validation prevents runtime mismatches when mixing components from different LTX-2 versions.

---

## How the Trainer Detects Layout Automatically

The LTX-2 trainer inspects checkpoint metadata to distinguish layouts without explicit user flags. As documented in [`packages/ltx-trainer/docs/dataset-preparation.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/docs/dataset-preparation.md), unified checkpoints are expected by default, but the training pipeline adapts when split components are detected.

This automatic detection means you generally only need to provide correct paths—the framework handles the rest.

---

## Summary

- **Monolith = one file**: Simplest configuration, fewer flags, preferred for standard usage
- **Split = multiple files**: Enables VAE swapping and granular storage control
- **Both use `ModelPaths`**: The abstraction in [`model_paths.py`](https://github.com/Lightricks/LTX-2/blob/main/model_paths.py) unifies access regardless of layout
- **Metadata drives detection**: Automatic layout inference eliminates manual configuration

---

## Frequently Asked Questions

### Can I convert a monolith checkpoint to split layout?

LTX-2 does not ship with an official conversion utility, but you can extract components programmatically by loading the monolith with `ModelPaths.from_monolith()` and saving individual state dicts. For production use, Lightricks provides pre-converted split checkpoints for major releases.

### Does split layout affect inference speed?

No meaningful performance difference exists between layouts. Both approaches load the same total parameter count into GPU memory; the split layout simply staggers file I/O. Latency differences are negligible compared to model computation time.

### Which layout should I choose for fine-tuning?

The trainer accepts both, but [`packages/ltx-trainer/docs/training-modes.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/docs/training-modes.md) notes that dataset-preparation scripts expect unified checkpoints by default. Use monolith unless you specifically need to freeze or swap the VAE during training.

### Is the audio VAE required for split layout?

No—audio VAE is optional in both layouts. For video-only generation, omit `--audio-vae-path` (split) or rely on the monolith's embedded video components. The pipeline skips audio processing when the component is unavailable.