Monolith vs Split Checkpoint Layouts in LTX-2: A Complete Guide

The monolith layout packages all LTX-2 components (transformer, video VAE, audio VAE) into a single .safetensors file, while the split layout separates them into individual checkpoint files for modular flexibility.

LTX-2 supports two distinct ways of organizing model weights and auxiliary components. Understanding the difference between monolith and split checkpoint layouts helps you optimize storage, simplify CLI usage, and enable advanced workflows like VAE swapping. This guide breaks down both approaches using the actual implementation in Lightricks/LTX-2.


Monolith (Unified) Checkpoint Layout

The monolith layout combines the transformer, video VAE, and optional audio VAE into one "fat" checkpoint file.

File Structure

A typical monolith checkpoint follows this naming convention:

  • ltx-2.x-checkpoint.safetensors (e.g., ltx-2.3-22b-dev.safetensors)

The Gemma text encoder remains in a separate folder by design, but all other model parts live inside the single file.

Code Implementation

In packages/ltx-pipelines/src/ltx_pipelines/utils/model_paths.py, the ModelPaths.from_monolith() static method handles this layout:

from ltx_pipelines.utils.model_paths import ModelPaths

paths = ModelPaths.from_monolith(
    ltx_model_path="/models/checkpoints/ltx-2.3-22b-dev.safetensors",
    gemma_root_path="/models/gemma3",
)

The method automatically discovers the internal VAE from checkpoint metadata—no separate path required.

YAML Configuration

model:
  model_path: "/models/checkpoints/ltx-2.3-22b-dev.safetensors"
  gemma_root_path: "/models/gemma3"

When to Use Monolith

  • Default choice for most LTX-2, LTX-2.3, and LTX-2.5 releases
  • When simplicity outweighs storage constraints
  • When you want minimal CLI flags (--model-path only)

Split Checkpoint Layout

The split layout decomposes the model into separate files, enabling component-level flexibility.

File Structure

Component Typical Filename
Transformer ltx-2-transformer.safetensors
Video VAE ltx-2-video_vae.safetensors
Audio VAE (optional) ltx-2-audio_vae.safetensors
Text encoder Separate Gemma folder

Code Implementation

Use ModelPaths constructor directly or explicit split methods:

from ltx_pipelines.utils.model_paths import ModelPaths

# Explicit construction

paths = ModelPaths(
    transformer_path="/models/checkpoints/ltx-2-transformer.safetensors",
    video_vae_path="/models/checkpoints/ltx-2-video_vae.safetensors",
    gemma_root_path="/models/gemma4-ltx-v1",
)

YAML Configuration

model:
  model_path: "/models/checkpoints/ltx-2-transformer.safetensors"
  video_vae_path: "/models/checkpoints/ltx-2-video_vae.safetensors"
  gemma_root_path: "/models/gemma4-ltx-v1"

When to Use Split

  • VAE experimentation: Swap the default VAE for a distilled DiffVAE
  • Storage constraints: Download only components you need
  • Version mixing: Upgrade individual parts without re-downloading everything

Key Differences at a Glance

Aspect Monolith Split
Files Single .safetensors Multiple separate files
CLI flags --model-path only --model-path + --video-vae-path (+ optional --audio-vae-path)
Code entry ModelPaths.from_monolith() ModelPaths() constructor or from_split()
Metadata Single model_type and vcd fields Per-component metadata with cross-validation
Flexibility Fixed bundle Modular replacement

Metadata and Validation

Both layouts rely on checkpoint metadata, but validation differs:

  • Monolith: The model_type field describes architecture; vcd specifies the decoder type (conv vs diffusion). The pipeline extracts VAE settings automatically.

  • Split: Each component carries independent metadata. The pipeline validates consistency— for example, checking that is_diffusion_video_vae matches between transformer and supplied VAE.

According to packages/ltx-pipelines/docs/installation.md, this validation prevents runtime mismatches when mixing components from different LTX-2 versions.


How the Trainer Detects Layout Automatically

The LTX-2 trainer inspects checkpoint metadata to distinguish layouts without explicit user flags. As documented in packages/ltx-trainer/docs/dataset-preparation.md, unified checkpoints are expected by default, but the training pipeline adapts when split components are detected.

This automatic detection means you generally only need to provide correct paths—the framework handles the rest.


Summary

  • Monolith = one file: Simplest configuration, fewer flags, preferred for standard usage
  • Split = multiple files: Enables VAE swapping and granular storage control
  • Both use ModelPaths: The abstraction in model_paths.py unifies access regardless of layout
  • Metadata drives detection: Automatic layout inference eliminates manual configuration

Frequently Asked Questions

Can I convert a monolith checkpoint to split layout?

LTX-2 does not ship with an official conversion utility, but you can extract components programmatically by loading the monolith with ModelPaths.from_monolith() and saving individual state dicts. For production use, Lightricks provides pre-converted split checkpoints for major releases.

Does split layout affect inference speed?

No meaningful performance difference exists between layouts. Both approaches load the same total parameter count into GPU memory; the split layout simply staggers file I/O. Latency differences are negligible compared to model computation time.

Which layout should I choose for fine-tuning?

The trainer accepts both, but packages/ltx-trainer/docs/training-modes.md notes that dataset-preparation scripts expect unified checkpoints by default. Use monolith unless you specifically need to freeze or swap the VAE during training.

Is the audio VAE required for split layout?

No—audio VAE is optional in both layouts. For video-only generation, omit --audio-vae-path (split) or rely on the monolith's embedded video components. The pipeline skips audio processing when the component is unavailable.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →