How to Configure SingleGPUModelBuilder for Loading LTX-2 Model Checkpoints

Use the immutable SingleGPUModelBuilder class to construct LTX-2 models on a single GPU by specifying a model_class_configurator, checkpoint path, and optional LoRA adapters, then calling build() to fuse weights and return a ready-to-use PyTorch module.

Configuring the SingleGPUModelBuilder for loading LTX-2 model checkpoints requires understanding its immutable builder pattern and core parameters. Located in packages/ltx-core/src/ltx_core/loader/single_gpu_model_builder.py within the Lightricks/LTX-2 repository, this class handles safetensors deserialization, meta-model instantiation, and LoRA weight fusion through a fluent API that returns a shallow copy on each configuration call.

Understanding the SingleGPUModelBuilder Architecture

The SingleGPUModelBuilder follows a strict builder pattern where every configuration method returns a new instance, leaving the original untouched. This immutability ensures thread safety and configuration reusability across different model loading scenarios.

Core Configuration Components

The builder relies on several key components defined during instantiation or chained configuration:

  • model_class_configurator: A class such as LtxModelConfigurator that transforms model configuration dictionaries into concrete PyTorch modules (used in meta_model() at lines 199-202).
  • model_path: File path or tuple of shard paths pointing to .safetensors checkpoint files containing base weights (accessed in _load_model_weights() at lines 44-47 and model_config() at lines 198-199).
  • module_ops: A sequence of module-level mutations applied to the meta model before weight loading, such as ReplaceModuleOp (stored in _module_ops and applied in build() at lines 202-203).
  • loras: Optional LoRA adapters specified as paths with strength values and optional sd_ops, fused into the base model during loading (managed by lora(), with_loras(), and consumed in _load_model_weights() at lines 70-84).
  • model_loader: Strategy for reading safetensors state dictionaries, defaulting to SafetensorsModelStateDictLoader (initialized at lines 115-126).
  • registry: Caches loaded state dictionaries to prevent redundant disk reads, defaulting to DummyRegistry (initialized at lines 126-128).
  • lora_load_device: Device for initial LoRA tensor loading, defaulting to cpu to conserve GPU memory (set at lines 127-129 and passed to _load_model_weights()).
  • fuse_rule: Policy defining LoRA weight merging mechanics, defaulting to bf16_fuse_rule (set at lines 128-130 and used in _load_model_weights() at lines 74-78).

The Build Execution Flow

When you invoke build(device, dtype), the builder executes a precise sequence of operations:

  1. Device Resolution: Defaults to the first available CUDA device if none is specified (lines 15-16).
  2. Configuration Reading: Parses the model config via read_model_config() (lines 198-199).
  3. Meta Model Creation: Instantiates an un-materialized model on the meta device via create_meta_model() (lines 202-203).
  4. Weight Loading: Populates parameters from safetensors files and fuses LoRA adapters if present (_load_model_weights() at lines 44-84).
  5. Verification: Ensures no parameters remain on the meta device (_check_uninitialized() at lines 32-41).
  6. Device Transfer: Moves the fully initialized model to the target GPU (meta_model.to(device) at line 36).

Basic Configuration Examples

Loading a Base Checkpoint

The simplest configuration requires only the model configurator and checkpoint path:

from ltx_core.loader.single_gpu_model_builder import SingleGPUModelBuilder
from ltx_core.model.model_protocol import LtxModelConfigurator

builder = SingleGPUModelBuilder(
    model_class_configurator=LtxModelConfigurator,
    model_path="models/ltx2_v1.safetensors"
)

model = builder.build()

This configuration loads the model onto the default CUDA device with bf16 precision, reading the architecture definition from the configurator and weights from the specified safetensors file.

Specifying Target Device and Precision

Override the default device and data type for specific hardware configurations:

builder = SingleGPUModelBuilder(
    model_class_configurator=LtxModelConfigurator,
    model_path="models/ltx2_v1.safetensors"
)

model = builder.build(device="cuda:1", dtype=torch.float16)

Use this pattern for mixed-precision inference or when distributing multiple models across different GPUs.

Advanced Configuration Patterns

Injecting Custom Module Operations

Apply structural modifications to the model before weight loading using with_module_ops():

from ltx_core.loader.module_ops import ReplaceModuleOp
from my_custom_modules import MyAttention

replace_attention = ReplaceModuleOp(
    pattern=".*attention.*",
    new_module=MyAttention
)

builder = (
    SingleGPUModelBuilder(LtxModelConfigurator, "models/ltx2_v1.safetensors")
    .with_module_ops((replace_attention,))
)

model = builder.build()

The ReplaceModuleOp matches modules using regex patterns and substitutes them with custom implementations before the checkpoint weights are Loaded.

Loading and Fusing LoRA Adapters

Chain multiple LoRA configurations to fuse style adapters at different strengths:

from ltx_core.loader.sd_ops import SDOps

builder = (
    SingleGPUModelBuilder(LtxModelConfigurator, "models/ltx2_v1.safetensors")
    .lora("loras/style.safetensors", strength=0.7, sd_ops=SDOps())
    .lora("loras/color.safetensors", strength=0.3, sd_ops=SDOps())
)

model = builder.build()

Each lora() call returns a new builder instance. During build(), the adapters are fused on-the-fly using the configured fuse_rule, ensuring only the final merged weights occupy GPU memory.

Optimizing LoRA Loading Performance

Configure the loading device and fusion rules to optimize memory usage:

builder = (
    SingleGPUModelBuilder(LtxModelConfigurator, "models/ltx2_v1.safetensors")
    .with_lora_load_device("cpu")
    .with_fuse_rule(custom_fuse_rule)
)

Setting lora_load_device to cpu keeps GPU memory available during the loading phase, while custom fuse rules control how adapters merge with base weights.

Key Implementation Files

Understanding these source files provides deeper insight into the loading mechanics:

  • single_gpu_model_builder.py: Contains the core immutable builder logic, including the build() method and weight loading orchestration (lines 1-210).
  • helpers.py: Implements create_meta_model() and read_model_config() for meta-model instantiation and configuration parsing.
  • module_ops.py: Defines base classes for module mutations like ReplaceModuleOp used in pre-loading transformations.
  • fuse_loras.py: Implements LoRA fusion policies including bf16_fuse_rule and the apply_loras function.
  • sft_loader.py: Provides SafetensorsModelStateDictLoader for efficient checkpoint deserialization.
  • model_protocol.py: Defines abstract interfaces for model configurators that translate config dictionaries to PyTorch modules.

Summary

  • SingleGPUModelBuilder uses an immutable builder pattern where each configuration method returns a shallow copy, enabling reusable configuration templates.
  • Core parameters include model_class_configurator for architecture definition, model_path for checkpoint location, and optional module_ops for structural modifications.
  • LoRA integration happens during the build phase through chained lora() calls, with weights fused according to the specified fuse_rule and loaded to lora_load_device before GPU transfer.
  • The build process automatically handles meta-model creation, weight verification, and device placement, returning a fully initialized PyTorch module ready for inference.

Frequently Asked Questions

What is the difference between SingleGPUModelBuilder and direct checkpoint loading?

SingleGPUModelBuilder provides a structured, configuration-driven approach that handles complex scenarios like LoRA fusion, module replacement, and memory optimization automatically. Direct loading requires manual implementation of state dict deserialization, device placement, and adapter merging that the builder handles internally through its _load_model_weights() and _check_uninitialized() methods.

How do I load multiple LoRA adapters with different strengths?

Chain multiple .lora() calls before invoking build(), specifying the strength parameter for each adapter. The builder fuses all specified adapters sequentially during the weight loading phase at lines 70-84 of single_gpu_model_builder.py, applying the strength multipliers as defined in the fuse_loras.py implementation.

Why does SingleGPUModelBuilder use the builder pattern?

The immutable builder pattern ensures that configuration objects remain thread-safe and reusable. Each method like with_module_ops() or lora() returns a shallow copy with the new setting, allowing you to create base configurations and derive specialized variants without side effects, as implemented in the constructor and chaining methods throughout the source file.

How do I verify that all model weights loaded correctly?

The builder automatically calls _check_uninitialized() (lines 32-41) after weight loading to verify that no parameters remain on the meta device. If any parameters were not populated from the checkpoint or LoRA fusion, this method raises an exception before the model moves to the target GPU, ensuring complete initialization.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →