How to Implement GuidedDenoiser vs SimpleDenoiser for Custom Denoisers in LTX-2

GuidedDenoiser applies classifier-free guidance (CFG) and self-teacher guidance (STG) through batched multi-pass transformer calls, while SimpleDenoiser executes a single conditional pass for maximum inference speed, and both implement the Denoiser protocol defined in ltx_pipelines/utils/types.py.

Lightricks/LTX-2 ships with three production-ready denoiser implementations that cater to different inference requirements. Understanding how to implement GuidedDenoiser vs SimpleDenoiser for custom denoisers in LTX-2 allows you to balance generation quality against computational overhead, whether you need unconditional fast sampling or fine-grained guidance control.

Understanding the Built-in Denoiser Architecture

LTX-2 organizes its denoising logic around a strict protocol-based design. All denoisers reside in ltx_pipelines/utils/denoisers.py and conform to the Denoiser protocol defined in ltx_pipelines/utils/types.py (lines 73-88).

The Denoiser Protocol

Every denoiser must implement the __call__ method with the following signature:

def __call__(
    self,
    transformer: X0Model,
    video_state: LatentState | None,
    audio_state: LatentState | None,
    sigmas: torch.Tensor,
    step_index: int,
) -> Tuple[DenoisedLatentResult | None, DenoisedLatentResult | None]:

This protocol ensures interoperability with LTX-2 diffusion pipelines such as ti2vid_two_stages.py.

SimpleDenoiser for Single-Pass Inference

SimpleDenoiser performs exactly one transformer call per denoising step. It builds conditional modalities using modality_from_latent_state (from ltx_pipelines/utils/helpers.py) and immediately returns the transformer output. This minimizes GPU kernel launches and provides the fastest inference path when you do not require classifier-free guidance.

GuidedDenoiser for Multi-Pass Guidance

GuidedDenoiser supports static guidance strategies including CFG, STG (self-teacher guidance), and isolated-modality guidance. Internally, it delegates to the shared _guided_denoise function (lines 57-80 in denoisers.py), which batches all guidance passes—conditional, unconditioned, and perturbed—into a single transformer call. This batching maintains throughput even when computing multiple guidance terms.

FactoryGuidedDenoiser for Dynamic Guidance

FactoryGuidedDenoiser instantiates guiders per-step based on the current sigma value. Use this when you need to vary cfg_scale or stg_scale throughout the diffusion schedule, such as ramping guidance strength from low to high noise levels.

Implementing a Custom Denoiser from Scratch

To create a custom denoiser that conforms to the protocol, implement the __call__ method and handle modality construction:

from typing import Tuple
import torch
from ltx_core.model.transformer import X0Model
from ltx_core.types import LatentState
from ltx_pipelines.utils.types import DenoisedLatentResult, Denoiser
from ltx_pipelines.utils.helpers import modality_from_latent_state

class MyCustomDenoiser:
    """Example custom denoiser that mixes user-defined perturbations
    with a standard conditional pass."""
    def __init__(self, v_context: torch.Tensor | None, a_context: torch.Tensor | None):
        self.v_context = v_context
        self.a_context = a_context

    def __call__(
        self,
        transformer: X0Model,
        video_state: LatentState | None,
        audio_state: LatentState | None,
        sigmas: torch.Tensor,
        step_index: int,
    ) -> Tuple[DenoisedLatentResult | None, DenoisedLatentResult | None]:
        sigma = sigmas[step_index]

        # Build the "cond" modality (same pattern as SimpleDenoiser)

        video_mod = (
            modality_from_latent_state(video_state, self.v_context, sigma)
            if video_state is not None else None
        )
        audio_mod = (
            modality_from_latent_state(audio_state, self.a_context, sigma)
            if audio_state is not None else None
        )

        # Apply custom perturbation before the transformer call

        if video_mod is not None:
            mask = torch.rand_like(video_mod.token_embeddings) > 0.5
            video_mod.token_embeddings = video_mod.token_embeddings * mask

        # Single transformer call (no guidance)

        denoised_video, denoised_audio = transformer(
            video=video_mod, audio=audio_mod, perturbations=None
        )

        return (
            DenoisedLatentResult.result_or_none(denoised=denoised_video),
            DenoisedLatentResult.result_or_none(denoised=denoised_audio),
        )

Key implementation details:

  • The constructor receives conditional contexts (v_context, a_context) that align with the latent states.
  • Apply perturbations before calling the transformer, as the model expects fully prepared Modality objects.
  • Return DenoisedLatentResult containers using the result_or_none factory method to handle null states safely.

Extending GuidedDenoiser vs SimpleDenoiser

When you need to customize initialization logic while preserving the optimized __call__ implementation, subclass the existing denoisers.

Extending GuidedDenoiser

Override __init__ to inject custom guiders with specific cfg_scale values:

from ltx_pipelines.utils.denoisers import GuidedDenoiser
from ltx_core.components.guiders import MultiModalGuider, MultiModalGuiderParams

class MyGuidedDenoiser(GuidedDenoiser):
    """Guided denoiser with custom CFG scale for video only."""
    def __init__(self, v_context, a_context):
        # Create a guider with cfg_scale=7.5 for video

        video_guider = MultiModalGuider(
            params=MultiModalGuiderParams(cfg_scale=7.5, stg_scale=0.0, modality_scale=1.0)
        )
        super().__init__(
            v_context=v_context,
            a_context=a_context,
            video_guider=video_guider,
            audio_guider=None,  # Falls back to positive-only guider

            force_uncond_pass=False,
        )

Passing audio_guider=None triggers the default _POSITIVE_ONLY_GUIDER (defined at lines 25-28 in denoisers.py), which performs a single conditional pass without guidance.

Extending SimpleDenoiser

Subclass SimpleDenoiser when you need to modify preprocessing or post-processing while keeping the single-pass logic:

from ltx_pipelines.utils.denoisers import SimpleDenoiser

class PreprocessingDenoiser(SimpleDenoiser):
    """SimpleDenoiser with custom sigma rescaling."""
    def __call__(self, transformer, video_state, audio_state, sigmas, step_index):
        # Modify sigma before processing

        modified_sigmas = sigmas * 0.9
        return super().__call__(transformer, video_state, audio_state, modified_sigmas, step_index)

Choosing the Right Denoiser for Your Use Case

Select your base class based on inference requirements and guidance complexity:

Scenario Recommended Denoiser Implementation Strategy
Fastest inference – unconditional or single-condition generation SimpleDenoiser Use directly or subclass for preprocessing
Static CFG/STG – constant guidance strength throughout sampling GuidedDenoiser Subclass and inject custom MultiModalGuider instances
Dynamic guidance – varying cfg_scale based on noise level FactoryGuidedDenoiser Provide a factory function that returns guiders per sigma
Custom perturbations – sigma-dependent masking or dropout Custom Denoiser protocol Delegate to _guided_denoise for multi-pass batching, or build modalities manually for single-pass

Integrating your denoiser into a pipeline follows the pattern shown in ti2vid_two_stages.py (lines 161-170):


# For guided generation

denoiser = GuidedDenoiser(
    v_context=video_cond,
    a_context=audio_cond,
    video_guider=video_guider,  # Custom cfg_scale=8.0

    audio_guider=None,
    force_uncond_pass=False,
)
video_res, audio_res = denoiser(transformer, video_state, audio_state, sigmas, step_i)

# For fast inference

simple_denoiser = SimpleDenoiser(v_context=video_cond, a_context=audio_cond)
video_res, audio_res = simple_denoiser(transformer, video_state, audio_state, sigmas, step_i)

Summary

  • SimpleDenoiser executes one transformer pass per step, providing the fastest inference for unconditional or purely conditional generation.
  • GuidedDenoiser batches conditional, unconditioned, and perturbed passes via _guided_denoise (lines 57-80) to support CFG and STG without excessive kernel launches.
  • FactoryGuidedDenoiser enables per-step guidance parameter variation by instantiating guiders from sigma values.
  • All denoisers implement the Denoiser protocol from ltx_pipelines/utils/types.py, ensuring compatibility with LTX-2 pipelines.
  • Custom implementations should use modality_from_latent_state from ltx_pipelines/utils/helpers.py to construct transformer-compatible modalities.
  • Subclass existing denoisers when you need customized initialization but want to preserve the optimized __call__ logic.

Frequently Asked Questions

What is the difference between GuidedDenoiser and SimpleDenoiser in LTX-2?

SimpleDenoiser performs a single transformer call using only the conditional input, making it optimal for speed when no guidance is needed. GuidedDenoiser computes multiple passes (conditional, unconditional, and perturbed) batched together via _guided_denoise, enabling classifier-free guidance and self-teacher guidance at the cost of increased memory usage. According to the source code in ltx_pipelines/utils/denoisers.py, both classes conform to the same Denoiser protocol but differ in their internal pass composition.

When should I use FactoryGuidedDenoiser instead of GuidedDenoiser?

Use FactoryGuidedDenoiser when your guidance parameters need to change based on the current noise level (sigma). Unlike GuidedDenoiser, which uses static guiders initialized once, FactoryGuidedDenoiser creates new guider instances for each step using the current sigma value. This is essential for techniques like dynamic CFG scaling that ramps from low to high strength during the diffusion process.

How do I add custom perturbations to a denoiser in LTX-2?

Apply custom perturbations to the Modality objects before passing them to the transformer. If using GuidedDenoiser, you can provide a custom PerturbationConfig in the passes list (see lines 94-119 in denoisers.py) to leverage the existing multi-pass batching infrastructure. For SimpleDenoiser or custom protocol implementations, manually modify the token_embeddings or mask tensors after calling modality_from_latent_state but before the transformer forward pass.

Can I combine SimpleDenoiser speed with guidance features?

No, guidance inherently requires multiple forward passes to compute the conditional and unconditional directions. However, you can minimize overhead by ensuring your GuidedDenoiser uses the _guided_denoise helper, which batches all passes into a single transformer call. This approach maintains near-optimal throughput while still supporting CFG and STG, as implemented in the base class at lines 57-80 of ltx_pipelines/utils/denoisers.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →