How to Implement GuidedDenoiser vs SimpleDenoiser for Custom Denoisers in LTX-2
GuidedDenoiser applies classifier-free guidance (CFG) and self-teacher guidance (STG) through batched multi-pass transformer calls, while SimpleDenoiser executes a single conditional pass for maximum inference speed, and both implement the Denoiser protocol defined in ltx_pipelines/utils/types.py.
Lightricks/LTX-2 ships with three production-ready denoiser implementations that cater to different inference requirements. Understanding how to implement GuidedDenoiser vs SimpleDenoiser for custom denoisers in LTX-2 allows you to balance generation quality against computational overhead, whether you need unconditional fast sampling or fine-grained guidance control.
Understanding the Built-in Denoiser Architecture
LTX-2 organizes its denoising logic around a strict protocol-based design. All denoisers reside in ltx_pipelines/utils/denoisers.py and conform to the Denoiser protocol defined in ltx_pipelines/utils/types.py (lines 73-88).
The Denoiser Protocol
Every denoiser must implement the __call__ method with the following signature:
def __call__(
self,
transformer: X0Model,
video_state: LatentState | None,
audio_state: LatentState | None,
sigmas: torch.Tensor,
step_index: int,
) -> Tuple[DenoisedLatentResult | None, DenoisedLatentResult | None]:
This protocol ensures interoperability with LTX-2 diffusion pipelines such as ti2vid_two_stages.py.
SimpleDenoiser for Single-Pass Inference
SimpleDenoiser performs exactly one transformer call per denoising step. It builds conditional modalities using modality_from_latent_state (from ltx_pipelines/utils/helpers.py) and immediately returns the transformer output. This minimizes GPU kernel launches and provides the fastest inference path when you do not require classifier-free guidance.
GuidedDenoiser for Multi-Pass Guidance
GuidedDenoiser supports static guidance strategies including CFG, STG (self-teacher guidance), and isolated-modality guidance. Internally, it delegates to the shared _guided_denoise function (lines 57-80 in denoisers.py), which batches all guidance passes—conditional, unconditioned, and perturbed—into a single transformer call. This batching maintains throughput even when computing multiple guidance terms.
FactoryGuidedDenoiser for Dynamic Guidance
FactoryGuidedDenoiser instantiates guiders per-step based on the current sigma value. Use this when you need to vary cfg_scale or stg_scale throughout the diffusion schedule, such as ramping guidance strength from low to high noise levels.
Implementing a Custom Denoiser from Scratch
To create a custom denoiser that conforms to the protocol, implement the __call__ method and handle modality construction:
from typing import Tuple
import torch
from ltx_core.model.transformer import X0Model
from ltx_core.types import LatentState
from ltx_pipelines.utils.types import DenoisedLatentResult, Denoiser
from ltx_pipelines.utils.helpers import modality_from_latent_state
class MyCustomDenoiser:
"""Example custom denoiser that mixes user-defined perturbations
with a standard conditional pass."""
def __init__(self, v_context: torch.Tensor | None, a_context: torch.Tensor | None):
self.v_context = v_context
self.a_context = a_context
def __call__(
self,
transformer: X0Model,
video_state: LatentState | None,
audio_state: LatentState | None,
sigmas: torch.Tensor,
step_index: int,
) -> Tuple[DenoisedLatentResult | None, DenoisedLatentResult | None]:
sigma = sigmas[step_index]
# Build the "cond" modality (same pattern as SimpleDenoiser)
video_mod = (
modality_from_latent_state(video_state, self.v_context, sigma)
if video_state is not None else None
)
audio_mod = (
modality_from_latent_state(audio_state, self.a_context, sigma)
if audio_state is not None else None
)
# Apply custom perturbation before the transformer call
if video_mod is not None:
mask = torch.rand_like(video_mod.token_embeddings) > 0.5
video_mod.token_embeddings = video_mod.token_embeddings * mask
# Single transformer call (no guidance)
denoised_video, denoised_audio = transformer(
video=video_mod, audio=audio_mod, perturbations=None
)
return (
DenoisedLatentResult.result_or_none(denoised=denoised_video),
DenoisedLatentResult.result_or_none(denoised=denoised_audio),
)
Key implementation details:
- The constructor receives conditional contexts (
v_context,a_context) that align with the latent states. - Apply perturbations before calling the transformer, as the model expects fully prepared
Modalityobjects. - Return
DenoisedLatentResultcontainers using theresult_or_nonefactory method to handle null states safely.
Extending GuidedDenoiser vs SimpleDenoiser
When you need to customize initialization logic while preserving the optimized __call__ implementation, subclass the existing denoisers.
Extending GuidedDenoiser
Override __init__ to inject custom guiders with specific cfg_scale values:
from ltx_pipelines.utils.denoisers import GuidedDenoiser
from ltx_core.components.guiders import MultiModalGuider, MultiModalGuiderParams
class MyGuidedDenoiser(GuidedDenoiser):
"""Guided denoiser with custom CFG scale for video only."""
def __init__(self, v_context, a_context):
# Create a guider with cfg_scale=7.5 for video
video_guider = MultiModalGuider(
params=MultiModalGuiderParams(cfg_scale=7.5, stg_scale=0.0, modality_scale=1.0)
)
super().__init__(
v_context=v_context,
a_context=a_context,
video_guider=video_guider,
audio_guider=None, # Falls back to positive-only guider
force_uncond_pass=False,
)
Passing audio_guider=None triggers the default _POSITIVE_ONLY_GUIDER (defined at lines 25-28 in denoisers.py), which performs a single conditional pass without guidance.
Extending SimpleDenoiser
Subclass SimpleDenoiser when you need to modify preprocessing or post-processing while keeping the single-pass logic:
from ltx_pipelines.utils.denoisers import SimpleDenoiser
class PreprocessingDenoiser(SimpleDenoiser):
"""SimpleDenoiser with custom sigma rescaling."""
def __call__(self, transformer, video_state, audio_state, sigmas, step_index):
# Modify sigma before processing
modified_sigmas = sigmas * 0.9
return super().__call__(transformer, video_state, audio_state, modified_sigmas, step_index)
Choosing the Right Denoiser for Your Use Case
Select your base class based on inference requirements and guidance complexity:
| Scenario | Recommended Denoiser | Implementation Strategy |
|---|---|---|
| Fastest inference – unconditional or single-condition generation | SimpleDenoiser |
Use directly or subclass for preprocessing |
| Static CFG/STG – constant guidance strength throughout sampling | GuidedDenoiser |
Subclass and inject custom MultiModalGuider instances |
Dynamic guidance – varying cfg_scale based on noise level |
FactoryGuidedDenoiser |
Provide a factory function that returns guiders per sigma |
| Custom perturbations – sigma-dependent masking or dropout | Custom Denoiser protocol |
Delegate to _guided_denoise for multi-pass batching, or build modalities manually for single-pass |
Integrating your denoiser into a pipeline follows the pattern shown in ti2vid_two_stages.py (lines 161-170):
# For guided generation
denoiser = GuidedDenoiser(
v_context=video_cond,
a_context=audio_cond,
video_guider=video_guider, # Custom cfg_scale=8.0
audio_guider=None,
force_uncond_pass=False,
)
video_res, audio_res = denoiser(transformer, video_state, audio_state, sigmas, step_i)
# For fast inference
simple_denoiser = SimpleDenoiser(v_context=video_cond, a_context=audio_cond)
video_res, audio_res = simple_denoiser(transformer, video_state, audio_state, sigmas, step_i)
Summary
- SimpleDenoiser executes one transformer pass per step, providing the fastest inference for unconditional or purely conditional generation.
- GuidedDenoiser batches conditional, unconditioned, and perturbed passes via
_guided_denoise(lines 57-80) to support CFG and STG without excessive kernel launches. - FactoryGuidedDenoiser enables per-step guidance parameter variation by instantiating guiders from sigma values.
- All denoisers implement the
Denoiserprotocol fromltx_pipelines/utils/types.py, ensuring compatibility with LTX-2 pipelines. - Custom implementations should use
modality_from_latent_statefromltx_pipelines/utils/helpers.pyto construct transformer-compatible modalities. - Subclass existing denoisers when you need customized initialization but want to preserve the optimized
__call__logic.
Frequently Asked Questions
What is the difference between GuidedDenoiser and SimpleDenoiser in LTX-2?
SimpleDenoiser performs a single transformer call using only the conditional input, making it optimal for speed when no guidance is needed. GuidedDenoiser computes multiple passes (conditional, unconditional, and perturbed) batched together via _guided_denoise, enabling classifier-free guidance and self-teacher guidance at the cost of increased memory usage. According to the source code in ltx_pipelines/utils/denoisers.py, both classes conform to the same Denoiser protocol but differ in their internal pass composition.
When should I use FactoryGuidedDenoiser instead of GuidedDenoiser?
Use FactoryGuidedDenoiser when your guidance parameters need to change based on the current noise level (sigma). Unlike GuidedDenoiser, which uses static guiders initialized once, FactoryGuidedDenoiser creates new guider instances for each step using the current sigma value. This is essential for techniques like dynamic CFG scaling that ramps from low to high strength during the diffusion process.
How do I add custom perturbations to a denoiser in LTX-2?
Apply custom perturbations to the Modality objects before passing them to the transformer. If using GuidedDenoiser, you can provide a custom PerturbationConfig in the passes list (see lines 94-119 in denoisers.py) to leverage the existing multi-pass batching infrastructure. For SimpleDenoiser or custom protocol implementations, manually modify the token_embeddings or mask tensors after calling modality_from_latent_state but before the transformer forward pass.
Can I combine SimpleDenoiser speed with guidance features?
No, guidance inherently requires multiple forward passes to compute the conditional and unconditional directions. However, you can minimize overhead by ensuring your GuidedDenoiser uses the _guided_denoise helper, which batches all passes into a single transformer call. This approach maintains near-optimal throughput while still supporting CFG and STG, as implemented in the base class at lines 57-80 of ltx_pipelines/utils/denoisers.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →