How Classifier-Free Guidance (CFG) Works to Blend Conditioned and Unconditioned Predictions in LTX‑2
Classifier-free guidance (CFG) steers a diffusion model toward a desired prompt by interpolating between a conditioned prediction (using the real prompt) and an unconditioned prediction (using a null or negative prompt), with the blend controlled by a cfg_scale parameter.
Classifier-free guidance is the standard technique for controlling prompt adherence in modern diffusion models, and the LTX‑2 video generation system implements an extended, multimodal variant. This article breaks down exactly how CFG blending works in the Lightricks/LTX‑2 codebase, covering the core formula, the classes that implement it, and how the framework handles multiple guidance modes simultaneously.
The CFG Formula in LTX‑2
The heart of CFG blending in LTX‑2 lives in MultiModalGuider.calculate, defined in guiders.py at lines 52–66:
pred = (
cond
+ (self.params.cfg_scale - 1) * (cond - uncond_text) # ← CFG term
+ self.params.stg_scale * (cond - uncond_perturbed) # ← optional STG term
+ (self.params.modality_scale - 1) * (cond - uncond_modality) # ← optional modality term
)
Key variables:
cond– the conditioned model output, computed with the full positive promptuncond_text– the unconditioned output, computed with an empty or negative text promptcfg_scale– controls the strength of text guidance (default 3.0 for video, 7.0 for audio)
When cfg_scale = 1.0, the CFG term (cfg_scale - 1) * (cond - uncond_text) evaluates to zero, and the prediction equals the plain conditioned output. Values above 1.0 amplify the directional push toward the prompt.
The MultiModalGuider Architecture
LTX‑2 extends classic CFG with a MultiModalGuider class that can blend guidance across text, spatial, and modality-specific signals. This design enables simultaneous control over video and audio generation.
Configuration Via MultiModalGuiderParams
The guidance parameters are encapsulated in MultiModalGuiderParams, also in guiders.py:
cfg_scale: float = 1.0 # "how strongly the model adheres to the prompt"
Default scales are defined in config.py (lines 9–15):
video_cfg_scale: float = Field(default=3.0, ge=1.0)
audio_cfg_scale: float = Field(default=7.0, ge=1.0)
Video generation uses a moderate 3.0 default, while audio generation uses 7.0—reflecting the greater difficulty of guiding audio synthesis with text alone.
Integration With the Denoising Loop
CFG blending occurs once per denoising step inside GuidedDenoiser and FactoryGuidedDenoiser (denoisers.py, lines 13–31). The denoiser orchestrates the conditional and unconditional forward passes:
- Calls the transformer twice (or up to four times when Self-Tuning Guidance (STG) or modality guidance is enabled)
- Stacks the results into tensors matching the guider's expected signature
- Invokes
MultiModalGuider.calculateto produce the blended prediction - Returns the result for the next diffusion step
This architecture keeps the core CFG formula isolated and testable while allowing the denoiser to handle batching, masking, and multi-GPU concerns.
Practical CFG Usage Examples
Manual CFG Blending
For debugging or research purposes, you can instantiate the guider directly:
from ltx_core.components.guiders import MultiModalGuider, MultiModalGuiderParams
# Configure CFG for video at default strength
params = MultiModalGuiderParams(cfg_scale=3.0)
guider = MultiModalGuider(params=params)
# Assuming cond and uncond_text are transformer outputs:
blended = guider.calculate(
cond,
uncond_text,
uncond_perturbed=0.0,
uncond_modality=0.0
)
Standard Inference Path
Typical inference uses GuidedDenoiser, which automates the conditional/unconditional split:
from ltx_pipelines.utils.denoisers import GuidedDenoiser
from ltx_core.components.guiders import MultiModalGuiderParams
video_params = MultiModalGuiderParams(cfg_scale=3.0)
audio_params = MultiModalGuiderParams(cfg_scale=7.0)
denoiser = GuidedDenoiser(
video_guider_params=video_params,
audio_guider_params=audio_params,
force_uncond_pass=True, # Ensures uncond tensor even if cfg_scale==1
negative_context=None,
)
# Inside the diffusion loop:
video_res, audio_res = denoiser(
transformer, video_state, audio_state, sigmas, step_index
)
# video_res.denoised contains the CFG-blended prediction
Runtime CFG Adjustment
Modify guidance strength dynamically between steps:
# Increase video guidance for stronger prompt adherence
denoiser.video_guider.params = MultiModalGuiderParams(cfg_scale=5.0)
Extension: CFG++ and Advanced Samplers
The base CFG formula can be modified for improved sample quality. LTX‑2 implements CFG++ in euler_cfg_pp_denoising_loop (samplers.py, line 35), which adjusts how the blended prediction feeds back into the ODE solver. This variant uses the same MultiModalGuider infrastructure but wraps it in a different update rule.
Summary
- CFG blends predictions via
cond + (cfg_scale - 1) * (cond - uncond_text)as implemented inMultiModalGuider.calculate - Default scales differ by modality: 3.0 for video, 7.0 for audio, configured in
MultiModalGuiderParams - Blending happens per step in
GuidedDenoiser, which manages conditional/unconditional model calls - Extended guidance (STG, modality scales) uses additive terms in the same formula without code duplication
- Runtime modification is supported by updating the
paramsattribute of active guiders
Frequently Asked Questions
What happens when CFG scale equals 1.0?
The guidance term vanishes. At cfg_scale = 1.0, the formula reduces to pred = cond, meaning the model ignores the unconditioned output and generates purely from the positive prompt. This disables steering but preserves computational cost since the uncond pass may still be computed depending on force_uncond_pass.
Why does audio use a higher CFG scale than video?
Audio generation from text is inherently more ambiguous—text descriptions map less precisely to audio features than to visual ones. The default audio_cfg_scale=7.0 versus video_cfg_scale=3.0 reflects this difficulty, providing stronger guidance to compensate for weaker text-to-audio conditioning.
Can I use negative prompts with CFG in LTX‑2?
Yes. The negative_context parameter in GuidedDenoiser accepts tensors that replace or augment the default "null" conditioning. Pass a computed negative embedding to steer generation away from undesired concepts while pulling toward the positive prompt via CFG.
How does CFG relate to Self-Tuning Guidance (STG)?
STG is an additive term stg_scale * (cond - uncond_perturbed) that uses a perturbed version of the conditioned input rather than a null prompt. Both terms share the same MultiModalGuider.calculate implementation and can operate simultaneously, though STG requires additional transformer evaluations per step.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →