# How Classifier-Free Guidance (CFG) Works to Blend Conditioned and Unconditioned Predictions in LTX‑2

> Discover how Classifier-Free Guidance (CFG) blends conditioned and unconditioned predictions in LTX-2. Learn how cfg_scale controls the blend to steer diffusion models toward your prompts.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-18

---

**Classifier-free guidance (CFG) steers a diffusion model toward a desired prompt by interpolating between a conditioned prediction (using the real prompt) and an unconditioned prediction (using a null or negative prompt), with the blend controlled by a `cfg_scale` parameter.**

Classifier-free guidance is the standard technique for controlling prompt adherence in modern diffusion models, and the LTX‑2 video generation system implements an extended, multimodal variant. This article breaks down exactly how CFG blending works in the Lightricks/LTX‑2 codebase, covering the core formula, the classes that implement it, and how the framework handles multiple guidance modes simultaneously.

## The CFG Formula in LTX‑2

The heart of CFG blending in LTX‑2 lives in `MultiModalGuider.calculate`, defined in [`guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/guiders.py) at lines 52–66:

```python
pred = (
    cond
    + (self.params.cfg_scale - 1) * (cond - uncond_text)   # ← CFG term

    + self.params.stg_scale * (cond - uncond_perturbed)   # ← optional STG term

    + (self.params.modality_scale - 1) * (cond - uncond_modality)  # ← optional modality term

)

```

**Key variables:**

- **`cond`** – the conditioned model output, computed with the full positive prompt
- **`uncond_text`** – the unconditioned output, computed with an empty or negative text prompt
- **`cfg_scale`** – controls the strength of text guidance (default 3.0 for video, 7.0 for audio)

When `cfg_scale = 1.0`, the CFG term `(cfg_scale - 1) * (cond - uncond_text)` evaluates to zero, and the prediction equals the plain conditioned output. Values above 1.0 amplify the directional push toward the prompt.

## The MultiModalGuider Architecture

LTX‑2 extends classic CFG with a `MultiModalGuider` class that can blend guidance across text, spatial, and modality-specific signals. This design enables simultaneous control over video and audio generation.

### Configuration Via MultiModalGuiderParams

The guidance parameters are encapsulated in `MultiModalGuiderParams`, also in [`guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/guiders.py):

```python
cfg_scale: float = 1.0   # "how strongly the model adheres to the prompt"

```

Default scales are defined in [`config.py`](https://github.com/Lightricks/LTX-2/blob/main/config.py) (lines 9–15):

```python
video_cfg_scale: float = Field(default=3.0, ge=1.0)
audio_cfg_scale: float = Field(default=7.0, ge=1.0)

```

Video generation uses a moderate 3.0 default, while audio generation uses 7.0—reflecting the greater difficulty of guiding audio synthesis with text alone.

## Integration With the Denoising Loop

CFG blending occurs **once per denoising step** inside `GuidedDenoiser` and `FactoryGuidedDenoiser` ([`denoisers.py`](https://github.com/Lightricks/LTX-2/blob/main/denoisers.py), lines 13–31). The denoiser orchestrates the conditional and unconditional forward passes:

1. **Calls the transformer twice** (or up to four times when Self-Tuning Guidance (STG) or modality guidance is enabled)
2. **Stacks the results** into tensors matching the guider's expected signature
3. **Invokes `MultiModalGuider.calculate`** to produce the blended prediction
4. **Returns the result** for the next diffusion step

This architecture keeps the core CFG formula isolated and testable while allowing the denoiser to handle batching, masking, and multi-GPU concerns.

## Practical CFG Usage Examples

### Manual CFG Blending

For debugging or research purposes, you can instantiate the guider directly:

```python
from ltx_core.components.guiders import MultiModalGuider, MultiModalGuiderParams

# Configure CFG for video at default strength

params = MultiModalGuiderParams(cfg_scale=3.0)
guider = MultiModalGuider(params=params)

# Assuming cond and uncond_text are transformer outputs:

blended = guider.calculate(
    cond, 
    uncond_text, 
    uncond_perturbed=0.0, 
    uncond_modality=0.0
)

```

### Standard Inference Path

Typical inference uses `GuidedDenoiser`, which automates the conditional/unconditional split:

```python
from ltx_pipelines.utils.denoisers import GuidedDenoiser
from ltx_core.components.guiders import MultiModalGuiderParams

video_params = MultiModalGuiderParams(cfg_scale=3.0)
audio_params = MultiModalGuiderParams(cfg_scale=7.0)

denoiser = GuidedDenoiser(
    video_guider_params=video_params,
    audio_guider_params=audio_params,
    force_uncond_pass=True,    # Ensures uncond tensor even if cfg_scale==1

    negative_context=None,
)

# Inside the diffusion loop:

video_res, audio_res = denoiser(
    transformer, video_state, audio_state, sigmas, step_index
)

# video_res.denoised contains the CFG-blended prediction

```

### Runtime CFG Adjustment

Modify guidance strength dynamically between steps:

```python

# Increase video guidance for stronger prompt adherence

denoiser.video_guider.params = MultiModalGuiderParams(cfg_scale=5.0)

```

## Extension: CFG++ and Advanced Samplers

The base CFG formula can be modified for improved sample quality. LTX‑2 implements **CFG++** in `euler_cfg_pp_denoising_loop` ([`samplers.py`](https://github.com/Lightricks/LTX-2/blob/main/samplers.py), line 35), which adjusts how the blended prediction feeds back into the ODE solver. This variant uses the same `MultiModalGuider` infrastructure but wraps it in a different update rule.

## Summary

- **CFG blends predictions** via `cond + (cfg_scale - 1) * (cond - uncond_text)` as implemented in `MultiModalGuider.calculate`
- **Default scales differ by modality**: 3.0 for video, 7.0 for audio, configured in `MultiModalGuiderParams`
- **Blending happens per step** in `GuidedDenoiser`, which manages conditional/unconditional model calls
- **Extended guidance** (STG, modality scales) uses additive terms in the same formula without code duplication
- **Runtime modification** is supported by updating the `params` attribute of active guiders

## Frequently Asked Questions

### What happens when CFG scale equals 1.0?

The guidance term vanishes. At `cfg_scale = 1.0`, the formula reduces to `pred = cond`, meaning the model ignores the unconditioned output and generates purely from the positive prompt. This disables steering but preserves computational cost since the uncond pass may still be computed depending on `force_uncond_pass`.

### Why does audio use a higher CFG scale than video?

Audio generation from text is inherently more ambiguous—text descriptions map less precisely to audio features than to visual ones. The default `audio_cfg_scale=7.0` versus `video_cfg_scale=3.0` reflects this difficulty, providing stronger guidance to compensate for weaker text-to-audio conditioning.

### Can I use negative prompts with CFG in LTX‑2?

Yes. The `negative_context` parameter in `GuidedDenoiser` accepts tensors that replace or augment the default "null" conditioning. Pass a computed negative embedding to steer generation away from undesired concepts while pulling toward the positive prompt via CFG.

### How does CFG relate to Self-Tuning Guidance (STG)?

STG is an additive term `stg_scale * (cond - uncond_perturbed)` that uses a perturbed version of the conditioned input rather than a null prompt. Both terms share the same `MultiModalGuider.calculate` implementation and can operate simultaneously, though STG requires additional transformer evaluations per step.