# LTX-2 Guidance Parameters: CFG, STG, and Modality Guidance Explained

> Master LTX-2's guidance parameters: CFG, STG, and modality guidance. Tune these signals independently for advanced video and audio generation. Learn more now.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-08-20

---

**LTX-2 supports three distinct guidance signals—Classifier-Free Guidance (CFG), Spatio-Temporal Guidance (STG), and modality-specific guidance—that you can independently tune for video and audio generation via the `MultiModalGuiderParams` dataclass in [`guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/guiders.py).**

The Lightricks LTX-2 diffusion model exposes fine-grained control over generation quality through its multimodal guider system. Each modality (video and audio) receives its own guidance configuration, allowing precise adjustment of prompt adherence, temporal consistency, and cross-modal synchronization. This article breaks down every LTX-2 guidance parameter based on the source code in the official repository.

## Classifier-Free Guidance (cfg_scale)

CFG is the standard technique for amplifying text prompt influence. In LTX-2, `cfg_scale` controls the strength of the `(conditional - unconditional)` direction computed during each denoising step.

| Aspect | Details |
|--------|---------|
| Typical range | 2.0 – 5.0 |
| Default | ~3.0 |
| Disable | Set to 1.0 |

Higher values force stricter prompt adherence but can reduce motion naturalness. The video and audio guiders accept independent `cfg_scale` values, so you might prioritize sharper audio generation while keeping video more fluid.

```python
from ltx_core.components.guiders import MultiModalGuiderParams

# Strong text guidance for audio, moderate for video

audio_guider = MultiModalGuiderParams(cfg_scale=7.0, ...)
video_guider = MultiModalGuiderParams(cfg_scale=3.0, ...)

```

## Spatio-Temporal Guidance (stg_scale and stg_blocks)

STG improves temporal coherence by perturbing specific transformer blocks and steering away from the perturbed prediction. This mechanism is unique to LTX-2's video generation pipeline.

### stg_scale

Controls the perturbation guidance strength:

| Aspect | Details |
|--------|---------|
| Typical range | 0.5 – 1.5 |
| Default | ~1.0 |
| Disable | Set to 0.0 |

### stg_blocks

Specifies which transformer blocks receive the perturbation. The default `[29]` targets the final block in the architecture.

```python

# Enable STG on the last transformer block

video_guider = MultiModalGuiderParams(
    stg_scale=1.0,
    stg_blocks=[29],  # Perturb block 29 only

)

```

To completely disable STG, provide an empty list: `stg_blocks=[]`.

## Modality Guidance (modality_scale)

The `modality_scale` parameter implements **modality-specific CFG** for joint audio-video generation. It pushes the two streams toward synchronized results, preventing audio-visual drift.

| Aspect | Details |
|--------|---------|
| Typical range | 1.0 – 5.0 |
| Default | ~3.0 (joint generation) |
| Disable | Set to 1.0 |

This parameter is particularly critical when generating matched video and audio simultaneously. Higher values enforce tighter coupling between modalities.

## Rescaling and Performance Controls

### rescale_scale

Prevents over-saturation by scaling guided predictions to match the variance of conditional predictions:

| Aspect | Details |
|--------|---------|
| Typical range | 0.5 – 0.7 |
| Default | ~0.7 |
| Disable | Set to 0.0 |

### skip_step

Accelerator for inference. Skips guidance computation every N steps:

| Aspect | Details |
|--------|---------|
| Default | 0 (no skipping) |
| Use case | Speed-critical applications |

## Complete Configuration Example

```python
from ltx_core.components.guiders import MultiModalGuiderParams
from ltx_pipelines.ti2vid_one_stage import run_pipeline

# Video: balanced guidance with temporal coherence

video_guider = MultiModalGuiderParams(
    cfg_scale=3.0,
    stg_scale=1.0,
    rescale_scale=0.7,
    modality_scale=3.0,
    stg_blocks=[29],
    skip_step=0,
)

# Audio: stronger prompt adherence

audio_guider = MultiModalGuiderParams(
    cfg_scale=7.0,
    stg_scale=1.0,
    rescale_scale=0.7,
    modality_scale=3.0,
    stg_blocks=[29],
    skip_step=0,
)

run_pipeline(
    video_guider_params=video_guider,
    audio_guider_params=audio_guider,
    # ... additional arguments

)

```

## Key Implementation Files

The guidance system spans several source files in the LTX-2 repository:

- **[`packages/ltx-core/src/ltx_core/components/guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py)** — Core `MultiModalGuiderParams` dataclass and guidance merging logic
- **[`packages/ltx-pipelines/docs/multimodal-guidance.md`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/docs/multimodal-guidance.md)** — Official parameter documentation
- **[`packages/ltx-trainer/src/ltx_trainer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py)** — CLI flags for training validation
- **[`packages/ltx-pipelines/src/ltx_pipelines/utils/args.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py)** — Command-line argument parsing
- **[`packages/ltx-pipelines/src/ltx_pipelines/ti2vid_one_stage.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/ti2vid_one_stage.py)** — Pipeline integration example

## Summary

- **CFG (`cfg_scale`)** controls text-prompt strength per modality—set to 1.0 to disable
- **STG (`stg_scale`, `stg_blocks`)** improves temporal coherence via block perturbation—disable with `stg_scale=0.0` or `stg_blocks=[]`
- **Modality guidance (`modality_scale`)** synchronizes audio-video generation—critical for joint generation tasks
- **Rescaling (`rescale_scale`)** prevents saturation artifacts
- **Skip steps (`skip_step`)** trades quality for inference speed

All parameters live in `MultiModalGuiderParams` as defined in [`guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/guiders.py), with independent instances for video and audio.

## Frequently Asked Questions

### What is the difference between CFG and modality guidance in LTX-2?

**CFG (`cfg_scale`)** amplifies text prompt influence by contrasting conditional and unconditional predictions. **Modality guidance (`modality_scale`)** specifically targets multi-modal synchronization, pushing video and audio streams toward coherent joint results. You need both: CFG for prompt fidelity, modality guidance for audio-visual alignment.

### How do I completely disable STG in LTX-2?

Set `stg_scale=0.0` or provide an empty list for `stg_blocks=[]`. The latter is more explicit and prevents any STG-related computation overhead. Both approaches are valid; the source code in [`guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/guiders.py) checks these conditions to skip STG application.

### Can I use different guidance scales for video and audio generation?

Yes—this is the default pattern. Instantiate separate `MultiModalGuiderParams` objects and pass them as `video_guider_params` and `audio_guider_params` to your pipeline. The LTX-2 architecture treats each modality independently, allowing you to prioritize audio clarity (`cfg_scale=7.0`) while keeping video motion natural (`cfg_scale=3.0`).

### Where are the default guidance values defined in the LTX-2 repository?

Default values reside in [`packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py). These serve as the reference configuration that the pipelines use when you don't override parameters explicitly.