LTX‑2 STG and Modality Guidance Scales Configuration: A Complete Guide

Set video_stg_scale and audio_stg_scale to 0.0 to disable STG, adjust video_modality_guidance_scale and audio_modality_guidance_scale to control cross‑modal isolation strength, and use guidance_rescale to scale the combined guidance output—each field is defined in ltx_trainer/config.py and applied during the validation denoising loop.

LTX‑2, the open‑source video‑audio generation model from Lightricks, exposes a unified configuration system for fine‑tuning generation behavior. Understanding how Spatio‑Temporal Guidance (STG) and modality guidance scales interact is essential for controlling trade‑offs between cross‑modal conditioning and output fidelity. This guide walks through every configuration field, its default value, and where it is implemented in the LTX‑2 source code.


What STG and Modality Guidance Scales Control

LTX‑2 applies guidance at three interconnected levels during generation:

  1. Spatio‑Temporal Guidance (STG) – perturbs self‑attention blocks to guide temporal coherence
  2. Modality Guidance – isolates or strengthens audio‑to‑video and video‑to‑audio conditioning
  3. Guidance Rescale – uniformly scales the combined effect of all guidance mechanisms

These settings live in the Config class at packages/ltx-trainer/src/ltx_trainer/config.py and are consumed by the validation runner and denoising pipelines.


STG Configuration: video_stg_scale, audio_stg_scale, and stg_blocks

STG perturbations are applied selectively to transformer blocks to control temporal consistency without destabilizing generation.

Core STG Fields

Field Location Default Purpose
video_stg_scale config.py:521 1.0 Scales STG perturbation strength for video self‑attention; 0.0 disables it
audio_stg_scale config.py:527 1.0 Scales STG perturbation strength for audio self‑attention; 0.0 disables it
stg_blocks config.py:533 [28] List of block indices receiving STG; None applies to all blocks

The stg_blocks parameter gives surgical control. By default, only block 28 is perturbed, which the LTX‑2 authors found to balance quality and guidance effectiveness.

How STG Is Applied Internally

In validation_runner.py (lines 928‑1290), the method _build_stg_perturbation_config constructs Perturbation objects with types SKIP_VIDEO_SELF_ATTN and SKIP_AUDIO_SELF_ATTN. These are passed as MultiModalGuiderParams to the denoiser pipeline.


# Conceptual flow inside validation_runner.py

stg_perturbation_config = self._build_stg_perturbation_config(
    video_stg_scale=cfg.video_stg_scale,
    audio_stg_scale=cfg.audio_stg_scale,
    stg_blocks=cfg.stg_blocks,
)

# Perturbations later injected into MultiModalGuiderParams

Modality Guidance Scales: Cross‑Modal Isolation Control

Modality guidance scales control how strongly one modality influences the other during joint generation. This is critical for tasks like lip‑sync (audio‑driven video) or sound effects generation (video‑driven audio).

Core Modality Fields

Field Location Default Purpose
video_modality_guidance_scale config.py:546 3.0 Audio‑to‑video isolation strength; higher values increase audio conditioning influence on video
audio_modality_guidance_scale config.py:552 3.0 Video‑to‑audio isolation strength; higher values increase video conditioning influence on audio

These values are read in validation_runner.py (lines 1229‑1251) and applied as modality_scale factors when assembling guidance parameters for each modality stream.


Guidance Rescale: The Final Output Scaling Factor

After CFG, STG, and modality guidance are combined, guidance_rescale applies a uniform scaling factor to the final logits before sampling.

  • Field: guidance_rescale (config.py:538)
  • Default: 0.7
  • Effect: 0.0 disables combined guidance entirely; values closer to 1.0 preserve full guidance strength

The rescaling is implemented in packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py at line 26, where the accumulated guidance is multiplied by this factor before the final sampling step.


Practical Configuration Examples

YAML Configuration Override

The standard way to adjust guidance scales is through the trainer YAML config, as shown in packages/ltx-trainer/configs/video_suffix_lora.yaml:


# packages/ltx-trainer/configs/video_suffix_lora.yaml

guidance_rescale: 0.7                # Scale combined guidance (0 → disable all)

video_modality_guidance_scale: 3.0   # Audio→Video isolation strength

audio_modality_guidance_scale: 3.0   # Video→Audio isolation strength

video_stg_scale: 1.0                 # Video STG strength (0 → disable STG)

audio_stg_scale: 1.0                 # Audio STG strength (0 → disable STG)

stg_blocks: [28]                     # Apply STG only on block 28

Programmatic Configuration

For dynamic experimentation, modify the Config object directly before trainer initialization:

from ltx_trainer.config import Config
from ltx_trainer.trainer import Trainer

cfg = Config.parse_yaml("configs/video_suffix_lora.yaml")

# Weaken overall guidance

cfg.guidance_rescale = 0.5

# Reduce cross‑modal isolation

cfg.video_modality_guidance_scale = 2.0
cfg.audio_modality_guidance_scale = 2.0

# Disable video STG, keep audio STG

cfg.video_stg_scale = 0.0
cfg.audio_stg_scale = 1.0

# Target multiple transformer blocks

cfg.stg_blocks = [20, 28, 30]

trainer = Trainer(cfg)
trainer.run()

Inspecting Applied Perturbations

During debugging, dump the perturbation configuration to verify which blocks receive STG:


# Inside validation_runner after building perturbations

print("STG perturbations:", stg_perturbation_config)

# Example output:

# STG perturbations: [

#   Perturbation(type=SKIP_VIDEO_SELF_ATTN, blocks=[20, 28, 30]),

#   Perturbation(type=SKIP_AUDIO_SELF_ATTN, blocks=[20, 28, 30])

# ]

Key Source Files Reference

File Lines Role
packages/ltx-trainer/src/ltx_trainer/config.py 521, 527, 533, 538, 546, 552 Defines all guidance scale fields and defaults
packages/ltx-trainer/docs/configuration-reference.md — User‑facing documentation of guidance parameters
packages/ltx-trainer/configs/video_suffix_lora.yaml — Example configuration with typical guidance values
packages/ltx-trainer/src/ltx_trainer/validation_runner.py 928‑1290, 1229‑1251 Builds STG perturbations and applies modality scales
packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py 26 Applies final guidance_rescale factor

Summary

  • Disable STG by setting video_stg_scale and/or audio_stg_scale to 0.0
  • Control block‑level STG with the stg_blocks list; default [28] targets a single mid‑depth transformer block
  • Adjust cross‑modal influence via video_modality_guidance_scale (audio→video) and audio_modality_guidance_scale (video→audio); defaults are 3.0
  • Scale final guidance output with guidance_rescale; 0.0 disables combined guidance, 0.7 is the conservative default
  • All fields are defined in ltx_trainer/config.py and consumed downstream in validation_runner.py and denoisers.py

Frequently Asked Questions

What happens if I set guidance_rescale to 0.0?

Setting guidance_rescale to 0.0 disables the combined classifier‑free guidance, STG, and modality guidance entirely. The model generates outputs without any explicit guidance conditioning, which typically produces lower‑quality but more "neutral" samples. This is defined at line 538 of config.py and applied at line 26 of denoisers.py.

How do I completely disable STG for video but keep it for audio?

Set video_stg_scale: 0.0 in your config while keeping audio_stg_scale at your desired positive value (default 1.0). The stg_blocks list still controls which blocks receive audio STG perturbation. This configuration is parsed in _build_stg_perturbation_config in validation_runner.py.

Why does the default stg_blocks only include block 28?

Block 28 represents a mid‑depth layer in the LTX‑2 transformer where spatio‑temporal features are well‑formed but not yet fully committed to fine details. Perturbing this single block provides effective temporal guidance without the computational cost or instability of applying STG to all layers. You can extend to multiple blocks (e.g., [20, 28, 30]) for stronger guidance at increased inference cost.

What is the difference between modality guidance scales and STG scales?

Modality guidance scales (video_modality_guidance_scale, audio_modality_guidance_scale) control cross‑modal isolation—how strongly audio conditions video generation and vice versa. STG scales (video_stg_scale, audio_stg_scale) control spatio‑temporal coherence within a single modality via self‑attention perturbations. They operate independently and can be tuned separately to achieve different generation behaviors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →