LTX‑2 STG and Modality Guidance Scales Configuration: A Complete Guide
Set video_stg_scale and audio_stg_scale to 0.0 to disable STG, adjust video_modality_guidance_scale and audio_modality_guidance_scale to control cross‑modal isolation strength, and use guidance_rescale to scale the combined guidance output—each field is defined in ltx_trainer/config.py and applied during the validation denoising loop.
LTX‑2, the open‑source video‑audio generation model from Lightricks, exposes a unified configuration system for fine‑tuning generation behavior. Understanding how Spatio‑Temporal Guidance (STG) and modality guidance scales interact is essential for controlling trade‑offs between cross‑modal conditioning and output fidelity. This guide walks through every configuration field, its default value, and where it is implemented in the LTX‑2 source code.
What STG and Modality Guidance Scales Control
LTX‑2 applies guidance at three interconnected levels during generation:
- Spatio‑Temporal Guidance (STG) – perturbs self‑attention blocks to guide temporal coherence
- Modality Guidance – isolates or strengthens audio‑to‑video and video‑to‑audio conditioning
- Guidance Rescale – uniformly scales the combined effect of all guidance mechanisms
These settings live in the Config class at packages/ltx-trainer/src/ltx_trainer/config.py and are consumed by the validation runner and denoising pipelines.
STG Configuration: video_stg_scale, audio_stg_scale, and stg_blocks
STG perturbations are applied selectively to transformer blocks to control temporal consistency without destabilizing generation.
Core STG Fields
| Field | Location | Default | Purpose |
|---|---|---|---|
video_stg_scale |
config.py:521 | 1.0 |
Scales STG perturbation strength for video self‑attention; 0.0 disables it |
audio_stg_scale |
config.py:527 | 1.0 |
Scales STG perturbation strength for audio self‑attention; 0.0 disables it |
stg_blocks |
config.py:533 | [28] |
List of block indices receiving STG; None applies to all blocks |
The stg_blocks parameter gives surgical control. By default, only block 28 is perturbed, which the LTX‑2 authors found to balance quality and guidance effectiveness.
How STG Is Applied Internally
In validation_runner.py (lines 928‑1290), the method _build_stg_perturbation_config constructs Perturbation objects with types SKIP_VIDEO_SELF_ATTN and SKIP_AUDIO_SELF_ATTN. These are passed as MultiModalGuiderParams to the denoiser pipeline.
# Conceptual flow inside validation_runner.py
stg_perturbation_config = self._build_stg_perturbation_config(
video_stg_scale=cfg.video_stg_scale,
audio_stg_scale=cfg.audio_stg_scale,
stg_blocks=cfg.stg_blocks,
)
# Perturbations later injected into MultiModalGuiderParams
Modality Guidance Scales: Cross‑Modal Isolation Control
Modality guidance scales control how strongly one modality influences the other during joint generation. This is critical for tasks like lip‑sync (audio‑driven video) or sound effects generation (video‑driven audio).
Core Modality Fields
| Field | Location | Default | Purpose |
|---|---|---|---|
video_modality_guidance_scale |
config.py:546 | 3.0 |
Audio‑to‑video isolation strength; higher values increase audio conditioning influence on video |
audio_modality_guidance_scale |
config.py:552 | 3.0 |
Video‑to‑audio isolation strength; higher values increase video conditioning influence on audio |
These values are read in validation_runner.py (lines 1229‑1251) and applied as modality_scale factors when assembling guidance parameters for each modality stream.
Guidance Rescale: The Final Output Scaling Factor
After CFG, STG, and modality guidance are combined, guidance_rescale applies a uniform scaling factor to the final logits before sampling.
- Field:
guidance_rescale(config.py:538) - Default:
0.7 - Effect:
0.0disables combined guidance entirely; values closer to 1.0 preserve full guidance strength
The rescaling is implemented in packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py at line 26, where the accumulated guidance is multiplied by this factor before the final sampling step.
Practical Configuration Examples
YAML Configuration Override
The standard way to adjust guidance scales is through the trainer YAML config, as shown in packages/ltx-trainer/configs/video_suffix_lora.yaml:
# packages/ltx-trainer/configs/video_suffix_lora.yaml
guidance_rescale: 0.7 # Scale combined guidance (0 → disable all)
video_modality_guidance_scale: 3.0 # Audio→Video isolation strength
audio_modality_guidance_scale: 3.0 # Video→Audio isolation strength
video_stg_scale: 1.0 # Video STG strength (0 → disable STG)
audio_stg_scale: 1.0 # Audio STG strength (0 → disable STG)
stg_blocks: [28] # Apply STG only on block 28
Programmatic Configuration
For dynamic experimentation, modify the Config object directly before trainer initialization:
from ltx_trainer.config import Config
from ltx_trainer.trainer import Trainer
cfg = Config.parse_yaml("configs/video_suffix_lora.yaml")
# Weaken overall guidance
cfg.guidance_rescale = 0.5
# Reduce cross‑modal isolation
cfg.video_modality_guidance_scale = 2.0
cfg.audio_modality_guidance_scale = 2.0
# Disable video STG, keep audio STG
cfg.video_stg_scale = 0.0
cfg.audio_stg_scale = 1.0
# Target multiple transformer blocks
cfg.stg_blocks = [20, 28, 30]
trainer = Trainer(cfg)
trainer.run()
Inspecting Applied Perturbations
During debugging, dump the perturbation configuration to verify which blocks receive STG:
# Inside validation_runner after building perturbations
print("STG perturbations:", stg_perturbation_config)
# Example output:
# STG perturbations: [
# Perturbation(type=SKIP_VIDEO_SELF_ATTN, blocks=[20, 28, 30]),
# Perturbation(type=SKIP_AUDIO_SELF_ATTN, blocks=[20, 28, 30])
# ]
Key Source Files Reference
| File | Lines | Role |
|---|---|---|
packages/ltx-trainer/src/ltx_trainer/config.py |
521, 527, 533, 538, 546, 552 | Defines all guidance scale fields and defaults |
packages/ltx-trainer/docs/configuration-reference.md |
— | User‑facing documentation of guidance parameters |
packages/ltx-trainer/configs/video_suffix_lora.yaml |
— | Example configuration with typical guidance values |
packages/ltx-trainer/src/ltx_trainer/validation_runner.py |
928‑1290, 1229‑1251 | Builds STG perturbations and applies modality scales |
packages/ltx-pipelines/src/ltx_pipelines/utils/denoisers.py |
26 | Applies final guidance_rescale factor |
Summary
- Disable STG by setting
video_stg_scaleand/oraudio_stg_scaleto0.0 - Control block‑level STG with the
stg_blockslist; default[28]targets a single mid‑depth transformer block - Adjust cross‑modal influence via
video_modality_guidance_scale(audio→video) andaudio_modality_guidance_scale(video→audio); defaults are3.0 - Scale final guidance output with
guidance_rescale;0.0disables combined guidance,0.7is the conservative default - All fields are defined in
ltx_trainer/config.pyand consumed downstream invalidation_runner.pyanddenoisers.py
Frequently Asked Questions
What happens if I set guidance_rescale to 0.0?
Setting guidance_rescale to 0.0 disables the combined classifier‑free guidance, STG, and modality guidance entirely. The model generates outputs without any explicit guidance conditioning, which typically produces lower‑quality but more "neutral" samples. This is defined at line 538 of config.py and applied at line 26 of denoisers.py.
How do I completely disable STG for video but keep it for audio?
Set video_stg_scale: 0.0 in your config while keeping audio_stg_scale at your desired positive value (default 1.0). The stg_blocks list still controls which blocks receive audio STG perturbation. This configuration is parsed in _build_stg_perturbation_config in validation_runner.py.
Why does the default stg_blocks only include block 28?
Block 28 represents a mid‑depth layer in the LTX‑2 transformer where spatio‑temporal features are well‑formed but not yet fully committed to fine details. Perturbing this single block provides effective temporal guidance without the computational cost or instability of applying STG to all layers. You can extend to multiple blocks (e.g., [20, 28, 30]) for stronger guidance at increased inference cost.
What is the difference between modality guidance scales and STG scales?
Modality guidance scales (video_modality_guidance_scale, audio_modality_guidance_scale) control cross‑modal isolation—how strongly audio conditions video generation and vice versa. STG scales (video_stg_scale, audio_stg_scale) control spatio‑temporal coherence within a single modality via self‑attention perturbations. They operate independently and can be tuned separately to achieve different generation behaviors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →