How Self-Targeted Guidance (STG) Functions in LTX-2 and Choosing Optimal `stg_scale` Values
Self-Targeted Guidance (STG) steers the diffusion denoising process in LTX-2 by perturbing self-attention layers in specified transformer blocks, with stg_scale controlling the blend between conditioned and perturbed predictions.
STG is a modality-specific guidance mechanism implemented in the Lightricks/LTX-2 repository that improves temporal consistency and cross-modal alignment. Unlike traditional classifier-free guidance, STG creates a "perturbed" unconditioned prediction by selectively masking self-attention, then blends this with the standard conditioned output. This article explains the internal mechanics based on the actual source code and provides practical guidance for tuning stg_scale values.
Where STG Lives in the LTX-2 Architecture
STG parameters are defined in MultiModalGuiderParams within packages/ltx-core/src/ltx_core/components/guiders.py:
stg_scale— the guidance strength multiplierstg_blocks— list of transformer block indices to perturb
The core blending formula appears at lines 62-66 of guiders.py:
# Simplified view of the calculation
guided_output = cond + cfg_scale * (cond - uncond_text) + stg_scale * (cond - uncond_perturbed)
The uncond_perturbed term is what distinguishes STG from standard CFG. It comes from a forward pass where self-attention is masked in the blocks specified by stg_blocks.
How the Self-Attention Perturbation Works
The Skip-Self-Attention Mechanism
The perturbation type is defined in packages/ltx-core/src/ltx_core/guidance/perturbations.py (lines 13-16):
class PerturbationType(Enum):
SKIP_VIDEO_SELF_ATTN = "skip_video_self_attn"
SKIP_AUDIO_SELF_ATTN = "skip_audio_self_attn"
ZERO_VIDEO_SELF_ATTN = "zero_video_self_attn" # Alternative perturbation
ZERO_AUDIO_SELF_ATTN = "zero_audio_self_attn"
During validation inference, the ValidationRunner constructs a BatchedPerturbationConfig that instructs the transformer to skip self-attention for the unconditioned pass. From packages/ltx-trainer/src/ltx_trainer/validation_runner.py (lines 1288-1291):
if video_enabled:
perturbations.append(
Perturbation(type=PerturbationType.SKIP_VIDEO_SELF_ATTN, blocks=stg_blocks)
)
if audio_enabled:
perturbations.append(
Perturbation(type=PerturbationType.SKIP_AUDIO_SELF_ATTN, blocks=stg_blocks)
)
This creates a batch where:
- Conditioned sample: Full self-attention in all blocks
- Unconditioned sample: Self-attention masked in
stg_blocks - Perturbed unconditioned sample: Same latent as unconditioned, but with degraded self-attention
The difference between conditioned and perturbed predictions reveals how much the model relies on self-attention for coherence—STG amplifies this signal.
Default STG Configuration Values
| Parameter | Default | Source Location |
|---|---|---|
video_stg_scale |
1.0 |
packages/ltx-trainer/src/ltx_trainer/config.py lines 521-525 |
audio_stg_scale |
1.0 |
packages/ltx-trainer/src/ltx_trainer/config.py lines 527-531 |
stg_blocks (LTX-2.3) |
[28] |
packages/ltx-pipelines/CLAUDE.md line 34 |
stg_blocks (original LTX-2) |
[29] |
packages/ltx-pipelines/CLAUDE.md line 34 |
CLI flags exposing these settings are defined in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 59-70):
--video-stg-guidance-scale--audio-stg-guidance-scale--video-stg-blocks--audio-stg-blocks
Optimal stg_scale Values: Practical Guidance
Default Behavior (stg_scale = 1.0)
The default value of 1.0 provides balanced guidance. According to the CLAUDE.md documentation and config.py defaults, this setting:
- Enhances temporal consistency in video generation
- Improves audio-visual alignment in multimodal outputs
- Avoids the artifacts common with excessive guidance
Lower Values (0.0 to 0.5)
Reduce stg_scale when:
- Flickering or temporal inconsistency appears — lower values reduce the "forced" coherence that can cause oscillation
- Running low-VRAM or fast inference — the perturbed pass requires additional computation; lowering the scale or setting to
0.0skips this overhead - Seeking more "natural" diffusion trajectories — less intervention preserves the model's native noise-to-data path
Setting stg_scale = 0.0 completely disables STG without affecting standard classifier-free guidance.
Higher Values (1.0 to 2.0)
Increase stg_scale when:
- Strong temporal anchoring is needed — audio-driven video synthesis benefits from aggressive STG to lock lip-sync and motion to audio cues
- Cross-modal consistency is critical — values of
1.5to2.0can reinforce alignment between video and audio streams
Avoid exceeding 2.0. Higher values typically introduce:
- Over-sharpened edges and plastic appearance in video
- Audio distortion or "metallic" artifacts in generated sound
- Reduced diversity in outputs
Special Case: HQ Pipelines
The High-Quality pipelines (*_hq variants) deliberately disable STG. From packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py (lines 101-113), these pipelines set:
video_stg_scale=0.0,
audio_stg_scale=0.0,
The second-order samplers used in HQ pipelines provide fine-grained control through their internal schedules; STG interference would disrupt this precision.
Practical Configuration Examples
YAML Configuration File
# configs/v2a_lora.yaml — video-to-audio with tuned STG
video_stg_scale: 1.0 # Enable video STG for motion coherence
audio_stg_scale: 0.0 # Disable audio STG (audio is the target, not source)
stg_blocks: [28] # Perturb block 28 (LTX-2.3 architecture)
Command-Line Override
uv run python scripts/train.py \
configs/t2v_lora.yaml \
--video-stg-guidance-scale 0.8 \
--video-stg-blocks 27 28 29
This applies reduced STG strength across three consecutive blocks for smoother transitions.
Programmatic MultiModalGuider Creation
from ltx_core.components.guiders import MultiModalGuiderParams, MultiModalGuider
params = MultiModalGuiderParams(
cfg_scale=3.0,
stg_scale=1.2, # Stronger STG than default
stg_blocks=[28], # Single-block perturbation
modality_scale=1.0,
)
guider = MultiModalGuider(params=params)
guided = guider.calculate(
cond=cond_pred,
uncond_text=uncond_pred,
uncond_perturbed=perturbed_pred,
uncond_modality=uncond_pred,
)
Debugging Perturbation Masks
from ltx_core.guidance.perturbations import (
PerturbationType,
PerturbationConfig,
BatchedPerturbationConfig,
Perturbation,
)
# Configure skip-self-attention on block 28 for video
pert = Perturbation(
type=PerturbationType.SKIP_VIDEO_SELF_ATTN,
blocks=[28]
)
cfg = PerturbationConfig(perturbations=[pert])
batched_cfg = BatchedPerturbationConfig([cfg], num_blocks=32)
# Inspect the actual mask tensor
mask = batched_cfg.mask(PerturbationType.SKIP_VIDEO_SELF_ATTN, block=28)
print(mask.shape) # (batch_size, 1, 1)
print(mask) # 0.0 where attention is masked, 1.0 otherwise
Key Source Files for STG Implementation
| File Path | Purpose |
|---|---|
packages/ltx-core/src/ltx_core/components/guiders.py |
MultiModalGuiderParams, MultiModalGuider, and the STG blending formula |
packages/ltx-core/src/ltx_core/guidance/perturbations.py |
PerturbationType enum and perturbation configuration classes |
packages/ltx-trainer/src/ltx_trainer/config.py |
User-facing video_stg_scale, audio_stg_scale, and stg_blocks defaults |
packages/ltx-trainer/src/ltx_trainer/validation_runner.py |
BatchedPerturbationConfig construction for inference |
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py |
CLI argument definitions for STG parameters |
packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py |
HQ pipeline defaults that disable STG |
Summary
- STG perturbs self-attention in specified transformer blocks to create a degraded "unconditioned" prediction, then blends this with the conditioned output using
stg_scale - Default
stg_scale = 1.0works well for most generation tasks in both video and audio modalities - Lower values (0.0–0.5) reduce artifacts and computation; 0.0 fully disables STG
- Higher values (1.0–2.0) strengthen temporal consistency but risk quality degradation above 2.0
- HQ pipelines disable STG entirely to avoid interference with second-order samplers
- Block selection matters:
[28]for LTX-2.3,[29]for original LTX-2
Frequently Asked Questions
What happens when stg_scale is set to 0.0?
Setting stg_scale = 0.0 removes the STG term from the guidance calculation, reducing the formula to standard classifier-free guidance. The perturbed pass still runs but contributes zero weight. This is how HQ pipelines operate according to constants.py lines 101-113.
Why do different LTX-2 versions use different stg_blocks?
The LTX-2.3 architecture places critical self-attention mechanisms at block 28, while the original LTX-2 uses block 29. These specific blocks were identified as optimal intervention points through architectural analysis. The CLAUDE.md documentation at line 34 documents these version-specific defaults.
Can STG be applied to only one modality?
Yes. The configuration system separates video_stg_scale and audio_stg_scale, allowing independent control. Set the unused modality to 0.0 while keeping the target modality at your desired strength. This is common in video-to-audio tasks where audio STG is disabled.
How does STG differ from standard classifier-free guidance?
Standard CFG compares conditioned and unconditioned predictions. STG introduces a third "perturbed unconditioned" prediction where self-attention is masked. The difference between unconditioned and perturbed reveals self-attention's contribution to coherence, which STG amplifies through stg_scale.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →