How Self-Targeted Guidance (STG) Functions in LTX-2 and Choosing Optimal `stg_scale` Values

Self-Targeted Guidance (STG) steers the diffusion denoising process in LTX-2 by perturbing self-attention layers in specified transformer blocks, with stg_scale controlling the blend between conditioned and perturbed predictions.

STG is a modality-specific guidance mechanism implemented in the Lightricks/LTX-2 repository that improves temporal consistency and cross-modal alignment. Unlike traditional classifier-free guidance, STG creates a "perturbed" unconditioned prediction by selectively masking self-attention, then blends this with the standard conditioned output. This article explains the internal mechanics based on the actual source code and provides practical guidance for tuning stg_scale values.

Where STG Lives in the LTX-2 Architecture

STG parameters are defined in MultiModalGuiderParams within packages/ltx-core/src/ltx_core/components/guiders.py:

  • stg_scale — the guidance strength multiplier
  • stg_blocks — list of transformer block indices to perturb

The core blending formula appears at lines 62-66 of guiders.py:


# Simplified view of the calculation

guided_output = cond + cfg_scale * (cond - uncond_text) + stg_scale * (cond - uncond_perturbed)

The uncond_perturbed term is what distinguishes STG from standard CFG. It comes from a forward pass where self-attention is masked in the blocks specified by stg_blocks.

How the Self-Attention Perturbation Works

The Skip-Self-Attention Mechanism

The perturbation type is defined in packages/ltx-core/src/ltx_core/guidance/perturbations.py (lines 13-16):

class PerturbationType(Enum):
    SKIP_VIDEO_SELF_ATTN = "skip_video_self_attn"
    SKIP_AUDIO_SELF_ATTN = "skip_audio_self_attn"
    ZERO_VIDEO_SELF_ATTN = "zero_video_self_attn"  # Alternative perturbation

    ZERO_AUDIO_SELF_ATTN = "zero_audio_self_attn"

During validation inference, the ValidationRunner constructs a BatchedPerturbationConfig that instructs the transformer to skip self-attention for the unconditioned pass. From packages/ltx-trainer/src/ltx_trainer/validation_runner.py (lines 1288-1291):

if video_enabled:
    perturbations.append(
        Perturbation(type=PerturbationType.SKIP_VIDEO_SELF_ATTN, blocks=stg_blocks)
    )
if audio_enabled:
    perturbations.append(
        Perturbation(type=PerturbationType.SKIP_AUDIO_SELF_ATTN, blocks=stg_blocks)
    )

This creates a batch where:

  • Conditioned sample: Full self-attention in all blocks
  • Unconditioned sample: Self-attention masked in stg_blocks
  • Perturbed unconditioned sample: Same latent as unconditioned, but with degraded self-attention

The difference between conditioned and perturbed predictions reveals how much the model relies on self-attention for coherence—STG amplifies this signal.

Default STG Configuration Values

Parameter Default Source Location
video_stg_scale 1.0 packages/ltx-trainer/src/ltx_trainer/config.py lines 521-525
audio_stg_scale 1.0 packages/ltx-trainer/src/ltx_trainer/config.py lines 527-531
stg_blocks (LTX-2.3) [28] packages/ltx-pipelines/CLAUDE.md line 34
stg_blocks (original LTX-2) [29] packages/ltx-pipelines/CLAUDE.md line 34

CLI flags exposing these settings are defined in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 59-70):

  • --video-stg-guidance-scale
  • --audio-stg-guidance-scale
  • --video-stg-blocks
  • --audio-stg-blocks

Optimal stg_scale Values: Practical Guidance

Default Behavior (stg_scale = 1.0)

The default value of 1.0 provides balanced guidance. According to the CLAUDE.md documentation and config.py defaults, this setting:

  • Enhances temporal consistency in video generation
  • Improves audio-visual alignment in multimodal outputs
  • Avoids the artifacts common with excessive guidance

Lower Values (0.0 to 0.5)

Reduce stg_scale when:

  • Flickering or temporal inconsistency appears — lower values reduce the "forced" coherence that can cause oscillation
  • Running low-VRAM or fast inference — the perturbed pass requires additional computation; lowering the scale or setting to 0.0 skips this overhead
  • Seeking more "natural" diffusion trajectories — less intervention preserves the model's native noise-to-data path

Setting stg_scale = 0.0 completely disables STG without affecting standard classifier-free guidance.

Higher Values (1.0 to 2.0)

Increase stg_scale when:

  • Strong temporal anchoring is needed — audio-driven video synthesis benefits from aggressive STG to lock lip-sync and motion to audio cues
  • Cross-modal consistency is critical — values of 1.5 to 2.0 can reinforce alignment between video and audio streams

Avoid exceeding 2.0. Higher values typically introduce:

  • Over-sharpened edges and plastic appearance in video
  • Audio distortion or "metallic" artifacts in generated sound
  • Reduced diversity in outputs

Special Case: HQ Pipelines

The High-Quality pipelines (*_hq variants) deliberately disable STG. From packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py (lines 101-113), these pipelines set:

video_stg_scale=0.0,
audio_stg_scale=0.0,

The second-order samplers used in HQ pipelines provide fine-grained control through their internal schedules; STG interference would disrupt this precision.

Practical Configuration Examples

YAML Configuration File


# configs/v2a_lora.yaml — video-to-audio with tuned STG

video_stg_scale: 1.0      # Enable video STG for motion coherence

audio_stg_scale: 0.0      # Disable audio STG (audio is the target, not source)

stg_blocks: [28]          # Perturb block 28 (LTX-2.3 architecture)

Command-Line Override

uv run python scripts/train.py \
    configs/t2v_lora.yaml \
    --video-stg-guidance-scale 0.8 \
    --video-stg-blocks 27 28 29

This applies reduced STG strength across three consecutive blocks for smoother transitions.

Programmatic MultiModalGuider Creation

from ltx_core.components.guiders import MultiModalGuiderParams, MultiModalGuider

params = MultiModalGuiderParams(
    cfg_scale=3.0,
    stg_scale=1.2,          # Stronger STG than default

    stg_blocks=[28],        # Single-block perturbation

    modality_scale=1.0,
)
guider = MultiModalGuider(params=params)

guided = guider.calculate(
    cond=cond_pred,
    uncond_text=uncond_pred,
    uncond_perturbed=perturbed_pred,
    uncond_modality=uncond_pred,
)

Debugging Perturbation Masks

from ltx_core.guidance.perturbations import (
    PerturbationType,
    PerturbationConfig,
    BatchedPerturbationConfig,
    Perturbation,
)

# Configure skip-self-attention on block 28 for video

pert = Perturbation(
    type=PerturbationType.SKIP_VIDEO_SELF_ATTN,
    blocks=[28]
)
cfg = PerturbationConfig(perturbations=[pert])
batched_cfg = BatchedPerturbationConfig([cfg], num_blocks=32)

# Inspect the actual mask tensor

mask = batched_cfg.mask(PerturbationType.SKIP_VIDEO_SELF_ATTN, block=28)
print(mask.shape)   # (batch_size, 1, 1)

print(mask)         # 0.0 where attention is masked, 1.0 otherwise

Key Source Files for STG Implementation

File Path Purpose
packages/ltx-core/src/ltx_core/components/guiders.py MultiModalGuiderParams, MultiModalGuider, and the STG blending formula
packages/ltx-core/src/ltx_core/guidance/perturbations.py PerturbationType enum and perturbation configuration classes
packages/ltx-trainer/src/ltx_trainer/config.py User-facing video_stg_scale, audio_stg_scale, and stg_blocks defaults
packages/ltx-trainer/src/ltx_trainer/validation_runner.py BatchedPerturbationConfig construction for inference
packages/ltx-pipelines/src/ltx_pipelines/utils/args.py CLI argument definitions for STG parameters
packages/ltx-pipelines/src/ltx_pipelines/utils/constants.py HQ pipeline defaults that disable STG

Summary

  • STG perturbs self-attention in specified transformer blocks to create a degraded "unconditioned" prediction, then blends this with the conditioned output using stg_scale
  • Default stg_scale = 1.0 works well for most generation tasks in both video and audio modalities
  • Lower values (0.0–0.5) reduce artifacts and computation; 0.0 fully disables STG
  • Higher values (1.0–2.0) strengthen temporal consistency but risk quality degradation above 2.0
  • HQ pipelines disable STG entirely to avoid interference with second-order samplers
  • Block selection matters: [28] for LTX-2.3, [29] for original LTX-2

Frequently Asked Questions

What happens when stg_scale is set to 0.0?

Setting stg_scale = 0.0 removes the STG term from the guidance calculation, reducing the formula to standard classifier-free guidance. The perturbed pass still runs but contributes zero weight. This is how HQ pipelines operate according to constants.py lines 101-113.

Why do different LTX-2 versions use different stg_blocks?

The LTX-2.3 architecture places critical self-attention mechanisms at block 28, while the original LTX-2 uses block 29. These specific blocks were identified as optimal intervention points through architectural analysis. The CLAUDE.md documentation at line 34 documents these version-specific defaults.

Can STG be applied to only one modality?

Yes. The configuration system separates video_stg_scale and audio_stg_scale, allowing independent control. Set the unused modality to 0.0 while keeping the target modality at your desired strength. This is common in video-to-audio tasks where audio STG is disabled.

How does STG differ from standard classifier-free guidance?

Standard CFG compares conditioned and unconditioned predictions. STG introduces a third "perturbed unconditioned" prediction where self-attention is masked. The difference between unconditioned and perturbed reveals self-attention's contribution to coherence, which STG amplifies through stg_scale.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →