How Spatio-Temporal Guidance (STG) Works in LTX-2: A Technical Deep Dive

Spatio-Temporal Guidance (STG) improves temporal coherence in LTX-2 video generation by computing a scaled delta between standard denoised predictions and perturbed predictions where self-attention is disabled in selected transformer blocks, then steering the diffusion process toward the unperturbed result.

Spatio-Temporal Guidance (STG) is a specialized guidance mechanism in the Lightricks/LTX-2 repository that reduces flickering and enhances frame-to-frame consistency in generated video and audio. Unlike Classifier-Free Guidance (CFG), which contrasts conditional and unconditional predictions, STG operates by internally perturbing the model's self-attention patterns to create a temporal consistency signal. The implementation follows a three-stage pipeline involving perturbation configuration, delta computation via the STGGuider class, and integration within the Euler denoising loop.

Core STG Guiding Logic

The mathematical foundation of STG resides in ltx_core/components/guiders.py, where the STGGuider dataclass implements the GuiderProtocol:


# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py#L55-L71

@dataclass(frozen=True)
class STGGuider(GuiderProtocol):
    """
    Calculates the STG delta between conditioned and perturbed denoised samples.
    """
    scale: float                     # 0.0 disables STG; >0 applies guidance

    def delta(self, pos_denoised: torch.Tensor, perturbed_denoised: torch.Tensor) -> torch.Tensor:
        return self.scale * (pos_denoised - perturbed_denoised)

    def enabled(self) -> bool:
        return self.scale != 0.0

The delta() method computes the STG delta as Δ = scale × (positive − perturbed). Here, pos_denoised represents the standard forward pass, while perturbed_denoised comes from a pass where self-attention is skipped in specified transformer blocks. The scale parameter controls guidance strength—when set to 0.0, enabled() returns False and STG is bypassed entirely.

Configuring Perturbations

Before computing deltas, the system must define which layers to perturb. In ltx_trainer/validation_runner.py, the _build_stg_perturbation_config() function constructs a BatchedPerturbationConfig that specifies which transformer blocks have their self-attention disabled:


# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/validation_runner.py#L74-L84

def _build_stg_perturbation_config(stg_blocks, stg_mode):
    perturbations = [
        Perturbation(type=PerturbationType.SKIP_VIDEO_SELF_ATTN, blocks=stg_blocks)
    ]
    if stg_mode == "stg_av":               # also affect audio

        perturbations.append(
            Perturbation(type=PerturbationType.SKIP_AUDIO_SELF_ATTN, blocks=stg_blocks)
        )
    return BatchedPerturbationConfig(perturbations=[PerturbationConfig(perturbations=perturbations)])

The perturbation configuration accepts two critical parameters:

  • stg_blocks: A list of transformer block indices (e.g., [29]) where self-attention will be skipped.
  • stg_mode: Either "stg_v" (perturb video only) or "stg_av" (perturb both video and audio self-attention).

This configuration is passed to the model during the perturbed forward pass to temporarily disable attention mechanisms in the specified blocks.

Integration in the Denoising Loop

During each diffusion step in validation_runner.py, the system performs three distinct forward passes:

  1. Positive pass: Standard prediction without perturbations (perturbations=None).
  2. Negative pass: Unconditioned prediction for CFG (only when CFG is enabled).
  3. Perturbed pass: Prediction with the STG perturbation config applied.

The STG delta is then applied to steer the denoised latent:


# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/validation_runner.py#L11-L18

if stg_perturbation_config is not None:
    ptb_video, ptb_audio = x0_model(video=video, audio=audio,
                                    perturbations=stg_perturbation_config)
    if not video_frozen and denoised_video is not None:
        denoised_video = denoised_video + stg_guider.delta(pos_video, ptb_video)
    if not audio_frozen and denoised_audio is not None and ptb_audio is not None:
        denoised_audio = denoised_audio + stg_guider.delta(pos_audio, ptb_audio)

The delta pulls the latent toward the positive direction while penalizing the temporal inconsistencies introduced by the skipped self-attention. After applying the delta, the updated latents proceed to the Euler stepper (stepper.step) for the next diffusion iteration.

Configuration Parameters

STG parameters are exposed in ltx_trainer/config.py as user-configurable arguments:


# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py#L525-L538

stg_scale = Float(
    default=0.0,
    description="STG (Spatio‑Temporal Guidance) scale. 0.0 disables STG. Recommended value is 1.0."
)
stg_blocks = List(Int, default=[29],
    description="Which transformer blocks to perturb for STG."
)
stg_mode = Enum(
    choices=["stg_av", "stg_v"],
    default="stg_av",
    description="STG mode: 'stg_av' skips both audio and video self‑attention, 'stg_v' skips video only."
)

Recommended settings from the LTX-2 source code include:

  • stg_scale: Set to 1.0 for active guidance, 0.0 to disable.
  • stg_blocks: The default [29] targets specific transformer layers for perturbation.
  • stg_mode: Use "stg_av" for audio-video synchronization or "stg_v" for video-only temporal coherence.

Implementation Example

To enable STG in a validation script, instantiate the configuration with the desired parameters:

from ltx_trainer.validation_runner import ValidationRunner
from ltx_trainer.config import ValidationConfig

cfg = ValidationConfig(
    stg_scale=1.0,
    stg_blocks=[29],
    stg_mode="stg_av",
    guidance_scale=7.0,       # CFG scale

)

runner = ValidationRunner(cfg)
video_latent, audio_latent = runner.run()   # STG applied automatically

The ValidationRunner internally constructs the BatchedPerturbationConfig, instantiates the STGGuider, and applies the guidance delta during each denoising step.

Summary

  • Spatio-Temporal Guidance (STG) improves temporal consistency by perturbing self-attention in selected transformer blocks and steering toward the unperturbed prediction.
  • The STGGuider class in ltx_core/components/guiders.py computes the guidance delta as scale × (positive − perturbed).
  • stg_blocks controls which layers are perturbed (default [29]), while stg_mode selects between video-only ("stg_v") or audio-video ("stg_av") perturbation.
  • The delta is applied to denoised latents in validation_runner.py before the Euler step, pulling the generation toward temporally coherent outputs.
  • Set stg_scale to 1.0 to enable STG or 0.0 to disable it entirely.

Frequently Asked Questions

What is the difference between STG and CFG in LTX-2?

Classifier-Free Guidance (CFG) computes the difference between conditional and unconditional predictions to enforce prompt adherence, while Spatio-Temporal Guidance (STG) computes the difference between standard predictions and predictions where self-attention is disabled in specific blocks to enforce temporal coherence. According to the LTX-2 source code in validation_runner.py, these mechanisms operate independently and can be applied simultaneously during inference.

Which transformer blocks should I specify for stg_blocks?

The default configuration in ltx_trainer/config.py uses [29], targeting a specific middle-to-late transformer block. While the repository recommends this default for general use, you can specify multiple blocks (e.g., [28, 29, 30]) to increase the perturbation strength. Note that perturbing more blocks or earlier blocks may significantly alter the temporal characteristics of the output.

Why does STG have separate modes for video and audio?

The stg_mode parameter accommodates different media requirements. When set to "stg_av", the perturbation config includes both SKIP_VIDEO_SELF_ATTN and SKIP_AUDIO_SELF_ATTN, ensuring temporal consistency across both modalities. When set to "stg_v", only video layers are perturbed, leaving audio generation unaffected. This flexibility allows users to optimize STG for video-only models or to prevent audio quality degradation when temporal coherence is less critical for sound.

How does the stg_scale parameter affect video quality?

The stg_scale acts as a multiplier for the guidance delta. A value of 0.0 disables STG entirely, while the recommended value of 1.0 provides balanced temporal smoothing without over-constraining the generation. Values significantly higher than 1.0 may over-penalize the perturbed predictions, potentially leading to over-smoothed or temporally rigid outputs. The STGGuider.enabled() method returns False when scale equals 0.0, efficiently bypassing the perturbed forward pass for computational savings.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →