How Spatio-Temporal Guidance (STG) Works in LTX-2: A Technical Deep Dive
Spatio-Temporal Guidance (STG) improves temporal coherence in LTX-2 video generation by computing a scaled delta between standard denoised predictions and perturbed predictions where self-attention is disabled in selected transformer blocks, then steering the diffusion process toward the unperturbed result.
Spatio-Temporal Guidance (STG) is a specialized guidance mechanism in the Lightricks/LTX-2 repository that reduces flickering and enhances frame-to-frame consistency in generated video and audio. Unlike Classifier-Free Guidance (CFG), which contrasts conditional and unconditional predictions, STG operates by internally perturbing the model's self-attention patterns to create a temporal consistency signal. The implementation follows a three-stage pipeline involving perturbation configuration, delta computation via the STGGuider class, and integration within the Euler denoising loop.
Core STG Guiding Logic
The mathematical foundation of STG resides in ltx_core/components/guiders.py, where the STGGuider dataclass implements the GuiderProtocol:
# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py#L55-L71
@dataclass(frozen=True)
class STGGuider(GuiderProtocol):
"""
Calculates the STG delta between conditioned and perturbed denoised samples.
"""
scale: float # 0.0 disables STG; >0 applies guidance
def delta(self, pos_denoised: torch.Tensor, perturbed_denoised: torch.Tensor) -> torch.Tensor:
return self.scale * (pos_denoised - perturbed_denoised)
def enabled(self) -> bool:
return self.scale != 0.0
The delta() method computes the STG delta as Δ = scale × (positive − perturbed). Here, pos_denoised represents the standard forward pass, while perturbed_denoised comes from a pass where self-attention is skipped in specified transformer blocks. The scale parameter controls guidance strength—when set to 0.0, enabled() returns False and STG is bypassed entirely.
Configuring Perturbations
Before computing deltas, the system must define which layers to perturb. In ltx_trainer/validation_runner.py, the _build_stg_perturbation_config() function constructs a BatchedPerturbationConfig that specifies which transformer blocks have their self-attention disabled:
# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/validation_runner.py#L74-L84
def _build_stg_perturbation_config(stg_blocks, stg_mode):
perturbations = [
Perturbation(type=PerturbationType.SKIP_VIDEO_SELF_ATTN, blocks=stg_blocks)
]
if stg_mode == "stg_av": # also affect audio
perturbations.append(
Perturbation(type=PerturbationType.SKIP_AUDIO_SELF_ATTN, blocks=stg_blocks)
)
return BatchedPerturbationConfig(perturbations=[PerturbationConfig(perturbations=perturbations)])
The perturbation configuration accepts two critical parameters:
stg_blocks: A list of transformer block indices (e.g.,[29]) where self-attention will be skipped.stg_mode: Either"stg_v"(perturb video only) or"stg_av"(perturb both video and audio self-attention).
This configuration is passed to the model during the perturbed forward pass to temporarily disable attention mechanisms in the specified blocks.
Integration in the Denoising Loop
During each diffusion step in validation_runner.py, the system performs three distinct forward passes:
- Positive pass: Standard prediction without perturbations (
perturbations=None). - Negative pass: Unconditioned prediction for CFG (only when CFG is enabled).
- Perturbed pass: Prediction with the STG perturbation config applied.
The STG delta is then applied to steer the denoised latent:
# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/validation_runner.py#L11-L18
if stg_perturbation_config is not None:
ptb_video, ptb_audio = x0_model(video=video, audio=audio,
perturbations=stg_perturbation_config)
if not video_frozen and denoised_video is not None:
denoised_video = denoised_video + stg_guider.delta(pos_video, ptb_video)
if not audio_frozen and denoised_audio is not None and ptb_audio is not None:
denoised_audio = denoised_audio + stg_guider.delta(pos_audio, ptb_audio)
The delta pulls the latent toward the positive direction while penalizing the temporal inconsistencies introduced by the skipped self-attention. After applying the delta, the updated latents proceed to the Euler stepper (stepper.step) for the next diffusion iteration.
Configuration Parameters
STG parameters are exposed in ltx_trainer/config.py as user-configurable arguments:
# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py#L525-L538
stg_scale = Float(
default=0.0,
description="STG (Spatio‑Temporal Guidance) scale. 0.0 disables STG. Recommended value is 1.0."
)
stg_blocks = List(Int, default=[29],
description="Which transformer blocks to perturb for STG."
)
stg_mode = Enum(
choices=["stg_av", "stg_v"],
default="stg_av",
description="STG mode: 'stg_av' skips both audio and video self‑attention, 'stg_v' skips video only."
)
Recommended settings from the LTX-2 source code include:
stg_scale: Set to1.0for active guidance,0.0to disable.stg_blocks: The default[29]targets specific transformer layers for perturbation.stg_mode: Use"stg_av"for audio-video synchronization or"stg_v"for video-only temporal coherence.
Implementation Example
To enable STG in a validation script, instantiate the configuration with the desired parameters:
from ltx_trainer.validation_runner import ValidationRunner
from ltx_trainer.config import ValidationConfig
cfg = ValidationConfig(
stg_scale=1.0,
stg_blocks=[29],
stg_mode="stg_av",
guidance_scale=7.0, # CFG scale
)
runner = ValidationRunner(cfg)
video_latent, audio_latent = runner.run() # STG applied automatically
The ValidationRunner internally constructs the BatchedPerturbationConfig, instantiates the STGGuider, and applies the guidance delta during each denoising step.
Summary
- Spatio-Temporal Guidance (STG) improves temporal consistency by perturbing self-attention in selected transformer blocks and steering toward the unperturbed prediction.
- The
STGGuiderclass inltx_core/components/guiders.pycomputes the guidance delta asscale × (positive − perturbed). stg_blockscontrols which layers are perturbed (default[29]), whilestg_modeselects between video-only ("stg_v") or audio-video ("stg_av") perturbation.- The delta is applied to denoised latents in
validation_runner.pybefore the Euler step, pulling the generation toward temporally coherent outputs. - Set
stg_scaleto1.0to enable STG or0.0to disable it entirely.
Frequently Asked Questions
What is the difference between STG and CFG in LTX-2?
Classifier-Free Guidance (CFG) computes the difference between conditional and unconditional predictions to enforce prompt adherence, while Spatio-Temporal Guidance (STG) computes the difference between standard predictions and predictions where self-attention is disabled in specific blocks to enforce temporal coherence. According to the LTX-2 source code in validation_runner.py, these mechanisms operate independently and can be applied simultaneously during inference.
Which transformer blocks should I specify for stg_blocks?
The default configuration in ltx_trainer/config.py uses [29], targeting a specific middle-to-late transformer block. While the repository recommends this default for general use, you can specify multiple blocks (e.g., [28, 29, 30]) to increase the perturbation strength. Note that perturbing more blocks or earlier blocks may significantly alter the temporal characteristics of the output.
Why does STG have separate modes for video and audio?
The stg_mode parameter accommodates different media requirements. When set to "stg_av", the perturbation config includes both SKIP_VIDEO_SELF_ATTN and SKIP_AUDIO_SELF_ATTN, ensuring temporal consistency across both modalities. When set to "stg_v", only video layers are perturbed, leaving audio generation unaffected. This flexibility allows users to optimize STG for video-only models or to prevent audio quality degradation when temporal coherence is less critical for sound.
How does the stg_scale parameter affect video quality?
The stg_scale acts as a multiplier for the guidance delta. A value of 0.0 disables STG entirely, while the recommended value of 1.0 provides balanced temporal smoothing without over-constraining the generation. Values significantly higher than 1.0 may over-penalize the perturbed predictions, potentially leading to over-smoothed or temporally rigid outputs. The STGGuider.enabled() method returns False when scale equals 0.0, efficiently bypassing the perturbed forward pass for computational savings.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →