# How Spatio-Temporal Guidance (STG) Works in LTX-2: A Technical Deep Dive

> Explore Spatio-Temporal Guidance (STG) in LTX-2. Learn how STG enhances video generation by steering the diffusion process for improved temporal coherence. Technical deep dive.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-21

---

**Spatio-Temporal Guidance (STG) improves temporal coherence in LTX-2 video generation by computing a scaled delta between standard denoised predictions and perturbed predictions where self-attention is disabled in selected transformer blocks, then steering the diffusion process toward the unperturbed result.**

Spatio-Temporal Guidance (STG) is a specialized guidance mechanism in the Lightricks/LTX-2 repository that reduces flickering and enhances frame-to-frame consistency in generated video and audio. Unlike Classifier-Free Guidance (CFG), which contrasts conditional and unconditional predictions, STG operates by internally perturbing the model's self-attention patterns to create a temporal consistency signal. The implementation follows a three-stage pipeline involving perturbation configuration, delta computation via the `STGGuider` class, and integration within the Euler denoising loop.

## Core STG Guiding Logic

The mathematical foundation of STG resides in [`ltx_core/components/guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/components/guiders.py), where the `STGGuider` dataclass implements the `GuiderProtocol`:

```python

# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/components/guiders.py#L55-L71

@dataclass(frozen=True)
class STGGuider(GuiderProtocol):
    """
    Calculates the STG delta between conditioned and perturbed denoised samples.
    """
    scale: float                     # 0.0 disables STG; >0 applies guidance

    def delta(self, pos_denoised: torch.Tensor, perturbed_denoised: torch.Tensor) -> torch.Tensor:
        return self.scale * (pos_denoised - perturbed_denoised)

    def enabled(self) -> bool:
        return self.scale != 0.0

```

The `delta()` method computes the **STG delta** as `Δ = scale × (positive − perturbed)`. Here, `pos_denoised` represents the standard forward pass, while `perturbed_denoised` comes from a pass where self-attention is skipped in specified transformer blocks. The `scale` parameter controls guidance strength—when set to `0.0`, `enabled()` returns `False` and STG is bypassed entirely.

## Configuring Perturbations

Before computing deltas, the system must define which layers to perturb. In [`ltx_trainer/validation_runner.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/validation_runner.py), the `_build_stg_perturbation_config()` function constructs a `BatchedPerturbationConfig` that specifies which transformer blocks have their self-attention disabled:

```python

# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/validation_runner.py#L74-L84

def _build_stg_perturbation_config(stg_blocks, stg_mode):
    perturbations = [
        Perturbation(type=PerturbationType.SKIP_VIDEO_SELF_ATTN, blocks=stg_blocks)
    ]
    if stg_mode == "stg_av":               # also affect audio

        perturbations.append(
            Perturbation(type=PerturbationType.SKIP_AUDIO_SELF_ATTN, blocks=stg_blocks)
        )
    return BatchedPerturbationConfig(perturbations=[PerturbationConfig(perturbations=perturbations)])

```

The **perturbation configuration** accepts two critical parameters:

- **`stg_blocks`**: A list of transformer block indices (e.g., `[29]`) where self-attention will be skipped.
- **`stg_mode`**: Either `"stg_v"` (perturb video only) or `"stg_av"` (perturb both video and audio self-attention).

This configuration is passed to the model during the perturbed forward pass to temporarily disable attention mechanisms in the specified blocks.

## Integration in the Denoising Loop

During each diffusion step in [`validation_runner.py`](https://github.com/Lightricks/LTX-2/blob/main/validation_runner.py), the system performs three distinct forward passes:

1. **Positive pass**: Standard prediction without perturbations (`perturbations=None`).
2. **Negative pass**: Unconditioned prediction for CFG (only when CFG is enabled).
3. **Perturbed pass**: Prediction with the STG perturbation config applied.

The STG delta is then applied to steer the denoised latent:

```python

# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/validation_runner.py#L11-L18

if stg_perturbation_config is not None:
    ptb_video, ptb_audio = x0_model(video=video, audio=audio,
                                    perturbations=stg_perturbation_config)
    if not video_frozen and denoised_video is not None:
        denoised_video = denoised_video + stg_guider.delta(pos_video, ptb_video)
    if not audio_frozen and denoised_audio is not None and ptb_audio is not None:
        denoised_audio = denoised_audio + stg_guider.delta(pos_audio, ptb_audio)

```

The delta **pulls the latent toward the positive direction** while penalizing the temporal inconsistencies introduced by the skipped self-attention. After applying the delta, the updated latents proceed to the Euler stepper (`stepper.step`) for the next diffusion iteration.

## Configuration Parameters

STG parameters are exposed in [`ltx_trainer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/config.py) as user-configurable arguments:

```python

# https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/config.py#L525-L538

stg_scale = Float(
    default=0.0,
    description="STG (Spatio‑Temporal Guidance) scale. 0.0 disables STG. Recommended value is 1.0."
)
stg_blocks = List(Int, default=[29],
    description="Which transformer blocks to perturb for STG."
)
stg_mode = Enum(
    choices=["stg_av", "stg_v"],
    default="stg_av",
    description="STG mode: 'stg_av' skips both audio and video self‑attention, 'stg_v' skips video only."
)

```

**Recommended settings** from the LTX-2 source code include:
- **`stg_scale`**: Set to `1.0` for active guidance, `0.0` to disable.
- **`stg_blocks`**: The default `[29]` targets specific transformer layers for perturbation.
- **`stg_mode`**: Use `"stg_av"` for audio-video synchronization or `"stg_v"` for video-only temporal coherence.

## Implementation Example

To enable STG in a validation script, instantiate the configuration with the desired parameters:

```python
from ltx_trainer.validation_runner import ValidationRunner
from ltx_trainer.config import ValidationConfig

cfg = ValidationConfig(
    stg_scale=1.0,
    stg_blocks=[29],
    stg_mode="stg_av",
    guidance_scale=7.0,       # CFG scale

)

runner = ValidationRunner(cfg)
video_latent, audio_latent = runner.run()   # STG applied automatically

```

The `ValidationRunner` internally constructs the `BatchedPerturbationConfig`, instantiates the `STGGuider`, and applies the guidance delta during each denoising step.

## Summary

- **Spatio-Temporal Guidance (STG)** improves temporal consistency by perturbing self-attention in selected transformer blocks and steering toward the unperturbed prediction.
- The **`STGGuider`** class in [`ltx_core/components/guiders.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_core/components/guiders.py) computes the guidance delta as `scale × (positive − perturbed)`.
- **`stg_blocks`** controls which layers are perturbed (default `[29]`), while **`stg_mode`** selects between video-only (`"stg_v"`) or audio-video (`"stg_av"`) perturbation.
- The delta is applied to denoised latents in [`validation_runner.py`](https://github.com/Lightricks/LTX-2/blob/main/validation_runner.py) before the Euler step, pulling the generation toward temporally coherent outputs.
- Set **`stg_scale`** to `1.0` to enable STG or `0.0` to disable it entirely.

## Frequently Asked Questions

### What is the difference between STG and CFG in LTX-2?

**Classifier-Free Guidance (CFG)** computes the difference between conditional and unconditional predictions to enforce prompt adherence, while **Spatio-Temporal Guidance (STG)** computes the difference between standard predictions and predictions where self-attention is disabled in specific blocks to enforce temporal coherence. According to the LTX-2 source code in [`validation_runner.py`](https://github.com/Lightricks/LTX-2/blob/main/validation_runner.py), these mechanisms operate independently and can be applied simultaneously during inference.

### Which transformer blocks should I specify for `stg_blocks`?

The default configuration in [`ltx_trainer/config.py`](https://github.com/Lightricks/LTX-2/blob/main/ltx_trainer/config.py) uses `[29]`, targeting a specific middle-to-late transformer block. While the repository recommends this default for general use, you can specify multiple blocks (e.g., `[28, 29, 30]`) to increase the perturbation strength. Note that perturbing more blocks or earlier blocks may significantly alter the temporal characteristics of the output.

### Why does STG have separate modes for video and audio?

The **`stg_mode`** parameter accommodates different media requirements. When set to `"stg_av"`, the perturbation config includes both `SKIP_VIDEO_SELF_ATTN` and `SKIP_AUDIO_SELF_ATTN`, ensuring temporal consistency across both modalities. When set to `"stg_v"`, only video layers are perturbed, leaving audio generation unaffected. This flexibility allows users to optimize STG for video-only models or to prevent audio quality degradation when temporal coherence is less critical for sound.

### How does the `stg_scale` parameter affect video quality?

The **`stg_scale`** acts as a multiplier for the guidance delta. A value of `0.0` disables STG entirely, while the recommended value of `1.0` provides balanced temporal smoothing without over-constraining the generation. Values significantly higher than `1.0` may over-penalize the perturbed predictions, potentially leading to over-smoothed or temporally rigid outputs. The `STGGuider.enabled()` method returns `False` when `scale` equals `0.0`, efficiently bypassing the perturbed forward pass for computational savings.