How to Use CFG and STG Guidance Scales for Better Generation Quality in LTX‑2
Use CFG (Classifier‑Free Guidance) scales to control prompt fidelity and STG (Spatio‑Temporal Guidance) scales to improve temporal consistency—tune them together with a guidance rescale factor to balance quality and artifacts in LTX‑2 video and audio generation.
LTX‑2, the open‑source video‑and‑audio generation model by Lightricks, exposes two independent guidance mechanisms that operate during diffusion sampling. Understanding how to configure CFG and STG guidance scales lets you trade off prompt adherence against temporal smoothness and artifact suppression. This guide walks through the configuration files, code paths, and practical tuning strategies based on the official LTX‑2 repository.
What Are CFG and STG Guidance Scales?
LTX‑2 separates guidance into two orthogonal controls defined in packages/ltx‑trainer/src/ltx_trainer/config.py:
CFG (Classifier‑Free Guidance) amplifies the difference between conditional and unconditional model predictions. Higher values force stricter prompt adherence but can introduce oversaturation or structural artifacts.
STG (Spatio‑Temporal Guidance) perturbs specific transformer blocks to improve temporal consistency in video and temporal‑frequency coherence in audio. Unlike CFG, STG operates locally on selected blocks rather than globally.
| Parameter | Purpose | Default (Validation) | Config Location |
|---|---|---|---|
video_cfg_scale |
Video prompt fidelity | 3.0 |
config.py lines 509‑511 |
audio_cfg_scale |
Audio prompt fidelity | 7.0 |
config.py lines 515‑517 |
video_stg_scale |
Video temporal smoothing | 1.0 |
config.py lines 521‑527 |
audio_stg_scale |
Audio coherence | 1.0 |
config.py lines 527‑531 |
stg_blocks |
Which blocks receive STG | [28] |
config.py lines 533‑535 |
guidance_rescale |
Variance interpolation post‑guidance | 0.7 |
config.py lines 38‑44 |
The guidance_rescale parameter deserves special attention: it interpolates the variance of the guided prediction with the original variance (1.0). Lower values soften the combined effect of CFG and STG to prevent over‑sharpening.
How Guidance Scales Are Applied During Sampling
During validation and inference, the trainer passes guidance parameters to the diffusion denoiser through the MultiModalGuiderParams structure. In packages/ltx‑trainer/src/ltx_trainer/validation_runner.py:
- Video path receives
cfg.video_cfg_scaleandcfg.video_stg_scale(lines 1232‑1234) - Audio path receives
cfg.audio_cfg_scaleandcfg.audio_stg_scale(lines 1254‑1256)
The underlying denoiser euler_cfg_pp_denoising_loop in packages/ltx‑pipelines/src/ltx_pipelines/utils/samplers.py validates that cfg_scale != 1 and stg_scale != 0 before applying each guidance type. This independence means you can enable CFG without STG, STG without CFG, or both together.
Tuning Strategy for CFG and STG Guidance Scales
The interaction between scales follows a predictable pattern:
- CFG provides global directional push toward conditioned latents
- STG adds local block‑level perturbation correcting temporal drift
- Guidance rescale smooths the mixture to prevent runaway variance
Follow this tuning workflow:
-
Start with defaults:
video_cfg_scale=3.0,audio_cfg_scale=7.0,video_stg_scale=1.0,audio_stg_scale=1.0 -
If motion is jittery despite good prompt match: increase
video_stg_scaleto1.5–2.0while keeping CFG modest -
If audio is noisy or misaligned: raise
audio_stg_scaleand consider loweringaudio_cfg_scaleto reduce over‑conditioning -
If colors oversaturate or ring artifacts appear: reduce
guidance_rescaleto0.5–0.6
Configuration Examples
YAML Configuration
For training runs or custom inference scripts:
validation:
video_cfg_scale: 3.0
audio_cfg_scale: 7.0
video_stg_scale: 1.5
audio_stg_scale: 1.0
stg_blocks: [28]
guidance_rescale: 0.6
Command‑Line Usage
Via the LTX‑Pipelines entry point with flags defined in packages/ltx‑pipelines/src/ltx_pipelines/utils/args.py:
python -m ltx_pipelines.ti2vid_two_stages \
--video-cfg-guidance-scale 4.0 \
--audio-cfg-guidance-scale 8.0 \
--video-stg-guidance-scale 1.8 \
--audio-stg-guidance-scale 1.0 \
--stg-blocks 28 \
--guidance-rescale 0.5 \
...
The argument definitions reside at lines 948‑962 for CFG defaults and lines 961‑966 for STG defaults in args.py.
Programmatic Construction
Building a MultiModalGuiderParams object directly:
from ltx_pipelines.utils.deniosers import MultiModalGuiderParams
guider = MultiModalGuiderParams(
cfg_scale=4.0,
stg_scale=1.8,
modality_scale=1.0,
)
# Pass guider to the sampler's denoise method
Core Implementation Files
| File | Role |
|---|---|
packages/ltx‑trainer/src/ltx_trainer/config.py |
Hyper‑parameter definitions |
packages/ltx‑trainer/src/ltx_trainer/validation_runner.py |
Scale application during inference |
packages/ltx‑pipelines/src/ltx_pipelines/utils/args.py |
CLI argument exposure |
packages/ltx‑core/src/ltx_core/components/guiders.py |
Mathematical combination logic |
packages/ltx‑pipelines/src/ltx_pipelines/utils/samplers.py |
Denoising loop with scale checks |
Summary
- CFG scales (
video_cfg_scale,audio_cfg_scale) control global prompt fidelity—higher values increase adherence but risk artifacts - STG scales (
video_stg_scale,audio_stg_scale) enable block‑level temporal correction via transformer perturbation - STG blocks (
stg_blocks) select which layers receive guidance—the default[28]works broadly - Guidance rescale (
guidance_rescale) softens combined guidance to prevent oversaturation - All parameters can be set via YAML, CLI flags in
args.py, orMultiModalGuiderParamsobjects - The denoiser in
samplers.pyapplies guidance independently, allowing flexible combinations
Frequently Asked Questions
What is the difference between CFG and STG in LTX‑2?
CFG operates globally on the entire model prediction, amplifying the gap between conditional and unconditional outputs. STG operates locally on specified transformer blocks, perturbing their activations to improve temporal structure. They address different failure modes: CFG improves prompt alignment, while STG reduces temporal jitter and frequency drift.
When should I increase STG scale instead of CFG scale?
Increase STG when your generation follows the prompt correctly but shows motion inconsistency, frame flicker, or audio desynchronization. STG specifically targets temporal coherence without the artifact risks of pushing CFG higher. The default stg_blocks = [28] targets late‑stage transformer layers where temporal structure crystallizes.
How does guidance rescale prevent artifacts?
After applying CFG and STG, the predicted noise variance can become artificially inflated, causing oversaturated colors and ringing. The guidance_rescale parameter interpolates between the guided variance and the original variance (1.0). Values below 0.7 progressively soften this effect, with 0.5–0.6 typically effective for high‑guidance configurations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →