How Modality CFG Improves Audio-Visual Synchronization in LTX-2
Modality CFG improves audio-visual synchronization by applying a cross-modal guidance scale that amplifies the difference between conditioned and unconditional modalities, penalizing the model when video and audio streams drift out of sync.
LTX-2 generates video and audio jointly through a diffusion transformer architecture that processes multiple modalities simultaneously. The Modality CFG technique extends traditional Classifier-Free Guidance to the cross-modal dimension, allowing users to enforce tighter alignment between visual content and corresponding audio tracks during the denoising process.
Understanding Modality CFG
LTX-2 represents media inputs as Modality objects that encapsulate latent tokens, timesteps, positional embeddings, and text-conditioning context. Each modality instance flows through the transformer independently while maintaining cross-attention paths between video and audio branches.
Traditional CFG operates by running the model twice—once with full conditioning (positive) and once with an unconditional prompt (negative)—then scaling the difference between these outputs. Modality CFG introduces a secondary guidance mechanism specifically for the cross-modal dimension, targeting the relationship between simultaneous video and audio generation.
Technical Implementation
The implementation centers on MultiModalGuiderParams and the CFGGuider class within the LTX-2 core components.
MultiModalGuiderParams Configuration
In packages/ltx-core/src/ltx_core/components/guiders.py at line 208, the MultiModalGuiderParams dataclass stores the modality-specific guidance factor:
@dataclass
class MultiModalGuiderParams:
cfg_scale: float = 1.0 # Classic CFG scale
modality_scale: float = 1.0 # Modality CFG scale (1.0 = disabled)
stg_scale: float = 0.0
When modality_scale exceeds 1.0, the guider computes an additional delta term that specifically targets audio-visual coherence.
The Delta Calculation
At line 265 in packages/ltx-core/src/ltx_core/components/guiders.py, the guider applies the following formula during each denoising step:
modality_delta = (modality_scale - 1) * (cond_modality - uncond_modality)
This delta is blended into the final latent update, creating a gradient that pushes the video and audio latents toward synchronization. The cond_modality represents the fully conditioned state where video and audio are prompted to align, while uncond_modality represents the deliberately unsynchronized baseline.
Achieving Audio-Visual Alignment
The synchronization mechanism relies on the contrast between conditioned and unconditional modality states.
Why Unconditional Modality Is Unsynchronized
During the negative pass, the model generates video and audio independently without cross-modal conditioning. This creates a baseline where lip movements, facial expressions, and rhythmic visual cues bear no relationship to the audio waveform. By amplifying the difference between this chaotic baseline and the conditioned state, Modality CFG forces the model to maintain coherent relationships between visual and auditory elements.
Synchronization Effects
When modality_scale is set above 1.0, the technique produces:
- Reduced lip-sync artifacts – Mouth shapes and phoneme generation align with spoken audio timing
- Rhythmic alignment – Visual beats, movements, and transitions synchronize with musical cues or sound effects
- Temporal coherence – Prevents drifting between video and audio streams over long generation sequences
Configuration and Usage
Activate Modality CFG through command-line interfaces, configuration files, or direct Python API calls.
CLI Configuration
Use the pipeline commands with modality-specific flags:
ltx-pipelines ti2vid_one_stage \
--prompt "A person dances to a drum beat" \
--video-guidance-scale 3.0 \
--audio-guidance-scale 3.0 \
--video-modality-scale 3.0 \
--audio-modality-scale 3.0
The --video-modality-scale and --audio-modality-scale arguments (defined in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py at lines 572 and 632) control the cross-modal guidance strength for each stream.
YAML Configuration
For training or batch generation, set the scale in YAML configs:
# packages/ltx-trainer/configs/video_suffix_lora.yaml
guidance:
cfg_scale: 4.0 # Classic CFG
modality_scale: 3.0 # Modality CFG for AV sync
stg_scale: 0.0
Python API
Construct modalities and apply the guider programmatically:
from ltx_core.model.transformer.modality import Modality
from ltx_core.components.guiders import (
CFGGuider, MultiModalGuiderParams, GuidedDenoiser
)
# Create conditioned and unconditional modalities
cond_mod = Modality(
latent=cond_latent,
sigma=cond_sigma,
timesteps=cond_ts,
positions=cond_pos,
context=cond_ctx,
enabled=True
)
uncond_mod = Modality(
latent=uncond_latent,
sigma=uncond_sigma,
timesteps=uncond_ts,
positions=uncond_pos,
context=uncond_ctx,
enabled=True
)
# Configure with Modality CFG
params = MultiModalGuiderParams(
cfg_scale=4.0,
modality_scale=3.0
)
guided = GuidedDenoiser(params=params)
next_latent = guided.denoise_step(cond_mod, uncond_mod)
Tuning Guidance Scales
Selecting appropriate values balances synchronization quality against creative flexibility.
Recommended Values
- 1.0 – Disables Modality CFG entirely; modalities generate independently
- 2.0–4.0 – Optimal range for strong audio-visual alignment without over-constraining the model
- Above 4.0 – May produce overly rigid synchronization that sacrifices natural motion or audio variation
Trade-offs
Higher modality_scale values improve synchronization metrics but can reduce the model's ability to generate creative interpretations of audio-visual relationships. For scenes requiring loose association (background music over unrelated visuals), maintain values near 1.0. For dialogue-heavy or music-video content, values between 3.0 and 4.0 typically yield the best results.
Summary
- Modality CFG extends Classifier-Free Guidance to cross-modal dimensions in
packages/ltx-core/src/ltx_core/components/guiders.py - The mechanism amplifies differences between synchronized (conditioned) and unsynchronized (unconditional) modality states
- Configure via
MultiModalGuiderParams.modality_scaleor CLI flags--video-modality-scaleand--audio-modality-scale - Optimal values range from 2.0 to 4.0 for tight audio-visual synchronization while preserving generation quality
- The default value of 1.0 disables the effect, allowing independent modality generation
Frequently Asked Questions
What is the difference between regular CFG and Modality CFG?
Regular CFG (controlled by cfg_scale) guides the generation toward the text prompt by comparing conditioned and unconditional outputs. Modality CFG (controlled by modality_scale) specifically compares how video and audio modalities relate to each other in both conditioned and unconditional states, directly targeting synchronization rather than prompt adherence.
Why does Modality CFG use unsynchronized unconditional modalities?
The unconditional pass generates video and audio independently without cross-modal attention. By amplifying the difference between this unsynchronized baseline and the conditioned state, the model learns to penalize drift between modalities. This contrastive signal forces the denoiser to maintain alignment or face strong gradient corrections.
Can I use different modality scales for video and audio?
Yes. LTX-2 supports independent scales via --video-modality-scale and --audio-modality-scale flags. This allows asymmetric guidance where video might receive stronger synchronization pressure than audio, or vice versa, depending on which stream requires stricter alignment with the other.
Does increasing Modality CFG affect generation speed?
No. Modality CFG requires the same two forward passes (conditioned and unconditional) already necessary for standard CFG. The additional computation of the modality_delta term is negligible compared to the transformer inference cost, meaning you can improve synchronization without increasing wall-clock time per step.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →