How Gemma 3 Produces Audio vs Video Embeddings in LTX-2's Text Encoder
LTX-2 uses Gemma 3's hidden states processed through an EmbeddingsProcessor to generate separate video embeddings (always produced) and optional audio embeddings (when configured with an audio connector), enabling synchronized multimodal generation from a single text prompt.
The LTX-2 video generation framework leverages Google's Gemma 3 as its text encoder to bridge natural language prompts with multimodal outputs. Understanding how Gemma 3 audio vs video embeddings work in LTX-2's text encoder is crucial for developers building synchronized audio-video generation pipelines. The implementation splits hidden-state tensors into distinct streams through a specialized processor architecture that handles both modalities with shared attention mechanisms.
The Two-Stage Embedding Pipeline
LTX-2's text encoding operates through a distinct separation between raw language model inference and modality-specific projection. The architecture first extracts transformer hidden states from Gemma 3, then processes these through separate connectors for video and audio generation paths.
Stage 1: Extracting Hidden States with GemmaTextEncoder
The process begins in packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py, where the GemmaTextEncoder class wraps the full Gemma3ForConditionalGeneration model. While the underlying model includes the language modeling head, the encoder skips final logits generation when only hidden representations are needed.
from ltx_core.text_encoders.gemma import GemmaTextEncoder
encoder = GemmaTextEncoder(
model=gemma_model,
tokenizer=gemma_tokenizer,
processor=gemma_processor
)
# Returns tuple of tensors (one per transformer layer)
hidden_states = encoder.encode(["A sunrise over the mountains"])
The encode() method returns a tuple containing hidden-state tensors for each transformer layer, providing the raw material for downstream modality-specific processing.
Stage 2: Processing Through EmbeddingsProcessor
The second stage occurs in packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_processor.py, where the EmbeddingsProcessor class transforms raw hidden states into structured embeddings. This class is typically configured via EmbeddingsProcessorConfigurator in packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/encoder_configurator.py.
from ltx_core.text_encoders.gemma import EmbeddingsProcessor
processor = EmbeddingsProcessor(
feature_extractor=feature_extractor,
video_connector=video_connector,
audio_connector=audio_connector, # None for video-only pipelines
)
output = processor.process_hidden_states(
hidden_states=hidden_states,
attention_mask=attention_mask,
padding_side="left"
)
The process_hidden_states() method returns an EmbeddingsProcessorOutput NamedTuple containing video_encoding, optional audio_encoding, and the binary attention_mask.
How Video and Audio Embeddings Diverge
The critical distinction between modalities happens inside the feature extraction and projection pipeline. While both streams originate from the same Gemma 3 hidden states, they undergo separate transformations through dedicated connectors.
Feature Extraction and Stream Splitting
The processor's internal feature extractor splits the hidden-state tuple into distinct video and audio feature sets. When an audio_connector is present (configured in packages/ltx-trainer/src/ltx_trainer/model_loader.py), the extractor produces both streams; otherwise, it returns None for audio features.
The Role of Embeddings1DConnector
Both modalities utilize Embeddings1DConnector defined in packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py to project features into the target embedding space. These 1-D connectors handle the linear transformation from Gemma 3's hidden dimensions to the specific dimensions required by LTX-2's diffusion scheduler.
Video embeddings are always produced and returned via video_encoding. Audio embeddings are only returned when the audio_connector parameter is supplied; otherwise, audio_encoding is None.
Attention Mask Alignment
The processor handles complex padding logic to ensure both modalities share compatible attention masks:
- Additive mask creation: The binary attention mask from the tokenizer is converted to an additive format using
convert_to_additive_mask()for Flash-Attention compatibility - Right-padding computation: The
_compute_right_pad_order()method reorders features so valid tokens precede padding tokens - Mask application: After connector projection, the video mask is converted back to binary format and applied to the video embeddings
This shared attention mechanism ensures downstream modules can treat audio and video symmetrically during the diffusion process.
Complete Implementation Example
The following workflow demonstrates the full pipeline from model loading to embedding extraction:
from pathlib import Path
from ltx_core.text_encoders.gemma import GemmaTextEncoder, EmbeddingsProcessor
from ltx_core.loader.module_ops import load_module_ops_from_gemma_root
from ltx_core.text_encoders.gemma.embeddings_connector import Embeddings1DConnector
# 1. Load Gemma-3 assets (contains tokenizer.model & preprocessor_config.json)
gemma_root = Path("/path/to/gemma-3")
tokenizer_ops, processor_ops = load_module_ops_from_gemma_root(str(gemma_root))
gemma_encoder = GemmaTextEncoder()
gemma_encoder = tokenizer_ops.apply(gemma_encoder) # Loads tokenizer
gemma_encoder = processor_ops.apply(gemma_encoder) # Loads processor
# 2. Encode prompt to hidden states
prompt = ["A cat playing piano in a concert hall"]
hidden_states = gemma_encoder.encode(prompt)
attention_mask = gemma_encoder.tokenizer.get_attention_mask(prompt)
# 3. Configure connectors (simplified stubs)
feature_extractor = MyFeatureExtractor()
video_connector = Embeddings1DConnector(...)
audio_connector = Embeddings1DConnector(...) # Set to None for video-only
emb_processor = EmbeddingsProcessor(
feature_extractor=feature_extractor,
video_connector=video_connector,
audio_connector=audio_connector,
)
# 4. Generate final embeddings
emb_out = emb_processor.process_hidden_states(
hidden_states=hidden_states,
attention_mask=attention_mask,
padding_side="left",
)
print(f"Video embedding shape: {emb_out.video_encoding.shape}")
if emb_out.audio_encoding is not None:
print(f"Audio embedding shape: {emb_out.audio_encoding.shape}")
Summary
- GemmaTextEncoder in
base_encoder.pyextracts raw hidden states from Gemma 3 without generating logits - EmbeddingsProcessor splits hidden states into video and optional audio streams through configurable feature extractors
- Video embeddings are always produced, while audio embeddings require an
audio_connectorconfiguration - Embeddings1DConnector handles the final projection of features into modality-specific embedding spaces
- Shared attention mask processing ensures both modalities align with Flash-Attention requirements and padding conventions
Frequently Asked Questions
Are audio embeddings always generated in LTX-2's text encoder?
No. Audio embeddings are only produced when the EmbeddingsProcessor is configured with an audio_connector parameter. For video-only generation pipelines, setting audio_connector=None results in emb_out.audio_encoding returning None, while video_encoding always contains valid tensors.
Which source files handle the embedding generation pipeline?
The key files are: packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py (GemmaTextEncoder), packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_processor.py (EmbeddingsProcessor), and packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py (Embeddings1DConnector). Model loading utilities are located in packages/ltx-trainer/src/ltx_trainer/model_loader.py.
How does the attention mask work for both audio and video embeddings?
The processor uses convert_to_additive_mask() to transform binary masks into additive format suitable for Flash-Attention, then applies _compute_right_pad_order() to ensure valid tokens precede padding. Both modalities share a common attention mask structure, enabling symmetric handling by downstream diffusion schedulers.
Can I use video-only generation without audio connectors?
Yes. Simply instantiate EmbeddingsProcessor with audio_connector=None. The processor will only generate video_encoding tensors, skipping audio feature extraction entirely. This reduces computational overhead when generating video without synchronized audio.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →