# How Gemma 3 Produces Audio vs Video Embeddings in LTX-2's Text Encoder

> Discover how Gemma 3 generates distinct audio and video embeddings within LTX-2's text encoder. Learn about synchronized multimodal generation from single prompts.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: deep-dive
- Published: 2026-06-20

---

**LTX-2 uses Gemma 3's hidden states processed through an EmbeddingsProcessor to generate separate video embeddings (always produced) and optional audio embeddings (when configured with an audio connector), enabling synchronized multimodal generation from a single text prompt.**

The LTX-2 video generation framework leverages Google's Gemma 3 as its text encoder to bridge natural language prompts with multimodal outputs. Understanding how Gemma 3 audio vs video embeddings work in LTX-2's text encoder is crucial for developers building synchronized audio-video generation pipelines. The implementation splits hidden-state tensors into distinct streams through a specialized processor architecture that handles both modalities with shared attention mechanisms.

## The Two-Stage Embedding Pipeline

LTX-2's text encoding operates through a distinct separation between raw language model inference and modality-specific projection. The architecture first extracts transformer hidden states from Gemma 3, then processes these through separate connectors for video and audio generation paths.

### Stage 1: Extracting Hidden States with GemmaTextEncoder

The process begins in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py), where the **GemmaTextEncoder** class wraps the full `Gemma3ForConditionalGeneration` model. While the underlying model includes the language modeling head, the encoder skips final logits generation when only hidden representations are needed.

```python
from ltx_core.text_encoders.gemma import GemmaTextEncoder

encoder = GemmaTextEncoder(
    model=gemma_model,
    tokenizer=gemma_tokenizer,
    processor=gemma_processor
)

# Returns tuple of tensors (one per transformer layer)

hidden_states = encoder.encode(["A sunrise over the mountains"])

```

The `encode()` method returns a tuple containing hidden-state tensors for each transformer layer, providing the raw material for downstream modality-specific processing.

### Stage 2: Processing Through EmbeddingsProcessor

The second stage occurs in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_processor.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_processor.py), where the **EmbeddingsProcessor** class transforms raw hidden states into structured embeddings. This class is typically configured via **EmbeddingsProcessorConfigurator** in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/encoder_configurator.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/encoder_configurator.py).

```python
from ltx_core.text_encoders.gemma import EmbeddingsProcessor

processor = EmbeddingsProcessor(
    feature_extractor=feature_extractor,
    video_connector=video_connector,
    audio_connector=audio_connector,  # None for video-only pipelines

)

output = processor.process_hidden_states(
    hidden_states=hidden_states,
    attention_mask=attention_mask,
    padding_side="left"
)

```

The `process_hidden_states()` method returns an **EmbeddingsProcessorOutput** NamedTuple containing `video_encoding`, optional `audio_encoding`, and the binary `attention_mask`.

## How Video and Audio Embeddings Diverge

The critical distinction between modalities happens inside the feature extraction and projection pipeline. While both streams originate from the same Gemma 3 hidden states, they undergo separate transformations through dedicated connectors.

### Feature Extraction and Stream Splitting

The processor's internal **feature extractor** splits the hidden-state tuple into distinct video and audio feature sets. When an `audio_connector` is present (configured in [`packages/ltx-trainer/src/ltx_trainer/model_loader.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/model_loader.py)), the extractor produces both streams; otherwise, it returns `None` for audio features.

### The Role of Embeddings1DConnector

Both modalities utilize **Embeddings1DConnector** defined in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py) to project features into the target embedding space. These 1-D connectors handle the linear transformation from Gemma 3's hidden dimensions to the specific dimensions required by LTX-2's diffusion scheduler.

**Video embeddings** are always produced and returned via `video_encoding`. **Audio embeddings** are only returned when the `audio_connector` parameter is supplied; otherwise, `audio_encoding` is `None`.

### Attention Mask Alignment

The processor handles complex padding logic to ensure both modalities share compatible attention masks:

1. **Additive mask creation**: The binary attention mask from the tokenizer is converted to an additive format using `convert_to_additive_mask()` for Flash-Attention compatibility
2. **Right-padding computation**: The `_compute_right_pad_order()` method reorders features so valid tokens precede padding tokens
3. **Mask application**: After connector projection, the video mask is converted back to binary format and applied to the video embeddings

This shared attention mechanism ensures downstream modules can treat audio and video symmetrically during the diffusion process.

## Complete Implementation Example

The following workflow demonstrates the full pipeline from model loading to embedding extraction:

```python
from pathlib import Path
from ltx_core.text_encoders.gemma import GemmaTextEncoder, EmbeddingsProcessor
from ltx_core.loader.module_ops import load_module_ops_from_gemma_root
from ltx_core.text_encoders.gemma.embeddings_connector import Embeddings1DConnector

# 1. Load Gemma-3 assets (contains tokenizer.model & preprocessor_config.json)

gemma_root = Path("/path/to/gemma-3")
tokenizer_ops, processor_ops = load_module_ops_from_gemma_root(str(gemma_root))

gemma_encoder = GemmaTextEncoder()
gemma_encoder = tokenizer_ops.apply(gemma_encoder)   # Loads tokenizer

gemma_encoder = processor_ops.apply(gemma_encoder)   # Loads processor

# 2. Encode prompt to hidden states

prompt = ["A cat playing piano in a concert hall"]
hidden_states = gemma_encoder.encode(prompt)
attention_mask = gemma_encoder.tokenizer.get_attention_mask(prompt)

# 3. Configure connectors (simplified stubs)

feature_extractor = MyFeatureExtractor()
video_connector = Embeddings1DConnector(...)
audio_connector = Embeddings1DConnector(...)  # Set to None for video-only

emb_processor = EmbeddingsProcessor(
    feature_extractor=feature_extractor,
    video_connector=video_connector,
    audio_connector=audio_connector,
)

# 4. Generate final embeddings

emb_out = emb_processor.process_hidden_states(
    hidden_states=hidden_states,
    attention_mask=attention_mask,
    padding_side="left",
)

print(f"Video embedding shape: {emb_out.video_encoding.shape}")
if emb_out.audio_encoding is not None:
    print(f"Audio embedding shape: {emb_out.audio_encoding.shape}")

```

## Summary

- **GemmaTextEncoder** in [`base_encoder.py`](https://github.com/Lightricks/LTX-2/blob/main/base_encoder.py) extracts raw hidden states from Gemma 3 without generating logits
- **EmbeddingsProcessor** splits hidden states into video and optional audio streams through configurable feature extractors
- **Video embeddings** are always produced, while **audio embeddings** require an `audio_connector` configuration
- **Embeddings1DConnector** handles the final projection of features into modality-specific embedding spaces
- Shared attention mask processing ensures both modalities align with Flash-Attention requirements and padding conventions

## Frequently Asked Questions

### Are audio embeddings always generated in LTX-2's text encoder?

No. Audio embeddings are only produced when the `EmbeddingsProcessor` is configured with an `audio_connector` parameter. For video-only generation pipelines, setting `audio_connector=None` results in `emb_out.audio_encoding` returning `None`, while `video_encoding` always contains valid tensors.

### Which source files handle the embedding generation pipeline?

The key files are: [`packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py) (GemmaTextEncoder), [`packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_processor.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_processor.py) (EmbeddingsProcessor), and [`packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/embeddings_connector.py) (Embeddings1DConnector). Model loading utilities are located in [`packages/ltx-trainer/src/ltx_trainer/model_loader.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-trainer/src/ltx_trainer/model_loader.py).

### How does the attention mask work for both audio and video embeddings?

The processor uses `convert_to_additive_mask()` to transform binary masks into additive format suitable for Flash-Attention, then applies `_compute_right_pad_order()` to ensure valid tokens precede padding. Both modalities share a common attention mask structure, enabling symmetric handling by downstream diffusion schedulers.

### Can I use video-only generation without audio connectors?

Yes. Simply instantiate `EmbeddingsProcessor` with `audio_connector=None`. The processor will only generate `video_encoding` tensors, skipping audio feature extraction entirely. This reduces computational overhead when generating video without synchronized audio.