# Best Practices for Prompt Engineering in LTX-2: A Complete Technical Guide

> Master prompt engineering for LTX-2 with our technical guide. Learn to create structured video prompts for Gemma text encoder, ensuring chronological flow and formatting constraints.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: best-practices
- Published: 2026-06-20

---

**LTX-2 leverages LLM-driven system prompts via the Gemma text encoder to transform raw user descriptions into structured, single-paragraph video-generation prompts that strictly enforce chronological flow, audio layer details, and formatting constraints.**

The Lightricks/LTX-2 repository implements a sophisticated prompt enhancement pipeline that relies on immutable system prompts to standardize video generation inputs. Understanding how these system prompts function within the `GemmaTextEncoder` class is essential for optimizing output quality and ensuring your prompts align with the model's architectural expectations.

## Core System Prompt Architecture

LTX-2 utilizes two distinct system prompt files that define the behavior of the text-to-video and image-to-video generation workflows. These files are located in the `ltx-core` package and are loaded as immutable resources by the encoder.

### Text-to-Video System Prompt

The **Text-to-Video system prompt** resides in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/prompts/gemma_t2v_system_prompt.txt`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/prompts/gemma_t2v_system_prompt.txt). This file establishes the style guidelines, chronological flow requirements, audio integration rules, and strict formatting constraints for generating video prompts. It specifically instructs the LLM to produce output as a single continuous paragraph without timestamps, markdown headings, or special characters.

### Image-to-Video System Prompt

The **Image-to-Video system prompt** is stored in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/prompts/gemma_i2v_system_prompt.txt`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/prompts/gemma_i2v_system_prompt.txt). This prompt governs how input images are described and how additional motion, sound, or camera instructions are incorporated into the final generation prompt. It ensures that static visual inputs are transformed into dynamic, temporally-aware descriptions suitable for video diffusion.

### Loading and Caching Mechanism

The `GemmaTextEncoder` class loads these system prompts through the `_load_system_prompt` method, which is decorated with `functools.cached_property` to ensure efficient caching. The encoder exposes these as `default_gemma_t2v_system_prompt` and `default_gemma_i2v_system_prompt` properties. According to the source code in [`packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py`](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py), this caching mechanism prevents redundant disk reads while maintaining prompt immutability throughout the generation process.

## Prompt Enhancement Workflow

The enhancement process automatically reformats user inputs to comply with LTX-2's strict output requirements. This workflow is implemented in the `enhance_t2v` and `enhance_i2v` methods of the Gemma encoder.

### The Enhancement Pipeline

When you call `enhance_t2v` or `enhance_i2v`, the encoder executes the `_enhance` method, which constructs a chat message structure consisting of a system message (containing the loaded system prompt) and a user message (containing your raw prompt). This chat is then processed by `self.model.generate` to produce the enhanced output. As implemented in [`base_encoder.py`](https://github.com/Lightricks/LTX-2/blob/main/base_encoder.py), this approach ensures that the LLM consistently applies the formatting rules defined in the system prompt files.

### Strict Output Formatting

The system prompts enforce rigid output constraints that downstream pipeline components expect. The LLM is explicitly instructed to return a single paragraph without markdown, titles, or special characters, and specifically to avoid narrative constructs like "The scene opens..." or timestamp markers. The `ltx_pipelines` module treats this returned string as the final prompt for audio-visual conditioning, making adherence to these constraints critical for successful generation.

## Writing Effective Raw Prompts

The repository's README provides a concise checklist for crafting raw prompts that work optimally with the enhancement system. Following these guidelines ensures the LLM can process your intent without losing critical details.

Start with a **chronological, single-sentence** description of the main action. Add concrete details covering motions, gestures, appearances, environment, camera angles, lighting, and colors. Describe the complete audio layer, including background sounds, speech (using exact quoted words), music, or foley effects. Keep the entire description as a single paragraph without headings or timestamps, and constrain the length to approximately 200 words to ensure compatibility with the LLM's context window.

## Implementation Examples

You can access the prompt enhancement functionality through both the Python API and the command-line interface.

### Python API Usage

The following example demonstrates loading the Gemma encoder and enhancing a raw prompt programmatically:

```python
from ltx_core.text_encoders.gemma.encoders.base_encoder import GemmaTextEncoder
from ltx_core.loader.module_ops import load_modules
import torch

# 1️⃣ Load a Gemma checkpoint (replace <path> with your model directory)

gemma_root = "<path>/gemma-3b"
tokenizer_ops, processor_ops = load_modules.module_ops_from_gemma_root(gemma_root)
model = GemmaTextEncoder()
model = tokenizer_ops.apply(model)
model = processor_ops.apply(model)

# 2️⃣ Raw user prompt (follow the checklist from the README)

raw_prompt = (
    "A woman at a coffee shop talking on the phone"
)

# 3️⃣ Enhance the prompt (uses the built‑in system prompt)

enhanced = model.enhance_t2v(
    prompt=raw_prompt,
    max_new_tokens=512,
    seed=42,
    # optional: override system prompt if you need a custom style

    # system_prompt="Style: cinematic, ..."

)

print("Enhanced prompt:")
print(enhanced)

```

### CLI Integration

When using the `ltx-pipelines` entry point, enable automatic enhancement with the `--enhance-prompt` flag:

```bash
ltx-pipelines ti2vid_one_stage \
  --prompt "A woman at a coffee shop talking on the phone" \
  --enhance-prompt \
  --output out.mp4

```

Passing `--enhance-prompt` invokes the same `enhance_t2v` routine, ensuring the final prompt adheres to all LTX-2 best-practice rules before being fed to the diffusion model.

## Summary

- **Immutable system prompts** in [`gemma_t2v_system_prompt.txt`](https://github.com/Lightricks/LTX-2/blob/main/gemma_t2v_system_prompt.txt) and [`gemma_i2v_system_prompt.txt`](https://github.com/Lightricks/LTX-2/blob/main/gemma_i2v_system_prompt.txt) define strict formatting and content rules for video generation.
- The `GemmaTextEncoder` caches these prompts and applies them via `enhance_t2v` and `enhance_i2v` methods to reformat user inputs.
- **Best practices** include chronological single-sentence starts, concrete visual and audio details, single-paragraph structure, and a 200-word limit.
- **Automatic enhancement** can be triggered via `enhance_prompt=True` in Python or `--enhance-prompt` in the CLI, ensuring consistent prompt quality.

## Frequently Asked Questions

### What is the maximum recommended length for LTX-2 prompts?

The documentation recommends keeping prompts under approximately 200 words. This constraint ensures the LLM can process the entire description within its context window while maintaining the coherence required for high-fidelity video generation.

### How does LTX-2 handle audio descriptions in prompts?

The system prompt explicitly instructs the LLM to describe the complete soundscape, including background sounds, speech with exact quoted words, music, and foley effects. This audio layer is integrated into the enhanced prompt as part of the single-paragraph output, allowing the diffusion pipeline to condition both visual and audio generation.

### Can I customize the system prompt for specific styles?

Yes, the `enhance_t2v` and `enhance_i2v` methods accept an optional `system_prompt` parameter that overrides the default system prompts loaded from the text files. This allows you to inject custom style guidelines while maintaining the required output format constraints.

### What happens if my prompt violates the formatting rules?

If the LLM cannot parse your input according to the system prompt constraints, the `enhance_t2v` or `enhance_i2v` methods may return the original user text unmodified, or the system prompt may strip invalid elements like timestamps or markdown. For optimal results, adhere to the chronological, single-paragraph structure defined in the README guidelines.