Best Practices for Prompt Engineering in LTX-2: A Complete Technical Guide

LTX-2 leverages LLM-driven system prompts via the Gemma text encoder to transform raw user descriptions into structured, single-paragraph video-generation prompts that strictly enforce chronological flow, audio layer details, and formatting constraints.

The Lightricks/LTX-2 repository implements a sophisticated prompt enhancement pipeline that relies on immutable system prompts to standardize video generation inputs. Understanding how these system prompts function within the GemmaTextEncoder class is essential for optimizing output quality and ensuring your prompts align with the model's architectural expectations.

Core System Prompt Architecture

LTX-2 utilizes two distinct system prompt files that define the behavior of the text-to-video and image-to-video generation workflows. These files are located in the ltx-core package and are loaded as immutable resources by the encoder.

Text-to-Video System Prompt

The Text-to-Video system prompt resides in packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/prompts/gemma_t2v_system_prompt.txt. This file establishes the style guidelines, chronological flow requirements, audio integration rules, and strict formatting constraints for generating video prompts. It specifically instructs the LLM to produce output as a single continuous paragraph without timestamps, markdown headings, or special characters.

Image-to-Video System Prompt

The Image-to-Video system prompt is stored in packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/prompts/gemma_i2v_system_prompt.txt. This prompt governs how input images are described and how additional motion, sound, or camera instructions are incorporated into the final generation prompt. It ensures that static visual inputs are transformed into dynamic, temporally-aware descriptions suitable for video diffusion.

Loading and Caching Mechanism

The GemmaTextEncoder class loads these system prompts through the _load_system_prompt method, which is decorated with functools.cached_property to ensure efficient caching. The encoder exposes these as default_gemma_t2v_system_prompt and default_gemma_i2v_system_prompt properties. According to the source code in packages/ltx-core/src/ltx_core/text_encoders/gemma/encoders/base_encoder.py, this caching mechanism prevents redundant disk reads while maintaining prompt immutability throughout the generation process.

Prompt Enhancement Workflow

The enhancement process automatically reformats user inputs to comply with LTX-2's strict output requirements. This workflow is implemented in the enhance_t2v and enhance_i2v methods of the Gemma encoder.

The Enhancement Pipeline

When you call enhance_t2v or enhance_i2v, the encoder executes the _enhance method, which constructs a chat message structure consisting of a system message (containing the loaded system prompt) and a user message (containing your raw prompt). This chat is then processed by self.model.generate to produce the enhanced output. As implemented in base_encoder.py, this approach ensures that the LLM consistently applies the formatting rules defined in the system prompt files.

Strict Output Formatting

The system prompts enforce rigid output constraints that downstream pipeline components expect. The LLM is explicitly instructed to return a single paragraph without markdown, titles, or special characters, and specifically to avoid narrative constructs like "The scene opens..." or timestamp markers. The ltx_pipelines module treats this returned string as the final prompt for audio-visual conditioning, making adherence to these constraints critical for successful generation.

Writing Effective Raw Prompts

The repository's README provides a concise checklist for crafting raw prompts that work optimally with the enhancement system. Following these guidelines ensures the LLM can process your intent without losing critical details.

Start with a chronological, single-sentence description of the main action. Add concrete details covering motions, gestures, appearances, environment, camera angles, lighting, and colors. Describe the complete audio layer, including background sounds, speech (using exact quoted words), music, or foley effects. Keep the entire description as a single paragraph without headings or timestamps, and constrain the length to approximately 200 words to ensure compatibility with the LLM's context window.

Implementation Examples

You can access the prompt enhancement functionality through both the Python API and the command-line interface.

Python API Usage

The following example demonstrates loading the Gemma encoder and enhancing a raw prompt programmatically:

from ltx_core.text_encoders.gemma.encoders.base_encoder import GemmaTextEncoder
from ltx_core.loader.module_ops import load_modules
import torch

# 1️⃣ Load a Gemma checkpoint (replace <path> with your model directory)

gemma_root = "<path>/gemma-3b"
tokenizer_ops, processor_ops = load_modules.module_ops_from_gemma_root(gemma_root)
model = GemmaTextEncoder()
model = tokenizer_ops.apply(model)
model = processor_ops.apply(model)

# 2️⃣ Raw user prompt (follow the checklist from the README)

raw_prompt = (
    "A woman at a coffee shop talking on the phone"
)

# 3️⃣ Enhance the prompt (uses the built‑in system prompt)

enhanced = model.enhance_t2v(
    prompt=raw_prompt,
    max_new_tokens=512,
    seed=42,
    # optional: override system prompt if you need a custom style

    # system_prompt="Style: cinematic, ..."

)

print("Enhanced prompt:")
print(enhanced)

CLI Integration

When using the ltx-pipelines entry point, enable automatic enhancement with the --enhance-prompt flag:

ltx-pipelines ti2vid_one_stage \
  --prompt "A woman at a coffee shop talking on the phone" \
  --enhance-prompt \
  --output out.mp4

Passing --enhance-prompt invokes the same enhance_t2v routine, ensuring the final prompt adheres to all LTX-2 best-practice rules before being fed to the diffusion model.

Summary

  • Immutable system prompts in gemma_t2v_system_prompt.txt and gemma_i2v_system_prompt.txt define strict formatting and content rules for video generation.
  • The GemmaTextEncoder caches these prompts and applies them via enhance_t2v and enhance_i2v methods to reformat user inputs.
  • Best practices include chronological single-sentence starts, concrete visual and audio details, single-paragraph structure, and a 200-word limit.
  • Automatic enhancement can be triggered via enhance_prompt=True in Python or --enhance-prompt in the CLI, ensuring consistent prompt quality.

Frequently Asked Questions

The documentation recommends keeping prompts under approximately 200 words. This constraint ensures the LLM can process the entire description within its context window while maintaining the coherence required for high-fidelity video generation.

How does LTX-2 handle audio descriptions in prompts?

The system prompt explicitly instructs the LLM to describe the complete soundscape, including background sounds, speech with exact quoted words, music, and foley effects. This audio layer is integrated into the enhanced prompt as part of the single-paragraph output, allowing the diffusion pipeline to condition both visual and audio generation.

Can I customize the system prompt for specific styles?

Yes, the enhance_t2v and enhance_i2v methods accept an optional system_prompt parameter that overrides the default system prompts loaded from the text files. This allows you to inject custom style guidelines while maintaining the required output format constraints.

What happens if my prompt violates the formatting rules?

If the LLM cannot parse your input according to the system prompt constraints, the enhance_t2v or enhance_i2v methods may return the original user text unmodified, or the system prompt may strip invalid elements like timestamps or markdown. For optimal results, adhere to the chronological, single-paragraph structure defined in the README guidelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →