LTX-2 Prompt Writing and Enhancement: Best Practices for High-Fidelity Video Generation

Write clear, chronologically structured prompts under 200 words and enable automatic prompt enhancement with --enhance-prompt to maximize LTX-2 video quality.

LTX-2 is Lightricks' open-source video generation model that produces high-fidelity video conditioned on text prompts. The quality of your output depends heavily on prompt structure and whether you leverage the built-in prompt enhancement system. This guide covers the exact methodology used in the LTX-2 codebase to craft effective prompts and optimize them programmatically.

The 7-Step LTX-2 Prompt Structure

According to the repository README (lines 34-45), effective prompts follow a specific narrative sequence that guides the diffusion model through temporal progression.

1. Start with the Main Action

Open with a single sentence stating the core event. This gives the model a clear primary focus.


A cat jumps onto a windowsill

2. Add Movement and Gestures

Describe specific motions, timings, and dynamics to drive temporal progression.


quickly arches its back, then lands gracefully

3. Specify Appearances

Detail characters and objects including color, clothing, and facial expressions to improve visual fidelity.

4. Describe Background and Environment

Include setting, props, and atmosphere for scene context and consistency.


sunlit kitchen with a vintage clock on the wall

5. Define Camera Angles and Movements

Mention perspective, zoom, pan, or tilt to guide LTX-2's implicit cinematographer.


close-up from a low angle, then a smooth pan upwards

6. State Lighting and Colors

Specify time of day, light sources, and color palette to influence mood and shading.


soft golden hour lighting, warm tones

7. Note Changes or Events

Capture sudden or evolving elements for dynamic scene updates.


a sudden gust of wind blows the curtains

Formatting rule: Combine all seven elements into one flowing paragraph under 200 words total.

Automatic Prompt Enhancement in LTX-2

LTX-2 can rewrite prompts before encoding using a separate Gemma text-encoder model. This enhancement adds structure, improves grammar, and inserts hidden tokens the diffusion model recognizes.

When to Enable Enhancement

  • Complex or noisy prompts requiring grammatical clarity
  • Multi-prompt scenarios (positive + negative) where shared KV-cache speeds up encoding
  • Batch processing where consistent prompt structure improves output reliability

CLI Configuration Options

The enhancement system is controlled through arguments defined in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 13-30):

Flag Purpose
--enhance-prompt Activates the enhancement pipeline
--prompt-enhancer-gemma-root Path to separate enhancer checkpoint (required when encoder isn't Gemma-3)
--enhance-static-cache Reuses KV-cache across prompts to reduce latency

How Enhancement Works

The generate_enhanced_prompt function in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py (lines 20-38) handles the enhancement workflow:

  1. Loads image input if provided (for image-to-video)
  2. Calls text_encoder.enhance_i2v or enhance_t2v depending on mode
  3. Logs the enhanced result
  4. Cleans stray Unicode characters

Complete CLI Workflow Example

python -m ltx_pipelines.ti2vid_one_stage \
    --prompt "A cat jumps onto a windowsill, quickly arches its back, then lands gracefully. Sunlight streams through a kitchen window, casting warm golden tones. Low-angle close-up with a smooth upward pan." \
    --negative-prompt "low quality, blurry" \
    --enhance-prompt \
    --prompt-enhancer-gemma-root /path/to/gemma4_enhancer \
    --enhance-static-cache \
    --num-frames 120 \
    --output-path result.mp4

This example demonstrates: the 7-step prompt structure, --enhance-prompt activation, separate Gemma-4 enhancer supply (when base encoder differs), and static caching for performance.

Python API Implementation

For programmatic control, use generate_enhanced_prompt directly:

from ltx_pipelines import TI2VidOneStagePipeline
from ltx_pipelines.utils.helpers import generate_enhanced_prompt

# Initialize pipeline with enhancement enabled

pipeline = TI2VidOneStagePipeline(
    model_path="models/ltx_v2.5.ckpt",
    text_encoder_path="models/gemma3_encoder",
    enhance_prompt=True,
    prompt_enhancer_gemma_root="models/gemma4_enhancer",
    enhance_static_cache=True,
)

# Manual enhancement for custom control

raw_prompt = (
    "A cat jumps onto a windowsill, quickly arches its back, then lands gracefully. "
    "Sunlight streams through a kitchen window, casting warm golden tones. "
    "Low-angle close-up with a smooth upward pan."
)

enhanced = generate_enhanced_prompt(
    pipeline.prompt_encoder._enhancer_text_encoder_builder.build(),
    raw_prompt,
    static_cache=True,
)

# Generate video

video = pipeline(prompt=enhanced, num_frames=120)
video.save("result.mp4")

Architectural Implementation Details

Separate Encoder Builders

The pipeline maintains two distinct builders as implemented in packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py (lines 706-714):

  • _text_encoder_builder – for the encoding model
  • _enhancer_text_encoder_builder – for the enhancement model

When both point to the same checkpoint, the builders share memory to avoid duplication.

System Prompt Integration

LTX-2 loads preset system prompts (e.g., gemma_t2v_system_prompt.txt) that prepend fixed instructions to every user prompt. These ensure consistent "role-playing" guidance. Users generally should not modify these files. Source: packages/ltx-core/README.md (lines 399-402).

Attention Mask Caching

After enhancement, the text encoder returns embeddings plus a prompt_attention_mask. This mask conditions each diffusion timestep and is cached for validation reuse to save compute. Source: packages/ltx-trainer/src/ltx_trainer/validation_runner.py (lines 321-327).

Quick Checklist for Optimal LTX-2 Prompts

  • ✅ Follow the 7-step structure: action → movement → appearance → background → camera → lighting → changes
  • ✅ Keep prompts under 200 words in a single paragraph
  • ✅ Use descriptive, concrete language; avoid vague adjectives
  • ✅ Include audio cues for audio-enabled generations
  • ✅ Enable --enhance-prompt for long or grammar-heavy prompts
  • ✅ Supply --prompt-enhancer-gemma-root when encoder isn't Gemma-3
  • ✅ Turn on --enhance-static-cache for batch runs

Summary

  • Structure matters: The 7-step prompt format (action through temporal changes) provides clear temporal guidance to LTX-2's diffusion model
  • Length discipline: Single-paragraph prompts under 200 words optimize model attention
  • Enhancement helps: --enhance-prompt rewrites prompts for grammar, structure, and hidden token insertion
  • Encoder compatibility: Use --prompt-enhancer-gemma-root when your base encoder differs from Gemma-3
  • Performance tuning: --enhance-static-cache reduces latency in multi-prompt and batch scenarios

Frequently Asked Questions

What makes a good LTX-2 prompt?

A good LTX-2 prompt follows the 7-step structure starting with main action, adding movement details, specifying appearances, describing environment, defining camera work, stating lighting, and noting temporal changes. Keep it under 200 words as one flowing paragraph with concrete, descriptive language rather than vague adjectives.

When should I use prompt enhancement in LTX-2?

Enable prompt enhancement with --enhance-prompt when your prompts are complex, grammatically noisy, or when running batch generations where consistent structure matters. Enhancement is especially valuable for multi-prompt scenarios where the shared KV-cache from --enhance-static-cache improves performance.

Do I need a separate model for prompt enhancement?

You need a separate enhancer checkpoint via --prompt-enhancer-gemma-root only when your base text encoder is not a Gemma-3 model. The pipeline in blocks.py automatically shares builders when encoder and enhancer point to the same checkpoint to conserve memory.

How does LTX-2 handle system prompts?

LTX-2 automatically prepends preset system prompts (like gemma_t2v_system_prompt.txt) to every user prompt before enhancement. These provide consistent role-playing instructions to the model. You generally should not modify these internal files as they are tuned for optimal generation behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →