How to Use Prompt Enhancement for Better Results in LTX-2

LTX-2 improves video generation quality by using a Gemma-based model to rewrite and enrich prompts before encoding, optionally conditioned on reference images for more faithful visual results.

Prompt enhancement in LTX-2 is a preprocessing step that transforms sparse or ambiguous text descriptions into detailed, visually rich prompts that better guide the diffusion model. This feature leverages Google's Gemma instruction-tuned model to expand prompts with relevant descriptors, lighting details, and scene composition—resulting in more coherent and visually accurate video outputs.

Enabling Prompt Enhancement in the CLI

The fastest way to activate prompt enhancement is through the --enhance-prompt flag exposed by the two-stage pipeline. In packages/ltx-pipelines/src/ltx_pipelines/ti2vid_two_stages.py (lines 74-78), the argument parser captures this boolean and forwards it to the internal PromptEncoder.

Key CLI flags for prompt enhancement include:

  • --enhance-prompt – activates the enhancement workflow
  • --enhance-static-cache – ensures deterministic output by caching the enhanced prompt
  • --enhance-prompt-image <path> – conditions enhancement on a reference image
  • --prompt-enhancer-gemma-root <path> – specifies a separate Gemma checkpoint for enhancement (if different from the main text encoder)

These arguments are defined in packages/ltx-pipelines/src/ltx_pipelines/utils/args.py (lines 613-620) and passed through the pipeline constructor.

CLI Example with Image Conditioning

python -m ltx_pipelines.ti2vid_two_stages \
    --model-paths path/to/encode_gemma_root \
    --prompt-enhancer-gemma-root path/to/enhance_gemma_root \
    --prompt "a futuristic city skyline at night" \
    --negative-prompt "" \
    --seed 42 \
    --height 1024 --width 1024 \
    --frame-rate 30 \
    --num-frames 120 \
    --enhance-prompt \
    --enhance-static-cache \
    --enhance-prompt-image path/to/reference.jpg

How Prompt Enhancement Works Programmatically

When using the TI2VidTwoStagesPipeline class directly, prompt enhancement is controlled through boolean parameters in the pipeline call. The implementation in ti2vid_two_stages.py (lines 94-100) handles four enhancement flags:

  • enhance_first_prompt
  • enhance_static_cache
  • enhance_prompt_image
  • enhance_prompt_seed

Only the first prompt in a batch is enhanced when enhance_prompt=True—subsequent prompts remain unchanged to preserve variation in batched generation.

Python API Example

from ltx_pipelines.ti2vid_two_stages import TI2VidTwoStagesPipeline

pipeline = TI2VidTwoStagesPipeline(
    model_paths="path/to/encode_gemma_root",
    prompt_enhancer_gemma_root="path/to/enhance_gemma_root",
)

video, audio, num_frames, tiling_cfg = pipeline(
    prompt="a medieval castle on a hill",
    negative_prompt="",
    seed=123,
    height=720,
    width=1280,
    frame_rate=24,
    num_frames=60,
    video_guider_params=…,
    audio_guider_params=…,
    images=[],
    enhance_prompt=True,           # enable Gemma-based enhancement

    enhance_static_cache=True,     # deterministic output

    enhance_prompt_image=None,     # no image conditioning

)

The Enhancement Implementation: generate_enhanced_prompt

The core enhancement logic lives in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py (lines 20-37). The generate_enhanced_prompt function performs:

  1. Model loading – instantiates LTXGemmaTextEncoder for enhancement
  2. Image preprocessing – optionally decodes and resizes enhance_prompt_image for conditioning
  3. Mode selection – routes to enhance_i2v (image-to-video) or enhance_t2v (text-to-video)
  4. Output sanitization – strips stray Unicode characters from the Gemma output

# Partial enhancement for batched prompts (first prompt only)

prompts = ["a sunrise over mountains", "a rainy street"]
videos, audios, _, _ = pipeline(
    prompt=prompts,
    enhance_prompt=True,  # enhances only prompts[0]

)

The function also supports separate model paths: if prompt_enhancer_gemma_root differs from the main encoder checkpoint, a dedicated Builder is created in helpers.py (lines 750-717) to keep both models resident independently.

When to Use Prompt Enhancement in LTX-2

Use prompt enhancement when:

  • Prompts are underspecified – short descriptions like "a sunny beach" expand to richer scenes with lighting, texture, and composition details
  • Visual fidelity matters – image-conditioned enhancement (enhance_i2v) incorporates reference visual cues for more accurate recreation
  • Reproducibility is required – combine enhance_static_cache with a fixed --seed for identical enhancement across runs

Performance considerations:

  • Enhancement adds one full forward pass through the Gemma model per prompt
  • Enable only when quality gains justify the compute overhead
  • Cache embeddings during validation to avoid redundant enhancement—see _cache_prompt_embeddings in packages/ltx-trainer/src/ltx_trainer/validation_runner.py (lines 321-328)

Model Compatibility and Configuration

The enhancer and main text encoder must both be Gemma-based. If checkpoints differ, explicitly set --prompt-enhancer-gemma-root or prompt_enhancer_gemma_root in Python. When unspecified, the pipeline reuses the main encoder checkpoint automatically.

Summary

  • Prompt enhancement in LTX-2 uses a Gemma instruction model to rewrite prompts before encoding, improving video quality and relevance
  • Enable via CLI with --enhance-prompt and related flags, or programmatically through enhance_prompt=True in pipeline calls
  • Image conditioning through --enhance-prompt-image or enhance_prompt_image parameter enables visually grounded enhancement
  • Deterministic results require enhance_static_cache combined with a fixed seed
  • Only the first prompt is enhanced in batched generation to preserve output diversity
  • Core implementation is in generate_enhanced_prompt within packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py

Frequently Asked Questions

Does prompt enhancement work with any text encoder in LTX-2?

No. Prompt enhancement requires a Gemma-based model architecture compatible with LTXGemmaTextEncoder. Both the enhancer and main text encoder must be Gemma variants. If you use a different encoder family, the enhancement pipeline will not function.

Can I enhance prompts for image-to-video generation?

Yes. Supply --enhance-prompt-image (CLI) or enhance_prompt_image (Python) with a reference image path. The generate_enhanced_prompt function automatically routes to enhance_i2v instead of enhance_t2v, conditioning the rewritten prompt on visual features from your source image.

Why is only the first prompt enhanced when I pass multiple prompts?

This is intentional behavior in packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py (lines 772-776). Enhancing every prompt in a batch would homogenize outputs and increase compute cost. The design preserves diversity: the first prompt sets the scene with detailed enhancement, while remaining prompts provide variation.

How do I make prompt enhancement deterministic across runs?

Pass --enhance-static-cache (CLI) or enhance_static_cache=True (Python) together with a fixed --seed or seed parameter. This combination caches the enhanced prompt embedding and controls all random operations, ensuring identical output for identical inputs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →