LTX-2 Prompt Enhancement Feature for Improved Generations

The LTX-2 video generation model includes an optional prompt enhancement pipeline that automatically rewrites user prompts into more detailed, model-optimized descriptions using a Gemma-based text enhancer before diffusion processing.

LTX-2 ships with a built-in prompt enhancement system that bridges the gap between concise user input and the rich descriptions needed for high-quality video generation. This feature, implemented in the ltx-pipelines package, leverages an optional Gemma text encoder to rewrite prompts while maintaining full compatibility with reproducibility controls and performance optimizations.

How Prompt Enhancement Works in LTX-2

The enhancement system operates as a preprocessing layer in the inference pipeline. When enabled, it intercepts the first prompt in a batch, processes it through a dedicated generative model, and substitutes the enhanced version before standard text encoding begins.

CLI Flags for Enabling Enhancement

All enhancement behavior is controlled through arguments defined in [packages/ltx-pipelines/src/ltx_pipelines/utils/args.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/args.py):

  • --enhance-prompt – Master flag that activates the enhancement pipeline.
  • --enhance-static-cache – Enables static KV-cache for faster repeated enhancement.
  • --prompt-enhancer-gemma-root – Path to a separate Gemma checkpoint used exclusively for enhancement.
  • --enhance-prompt-seed – Fixes the random seed for deterministic prompt rewriting.
python -m ltx_pipelines.run \
    --pipeline ti2vid_two_stages \
    --model-dir /path/to/ltx-checkpoint \
    --prompt "A surfer rides a massive wave at sunrise" \
    --enhance-prompt \
    --prompt-enhancer-gemma-root /path/to/gemma4-instruct-checkpoint \
    --enhance-static-cache \
    --enhance-prompt-seed 12345

Builder Logic and Model Selection

The core orchestration happens in [packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/blocks.py) at lines 66-81. The pipeline builder evaluates whether to trigger enhancement based on two conditions:

  1. --enhance-prompt must be true.
  2. A separate enhancer model must be available (self._enhancer_text_encoder_builder != self._text_encoder_builder).

When both conditions are met, the enhancer loads via gpu_model() and transforms the prompt through generate_enhanced_prompt. If the primary encoder is not Gemma-3 and no separate enhancer is configured, the pipeline raises a ValueError directing users to supply --prompt-enhancer-gemma-root.

The Enhancement Routine: generate_enhanced_prompt

The actual prompt rewriting logic resides in [packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py](https://github.com/Lightricks/LTX-2/blob/main/packages/ltx-pipelines/src/ltx_pipelines/utils/helpers.py) at lines 20-38. The generate_enhanced_prompt function handles both image-guided and text-only scenarios:

  • Image-to-video (I2V): Decodes and resizes the reference image, then calls text_encoder.enhance_i2v with the raw prompt.
  • Text-to-video (T2V): Invokes text_encoder.enhance_t2v for pure text enhancement.
from ltx_pipelines.utils.helpers import generate_enhanced_prompt

enhanced = generate_enhanced_prompt(
    text_encoder=text_encoder,  # Loaded Gemma model (enhancer or primary)

    prompt="A surfer rides a massive wave at sunrise",
    seed=12345,
    static_cache=True,
)
print(enhanced)

# → "A cinematic shot of a lone surfer skillfully riding a towering ocean wave at dawn, sunlight glittering on the water, vivid colors, high-speed motion blur, dynamic camera tracking."

Post-processing through clean_response strips common Gemma artifacts—curly quotes, leading non-alphabetic characters, and formatting debris—ensuring clean prompt output.

Integration with the Main Pipeline

After enhancement completes, the modified prompt list flows into the standard encoding path through self._text_encoder_ctx(), producing hidden-state embeddings for the diffusion model. This design keeps enhancement orthogonal to the core generation process: the diffusion model receives identical inputs regardless of whether enhancement was applied.

Performance and Reproducibility Features

Static KV-cache (--enhance-static-cache): Reduces latency for batch processing by reusing cached key-value attention states after the first warm-up pass.

Deterministic seeding (--enhance-prompt-seed): Locks Gemma's generation process, enabling byte-identical prompt enhancement across repeated runs—critical for A/B testing and research reproducibility.

Flexible model selection: Users may use Gemma-3 for both encoding and enhancement, or deploy a separate generative checkpoint (e.g., Gemma-4 instruct) for richer creative rewriting without affecting the primary encoder's embeddings.

Summary

  • LTX-2 prompt enhancement rewrites concise prompts into detailed descriptions using an optional Gemma-based enhancer.
  • Enable via --enhance-prompt with optional flags for caching (--enhance-static-cache), seeding (--enhance-prompt-seed), and separate model paths (--prompt-enhancer-gemma-root).
  • Core files: args.py (CLI), blocks.py (orchestration), helpers.py (generation logic).
  • The pipeline validates configuration and raises explicit errors when enhancer models are missing.
  • Enhanced prompts feed into standard text encoding; the diffusion model remains unaware of the preprocessing step.

Frequently Asked Questions

What happens if I enable --enhance-prompt but don't provide a separate enhancer model?

If your primary text encoder is Gemma-3, the pipeline uses it for both enhancement and standard encoding. For non-Gemma-3 encoders, blocks.py raises a ValueError instructing you to set --prompt-enhancer-gemma-root to a valid Gemma checkpoint path.

Can I use the same seed for both prompt enhancement and video generation?

Yes. --enhance-prompt-seed controls only the enhancement step's randomness. Set it independently of any video generation seed. For fully deterministic pipelines, configure both seeds explicitly.

Does static KV-cache affect output quality?

No. --enhance-static-cache is a pure performance optimization that reuses attention state caches. The enhanced prompt output remains identical to non-cached runs with the same seed.

Is prompt enhancement available for all LTX-2 pipelines?

The enhancement logic lives in the shared pipeline utilities, but individual pipelines must implement support for the --enhance-prompt flag. The ti2vid_two_stages pipeline supports it; check specific pipeline documentation for others.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →