Custom CUDA Kernels in LTX-2: Complete Build Guide and Kernel Reference
LTX-2 provides four compiled CUDA kernel extensions (ops_cpp, nvfp4_cpp, blockwise_cpp, all2all_cpp) plus two JIT-compiled CuTe DSL kernels, all built from the packages/ltx-kernels directory using uv workspace groups or manual pip installation.
This guide walks through every custom CUDA kernel shipped with Lightricks/LTX-2 and the exact commands needed to compile them. Whether you're optimizing inference on Ada, Hopper, or Blackwell architectures, you'll find the specific source paths and build flags required.
Overview of Custom CUDA Kernels
The LTX-2 inference engine relies on hand-optimized CUDA kernels to accelerate quantization, attention, and multi-GPU communication. These kernels are organized into four Python extension modules, each targeting a distinct computational pattern.
Extension Modules and Their Kernels
| Extension | Purpose | Source Kernels |
|---|---|---|
ops_cpp |
Fused element-wise operations for blockwise quantization | RMS-Norm + RoPE, RMS-Norm + Split-RoPE, FP6 pack/unpack |
nvfp4_cpp |
NVFP4 quantization and cuBLASLt block-scaled GEMM | Quantize kernels |
blockwise_cpp |
Blockwise FP8 GEMM for GeForce/Ada (SM89) and Hopper/Blackwell (SM90) | Architecture-specific GEMM implementations |
all2all_cpp |
Multi-GPU All-2-All communication for sequence-parallel inference | All-Gather and head redistribution |
Two additional JIT-compiled kernels—na_attn_dsl and block_fna_dsl—power the diffusion VAE decoder using the CuTe DSL, but these are not pre-compiled C++ extensions.
Kernel-by-Kernel Breakdown
ops_cpp: Fused Normalization and Quantization
The ops_cpp extension bundles three critical fused kernels found in packages/ltx-kernels/csrc/ops/:
RMS-Norm + RoPE (rms_norm_rope_cuda.cu)
Combines root-mean-square normalization with rotary positional embedding in a single kernel launch. This fusion eliminates intermediate memory traffic for transformer blocks.
RMS-Norm + Split-RoPE (rms_norm_split_rope_cuda.cu)
A variant for split attention patterns, applying RoPE with different rotational frequencies to query and key tensors.
FP6 Pack/Unpack (fp6_pack.cu)
Converts between high-precision activations and 6-bit floating-point storage used in aggressive memory-constrained quantization schemes.
nvfp4_cpp: NVIDIA FP4 Quantization
Located in packages/ltx-kernels/csrc/nvfp4/, this extension implements the emerging NVFP4 standard combining FP4 E2M1 weights with FP8 E4M3 activations.
Quantize kernel (quantize.cu)
Performs block-scaled quantization to NVFP4 format, preparing tensors for cuBLASLt block-scaled GEMM operations on supported hardware.
blockwise_cpp: Architecture-Specific FP8 GEMM
The blockwise GEMM kernels live in packages/ltx-kernels/csrc/blockwise/kernels/ with distinct implementations per GPU generation:
GeForce/Ada (SM89) (geforce/gemm.cu)
Optimized CUTLASS-based kernels for consumer and Ada Lovelace datacenter GPUs.
Hopper/Blackwell Deep-GEMM (SM90) (deep_gemm/include/deep_gemm/impls/sm90_fp8_gemm_1d2d_bias.cu)
Advanced FP8 GEMM with 1D/2D block scaling and bias fusion, targeting H100 and newer datacenter accelerators.
all2all_cpp: Multi-GPU Communication
Sequence-parallel inference requires efficient tensor redistribution across GPUs. The kernels in packages/ltx-kernels/csrc/all2all/cuda/ provide:
allgather.cu— All-Gather collective for concatenating sequence shardsall2all_heads.cu— Specialized head redistribution for attention tensor parallelism
How to Build the CUDA Kernels
The kernel package packages/ltx-kernels is excluded from LTX-2's default uv workspace. You must explicitly opt-in during synchronization.
Prerequisites
- CUDA toolkit matching your target GPU architecture (CUDA 12 recommended for SM 10/12)
- PyTorch with CUDA support pre-installed in your active environment
Build Methods
Option 1: UV Workspace Group (Recommended)
uv sync --group kernels
This pulls dependencies and compiles all four extensions in one command.
Option 2: Direct Package Installation
uv pip install -e packages/ltx-kernels --no-build-isolation
Use --no-build-isolation to ensure PyTorch's CUDA headers are accessible during compilation.
Targeting Specific GPU Architectures
Compilation time scales with the number of architectures targeted. Restrict TORCH_CUDA_ARCH_LIST to your deployment hardware:
Single architecture (H100/SM90):
TORCH_CUDA_ARCH_LIST="9.0" uv pip install -e packages/ltx-kernels --no-build-isolation
Multiple architectures (Ada, Hopper, Blackwell):
TORCH_CUDA_ARCH_LIST="9.0 9.0a 10.0 12.0" uv pip install -e packages/ltx-kernels --no-build-isolation
When unset, the build compiles for every architecture in the source, ensuring portability at the cost of longer build times.
CUTLASS Integration
The blockwise GEMM kernels require NVIDIA CUTLASS headers. The build system handles this automatically:
| Configuration | Command/ Location |
|---|---|
| Default cache | ~/.cache/ltx-kernels/ |
| Custom location | export LTX_KERNELS_CACHE_DIR=/custom/path |
| Local CUTLASS checkout | export CUTLASS_DIR=/path/to/cutlass |
CUTLASS is fetched on first build and cached for subsequent compilations.
Verifying the Build
Run the kernel test suite to confirm correct compilation:
uv run pytest packages/ltx-kernels/tests/ -v
Note: NVFP4 and VAE kernel tests require a datacenter Blackwell GPU and automatically skip on incompatible hardware.
Complete Source File Reference
| File Path | Kernel Function |
|---|---|
packages/ltx-kernels/csrc/ops/rms_norm_rope_cuda.cu |
RMS-Norm + RoPE fusion |
packages/ltx-kernels/csrc/ops/rms_norm_split_rope_cuda.cu |
RMS-Norm + Split-RoPE fusion |
packages/ltx-kernels/csrc/ops/fp6_pack.cu |
6-bit float packing |
packages/ltx-kernels/csrc/nvfp4/quantize.cu |
NVFP4 block quantization |
packages/ltx-kernels/csrc/blockwise/kernels/geforce/gemm.cu |
SM89 FP8 GEMM |
packages/ltx-kernels/csrc/blockwise/kernels/deep_gemm/include/deep_gemm/impls/sm90_fp8_gemm_1d2d_bias.cu |
SM90 deep GEMM with bias |
packages/ltx-kernels/csrc/all2all/cuda/allgather.cu |
All-Gather collective |
packages/ltx-kernels/csrc/all2all/cuda/all2all_heads.cu |
Head redistribution |
Summary
- Four compiled extensions provide quantized inference, GEMM, and multi-GPU communication:
ops_cpp,nvfp4_cpp,blockwise_cpp,all2all_cpp - Two JIT kernels (
na_attn_dsl,block_fna_dsl) serve the VAE decoder via CuTe - Source location:
packages/ltx-kernels/csrc/with architecture-specific subdirectories - Build command:
uv sync --group kernelsoruv pip install -e packages/ltx-kernels --no-build-isolation - Architecture targeting: Use
TORCH_CUDA_ARCH_LISTto reduce compile times - CUTLASS handling: Automatic caching with override via
CUTLASS_DIRorLTX_KERNELS_CACHE_DIR
Frequently Asked Questions
Which GPUs are supported by LTX-2's custom CUDA kernels?
The kernels compile and run on NVIDIA GPUs from Ampere through Blackwell. Specific features require newer architectures: NVFP4 quantization needs Blackwell datacenter GPUs, Hopper/Blackwell deep-GEMM requires SM90+, and GeForce/Ada kernels target SM89. The build system automatically selects appropriate code paths based on detected hardware.
Why does kernel installation require --no-build-isolation?
The PyTorch C++ extensions need access to the same CUDA headers and libraries used to compile PyTorch itself. --no-build-isolation allows the build process to inherit these from the active environment rather than creating a clean isolated environment that lacks CUDA tooling.
How can I speed up compilation for development iteration?
Set TORCH_CUDA_ARCH_LIST to your single development GPU architecture rather than building for all supported targets. For H100 development, use TORCH_CUDA_ARCH_LIST="9.0". Also ensure CUTLASS is cached locally to avoid repeated downloads.
What is the difference between compiled extensions and JIT DSL kernels?
The compiled extensions (ops_cpp, nvfp4_cpp, blockwise_cpp, all2all_cpp) are built ahead-of-time into shared libraries with torch.utils.cpp_extension. The JIT DSL kernels (na_attn_dsl, block_fna_dsl) use CUDA's CuTe DSL and are compiled just-in-time during first execution, allowing more aggressive fusion specialization for the VAE decoder's specific tensor shapes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →