# Custom CUDA Kernels in LTX-2: Complete Build Guide and Kernel Reference

> Explore LTX-2's custom CUDA kernels, including compiled and JIT-compiled options. Get a complete build guide and reference for efficient GPU acceleration. Learn how to build and integrate them.

- Repository: [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2)
- Tags: how-to-guide
- Published: 2026-08-20

---

**LTX-2 provides four compiled CUDA kernel extensions (`ops_cpp`, `nvfp4_cpp`, `blockwise_cpp`, `all2all_cpp`) plus two JIT-compiled CuTe DSL kernels, all built from the `packages/ltx-kernels` directory using `uv` workspace groups or manual pip installation.**

This guide walks through every custom CUDA kernel shipped with [Lightricks/LTX-2](https://github.com/Lightricks/LTX-2) and the exact commands needed to compile them. Whether you're optimizing inference on Ada, Hopper, or Blackwell architectures, you'll find the specific source paths and build flags required.

---

## Overview of Custom CUDA Kernels

The LTX-2 inference engine relies on hand-optimized CUDA kernels to accelerate quantization, attention, and multi-GPU communication. These kernels are organized into four Python extension modules, each targeting a distinct computational pattern.

### Extension Modules and Their Kernels

| Extension | Purpose | Source Kernels |
|-----------|---------|----------------|
| `ops_cpp` | Fused element-wise operations for blockwise quantization | RMS-Norm + RoPE, RMS-Norm + Split-RoPE, FP6 pack/unpack |
| `nvfp4_cpp` | NVFP4 quantization and cuBLASLt block-scaled GEMM | Quantize kernels |
| `blockwise_cpp` | Blockwise FP8 GEMM for GeForce/Ada (SM89) and Hopper/Blackwell (SM90) | Architecture-specific GEMM implementations |
| `all2all_cpp` | Multi-GPU All-2-All communication for sequence-parallel inference | All-Gather and head redistribution |

Two additional JIT-compiled kernels—`na_attn_dsl` and `block_fna_dsl`—power the diffusion VAE decoder using the CuTe DSL, but these are not pre-compiled C++ extensions.

---

## Kernel-by-Kernel Breakdown

### ops_cpp: Fused Normalization and Quantization

The `ops_cpp` extension bundles three critical fused kernels found in `packages/ltx-kernels/csrc/ops/`:

**RMS-Norm + RoPE** (`rms_norm_rope_cuda.cu`)

Combines root-mean-square normalization with rotary positional embedding in a single kernel launch. This fusion eliminates intermediate memory traffic for transformer blocks.

**RMS-Norm + Split-RoPE** (`rms_norm_split_rope_cuda.cu`)

A variant for split attention patterns, applying RoPE with different rotational frequencies to query and key tensors.

**FP6 Pack/Unpack** (`fp6_pack.cu`)

Converts between high-precision activations and 6-bit floating-point storage used in aggressive memory-constrained quantization schemes.

---

### nvfp4_cpp: NVIDIA FP4 Quantization

Located in `packages/ltx-kernels/csrc/nvfp4/`, this extension implements the emerging NVFP4 standard combining **FP4 E2M1** weights with **FP8 E4M3** activations.

**Quantize kernel** (`quantize.cu`)

Performs block-scaled quantization to NVFP4 format, preparing tensors for cuBLASLt block-scaled GEMM operations on supported hardware.

---

### blockwise_cpp: Architecture-Specific FP8 GEMM

The blockwise GEMM kernels live in `packages/ltx-kernels/csrc/blockwise/kernels/` with distinct implementations per GPU generation:

**GeForce/Ada (SM89)** (`geforce/gemm.cu`)

Optimized CUTLASS-based kernels for consumer and Ada Lovelace datacenter GPUs.

**Hopper/Blackwell Deep-GEMM (SM90)** (`deep_gemm/include/deep_gemm/impls/sm90_fp8_gemm_1d2d_bias.cu`)

Advanced FP8 GEMM with 1D/2D block scaling and bias fusion, targeting H100 and newer datacenter accelerators.

---

### all2all_cpp: Multi-GPU Communication

Sequence-parallel inference requires efficient tensor redistribution across GPUs. The kernels in `packages/ltx-kernels/csrc/all2all/cuda/` provide:

- `allgather.cu` — All-Gather collective for concatenating sequence shards
- `all2all_heads.cu` — Specialized head redistribution for attention tensor parallelism

---

## How to Build the CUDA Kernels

The kernel package `packages/ltx-kernels` is excluded from LTX-2's default `uv` workspace. You must explicitly opt-in during synchronization.

### Prerequisites

- CUDA toolkit matching your target GPU architecture (CUDA 12 recommended for SM 10/12)
- PyTorch with CUDA support pre-installed in your active environment

### Build Methods

#### Option 1: UV Workspace Group (Recommended)

```bash
uv sync --group kernels

```

This pulls dependencies and compiles all four extensions in one command.

#### Option 2: Direct Package Installation

```bash
uv pip install -e packages/ltx-kernels --no-build-isolation

```

Use `--no-build-isolation` to ensure PyTorch's CUDA headers are accessible during compilation.

---

## Targeting Specific GPU Architectures

Compilation time scales with the number of architectures targeted. Restrict `TORCH_CUDA_ARCH_LIST` to your deployment hardware:

**Single architecture (H100/SM90):**

```bash
TORCH_CUDA_ARCH_LIST="9.0" uv pip install -e packages/ltx-kernels --no-build-isolation

```

**Multiple architectures (Ada, Hopper, Blackwell):**

```bash
TORCH_CUDA_ARCH_LIST="9.0 9.0a 10.0 12.0" uv pip install -e packages/ltx-kernels --no-build-isolation

```

When unset, the build compiles for every architecture in the source, ensuring portability at the cost of longer build times.

---

## CUTLASS Integration

The blockwise GEMM kernels require NVIDIA CUTLASS headers. The build system handles this automatically:

| Configuration | Command/ Location |
|-------------|-------------------|
| Default cache | `~/.cache/ltx-kernels/` |
| Custom location | `export LTX_KERNELS_CACHE_DIR=/custom/path` |
| Local CUTLASS checkout | `export CUTLASS_DIR=/path/to/cutlass` |

CUTLASS is fetched on first build and cached for subsequent compilations.

---

## Verifying the Build

Run the kernel test suite to confirm correct compilation:

```bash
uv run pytest packages/ltx-kernels/tests/ -v

```

**Note:** NVFP4 and VAE kernel tests require a datacenter Blackwell GPU and automatically skip on incompatible hardware.

---

## Complete Source File Reference

| File Path | Kernel Function |
|-----------|-----------------|
| `packages/ltx-kernels/csrc/ops/rms_norm_rope_cuda.cu` | RMS-Norm + RoPE fusion |
| `packages/ltx-kernels/csrc/ops/rms_norm_split_rope_cuda.cu` | RMS-Norm + Split-RoPE fusion |
| `packages/ltx-kernels/csrc/ops/fp6_pack.cu` | 6-bit float packing |
| `packages/ltx-kernels/csrc/nvfp4/quantize.cu` | NVFP4 block quantization |
| `packages/ltx-kernels/csrc/blockwise/kernels/geforce/gemm.cu` | SM89 FP8 GEMM |
| `packages/ltx-kernels/csrc/blockwise/kernels/deep_gemm/include/deep_gemm/impls/sm90_fp8_gemm_1d2d_bias.cu` | SM90 deep GEMM with bias |
| `packages/ltx-kernels/csrc/all2all/cuda/allgather.cu` | All-Gather collective |
| `packages/ltx-kernels/csrc/all2all/cuda/all2all_heads.cu` | Head redistribution |

---

## Summary

- **Four compiled extensions** provide quantized inference, GEMM, and multi-GPU communication: `ops_cpp`, `nvfp4_cpp`, `blockwise_cpp`, `all2all_cpp`
- **Two JIT kernels** (`na_attn_dsl`, `block_fna_dsl`) serve the VAE decoder via CuTe
- **Source location**: `packages/ltx-kernels/csrc/` with architecture-specific subdirectories
- **Build command**: `uv sync --group kernels` or `uv pip install -e packages/ltx-kernels --no-build-isolation`
- **Architecture targeting**: Use `TORCH_CUDA_ARCH_LIST` to reduce compile times
- **CUTLASS handling**: Automatic caching with override via `CUTLASS_DIR` or `LTX_KERNELS_CACHE_DIR`

---

## Frequently Asked Questions

### Which GPUs are supported by LTX-2's custom CUDA kernels?

The kernels compile and run on NVIDIA GPUs from Ampere through Blackwell. Specific features require newer architectures: NVFP4 quantization needs Blackwell datacenter GPUs, Hopper/Blackwell deep-GEMM requires SM90+, and GeForce/Ada kernels target SM89. The build system automatically selects appropriate code paths based on detected hardware.

### Why does kernel installation require `--no-build-isolation`?

The PyTorch C++ extensions need access to the same CUDA headers and libraries used to compile PyTorch itself. `--no-build-isolation` allows the build process to inherit these from the active environment rather than creating a clean isolated environment that lacks CUDA tooling.

### How can I speed up compilation for development iteration?

Set `TORCH_CUDA_ARCH_LIST` to your single development GPU architecture rather than building for all supported targets. For H100 development, use `TORCH_CUDA_ARCH_LIST="9.0"`. Also ensure CUTLASS is cached locally to avoid repeated downloads.

### What is the difference between compiled extensions and JIT DSL kernels?

The **compiled extensions** (`ops_cpp`, `nvfp4_cpp`, `blockwise_cpp`, `all2all_cpp`) are built ahead-of-time into shared libraries with `torch.utils.cpp_extension`. The **JIT DSL kernels** (`na_attn_dsl`, `block_fna_dsl`) use CUDA's CuTe DSL and are compiled just-in-time during first execution, allowing more aggressive fusion specialization for the VAE decoder's specific tensor shapes.