# How the LuisaCompute Backend Code Generation Pipeline Translates AST to PTX, HLSL, and MSL

> Explore the LuisaCompute backend pipeline translating AST to PTX, HLSL, and MSL. Discover XIR, LLVM PTX generation, custom HLSL emitters, and SPIR-V for Metal.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: internals
- Published: 2026-03-06

---

**The LuisaCompute backend code generation pipeline converts high-level AST nodes into an intermediate XIR representation, then uses LLVM for PTX generation, custom HLSL emitters for DirectX, and SPIR-V cross-compilation for Metal.**

The `luisagroup/luisacompute` repository implements a unified GPU computing framework that compiles user kernels written in the Luisa DSL into device-specific machine code. Central to this system is the backend code generation pipeline, which bridges the gap between the high-level abstract syntax tree (AST) and low-level GPU languages such as PTX, HLSL, and MSL.

## Stage 1: From DSL to Intermediate Representation (AST → XIR)

Before any GPU code is emitted, the front-end transforms the user’s kernel into a backend-agnostic intermediate representation called **XIR** (eXtended Intermediate Representation).

1. **AST Construction** – The DSL parser builds a concrete AST from the C++ lambda that defines the kernel.
2. **AST → XIR Conversion** – The header [`include/luisa/ir/ast2ir.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/ir/ast2ir.h) declares the visitor that walks the AST and constructs the XIR graph. Each high-level construct (loops, buffers, textures) maps to a corresponding XIR node.
3. **Reverse Mapping** – For debugging and introspection, [`include/luisa/ir/ir2ast.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/ir/ir2ast.h) provides the inverse transformation (IR back to AST).

*Key source files*  
- **AST → IR interface** – [[`include/luisa/ir/ast2ir.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/ir/ast2ir.h)](https://github.com/luisagroup/luisacompute/blob/stable/include/luisa/ir/ast2ir.h)  
- **IR → AST interface** – [[`include/luisa/ir/ir2ast.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/ir/ir2ast.h)](https://github.com/luisagroup/luisacompute/blob/stable/include/luisa/ir/ir2ast.h)  

## Stage 2: Backend-Specific Code Generation Paths

Once the XIR graph is ready, the pipeline forks into three distinct paths. Each backend implements its own translation layer that converts XIR into the target shading language.

### CUDA Backend: XIR to PTX via LLVM

The CUDA path leverages the LLVM NVPTX target to generate PTX assembly.

1. **XIR → LLVM IR** – [`src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp) implements the `CudaCodegenLLVMImpl` visitor. It traverses the XIR graph and emits equivalent LLVM IR instructions, handling memory spaces (`global`, `shared`, `local`) and intrinsic mappings.
2. **LLVM Target Machine** – The file initialises the NVPTX target (`LLVMInitializeNVPTXTargetInfo`, `LLVMInitializeNVPTXTarget`, `LLVMInitializeNVPTXTargetMC`) and configures the target machine for the detected compute capability.
3. **LLVM → PTX** – The `TargetMachine::emitToFile` method (or in-memory buffer) produces the final PTX string.
4. **PTX Post‑Processing** – [`src/backends/cuda/cuda_shader.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_shader.cpp) performs driver‑specific patches (e.g., adjusting version headers for older CUDA drivers) before the PTX is handed to `cuModuleLoadData`.

*Key source files*  
- LLVM IR emitter – [[`src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp)](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp)  
- PTX handling – [[`src/backends/cuda/cuda_shader.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_shader.cpp)](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/cuda/cuda_shader.cpp)  

### DirectX Backend: XIR to HLSL

The DirectX path generates human‑readable HLSL source and then compiles it with the Microsoft DXC compiler.

1. **XIR → HLSL Source** – [`src/backends/common/hlsl/hlsl_codegen_util.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/common/hlsl/hlsl_codegen_util.cpp) contains the `HLSLCodegenUtil` class. It walks the XIR graph and prints HLSL constructs (e.g., `RWStructuredBuffer`, `Texture2D`, compute shader entry points). The utility uses internal templates such as `hlsl_header` and `accel_process` to inject common boilerplate.
2. **DXC Compilation** – [`src/backends/dx/Shader/ComputeShader.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/dx/Shader/ComputeShader.cpp) receives the HLSL string, creates a `dxc::Compiler` instance, and invokes `Compile` with the `cs_6_0` (or higher) target profile. The resulting DXIL blob is stored in a `ComputeShader` object.
3. **Runtime Upload** – The DXIL blob is uploaded to the GPU via `ID3D12Device::CreateComputePipelineState`.

*Key source files*  
- HLSL emitter – [[`src/backends/common/hlsl/hlsl_codegen_util.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/common/hlsl/hlsl_codegen_util.cpp)](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/common/hlsl/hlsl_codegen_util.cpp)  
- DXC driver – [[`src/backends/dx/Shader/ComputeShader.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/dx/Shader/ComputeShader.cpp)](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/dx/Shader/ComputeShader.cpp)  

### Vulkan and Metal Backends: XIR to SPIR‑V to MSL

The Vulkan and Metal backends share a common SPIR‑V path. Metal requires an additional translation step because Apple’s runtime consumes MSL source, not SPIR‑V directly.

1. **XIR → SPIR‑V** – The Vulkan backend reuses the LLVM infrastructure but targets the SPIR‑V backend (`LLVMInitializeSPIRVTarget`). The translation logic resides in [`src/backends/vk/device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/device.cpp), which configures the LLVM target machine for SPIR‑V generation.
2. **SPIR‑V → MSL** – [`src/backends/vk/shader_serializer.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/shader_serializer.cpp) implements the `ShaderSerializer` class. It takes the SPIR‑V binary and invokes **SPIRV‑Cross** (or an internal MSL serializer) to produce Metal Shading Language source. The serializer handles resource binding remapping (e.g., SPIR‑V descriptor sets → Metal argument buffers).
3. **Metal Compilation** – The resulting MSL string is passed to the Apple Metal compiler (`MTLDevice::newLibraryWithSource`) at runtime, producing a `MTLComputePipelineState`.

*Key source files*  
- SPIR‑V emitter – [[`src/backends/vk/device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/device.cpp)](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/vk/device.cpp)  
- MSL serializer – [[`src/backends/vk/shader_serializer.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/shader_serializer.cpp)](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/vk/shader_serializer.cpp)  

## Stage 3: Runtime Kernel Loading and Caching

After the backend produces the final binary (PTX, DXIL, or MSL), the runtime must load it onto the device.

- **CUDA** – [`src/backends/cuda/cuda_device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_device.cpp) implements `add_kernel`, which calls `cuModuleLoadData` on the PTX string and caches the resulting `CUmodule` in the device’s kernel table.
- **DirectX** – [`src/backends/dx/DXRuntime/Device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/dx/DXRuntime/Device.cpp) creates a `ComputeShader` object, stores the DXIL blob, and constructs a `D3D12_COMPUTE_PIPELINE_STATE_DESC` for execution.
- **Vulkan / Metal** – [`src/backends/vk/device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/device.cpp) caches the `vk::ShaderModule` (for Vulkan) or the `MTLComputePipelineState` (for Metal) in an internal hash map keyed by the kernel’s unique identifier.

## Code Example: Tracing a Kernel Through the Pipeline

```cpp
#include <luisa/luisa.h>

using namespace luisa::compute;

int main() {
    // 1. Define a kernel in the DSL
    auto kernel = Kernel([](Buffer<float> buffer) {
        auto idx = dispatch_id().x;
        buffer.write(idx, buffer.read(idx) + 1.0f);
    });

    // 2. Create a context (backend selected by environment or explicit flag)
    Context ctx = Context::create(); // or Context::create("cuda"), "dx", "metal"

    // 3. Allocate a buffer and dispatch
    auto buffer = ctx.create_buffer<float>(1024);
    kernel.dispatch(ctx, buffer);

    // Under the hood:
    // - AST → XIR (include/luisa/ir/ast2ir.h)
    // - XIR → PTX (src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp)
    //   OR XIR → HLSL (src/backends/common/hlsl/hlsl_codegen_util.cpp)
    //   OR XIR → SPIR-V → MSL (src/backends/vk/shader_serializer.cpp)
}

```

This example demonstrates how a single DSL kernel is transparently lowered through the AST → XIR → backend-specific code generation pipeline, regardless of whether the target is NVIDIA (PTX), DirectX (HLSL), or Apple Silicon (MSL).

## Summary

- **Unified Front-End** – All kernels start as Luisa DSL lambdas that are parsed into an AST and immediately converted to the backend-agnostic **XIR** via [`include/luisa/ir/ast2ir.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/ir/ast2ir.h).
- **LLVM-Based CUDA Path** – The CUDA backend uses [`src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp) to emit LLVM IR, targets the NVPTX backend, and produces PTX assembly that is loaded by [`src/backends/cuda/cuda_shader.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_shader.cpp).
- **Direct HLSL Emission** – The DirectX backend bypasses LLVM for HLSL, using [`src/backends/common/hlsl/hlsl_codegen_util.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/common/hlsl/hlsl_codegen_util.cpp) to print HLSL source that is compiled offline by DXC in [`src/backends/dx/Shader/ComputeShader.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/dx/Shader/ComputeShader.cpp).
- **SPIR-V Cross-Compilation for Metal** – The Vulkan/Metal backend first generates SPIR-V via LLVM (similar to CUDA but with the SPIR-V target), then uses [`src/backends/vk/shader_serializer.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/shader_serializer.cpp) to translate SPIR-V into MSL for Apple’s Metal driver.

## Frequently Asked Questions

### What intermediate representation does LuisaCompute use before targeting GPU languages?

LuisaCompute uses an internal representation called **XIR** (eXtended Intermediate Representation). After the DSL parser builds an AST, the header [`include/luisa/ir/ast2ir.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/ir/ast2ir.h) defines the visitor that lowers the AST into an XIR graph. This graph is backend-agnostic and serves as the single source of truth for all subsequent optimization and code generation passes.

### How does the CUDA backend generate PTX from the IR?

The CUDA backend translates XIR into LLVM IR using the implementation in [`src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp). It initializes the NVPTX target (`LLVMInitializeNVPTXTarget`), configures the target machine for the detected compute capability, and emits PTX assembly via LLVM’s `TargetMachine::emitToFile`. The resulting PTX string is then post-processed in [`src/backends/cuda/cuda_shader.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_shader.cpp) to handle driver version mismatches before being loaded into the CUDA driver.

### Why does the Metal backend use SPIR-V as an intermediate step?

Apple’s Metal runtime consumes **Metal Shading Language (MSL)** source, not SPIR-V binaries. However, LuisaCompute’s Vulkan and Metal backends share a common SPIR-V generation path to avoid duplicating translation logic. The backend first generates SPIR-V using LLVM’s SPIR-V target (configured in [`src/backends/vk/device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/device.cpp)), then invokes the serializer in [`src/backends/vk/shader_serializer.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/shader_serializer.cpp)—which uses SPIRV-Cross—to convert the SPIR-V module into MSL source that the Metal driver can compile at runtime.

### Where can I find the entry points for kernel compilation in each backend?

Each backend exposes a device-specific `add_kernel` or equivalent method that triggers the code generation pipeline:

- **CUDA** – [`src/backends/cuda/cuda_device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/cuda/cuda_device.cpp) implements `add_kernel`, which calls the LLVM codegen and caches the resulting `CUmodule`.
- **DirectX** – [`src/backends/dx/DXRuntime/Device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/dx/DXRuntime/Device.cpp) creates a `ComputeShader` object and stores the DXIL blob produced by the HLSL emitter.
- **Vulkan / Metal** – [`src/backends/vk/device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/device.cpp) caches the `vk::ShaderModule` (Vulkan) or the `MTLComputePipelineState` (Metal) after the SPIR-V → MSL conversion.