How the LuisaCompute Backend Code Generation Pipeline Translates AST to PTX, HLSL, and MSL
The LuisaCompute backend code generation pipeline converts high-level AST nodes into an intermediate XIR representation, then uses LLVM for PTX generation, custom HLSL emitters for DirectX, and SPIR-V cross-compilation for Metal.
The luisagroup/luisacompute repository implements a unified GPU computing framework that compiles user kernels written in the Luisa DSL into device-specific machine code. Central to this system is the backend code generation pipeline, which bridges the gap between the high-level abstract syntax tree (AST) and low-level GPU languages such as PTX, HLSL, and MSL.
Stage 1: From DSL to Intermediate Representation (AST → XIR)
Before any GPU code is emitted, the front-end transforms the user’s kernel into a backend-agnostic intermediate representation called XIR (eXtended Intermediate Representation).
- AST Construction – The DSL parser builds a concrete AST from the C++ lambda that defines the kernel.
- AST → XIR Conversion – The header
include/luisa/ir/ast2ir.hdeclares the visitor that walks the AST and constructs the XIR graph. Each high-level construct (loops, buffers, textures) maps to a corresponding XIR node. - Reverse Mapping – For debugging and introspection,
include/luisa/ir/ir2ast.hprovides the inverse transformation (IR back to AST).
Key source files
- AST → IR interface – [
include/luisa/ir/ast2ir.h](https://github.com/luisagroup/luisacompute/blob/stable/include/luisa/ir/ast2ir.h) - IR → AST interface – [
include/luisa/ir/ir2ast.h](https://github.com/luisagroup/luisacompute/blob/stable/include/luisa/ir/ir2ast.h)
Stage 2: Backend-Specific Code Generation Paths
Once the XIR graph is ready, the pipeline forks into three distinct paths. Each backend implements its own translation layer that converts XIR into the target shading language.
CUDA Backend: XIR to PTX via LLVM
The CUDA path leverages the LLVM NVPTX target to generate PTX assembly.
- XIR → LLVM IR –
src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cppimplements theCudaCodegenLLVMImplvisitor. It traverses the XIR graph and emits equivalent LLVM IR instructions, handling memory spaces (global,shared,local) and intrinsic mappings. - LLVM Target Machine – The file initialises the NVPTX target (
LLVMInitializeNVPTXTargetInfo,LLVMInitializeNVPTXTarget,LLVMInitializeNVPTXTargetMC) and configures the target machine for the detected compute capability. - LLVM → PTX – The
TargetMachine::emitToFilemethod (or in-memory buffer) produces the final PTX string. - PTX Post‑Processing –
src/backends/cuda/cuda_shader.cppperforms driver‑specific patches (e.g., adjusting version headers for older CUDA drivers) before the PTX is handed tocuModuleLoadData.
Key source files
- LLVM IR emitter – [
src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp) - PTX handling – [
src/backends/cuda/cuda_shader.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/cuda/cuda_shader.cpp)
DirectX Backend: XIR to HLSL
The DirectX path generates human‑readable HLSL source and then compiles it with the Microsoft DXC compiler.
- XIR → HLSL Source –
src/backends/common/hlsl/hlsl_codegen_util.cppcontains theHLSLCodegenUtilclass. It walks the XIR graph and prints HLSL constructs (e.g.,RWStructuredBuffer,Texture2D, compute shader entry points). The utility uses internal templates such ashlsl_headerandaccel_processto inject common boilerplate. - DXC Compilation –
src/backends/dx/Shader/ComputeShader.cppreceives the HLSL string, creates adxc::Compilerinstance, and invokesCompilewith thecs_6_0(or higher) target profile. The resulting DXIL blob is stored in aComputeShaderobject. - Runtime Upload – The DXIL blob is uploaded to the GPU via
ID3D12Device::CreateComputePipelineState.
Key source files
- HLSL emitter – [
src/backends/common/hlsl/hlsl_codegen_util.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/common/hlsl/hlsl_codegen_util.cpp) - DXC driver – [
src/backends/dx/Shader/ComputeShader.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/dx/Shader/ComputeShader.cpp)
Vulkan and Metal Backends: XIR to SPIR‑V to MSL
The Vulkan and Metal backends share a common SPIR‑V path. Metal requires an additional translation step because Apple’s runtime consumes MSL source, not SPIR‑V directly.
- XIR → SPIR‑V – The Vulkan backend reuses the LLVM infrastructure but targets the SPIR‑V backend (
LLVMInitializeSPIRVTarget). The translation logic resides insrc/backends/vk/device.cpp, which configures the LLVM target machine for SPIR‑V generation. - SPIR‑V → MSL –
src/backends/vk/shader_serializer.cppimplements theShaderSerializerclass. It takes the SPIR‑V binary and invokes SPIRV‑Cross (or an internal MSL serializer) to produce Metal Shading Language source. The serializer handles resource binding remapping (e.g., SPIR‑V descriptor sets → Metal argument buffers). - Metal Compilation – The resulting MSL string is passed to the Apple Metal compiler (
MTLDevice::newLibraryWithSource) at runtime, producing aMTLComputePipelineState.
Key source files
- SPIR‑V emitter – [
src/backends/vk/device.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/vk/device.cpp) - MSL serializer – [
src/backends/vk/shader_serializer.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/vk/shader_serializer.cpp)
Stage 3: Runtime Kernel Loading and Caching
After the backend produces the final binary (PTX, DXIL, or MSL), the runtime must load it onto the device.
- CUDA –
src/backends/cuda/cuda_device.cppimplementsadd_kernel, which callscuModuleLoadDataon the PTX string and caches the resultingCUmodulein the device’s kernel table. - DirectX –
src/backends/dx/DXRuntime/Device.cppcreates aComputeShaderobject, stores the DXIL blob, and constructs aD3D12_COMPUTE_PIPELINE_STATE_DESCfor execution. - Vulkan / Metal –
src/backends/vk/device.cppcaches thevk::ShaderModule(for Vulkan) or theMTLComputePipelineState(for Metal) in an internal hash map keyed by the kernel’s unique identifier.
Code Example: Tracing a Kernel Through the Pipeline
#include <luisa/luisa.h>
using namespace luisa::compute;
int main() {
// 1. Define a kernel in the DSL
auto kernel = Kernel([](Buffer<float> buffer) {
auto idx = dispatch_id().x;
buffer.write(idx, buffer.read(idx) + 1.0f);
});
// 2. Create a context (backend selected by environment or explicit flag)
Context ctx = Context::create(); // or Context::create("cuda"), "dx", "metal"
// 3. Allocate a buffer and dispatch
auto buffer = ctx.create_buffer<float>(1024);
kernel.dispatch(ctx, buffer);
// Under the hood:
// - AST → XIR (include/luisa/ir/ast2ir.h)
// - XIR → PTX (src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp)
// OR XIR → HLSL (src/backends/common/hlsl/hlsl_codegen_util.cpp)
// OR XIR → SPIR-V → MSL (src/backends/vk/shader_serializer.cpp)
}
This example demonstrates how a single DSL kernel is transparently lowered through the AST → XIR → backend-specific code generation pipeline, regardless of whether the target is NVIDIA (PTX), DirectX (HLSL), or Apple Silicon (MSL).
Summary
- Unified Front-End – All kernels start as Luisa DSL lambdas that are parsed into an AST and immediately converted to the backend-agnostic XIR via
include/luisa/ir/ast2ir.h. - LLVM-Based CUDA Path – The CUDA backend uses
src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cppto emit LLVM IR, targets the NVPTX backend, and produces PTX assembly that is loaded bysrc/backends/cuda/cuda_shader.cpp. - Direct HLSL Emission – The DirectX backend bypasses LLVM for HLSL, using
src/backends/common/hlsl/hlsl_codegen_util.cppto print HLSL source that is compiled offline by DXC insrc/backends/dx/Shader/ComputeShader.cpp. - SPIR-V Cross-Compilation for Metal – The Vulkan/Metal backend first generates SPIR-V via LLVM (similar to CUDA but with the SPIR-V target), then uses
src/backends/vk/shader_serializer.cppto translate SPIR-V into MSL for Apple’s Metal driver.
Frequently Asked Questions
What intermediate representation does LuisaCompute use before targeting GPU languages?
LuisaCompute uses an internal representation called XIR (eXtended Intermediate Representation). After the DSL parser builds an AST, the header include/luisa/ir/ast2ir.h defines the visitor that lowers the AST into an XIR graph. This graph is backend-agnostic and serves as the single source of truth for all subsequent optimization and code generation passes.
How does the CUDA backend generate PTX from the IR?
The CUDA backend translates XIR into LLVM IR using the implementation in src/backends/cuda/llvm_codegen/cuda_codegen_llvm_impl.cpp. It initializes the NVPTX target (LLVMInitializeNVPTXTarget), configures the target machine for the detected compute capability, and emits PTX assembly via LLVM’s TargetMachine::emitToFile. The resulting PTX string is then post-processed in src/backends/cuda/cuda_shader.cpp to handle driver version mismatches before being loaded into the CUDA driver.
Why does the Metal backend use SPIR-V as an intermediate step?
Apple’s Metal runtime consumes Metal Shading Language (MSL) source, not SPIR-V binaries. However, LuisaCompute’s Vulkan and Metal backends share a common SPIR-V generation path to avoid duplicating translation logic. The backend first generates SPIR-V using LLVM’s SPIR-V target (configured in src/backends/vk/device.cpp), then invokes the serializer in src/backends/vk/shader_serializer.cpp—which uses SPIRV-Cross—to convert the SPIR-V module into MSL source that the Metal driver can compile at runtime.
Where can I find the entry points for kernel compilation in each backend?
Each backend exposes a device-specific add_kernel or equivalent method that triggers the code generation pipeline:
- CUDA –
src/backends/cuda/cuda_device.cppimplementsadd_kernel, which calls the LLVM codegen and caches the resultingCUmodule. - DirectX –
src/backends/dx/DXRuntime/Device.cppcreates aComputeShaderobject and stores the DXIL blob produced by the HLSL emitter. - Vulkan / Metal –
src/backends/vk/device.cppcaches thevk::ShaderModule(Vulkan) or theMTLComputePipelineState(Metal) after the SPIR-V → MSL conversion.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →