# How to Debug Shaders Using Backend-Specific Profiling Tools in Luisa Compute

> Debug shaders efficiently using backend-specific profiling tools like Nsight, PIX, and Xcode. Enable validation for automatic markers in GPU captures. Learn more with Luisa Compute.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: how-to-guide
- Published: 2026-03-06

---

**Enable validation mode with `LUISA_ENABLE_VALIDATION=1` to automatically inject backend-specific markers (NVTX, PIX, Metal debug groups) that appear in Nsight, PIX, and Xcode GPU captures.**

Luisa Compute is a high-performance GPU computing framework that abstracts CUDA, Vulkan, Metal, and DirectX 12 behind a unified kernel syntax. When you need to debug shaders using backend-specific profiling tools, the runtime provides a validation-driven profiling layer that translates high-level dispatch commands into native markers visible in Nsight Systems, Microsoft PIX, and Xcode GPU Debugger.

## Understanding Backend-Specific Profiling Integration

Luisa Compute implements **backend-agnostic debug markers** that map to each GPU vendor's native profiling API. When you compile a kernel with `device.compile(...)` and submit it to a stream, the runtime records commands into a backend-specific command buffer. If validation is enabled, the runtime automatically wraps these commands with profiling markers.

The mapping between backends and their native tools works as follows:

- **CUDA** → **NVTX ranges** (`nvtxRangePush`/`nvtxRangePop`) captured by **Nsight Systems** and **Nsight Compute**
- **DirectX 12** → **PIX markers** (`PIXBeginEvent`/`PIXEndEvent`) captured by **Microsoft PIX**
- **Metal** → **Debug groups** (`pushDebugGroup`/`popDebugGroup` on `MTLCommandBuffer`) captured by **Xcode GPU Debugger**
- **Vulkan** → **Debug markers** (`vkCmdDebugMarkerBeginEXT`/`vkCmdDebugMarkerEndEXT`) captured by **RenderDoc** and **Nsight Systems**

According to the architecture documentation in [`docs/source/architecture.md`](https://github.com/luisagroup/luisacompute/blob/main/docs/source/architecture.md), the backend-specific implementations of `DeviceInterface::record_debug_*` inject these calls when validation is active.

## Enabling Debug Markers in Your Build

To expose shader boundaries and custom regions in profiling captures, you must build the library with debug information and enable runtime validation.

First, configure the build with debug symbols so that shader SPIR-V, PTX, or Metal IR contains source-level information:

```bash
xmake f --debug
xmake

```

Next, enable the validation layer before running your application. This switches the runtime into "profiling-enabled" mode and activates marker injection:

```bash
export LUISA_ENABLE_VALIDATION=1   # Linux/macOS

set LUISA_ENABLE_VALIDATION=1      # Windows CMD

```

Alternatively, enable validation programmatically in C++ before creating a device:

```cpp
luisa::runtime::set_validation(true);

```

## Capturing Shader Execution in Native Profilers

Once validation is enabled, launch your application while the profiling tool is attached. Each backend integration follows a specific capture workflow.

**Nsight Systems/Compute (CUDA)**
Start a capture from the Nsight GUI or CLI, run your Luisa Compute application, and view the timeline. Each kernel appears as a labeled range in the **CUDA** → **Kernels** section, showing the kernel name and any user-defined debug groups.

**Microsoft PIX (DirectX 12)**
Launch your application through PIX, press **F12** (or click **Capture**) while the shader is running, then open the captured frame. Navigate to the **Event List** to see PIX markers delimiting kernel dispatches.

**Xcode GPU Debugger (Metal)**
Enable **GPU Frame Capture** (Debug → Capture GPU Frame), run your Metal-backed Luisa Compute app, and trigger a capture. The debug navigator displays Metal debug groups as collapsible hierarchical sections surrounding command buffer submissions.

**RenderDoc (Vulkan)**
Launch your application through RenderDoc, capture a frame, and inspect the **Event Browser**. The Vulkan backend emits debug markers via `VK_EXT_debug_marker`, making each kernel dispatch visible as a labeled event in the command buffer hierarchy.

## Adding Custom Debug Markers to Shader Code

Beyond automatic kernel wrappers, you can delimit arbitrary code regions using the `push_debug_group` and `pop_debug_group` methods on **Stream** objects. These calls are backend-agnostic and forward to the appropriate native API.

```cpp
#include <luisa/runtime/device.h>
#include <luisa/runtime/stream.h>

int main() {
    // Enable validation (if not using env var)
    luisa::runtime::set_validation(true);
    
    auto device = luisa::runtime::Device::create();
    auto stream = device.create_stream();
    
    // Begin custom profiling region
    stream.push_debug_group("Particle-Advection");
    
    // Compile and dispatch kernel
    auto advect = device.compile([](Float3* pos, Float dt) noexcept {
        // shader logic here
    });
    stream << advect(particle_buffer, dt).dispatch(particle_count);
    
    // End region
    stream.pop_debug_group();
    stream.synchronize();
}

```

The string passed to `push_debug_group` appears verbatim in the profiler timeline, allowing you to correlate GPU work with specific sections of your C++ source code.

## Verifying Marker Visibility in Capture Tools

After capturing a frame or timeline segment, verify that markers appear in the expected locations:

- **Nsight Compute**: Check **Kernel** → **User-Defined Ranges** for hierarchical nodes named after your debug groups
- **PIX**: Inspect the **Event List** → **Markers** column for labeled sections surrounding dispatch calls
- **Xcode**: Look under **GPU Frame Capture** → **Debug Groups** for collapsible sections containing command buffer work

If markers are missing, confirm that `LUISA_ENABLE_VALIDATION` is set in the environment where the application launches, and verify that you linked against the debug build of the Luisa Compute runtime.

## Key Implementation Files

The profiling integration spans several core files in the `luisagroup/luisacompute` repository:

| File | Purpose |
|------|---------|
| [`docs/source/architecture.md`](https://github.com/luisagroup/luisacompute/blob/main/docs/source/architecture.md) | Documents the validation-driven profiling architecture and backend-specific marker APIs |
| [`src/runtime/stream.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/stream.cpp) | Implements `push_debug_group` and `pop_debug_group` wrappers used by application code |
| [`src/runtime/device.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/device.cpp) | Handles `set_validation` flag propagation to the RHI layer |
| [`src/runtime/rhi/cuda_backend.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/rhi/cuda_backend.cpp) | Emits NVTX ranges for CUDA kernels when validation is enabled |
| [`src/runtime/rhi/dx12_backend.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/rhi/dx12_backend.cpp) | Wraps DirectX 12 command lists with PIX event macros |
| [`src/runtime/rhi/metal_backend.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/rhi/metal_backend.cpp) | Adds Metal debug group calls to `MTLCommandBuffer` |
| [`src/runtime/rhi/vulkan_backend.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/rhi/vulkan_backend.cpp) | Implements Vulkan debug marker extensions |

## Summary

- **Enable validation** via `LUISA_ENABLE_VALIDATION=1` or `set_validation(true)` to activate backend-specific profiling markers
- **Build with debug symbols** (`xmake f --debug`) to ensure shaders contain source-level debug information
- **Use `push_debug_group`/`pop_debug_group`** on streams to create labeled regions visible in Nsight, PIX, and Xcode
- **Verify captures** in the native tool's event list or timeline view to confirm markers surround kernel dispatches
- **Reference implementation files** like [`src/runtime/rhi/cuda_backend.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/rhi/cuda_backend.cpp) to understand how Luisa Compute maps abstract debug calls to NVTX, PIX, or Metal APIs

## Frequently Asked Questions

### How do I know if validation mode is actually enabled at runtime?

Check the return value of `luisa::runtime::set_validation(true)` or inspect the standard output when the application starts. The runtime logs the validation state during device creation, and the presence of debug markers in your profiler capture confirms the mode is active.

### Can I use these debug markers in release builds for production profiling?

While possible, release builds typically strip debug symbols and may optimize kernel code in ways that obscure source correlation. For accurate profiling with backend-specific tools, maintain a **RelWithDebInfo** or dedicated profiling build that preserves symbol information while enabling optimizations.

### Why don't I see debug markers in RenderDoc for Vulkan backends?

Ensure your Vulkan driver supports the `VK_EXT_debug_marker` extension. Luisa Compute checks for this extension in [`src/runtime/rhi/vulkan_backend.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/runtime/rhi/vulkan_backend.cpp) before emitting markers. If the extension is missing, upgrade your graphics driver or use Nsight Systems instead, which can capture Vulkan work without explicit debug markers.

### Do custom debug groups impact GPU performance?

Yes, but minimally. The `push_debug_group` and `pop_debug_group` calls translate to lightweight CPU-side annotations that add negligible overhead to command buffer recording. However, excessive nesting (hundreds of levels deep) can increase command buffer memory usage slightly. Disable validation (`LUISA_ENABLE_VALIDATION=0`) for final performance benchmarks.