How to Debug Shaders Using Backend-Specific Profiling Tools in Luisa Compute
Enable validation mode with LUISA_ENABLE_VALIDATION=1 to automatically inject backend-specific markers (NVTX, PIX, Metal debug groups) that appear in Nsight, PIX, and Xcode GPU captures.
Luisa Compute is a high-performance GPU computing framework that abstracts CUDA, Vulkan, Metal, and DirectX 12 behind a unified kernel syntax. When you need to debug shaders using backend-specific profiling tools, the runtime provides a validation-driven profiling layer that translates high-level dispatch commands into native markers visible in Nsight Systems, Microsoft PIX, and Xcode GPU Debugger.
Understanding Backend-Specific Profiling Integration
Luisa Compute implements backend-agnostic debug markers that map to each GPU vendor's native profiling API. When you compile a kernel with device.compile(...) and submit it to a stream, the runtime records commands into a backend-specific command buffer. If validation is enabled, the runtime automatically wraps these commands with profiling markers.
The mapping between backends and their native tools works as follows:
- CUDA → NVTX ranges (
nvtxRangePush/nvtxRangePop) captured by Nsight Systems and Nsight Compute - DirectX 12 → PIX markers (
PIXBeginEvent/PIXEndEvent) captured by Microsoft PIX - Metal → Debug groups (
pushDebugGroup/popDebugGrouponMTLCommandBuffer) captured by Xcode GPU Debugger - Vulkan → Debug markers (
vkCmdDebugMarkerBeginEXT/vkCmdDebugMarkerEndEXT) captured by RenderDoc and Nsight Systems
According to the architecture documentation in docs/source/architecture.md, the backend-specific implementations of DeviceInterface::record_debug_* inject these calls when validation is active.
Enabling Debug Markers in Your Build
To expose shader boundaries and custom regions in profiling captures, you must build the library with debug information and enable runtime validation.
First, configure the build with debug symbols so that shader SPIR-V, PTX, or Metal IR contains source-level information:
xmake f --debug
xmake
Next, enable the validation layer before running your application. This switches the runtime into "profiling-enabled" mode and activates marker injection:
export LUISA_ENABLE_VALIDATION=1 # Linux/macOS
set LUISA_ENABLE_VALIDATION=1 # Windows CMD
Alternatively, enable validation programmatically in C++ before creating a device:
luisa::runtime::set_validation(true);
Capturing Shader Execution in Native Profilers
Once validation is enabled, launch your application while the profiling tool is attached. Each backend integration follows a specific capture workflow.
Nsight Systems/Compute (CUDA) Start a capture from the Nsight GUI or CLI, run your Luisa Compute application, and view the timeline. Each kernel appears as a labeled range in the CUDA → Kernels section, showing the kernel name and any user-defined debug groups.
Microsoft PIX (DirectX 12) Launch your application through PIX, press F12 (or click Capture) while the shader is running, then open the captured frame. Navigate to the Event List to see PIX markers delimiting kernel dispatches.
Xcode GPU Debugger (Metal) Enable GPU Frame Capture (Debug → Capture GPU Frame), run your Metal-backed Luisa Compute app, and trigger a capture. The debug navigator displays Metal debug groups as collapsible hierarchical sections surrounding command buffer submissions.
RenderDoc (Vulkan)
Launch your application through RenderDoc, capture a frame, and inspect the Event Browser. The Vulkan backend emits debug markers via VK_EXT_debug_marker, making each kernel dispatch visible as a labeled event in the command buffer hierarchy.
Adding Custom Debug Markers to Shader Code
Beyond automatic kernel wrappers, you can delimit arbitrary code regions using the push_debug_group and pop_debug_group methods on Stream objects. These calls are backend-agnostic and forward to the appropriate native API.
#include <luisa/runtime/device.h>
#include <luisa/runtime/stream.h>
int main() {
// Enable validation (if not using env var)
luisa::runtime::set_validation(true);
auto device = luisa::runtime::Device::create();
auto stream = device.create_stream();
// Begin custom profiling region
stream.push_debug_group("Particle-Advection");
// Compile and dispatch kernel
auto advect = device.compile([](Float3* pos, Float dt) noexcept {
// shader logic here
});
stream << advect(particle_buffer, dt).dispatch(particle_count);
// End region
stream.pop_debug_group();
stream.synchronize();
}
The string passed to push_debug_group appears verbatim in the profiler timeline, allowing you to correlate GPU work with specific sections of your C++ source code.
Verifying Marker Visibility in Capture Tools
After capturing a frame or timeline segment, verify that markers appear in the expected locations:
- Nsight Compute: Check Kernel → User-Defined Ranges for hierarchical nodes named after your debug groups
- PIX: Inspect the Event List → Markers column for labeled sections surrounding dispatch calls
- Xcode: Look under GPU Frame Capture → Debug Groups for collapsible sections containing command buffer work
If markers are missing, confirm that LUISA_ENABLE_VALIDATION is set in the environment where the application launches, and verify that you linked against the debug build of the Luisa Compute runtime.
Key Implementation Files
The profiling integration spans several core files in the luisagroup/luisacompute repository:
| File | Purpose |
|---|---|
docs/source/architecture.md |
Documents the validation-driven profiling architecture and backend-specific marker APIs |
src/runtime/stream.cpp |
Implements push_debug_group and pop_debug_group wrappers used by application code |
src/runtime/device.cpp |
Handles set_validation flag propagation to the RHI layer |
src/runtime/rhi/cuda_backend.cpp |
Emits NVTX ranges for CUDA kernels when validation is enabled |
src/runtime/rhi/dx12_backend.cpp |
Wraps DirectX 12 command lists with PIX event macros |
src/runtime/rhi/metal_backend.cpp |
Adds Metal debug group calls to MTLCommandBuffer |
src/runtime/rhi/vulkan_backend.cpp |
Implements Vulkan debug marker extensions |
Summary
- Enable validation via
LUISA_ENABLE_VALIDATION=1orset_validation(true)to activate backend-specific profiling markers - Build with debug symbols (
xmake f --debug) to ensure shaders contain source-level debug information - Use
push_debug_group/pop_debug_groupon streams to create labeled regions visible in Nsight, PIX, and Xcode - Verify captures in the native tool's event list or timeline view to confirm markers surround kernel dispatches
- Reference implementation files like
src/runtime/rhi/cuda_backend.cppto understand how Luisa Compute maps abstract debug calls to NVTX, PIX, or Metal APIs
Frequently Asked Questions
How do I know if validation mode is actually enabled at runtime?
Check the return value of luisa::runtime::set_validation(true) or inspect the standard output when the application starts. The runtime logs the validation state during device creation, and the presence of debug markers in your profiler capture confirms the mode is active.
Can I use these debug markers in release builds for production profiling?
While possible, release builds typically strip debug symbols and may optimize kernel code in ways that obscure source correlation. For accurate profiling with backend-specific tools, maintain a RelWithDebInfo or dedicated profiling build that preserves symbol information while enabling optimizations.
Why don't I see debug markers in RenderDoc for Vulkan backends?
Ensure your Vulkan driver supports the VK_EXT_debug_marker extension. Luisa Compute checks for this extension in src/runtime/rhi/vulkan_backend.cpp before emitting markers. If the extension is missing, upgrade your graphics driver or use Nsight Systems instead, which can capture Vulkan work without explicit debug markers.
Do custom debug groups impact GPU performance?
Yes, but minimally. The push_debug_group and pop_debug_group calls translate to lightweight CPU-side annotations that add negligible overhead to command buffer recording. However, excessive nesting (hundreds of levels deep) can increase command buffer memory usage slightly. Disable validation (LUISA_ENABLE_VALIDATION=0) for final performance benchmarks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →