# How to Profile and Optimize Rendering Performance on Different LuisaRender Backends

> Optimize LuisaRender performance across backends. Profile kernel bottlenecks using the Luisa Compute profiler and tune thread pools or tile granularity for faster rendering.

- Repository: [LuisaGroup/luisarender](https://github.com/luisagroup/luisarender)
- Tags: performance
- Published: 2026-03-06

---

**Enable the Luisa Compute profiler with `LUISA_PROFILER=1`, analyze the generated `luisa_profiler.csv` to identify kernel bottlenecks, and apply back‑end‑specific optimizations such as tuning thread‑pool sizes for CPU or increasing tile granularity for CUDA.**

LuisaRender is a physically‑based renderer built on the **Luisa Compute** abstraction layer, supporting multiple hardware back‑ends including CUDA, CPU, Metal, DirectX 12, and ISPC. To extract maximum throughput from each target, you must **profile and optimize rendering performance on different LuisaRender backends** using the built‑in profiler and back‑end‑specific tuning parameters.

## Selecting the Right Back‑end for Your Hardware

The command‑line interface in [`src/apps/cli.cpp`](https://github.com/luisagroup/luisarender/blob/main/src/apps/cli.cpp) exposes a `-b/--backend` option that maps directly to Luisa Compute’s device factory. Lines 59‑63 define the option:

```cpp
cli.add_option("", "b", "backend", backend_hints,
               cxxopts::value<luisa::string>(), "<backend>");

```

At line 176, the CLI instantiates the device by passing the user‑supplied string to `context.create_device()`:

```cpp
auto backend = options["backend"].as<luisa::string>();
auto device  = context.create_device(backend, &config);

```

Valid values are enumerated by `luisa::compute::Context::installed_backends()` (see lines 160‑168 in the same file). Choose **CUDA** for NVIDIA GPUs, **Metal** for Apple silicon, **CPU** for portability, **DirectX 12** for Windows‑integrated GPUs, and **ISPC** for SIMD‑optimized CPU vectorization.

## Enabling Luisa Compute Profiling

Luisa Compute ships with a high‑resolution profiler that records kernel launch times, memory transfers, and CPU‑side work. You can activate it via environment variable, CLI extension, or programmatically.

### Environment Variable Method

Set `LUISA_PROFILER` before launching the executable:

```bash
export LUISA_PROFILER=1

# or for detailed traces

export LUISA_PROFILER=trace

./luisa-render-cli -b cuda -d 0 scenes/cornell_box.luisa

```

The runtime writes `luisa_profiler.csv` on exit, containing per‑kernel timings and throughput metrics.

### Programmatic API

For fine‑grained control—profiling only a specific integrator pass—use the C++ API:

```cpp
#include <luisa/runtime/profiler.h>

int main() {
    luisa::compute::Context ctx;
    auto device = ctx.create_device("cpu");
    
    luisa::compute::Profiler::enable();  // start collection
    
    // ... build pipeline ...
    pipeline->render();
    
    luisa::compute::Profiler::dump("profile.csv");  // write results
}

```

## Back‑end‑Specific Optimization Strategies

Once profiling identifies bottlenecks, apply the following targeted tweaks.

### CUDA Back‑end Optimization

In [`src/apps/cli.cpp`](https://github.com/luisagroup/luisarender/blob/main/src/apps/cli.cpp), the CUDA device is created with default stream semantics. To maximize throughput:

- **Increase tile granularity** in integrators to amortize kernel launch overhead.
- **Minimize host‑side synchronization**—avoid calling `device.synchronize()` inside sampling loops.
- **Enable L2 cache persistence** if the Luisa Compute version exposes `device.set_cache_mode(luisa::CacheMode::L2)`.

### CPU Back‑end Optimization

The CPU back‑end relies on a thread pool defined in [`src/util/thread_pool.cpp`](https://github.com/luisagroup/luisarender/blob/main/src/util/thread_pool.cpp). Key optimizations:

- **Tune thread‑pool size**—the default uses `std::thread::hardware_concurrency()`, but reducing it to physical core count (excluding hyper‑threads) can improve cache locality.
- **Align data structures** to 64‑byte boundaries using `alignas(64)` to prevent false sharing.
- **Leverage Embree**—when available, Luisa Compute dispatches ray‑geometry queries to Intel Embree on CPU, which is significantly faster than scalar traversal.

### Metal Back‑end Optimization

For Apple silicon (M1/M2/M3):

- **Reduce descriptor set churn**—bundle textures into arrays rather than rebinding per draw.
- **Match threadgroup size** to the GPU’s execution width (typically 32 or 64 threads) using `device.set_preferred_threads(N)`.
- **Profile with Xcode Instruments** alongside the Luisa CSV to catch driver‑level latency.

### DirectX 12 Back‑end Optimization

On Windows:

- **Reuse `luisa::compute::Buffer` objects** across frames to avoid GPU memory allocation fragmentation.
- **Batch command list submissions**—call `CommandBuffer::commit()` once per frame rather than per kernel.
- **Place resource barriers carefully**—unnecessary transitions between `COPY` and `UNORDERED_ACCESS` states stall the graphics queue.

### ISPC Back‑end Optimization

The ISPC target vectorizes code for AVX‑512/NEON:

- **Compile with ISPC enabled**: `cmake -B build -DLUISA_COMPUTE_ENABLE_ISPC=ON -DCMAKE_BUILD_TYPE=Release`.
- **Set native architecture flags**: `-mavx2` or `-march=native` in `CMAKE_CXX_FLAGS` to match your CPU.
- **Align spectra buffers** to 32‑byte boundaries—ISPC loads/stores are fastest with aligned SIMD registers.

## Practical Profiling Examples

### Profiling a CUDA Render from the Command Line

```bash

# Enable profiler via environment

export LUISA_PROFILER=1

# Run with CUDA backend on device 0

./luisa-render-cli -b cuda -d 0 scenes/cornell_box.luisa

# Inspect the output

head luisa_profiler.csv

```

### Programmatic Profiling of a Specific Integrator Pass

```cpp
#include <luisa/runtime/context.h>
#include <luisa/runtime/device.h>
#include <luisa/runtime/profiler.h>
#include <luisa/render/pipeline.h>

int main() {
    // Create context and device
    luisa::compute::Context ctx;
    auto device = ctx.create_device("metal");
    
    // Enable profiling
    luisa::compute::Profiler::enable();
    
    // Build pipeline (simplified)
    auto pipeline = luisa::make_unique<luisa::render::Pipeline>(device);
    // ... set camera, integrator, film ...
    
    // Render and capture
    pipeline->render();
    
    // Dump results
    luisa::compute::Profiler::dump("metal_profile.csv");
    return 0;
}

```

### Adjusting CPU Thread‑Pool Size for Cache Efficiency

Edit [`src/util/thread_pool.cpp`](https://github.com/luisagroup/luisarender/blob/main/src/util/thread_pool.cpp) around line 30:

```cpp
// Default uses all hardware threads
// auto max_threads = std::thread::hardware_concurrency();

// Optimized for physical cores (e.g., 8 cores, 16 threads -> use 8)
auto max_threads = 8u;
ThreadPool pool(max_threads);

```

Rebuild and profile to verify reduced cache thrashing.

### Switching to the ISPC Back‑end via CMake

```bash

# Configure with ISPC support

cmake -B build \
  -DLUISA_COMPUTE_ENABLE_ISPC=ON \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_CXX_FLAGS="-march=native"

# Build

cmake --build build -j$(nproc)

# Run with ISPC backend

./build/luisa-render-cli -b ispc scene.luisa

```

## Summary

- **Select the appropriate back‑end** via the `-b/--backend` flag in [`src/apps/cli.cpp`](https://github.com/luisagroup/luisarender/blob/main/src/apps/cli.cpp) to match your hardware (CUDA for NVIDIA, Metal for Apple Silicon, CPU/ISPC for cross‑platform CPU).
- **Enable profiling** by setting `LUISA_PROFILER=1` or calling `luisa::compute::Profiler::enable()` to generate `luisa_profiler.csv` with per‑kernel timings.
- **Optimize CUDA** by increasing tile sizes and minimizing `device.synchronize()` calls.
- **Optimize CPU** by tuning the thread‑pool size in [`src/util/thread_pool.cpp`](https://github.com/luisagroup/luisarender/blob/main/src/util/thread_pool.cpp) and aligning data to 64‑byte boundaries.
- **Optimize Metal** by bundling textures and matching threadgroup sizes to Apple Silicon execution widths.
- **Optimize DirectX 12** by reusing buffers and batching command list submissions.
- **Optimize ISPC** by compiling with `-DLUISA_COMPUTE_ENABLE_ISPC=ON` and aligning buffers to 32 bytes for SIMD efficiency.

## Frequently Asked Questions

### How do I enable profiling without modifying source code?

Set the environment variable `LUISA_PROFILER=1` before running the CLI. The runtime will automatically collect kernel timings and write `luisa_profiler.csv` when the process exits. For detailed traces, use `LUISA_PROFILER=trace`.

### Can I limit the CPU back‑end to a specific number of threads?

Yes. Modify [`src/util/thread_pool.cpp`](https://github.com/luisagroup/luisarender/blob/main/src/util/thread_pool.cpp) around line 30 where the pool is constructed. Replace `std::thread::hardware_concurrency()` with your desired thread count (e.g., `8u` for 8 threads), then rebuild. This reduces context switching and improves cache locality on CPUs with hyper‑threading.

### Why does the Metal back‑end show high driver latency in the profiler?

Metal exhibits higher CPU‑side driver overhead when descriptor sets are rebound frequently. Bundle textures into arrays and minimize state changes. Additionally, ensure your threadgroup size matches the Apple Silicon GPU’s execution width (typically 32 or 64 threads) by setting `device.set_preferred_threads(N)` if available.

### What CMake flags are required to build the ISPC back‑end?

Pass `-DLUISA_COMPUTE_ENABLE_ISPC=ON` during configuration. For maximum vectorization, also set `-mavx2` or `-march=native` in `CMAKE_CXX_FLAGS`. Ensure your spectra buffers are aligned to 32‑byte boundaries to allow ISPC’s SIMD loads to operate at peak throughput.