Best Practices for Kernel Optimization in LuisaCompute: Memory Coalescing, Occupancy, and Divergence

Optimize LuisaCompute kernels by ensuring memory coalescing for contiguous data access, maximizing occupancy through power-of-two block sizes, and minimizing branch divergence using uniform control flow and the $if DSL construct.

Kernel optimization is critical for achieving peak performance in GPU compute workloads. In the luisagroup/luisacompute repository, the architecture documentation establishes three fundamental pillars for efficient kernel execution: memory coalescing, occupancy management, and branch divergence control. Mastering these concepts allows developers to write high-performance compute shaders that fully utilize modern GPU hardware.

The Three Pillars of Kernel Optimization

According to the docs/source/architecture.md file, LuisaCompute follows the same optimization principles as CUDA, Vulkan Compute, and Metal GPU kernels. The documentation explicitly identifies three pillars that govern kernel performance:

  1. Memory coalescing – arranging data accesses so consecutive threads read or write consecutive memory locations, allowing the GPU to combine transactions.
  2. Occupancy – selecting work-group sizes that balance register pressure, shared memory consumption, and hardware limits to maximize active warps.
  3. Branch divergence – avoiding code paths that force threads within the same warp to execute different branches serially.

Memory Coalescing for Contiguous Data Access

Why Coalesced Access Matters

GPUs fetch memory in 32- or 64-byte segments per warp. When threads access non-contiguous memory locations, the hardware generates multiple separate fetch transactions, severely reducing effective bandwidth. Coalesced access ensures that adjacent threads pull adjacent data, allowing the memory controller to service the entire warp with a single transaction.

Implementing Coalesced Reads in LuisaCompute

In LuisaCompute, you achieve coalescing by mapping dispatch_id() coordinates to linear buffer indices that increase by one per thread. The repository's matrix multiplication test in src/tests/test_matrix_multiply.cpp demonstrates this pattern in lines 92-100:

// Inside the GEMM kernel – each thread processes one output element
UInt3 id = dispatch_id().zyx();               // (batch, row, col)
Float r = 0.f;

// Coalesced reads: consecutive `i` values are read by consecutive threads
for (auto i : dynamic_range(iterate_size)) {
    auto lhs_val = lhs.read(lhs_global_offset + lhs_size.y * i + id.y);
    auto rhs_val = rhs.read(rhs_global_offset + rhs_size.y * id.x + i);
    r += Float(lhs_val * rhs_val);
}

All reads walk linearly through the buffers, guaranteeing that adjacent threads pull adjacent memory locations. This pattern ensures that BufferVar<T> reads coalesce into single memory transactions.

Maximizing Occupancy Through Block Size Tuning

Balancing GPU Resources

Occupancy measures the ratio of active warps to the maximum warps supported by each Streaming Multiprocessor (SM). High occupancy helps the warp scheduler hide memory latency by switching to ready warps when others stall. However, occupancy is constrained by register file usage, shared memory consumption, and the hardware's thread block limits.

Dynamic Block Size Selection with get_proper_dispatch_size()

LuisaCompute provides utilities to select block sizes that respect warp granularity. In src/tests/test_matrix_multiply.cpp, lines 27-36 define get_proper_dispatch_size(), which calculates a power-of-two block size that divides the matrix dimension evenly:

// Helper that picks a block size which is a power of two divisor of group_size
static uint2 get_proper_dispatch_size(uint group_size) {
    uint2 block_size(1);
    for (uint i = 7; i >= 0; --i) {
        if (group_size % (1 << i) == 0) {
            block_size.y = (1 << i);
            break;
        }
    }
    block_size.x = 128 / block_size.y;
    return block_size;
}

Lines 40-62 of the same file demonstrate applying this helper via set_block_size():

set_block_size(get_proper_dispatch_size(end_node_size));

By ensuring block dimensions are multiples of the warp size (32) and powers of two, this approach maximizes the number of active warps per SM, improving throughput for memory-bound kernels.

Eliminating Branch Divergence in Warps

The Cost of Divergent Execution

Modern GPUs execute instructions in warps (groups of 32 threads) in lockstep. When threads within a warp take different code paths—such as when an if condition evaluates differently per thread—the hardware serializes execution, running each path sequentially and masking inactive threads. This branch divergence effectively divides parallelism and halves (or worse) effective throughput.

Using $if for Uniform Control Flow

LuisaCompute's DSL distinguishes between host-side control flow and device-side control flow. The $if construct evaluates conditions at DSL-trace time, generating static shader code. When the condition is uniform across the warp (identical for all threads), the generated code contains no divergent branches.

The Mandelbrot example in docs/source/tutorials.md (lines 119-122) demonstrates minimizing divergence with early exit:

$for (i, MAX_ITER) {
    // ... compute z ...
    $if (dot(z, z) > 4.0f) {
        $break;               // early exit – divergent but short
    };
    n += 1.0f;
}

While the $break introduces divergence (threads exit at different iterations), keeping the divergent section short limits the performance penalty. For truly uniform conditions, use host-side if statements to avoid generating device branches entirely.

The src/tests/test_warp.cpp file (line 50) provides a minimal test case for verifying warp execution patterns and divergence behavior.

Key Implementation Files for Kernel Optimization

File Optimization Relevance
docs/source/architecture.md Documents the three pillars: memory coalescing, occupancy, and divergence.
src/tests/test_matrix_multiply.cpp Demonstrates coalesced memory access (lines 92-100), dynamic block sizing (lines 27-36), and occupancy tuning (lines 40-62).
src/tests/test_warp.cpp Validates warp execution patterns and divergence behavior (line 50).
docs/source/tutorials.md Contains the Mandelbrot kernel demonstrating $if usage for branch control (lines 119-122).
src/backends/cuda/cuda_builtin/cuda_builtin_kernels.cu Shows low-level CUDA code generation from DSL constructs.

Summary

  • Memory coalescing requires arranging buffer accesses so consecutive threads read consecutive memory locations, as demonstrated in src/tests/test_matrix_multiply.cpp using linear index calculations with dispatch_id().
  • Occupancy depends on selecting block sizes that are multiples of the warp size (32) and powers of two; use get_proper_dispatch_size() and set_block_size() to balance register pressure and active warps.
  • Branch divergence degrades performance when threads in a warp execute different paths; minimize divergence using host-side if for uniform conditions and keep device-side $if branches short when unavoidable.
  • Reference the docs/source/architecture.md guide and the matrix multiplication test case for concrete implementation patterns that adhere to CUDA, Vulkan, and Metal optimization principles.

Frequently Asked Questions

What is memory coalescing in GPU programming?

Memory coalescing is a hardware optimization where the GPU combines multiple memory requests from threads in the same warp into a single transaction. When consecutive threads access consecutive memory locations—such as adjacent elements in a BufferVar<T>—the memory controller fetches a single wide cache line rather than separate addresses, significantly increasing effective bandwidth and reducing latency.

How does block size affect kernel occupancy?

Block size directly determines how many warps can reside simultaneously on a Streaming Multiprocessor (SM). If a block uses too many registers or shared memory, fewer blocks can fit, reducing occupancy. Conversely, choosing a block size that is a multiple of the warp size (32) and a power of two—such as 64, 128, or 256—typically maximizes active warps and allows the scheduler to hide memory latency by switching between ready warps.

What causes branch divergence in GPU kernels?

Branch divergence occurs when threads within the same warp follow different execution paths due to data-dependent conditional statements. For example, if an if statement evaluates to true for some threads and false for others within a warp, the GPU must serialize execution, running each path separately while disabling threads on the opposite path. This effectively divides the warp's parallelism and reduces throughput proportionally to the number of divergent paths taken.

How does LuisaCompute handle divergent branches?

LuisaCompute provides the $if and $else DSL constructs to distinguish between host-side and device-side control flow. When conditions are uniform across the warp, $if generates branch-free shader code at trace time. For data-dependent conditions that must diverge—such as early exit in iterative algorithms—LuisaCompute compiles the branch normally, but developers should keep the divergent section minimal to limit performance penalties, as demonstrated in the Mandelbrot example in docs/source/tutorials.md.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →