# How LuisaCompute Implements Automatic Dependency Tracking and Command Reordering

> LuisaCompute simplifies GPU programming with automatic dependency tracking and command reordering. Eliminate manual barriers and optimize execution layers for efficient submission.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: internals
- Published: 2026-03-06

---

**LuisaCompute eliminates manual GPU barrier management by statically analyzing command dependencies and reordering operations into execution layers before submission.**

The command-based execution model in LuisaCompute automatically maximizes GPU parallelism while guaranteeing correctness through compile-time dependency analysis. Every GPU operation is recorded as a command derived from `luisa::runtime::rhi::Command`, which the system inspects to discover read/write dependencies between resources. This enables automatic dependency tracking and command reordering without runtime synchronization overhead, allowing independent kernels to execute concurrently while dependent operations wait in later execution layers.

## Static Dependency Analysis in the Command Reorder Visitor

The core of LuisaCompute's scheduling logic lives in [`src/backends/common/command_reorder_visitor.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/common/command_reorder_visitor.h). The **Command Reorder Visitor** traverses each command in a submitted list and determines exactly when each resource access can safely occur relative to other commands.

### Resource Handle Tracking

Each command reports the buffers, textures, bindless arrays, and acceleration structures it accesses. The visitor maintains a **`ResourceHandle`** for every tracked resource, storing the highest read and write layers already assigned (`RangeHandle::read_range`, `write_range`, and `max_view`). When the visitor encounters a command, it invokes helper methods—`set_read`, `set_write`, or `set_rw`—to query the current resource state and calculate the earliest safe **layer** for execution.

```cpp
// From src/backends/common/command_reorder_visitor.h
int64_t set_write(ResourceHandle *dst_handle, Range range) {
    int64_t layer = 0;
    switch (dst_handle->type) {
        case ResourceType::Texture_Buffer: {
            auto handle = static_cast<RangeHandle *>(dst_handle);
            layer = get_last_layer_write(handle, range);
            handle->emplace_write_layer(range, layer);
            _write_res_map.emplace(dst_handle->handle);
        } break;
        // Mesh / Accel / Bindless handled similarly …
    }
    return layer;  // layer = earliest safe execution depth
}

```

### Layer Assignment Logic

A **layer** represents a logical time-step in the execution timeline. Commands that access disjoint resources share the same layer and can run concurrently; commands that depend on a previous write are forced to a higher layer number. This is purely static analysis—no runtime synchronization primitives are emitted between commands within the same layer. The visitor guarantees that every read sees the most recent write by ensuring dependent commands occupy strictly higher layer indices than their dependencies.

## Building Execution Layers Through Command Reordering

After determining the target layer for each command, the visitor reconstructs the command sequence to respect these dependencies while maximizing parallelism.

### Per-Layer Linked Lists

The visitor stores reordered commands in `_cmd_lists`, a vector indexed by layer number. The `add_command` method inserts commands into per-layer linked lists (`CommandLink`) using a memory arena for efficient allocation.

```cpp
void add_command(Command const *cmd, int64_t layer) {
    if (_cmd_lists.size() <= layer) { _cmd_lists.resize(layer + 1); }
    auto &v = _cmd_lists[layer];
    auto new_cmd_list = _arena.allocate_memory<CommandLink, false>();
    new_cmd_list->cmd = cmd;
    new_cmd_list->p_next = v;
    v = new_cmd_list;
}

```

When the stream finally executes, it iterates the vector from layer 0 upward, dispatching all commands in layer *n* before proceeding to layer *n+1*. This ensures strict dependency ordering without explicit barrier insertion between independent operations.

### Hardware-Aware Dispatch Splitting

The visitor also respects hardware limitations through `_max_dispatch_blocks` and `max_allowed_dispatch_size`. Extremely large dispatches are automatically split across multiple layers to prevent exceeding GPU thread-block limits, ensuring the reordered command stream remains valid for the target backend's physical constraints.

## Runtime Integration with Stream Dispatch

The reordering pipeline integrates seamlessly with the public API through `CommandList` and `Stream` classes defined in [`include/luisa/runtime/command_list.h`](https://github.com/luisagroup/luisacompute/blob/main/include/luisa/runtime/command_list.h). When you submit work:

1. **User code** appends commands to a `CommandList` using the `<<` operator.
2. **Commit** wraps the list via `CommandList::commit()`, preparing it for move-only semantics.
3. **Stream::dispatch** passes the committed list to the `CommandReorderVisitor` along with a device-specific `FuncTable`.
4. The visitor returns a `span<CommandLink const *>` representing the ordered layers.
5. The **backend** (CUDA, DirectX 12, Metal) receives these layers and issues actual GPU calls, inserting hardware barriers only when crossing read-after-write or write-after-read boundaries between layers.

This architecture decouples the high-level command recording API from backend-specific synchronization primitives, allowing the same user code to run optimally across different GPU vendors.

## Practical Implementation Example

The following pattern demonstrates how automatic dependency tracking and command reordering operate transparently during typical usage:

```cpp
// ---------------------------------------------------
// 1️⃣ Build a command list
// ---------------------------------------------------
luisa::compute::CommandList cmd_list = luisa::compute::CommandList::create();
cmd_list << kernel(buffer_a, 1024u).dispatch();
cmd_list << kernel(buffer_b, 1024u).dispatch();   // independent, can be reordered

// ---------------------------------------------------
// 2️⃣ Submit to a stream (automatic reordering)
// ---------------------------------------------------
auto commit = cmd_list.commit();                    // wrap for move-only semantics
stream << std::move(commit);                       // stream internally runs CommandReorderVisitor

// ---------------------------------------------------
// 3️⃣ The visitor produces layers – user never sees them
// ---------------------------------------------------
// Internally: CommandReorderVisitor visits each command,
// calls set_read / set_write, builds _cmd_lists[layer],
// and the stream iterates these layers in order.

```

Even though the user submits `buffer_a` and `buffer_b` kernels sequentially, the visitor detects they access disjoint resources and assigns them to the same layer, allowing the backend to execute them concurrently if hardware permits.

## Summary

- **Static analysis drives scheduling**: The `CommandReorderVisitor` in [`src/backends/common/command_reorder_visitor.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/common/command_reorder_visitor.h) analyzes resource dependencies at submission time, eliminating manual barrier management.
- **Layer-based execution**: Commands are assigned to integer layers based on read/write dependencies; commands in the same layer run concurrently while higher layers wait for lower ones.
- **Zero runtime overhead**: Dependency tracking occurs before GPU submission, meaning no runtime synchronization is required between independent commands.
- **Hardware-aware splitting**: The system automatically splits oversized dispatches across layers to respect `max_allowed_dispatch_size` limits.
- **Backend agnostic**: The same dependency analysis works across CUDA, DirectX 12, and Metal backends through the `CommandList` and `Stream` API.

## Frequently Asked Questions

### What is the CommandReorderVisitor in LuisaCompute?

The **CommandReorderVisitor** is a static analyzer located in [`src/backends/common/command_reorder_visitor.h`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/common/command_reorder_visitor.h) that traverses GPU command lists before execution. It implements automatic dependency tracking and command reordering by examining which resources each command reads or writes, then reorganizing commands into dependency-respecting layers to maximize parallelism.

### How does automatic dependency tracking work without runtime synchronization?

LuisaCompute performs **static dependency analysis** during command list submission rather than at runtime. The visitor precomputes execution layers by tracking resource access ranges through `ResourceHandle` objects. Since the system knows the exact read/write dependencies before the GPU executes any work, it can reorder commands into layers that guarantee correctness without inserting runtime barriers between concurrent operations.

### What are layers in LuisaCompute's command execution model?

**Layers** are integer indices representing logical execution steps in the reordered command stream. Commands assigned to layer 0 execute first, followed by layer 1, and so on. Independent commands accessing different resources occupy the same layer and may run concurrently, while dependent commands occupy higher layers to ensure they observe the results of previous writes. The backend iterates through these layers sequentially in `_cmd_lists`.

### How does command reordering improve GPU performance?

Command reordering improves performance by exposing **instruction-level parallelism** to the GPU. Without reordering, commands execute serially in submission order, potentially leaving execution units idle. By analyzing dependencies and placing independent commands in the same layer, LuisaCompute allows the backend to submit multiple kernels simultaneously or overlap memory transfers with computation, maximizing hardware utilization while maintaining correct resource access ordering.