# How Stream Synchronization Works with Events for Cross-Stream Dependencies in LuisaCompute

> Learn how LuisaCompute uses Vulkan Events and timeline semaphores for efficient stream synchronization, enabling seamless cross-stream dependencies without CPU blocking.

- Repository: [LuisaGroup/luisacompute](https://github.com/luisagroup/luisacompute)
- Tags: internals
- Published: 2026-03-06

---

**LuisaCompute uses Vulkan timeline semaphores wrapped in Event objects to synchronize work across different GPU streams, allowing compute and graphics queues to signal completion and wait on dependencies without CPU blocking.**

The `luisagroup/luisacompute` repository implements a backend-agnostic GPU compute framework where multiple command streams often need to coordinate execution order. Stream synchronization with Events for cross-stream dependencies provides a zero-copy, low-overhead mechanism to ensure that consumer streams begin work only after producer streams have completed specific operations.

## Core Concepts of Stream Synchronization

### Streams as Command Queues

In LuisaCompute, a **Stream** represents a GPU command queue (e.g., compute or graphics). Each stream records commands independently, but dependencies between streams require explicit synchronization primitives to prevent race conditions.

### Events and Timeline Semaphores

An **Event** wraps a Vulkan timeline semaphore (`VK_SEMAPHORE_TYPE_TIMELINE`). It tracks a monotonically increasing 64-bit fence value:

- `Event::signal` records a signal operation that increments the semaphore value
- `Event::wait` records a wait operation that blocks until the semaphore reaches a specific value

Because timeline semaphores are shareable across queues, they naturally support cross-stream dependencies.

### TimelineEvent for Host-Side Control

A **TimelineEvent** extends `Event` with host-side tracking. It allows the CPU to query completion status (`is_completed`) or block until a specific fence value is reached (`synchronize`). This is essential for frame pacing and triple-buffering scenarios.

## Implementation Details in the Vulkan Backend

### Event Creation and Initialization

In [`src/backends/vk/event.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/event.cpp), `Device::create_event()` constructs a timeline semaphore with initial value 0:

```cpp
// src/backends/vk/event.cpp
void Event::signal(Stream &stream, uint64_t value, VkCommandBuffer *cmdbuffer) {
    {
        std::lock_guard lck(eventMtx);
        lastFence = std::max(lastFence, value);
    }
    auto timelineInfo1 = get_timeline_submit(&value);
    VkSubmitInfo info1{};
    info1.sType = VK_STRUCTURE_TYPE_SUBMIT_INFO;
    info1.pNext = &timelineInfo1;
    info1.signalSemaphoreCount = 1;
    info1.pSignalSemaphores = &_semaphore;
    info1.commandBufferCount = cmdbuffer ? 1 : 0;
    info1.pCommandBuffers = cmdbuffer;
    stream.queue_mtx().lock();
    vkQueueSubmit(stream.queue(), 1, &info1, VK_NULL_HANDLE);
    stream.queue_mtx().unlock();
    mark_signal_fence(value);
}

```

### Signaling Events from a Stream

When a stream signals an event, it submits a `VkTimelineSemaphoreSubmitInfo` structure with `signalSemaphoreValueCount = 1`. This records the signal operation in the queue without blocking.

### Waiting on Events in a Stream

The `Event::wait` method in [`src/backends/vk/event.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/event.cpp) validates that the requested value has been signaled, then submits a wait operation:

```cpp
// src/backends/vk/event.cpp
void Event::wait(Stream &stream, uint64_t value) {
    auto evt_value = signaledEvent.load();
    if (evt_value < value)
        LUISA_ERROR("Waiting for fence {} greater than last signaled-fence {}", value, evt_value);
    VkTimelineSemaphoreSubmitInfo timelineInfo1{};
    timelineInfo1.sType = VK_STRUCTURE_TYPE_TIMELINE_SEMAPHORE_SUBMIT_INFO;
    timelineInfo1.waitSemaphoreValueCount = 1;
    timelineInfo1.pWaitSemaphoreValues = &value;
    VkSubmitInfo info1{};
    info1.sType = VK_STRUCTURE_TYPE_SUBMIT_INFO;
    info1.pNext = &timelineInfo1;
    info1.waitSemaphoreCount = 1;
    info1.pWaitSemaphores = &_semaphore;
    stream.queue_mtx().lock();
    vkQueueSubmit(stream.queue(), 1, &info1, VK_NULL_HANDLE);
    stream.queue_mtx().unlock();
}

```

### Stream-Level Glue Code

In [`src/backends/vk/stream.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/stream.cpp), the `Stream` class forwards signal and wait calls to the underlying `Event` while holding the stream's dispatch mutex:

```cpp
// src/backends/vk/stream.cpp
void Stream::signal(Event *event, uint64_t value) {
    std::lock_guard lck{_dispatch_mtx};
    event->signal(*this, value);
    // additional bookkeeping …
}
void Stream::wait(Event *event, uint64_t value) {
    std::lock_guard lck{_dispatch_mtx};
    event->wait(*this, value);
}

```

## Practical Cross-Stream Synchronization Example

The test suite in [`src/tests/test_runtime.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/tests/test_runtime.cpp) demonstrates a typical producer-consumer pattern between compute and graphics streams:

```cpp
// create two streams
Stream graphics = device.create_stream(StreamTag::GRAPHICS);
Stream compute  = device.create_stream(StreamTag::COMPUTE);

// event used to synchronize the streams
Event compute_event = device.create_event();

// ---- Compute side -------------------------------------------------
compute << shader(...).dispatch(...)
        << compute_event.signal();          // signal when compute work finishes

// ---- Graphics side ------------------------------------------------
graphics << compute_event.wait()           // wait for the compute signal
         << shader(...).dispatch(...)
         << swap_chain.present(image);

```

Here, the compute stream signals `compute_event` after its kernel completes. The graphics stream waits on this event before executing its commands, ensuring the GPU driver schedules the graphics work only after the compute work finishes.

## Host-Side Frame Pacing with TimelineEvent

For CPU-GPU synchronization (e.g., triple-buffering), `TimelineEvent` provides host-side blocking:

```cpp
TimelineEvent frame_fence = device.create_timeline_event();

// In the render loop:
if (frame_index >= framebuffer_count) {
    frame_fence.synchronize(frame_index - (framebuffer_count - 1));
}

// After graphics work:
graphics << frame_fence.signal(frame_index);

```

The `synchronize` method blocks the CPU until the GPU signals the specified frame value, preventing the CPU from running too far ahead of the GPU.

## Key Source Files

| File | Description |
|------|-------------|
| [src/backends/vk/event.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/vk/event.cpp) | Implements `Event` using Vulkan timeline semaphores; contains `signal()` and `wait()` methods. |
| [src/backends/vk/stream.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/backends/vk/stream.cpp) | `Stream` class forwarding signal/wait calls to events with mutex protection. |
| [src/runtime/event.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/runtime/event.cpp) | Public API for `Device::create_event()` and `Device::create_timeline_event()`. |
| [src/tests/test_runtime.cpp](https://github.com/luisagroup/luisacompute/blob/stable/src/tests/test_runtime.cpp) | Integration tests demonstrating cross-stream dependencies between compute and graphics. |

## Summary

- **Stream synchronization** in LuisaCompute relies on Vulkan timeline semaphores wrapped in `Event` objects to coordinate cross-stream dependencies.
- **Signal operations** record a semaphore increment in a stream's command queue, marking completion of work.
- **Wait operations** insert a semaphore wait command that blocks the stream until the specified value is signaled.
- **TimelineEvent** extends this mechanism to the host, allowing CPU-side synchronization for frame pacing.
- The implementation in [`src/backends/vk/event.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/event.cpp) and [`src/backends/vk/stream.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/stream.cpp) ensures thread-safe submission through mutex protection.

## Frequently Asked Questions

### What underlying GPU primitive does LuisaCompute use for cross-stream synchronization?

LuisaCompute uses **Vulkan timeline semaphores** (`VK_SEMAPHORE_TYPE_TIMELINE`). These are 64-bit monotonic counters that can be signaled and waited upon across different queues (streams), providing a lightweight, GPU-native synchronization primitive that avoids CPU intervention.

### How does a Stream signal an Event in LuisaCompute?

A `Stream` calls `event.signal()` via `Stream::signal`, which forwards to `Event::signal` in [`src/backends/vk/event.cpp`](https://github.com/luisagroup/luisacompute/blob/main/src/backends/vk/event.cpp). This method submits a `VkSubmitInfo` with a `VkTimelineSemaphoreSubmitInfo` specifying the signal value. The operation records a semaphore signal command in the stream's Vulkan queue without blocking.

### What is the difference between Event and TimelineEvent?

An **Event** is a GPU-side-only synchronization primitive used for cross-stream dependencies within the GPU. A **TimelineEvent** extends `Event` with host-side tracking, storing a fence value visible to the CPU. This allows the host to call `synchronize()` to block until the GPU reaches a specific timeline value, essential for frame pacing and triple-buffering.

### Can TimelineEvent be used for triple-buffering scenarios?

Yes. `TimelineEvent` is specifically designed for host-GPU synchronization patterns like triple-buffering. The host can maintain a running frame index and call `frame_fence.synchronize(frame_index - (framebuffer_count - 1))` to ensure it stays within the buffering limit, preventing the CPU from running too far ahead of the GPU while allowing parallel execution of multiple frames.