How Stream Synchronization Works with Events for Cross-Stream Dependencies in LuisaCompute

LuisaCompute uses Vulkan timeline semaphores wrapped in Event objects to synchronize work across different GPU streams, allowing compute and graphics queues to signal completion and wait on dependencies without CPU blocking.

The luisagroup/luisacompute repository implements a backend-agnostic GPU compute framework where multiple command streams often need to coordinate execution order. Stream synchronization with Events for cross-stream dependencies provides a zero-copy, low-overhead mechanism to ensure that consumer streams begin work only after producer streams have completed specific operations.

Core Concepts of Stream Synchronization

Streams as Command Queues

In LuisaCompute, a Stream represents a GPU command queue (e.g., compute or graphics). Each stream records commands independently, but dependencies between streams require explicit synchronization primitives to prevent race conditions.

Events and Timeline Semaphores

An Event wraps a Vulkan timeline semaphore (VK_SEMAPHORE_TYPE_TIMELINE). It tracks a monotonically increasing 64-bit fence value:

  • Event::signal records a signal operation that increments the semaphore value
  • Event::wait records a wait operation that blocks until the semaphore reaches a specific value

Because timeline semaphores are shareable across queues, they naturally support cross-stream dependencies.

TimelineEvent for Host-Side Control

A TimelineEvent extends Event with host-side tracking. It allows the CPU to query completion status (is_completed) or block until a specific fence value is reached (synchronize). This is essential for frame pacing and triple-buffering scenarios.

Implementation Details in the Vulkan Backend

Event Creation and Initialization

In src/backends/vk/event.cpp, Device::create_event() constructs a timeline semaphore with initial value 0:

// src/backends/vk/event.cpp
void Event::signal(Stream &stream, uint64_t value, VkCommandBuffer *cmdbuffer) {
    {
        std::lock_guard lck(eventMtx);
        lastFence = std::max(lastFence, value);
    }
    auto timelineInfo1 = get_timeline_submit(&value);
    VkSubmitInfo info1{};
    info1.sType = VK_STRUCTURE_TYPE_SUBMIT_INFO;
    info1.pNext = &timelineInfo1;
    info1.signalSemaphoreCount = 1;
    info1.pSignalSemaphores = &_semaphore;
    info1.commandBufferCount = cmdbuffer ? 1 : 0;
    info1.pCommandBuffers = cmdbuffer;
    stream.queue_mtx().lock();
    vkQueueSubmit(stream.queue(), 1, &info1, VK_NULL_HANDLE);
    stream.queue_mtx().unlock();
    mark_signal_fence(value);
}

Signaling Events from a Stream

When a stream signals an event, it submits a VkTimelineSemaphoreSubmitInfo structure with signalSemaphoreValueCount = 1. This records the signal operation in the queue without blocking.

Waiting on Events in a Stream

The Event::wait method in src/backends/vk/event.cpp validates that the requested value has been signaled, then submits a wait operation:

// src/backends/vk/event.cpp
void Event::wait(Stream &stream, uint64_t value) {
    auto evt_value = signaledEvent.load();
    if (evt_value < value)
        LUISA_ERROR("Waiting for fence {} greater than last signaled-fence {}", value, evt_value);
    VkTimelineSemaphoreSubmitInfo timelineInfo1{};
    timelineInfo1.sType = VK_STRUCTURE_TYPE_TIMELINE_SEMAPHORE_SUBMIT_INFO;
    timelineInfo1.waitSemaphoreValueCount = 1;
    timelineInfo1.pWaitSemaphoreValues = &value;
    VkSubmitInfo info1{};
    info1.sType = VK_STRUCTURE_TYPE_SUBMIT_INFO;
    info1.pNext = &timelineInfo1;
    info1.waitSemaphoreCount = 1;
    info1.pWaitSemaphores = &_semaphore;
    stream.queue_mtx().lock();
    vkQueueSubmit(stream.queue(), 1, &info1, VK_NULL_HANDLE);
    stream.queue_mtx().unlock();
}

Stream-Level Glue Code

In src/backends/vk/stream.cpp, the Stream class forwards signal and wait calls to the underlying Event while holding the stream's dispatch mutex:

// src/backends/vk/stream.cpp
void Stream::signal(Event *event, uint64_t value) {
    std::lock_guard lck{_dispatch_mtx};
    event->signal(*this, value);
    // additional bookkeeping …
}
void Stream::wait(Event *event, uint64_t value) {
    std::lock_guard lck{_dispatch_mtx};
    event->wait(*this, value);
}

Practical Cross-Stream Synchronization Example

The test suite in src/tests/test_runtime.cpp demonstrates a typical producer-consumer pattern between compute and graphics streams:

// create two streams
Stream graphics = device.create_stream(StreamTag::GRAPHICS);
Stream compute  = device.create_stream(StreamTag::COMPUTE);

// event used to synchronize the streams
Event compute_event = device.create_event();

// ---- Compute side -------------------------------------------------
compute << shader(...).dispatch(...)
        << compute_event.signal();          // signal when compute work finishes

// ---- Graphics side ------------------------------------------------
graphics << compute_event.wait()           // wait for the compute signal
         << shader(...).dispatch(...)
         << swap_chain.present(image);

Here, the compute stream signals compute_event after its kernel completes. The graphics stream waits on this event before executing its commands, ensuring the GPU driver schedules the graphics work only after the compute work finishes.

Host-Side Frame Pacing with TimelineEvent

For CPU-GPU synchronization (e.g., triple-buffering), TimelineEvent provides host-side blocking:

TimelineEvent frame_fence = device.create_timeline_event();

// In the render loop:
if (frame_index >= framebuffer_count) {
    frame_fence.synchronize(frame_index - (framebuffer_count - 1));
}

// After graphics work:
graphics << frame_fence.signal(frame_index);

The synchronize method blocks the CPU until the GPU signals the specified frame value, preventing the CPU from running too far ahead of the GPU.

Key Source Files

File Description
src/backends/vk/event.cpp Implements Event using Vulkan timeline semaphores; contains signal() and wait() methods.
src/backends/vk/stream.cpp Stream class forwarding signal/wait calls to events with mutex protection.
src/runtime/event.cpp Public API for Device::create_event() and Device::create_timeline_event().
src/tests/test_runtime.cpp Integration tests demonstrating cross-stream dependencies between compute and graphics.

Summary

  • Stream synchronization in LuisaCompute relies on Vulkan timeline semaphores wrapped in Event objects to coordinate cross-stream dependencies.
  • Signal operations record a semaphore increment in a stream's command queue, marking completion of work.
  • Wait operations insert a semaphore wait command that blocks the stream until the specified value is signaled.
  • TimelineEvent extends this mechanism to the host, allowing CPU-side synchronization for frame pacing.
  • The implementation in src/backends/vk/event.cpp and src/backends/vk/stream.cpp ensures thread-safe submission through mutex protection.

Frequently Asked Questions

What underlying GPU primitive does LuisaCompute use for cross-stream synchronization?

LuisaCompute uses Vulkan timeline semaphores (VK_SEMAPHORE_TYPE_TIMELINE). These are 64-bit monotonic counters that can be signaled and waited upon across different queues (streams), providing a lightweight, GPU-native synchronization primitive that avoids CPU intervention.

How does a Stream signal an Event in LuisaCompute?

A Stream calls event.signal() via Stream::signal, which forwards to Event::signal in src/backends/vk/event.cpp. This method submits a VkSubmitInfo with a VkTimelineSemaphoreSubmitInfo specifying the signal value. The operation records a semaphore signal command in the stream's Vulkan queue without blocking.

What is the difference between Event and TimelineEvent?

An Event is a GPU-side-only synchronization primitive used for cross-stream dependencies within the GPU. A TimelineEvent extends Event with host-side tracking, storing a fence value visible to the CPU. This allows the host to call synchronize() to block until the GPU reaches a specific timeline value, essential for frame pacing and triple-buffering.

Can TimelineEvent be used for triple-buffering scenarios?

Yes. TimelineEvent is specifically designed for host-GPU synchronization patterns like triple-buffering. The host can maintain a running frame index and call frame_fence.synchronize(frame_index - (framebuffer_count - 1)) to ensure it stays within the buffering limit, preventing the CPU from running too far ahead of the GPU while allowing parallel execution of multiple frames.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →