How to Use BindlessArray to Reduce Binding Overhead in Complex Shader Pipelines
LuisaCompute's BindlessArray abstraction consolidates thousands of individual texture and buffer bindings into a single descriptor set, allowing shaders to access resources via integer indices rather than separate per-resource binding calls.
Complex GPU pipelines—such as path tracers or deferred renderers—often allocate hundreds of temporary textures and buffers per frame, causing severe CPU overhead from descriptor set churn. The luisagroup/luisacompute repository solves this through the BindlessArray API, which collapses individual resource bindings into one heap-like structure that requires only a single bind operation per dispatch.
Creating a BindlessArray from the Device
You initialize a bindless array through the Device interface, specifying the maximum slot count and reference type. In include/luisa/runtime/device.h (lines 78–82), the factory method signature is:
auto create_bindless_array(size_t slot_count = 65536u,
BindlessSlotType type = BindlessSlotType::MULTIPLE) -> BindlessArray;
The slot count determines how many resources the heap can reference simultaneously. The slot type controls reference semantics: BindlessSlotType::SINGLE allows one resource per slot (simpler bookkeeping), while MULTIPLE permits multiple references for advanced aliasing scenarios.
#include <luisa/runtime/device.h>
#include <luisa/runtime/bindless_array.h>
// Create a heap with 64,384 slots, supporting multiple references per slot
auto heap = device.create_bindless_array(64384, luisa::compute::BindlessSlotType::MULTIPLE);
Populating Slots with Resource Updates
Instead of binding resources individually, you populate the array by recording modification commands. The backend maintains an internal buffer that stores a descriptor index per slot. In the Vulkan backend (src/backends/vk/bindless_array.h lines 16–29), this is represented by BindlessStruct, which holds indices for buffers, 2D textures, 3D textures, and volumes.
To update the heap, construct a vector of modifications and dispatch an update command:
#include <luisa/runtime/command.h>
// Create resources
auto tex0 = device.create_image<float>(PixelStorage::FLOAT4, 512, 512);
auto tex1 = device.create_image<float>(PixelStorage::FLOAT4, 1024, 1024);
// Record which slots receive which resources
std::vector<luisa::compute::BindlessArrayUpdateCommand::Texture2DModification> mods;
mods.push_back({0, tex0.handle()}); // slot 0 ← tex0
mods.push_back({1, tex1.handle()}); // slot 1 ← tex1
// Dispatch the update; the backend writes descriptor indices into the heap's buffer
device.stream().command(luisa::compute::BindlessArrayUpdateCommand{heap, std::move(mods)})
.dispatch();
The BindlessArray::update() and BindlessArray::bind() overloads (lines 68–92 in src/backends/vk/bindless_array.h) handle the actual descriptor-set write commands, ensuring the GPU-visible buffer containing descriptor indices stays synchronized.
Accessing Resources in Kernel DSL
Inside a kernel or raster shader, you declare a BindlessVar argument. The DSL provides indexed access methods such as heap.texture2d(index) or heap[index] for buffers. The compiler translates these into bindless fetch instructions that read the descriptor index from the heap's internal buffer before accessing the actual resource.
#include <luisa/dsl/syntax.h>
using namespace luisa::compute;
Kernel2D kernel = [&](ImageVar<float4> out,
BindlessVar heap,
UInt2 dispatch_id) noexcept {
// Retrieve textures by their slot indices
auto tex0 = heap.texture2d(0_u); // slot 0
auto tex1 = heap.texture2d(1_u); // slot 1
Float2 uv = make_float2(dispatch_id) / make_float2(out.width(), out.height());
auto c0 = tex0.sample(uv);
auto c1 = tex1.sample(uv);
out[dispatch_id] = lerp(c0, c1, 0.5f);
};
auto shader = device.compile(kernel);
shader(out_image, heap, dispatch_size).dispatch();
Real-world usage patterns appear in src/tests/test_bindless.cpp, which demonstrates end-to-end creation, binding, and shader consumption. Path tracing implementations in src/tests/test_path_tracing.cpp further illustrate how bindless arrays manage thousands of scene textures.
Backend Implementation and Descriptor Binding
The performance gain stems from collapsing resource visibility into a single descriptor set. In the Vulkan backend (src/backends/vk/bindless_array.h lines 63–70), the pre_update and copy_index methods ensure the VK_DESCRIPTOR_SET is updated only when the heap changes. At dispatch time, only this single bindless array descriptor set is bound; all subsequent resource accesses are indirect via the heap buffer, eliminating per-resource vkCmdBindDescriptorSets calls.
The DirectX 12 backend follows an identical conceptual design in src/backends/dx/Resource/BindlessArray.h, using GPU-visible descriptor heaps indexed by integer offsets.
Python API for Rapid Prototyping
LuisaCompute’s Python bindings expose the same workflow for dynamic resource management:
import luisa
device = luisa.Device()
heap = device.create_bindless_array() # defaults to 65536 slots
# Create images
img0 = device.create_image(luisa.PixelStorage.FLOAT4, 256, 256)
img1 = device.create_image(luisa.PixelStorage.FLOAT4, 512, 512)
# Populate slots 0 and 1
heap.add_image(0, img0)
heap.add_image(1, img1)
@device.kernel
def blend(out: luisa.ImageFloat, heap: luisa.BindlessArray):
i = luisa.dispatch_id().x
uv = luisa.float2(i) / out.size()
c0 = heap.texture2d(0).sample(uv)
c1 = heap.texture2d(1).sample(uv)
out[i] = (c0 + c1) * 0.5
blend(out_img, heap, dispatch_size=(256, 1, 1))
Summary
- BindlessArray replaces per-resource descriptor bindings with a single heap object containing integer-indexed resource references.
- Create the array via
Device::create_bindless_array()with configurable slot counts up to 65,536 and selectable slot types (SINGLEorMULTIPLE). - Update resources using
BindlessArrayUpdateCommandwith modification structures (Texture2DModification,BufferModification, etc.) to batch descriptor index writes. - Shaders access resources through
BindlessVarusing index-based methods liketexture2d(slot); the compiler generates bindless fetch instructions. - The Vulkan backend stores descriptor indices in a
BindlessStructbuffer and binds only one descriptor set per dispatch, minimizing CPU overhead insrc/backends/vk/bindless_array.h.
Frequently Asked Questions
What is the maximum number of slots supported in a BindlessArray?
The default capacity is 65,536 slots, defined by the slot_count parameter in Device::create_bindless_array(). You can request fewer slots to reduce memory footprint, or request more if the backend supports larger descriptor heaps, though 65,536 is the standard tested limit across Vulkan and DirectX 12 backends.
How does BindlessArray differ from traditional descriptor set binding?
Traditional pipelines bind each texture or buffer to a specific descriptor set slot at the API level (e.g., vkCmdBindDescriptorSets per resource). BindlessArray binds one descriptor set containing a buffer of indices; the shader reads the index then accesses the resource. This reduces CPU-side binding calls from O(N) per dispatch to O(1), as implemented in the Vulkan backend's single-set binding logic in src/backends/vk/bindless_array.h.
Can I mix different resource types in a single BindlessArray?
Yes. The BindlessStruct in the Vulkan backend (lines 16–29 of src/backends/vk/bindless_array.h) reserves separate index fields for buffers, 2D textures, 3D textures, and volumes. You can store a texture in slot 0, a buffer in slot 1, and a volume in slot 2, provided you use the appropriate modification type (Texture2DModification, BufferModification, etc.) when updating the heap.
Is there a performance cost for the indirection introduced by bindless access?
There is a minor GPU-side indirection cost: the shader reads the descriptor index from the heap buffer before accessing the resource. However, this is typically cache-friendly and far outweighed by the CPU binding overhead savings, especially in complex pipelines issuing thousands of resource transitions per frame. The architecture is designed to keep the indirection buffer resident in GPU-visible memory to minimize latency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →