GPU-Based Splat Data Processors in SuperSplat: Implementation and Architecture

SuperSplat implements GPU-based splat data processors through a DataProcessor façade that orchestrates three specialized workers—Intersect, CalcBound, and CalcPositions—to perform selection tests, bound calculations, and position extraction on the GPU, serializing asynchronous operations via an internal promise chain to prevent race conditions.

SuperSplat is an open-source 3D Gaussian Splatting editor built on the PlayCanvas engine. To maintain interactive frame rates while handling millions of splats, the library offloads heavy computational workloads to the GPU using a dedicated processing architecture centered in the src/data-processor/ directory.

The DataProcessor Façade

The entry point for all GPU-based splat processing is the DataProcessor class defined in src/data-processor/index.ts. This class acts as a centralized coordinator that owns three specialized GPU worker instances: Intersect, CalcBound, and CalcPositions. Rather than exposing the complexity of shader management and texture binding directly, DataProcessor provides a simplified asynchronous API.

To prevent resource conflicts when multiple operations target the same textures, the class maintains an internal processingPromise chain. Every public method (intersect, calcBound, calcPositions, copyRt) appends its GPU workload to this chain, ensuring that GPU operations execute one at a time in the order they were called.

The Three Specialized GPU Workers

Each worker follows a consistent pattern: create shaders lazily from GLSL sources, cache render targets, bind uniforms via PlayCanvas ScopeSpace, draw a full-screen quad, and read back results asynchronously.

Intersect (src/data-processor/intersect.ts)

The Intersect processor computes per-splat intersection tests against geometric primitives including masks, rectangles, spheres, and axis-aligned boxes. It loads GLSL from src/shaders/intersection-shader.ts to execute tests in parallel across all splats.

The method returns a Uint8Array where each byte is 1 (selected) or 0 (rejected), enabling CPU-side logic to determine which splats fall within a user-drawn selection mask.

CalcBound (src/data-processor/calc-bound.ts)

The CalcBound processor calculates both selected and visible axis-aligned bounding boxes (AABBs) in a single render pass. It writes results into four 1-pixel textures representing selected-min, selected-max, visible-min, and visible-max coordinates.

Using src/shaders/bound-shader.ts, this processor performs min/max reduction operations on the GPU, avoiding expensive CPU iterations over millions of splats. The results populate BoundingBox objects passed by reference.

CalcPositions (src/data-processor/calc-positions.ts)

The CalcPositions processor generates world-space coordinates for every splat by applying model-view transformations on the GPU. It outputs to a floating-point texture via src/shaders/position-shader.ts and returns the data as a Float32Array containing [x, y, z, w] quads for each splat.

This is particularly useful for gizmo rendering, hit-testing, and exporting splat data to external formats.

Core Implementation Patterns

All three processors share four architectural patterns that ensure performance and reliability:

Lazy Shader Compilation. Each processor builds its custom shader from GLSL sources only when first invoked, storing the result to avoid recompilation overhead.

Resource Pooling. Textures and render targets are cached and recreated only when their dimensions change. This eliminates costly GPU allocations during animation or user interaction.

Scope-Based Uniform Binding. Processors use PlayCanvas ScopeSpace to bind textures, transformation matrices, and calculation parameters before invoking drawQuadWithShader to execute the compute pass.

Asynchronous GPU Read-Back. After rendering, processors call Texture.read() with immediate: false to fetch results asynchronously. CalcBound optimizes this by reading four textures in parallel using Promise.all, while the façade guarantees sequential execution through its promise chain.

Practical Usage Example

The following pattern demonstrates how to instantiate the processor and execute the three primary workflows:

import { DataProcessor } from './data-processor/index';
import { MaskOptions } from './data-processor/intersect';
import { BoundingBox } from 'playcanvas';

// Initialize with the PlayCanvas GraphicsDevice
const processor = new DataProcessor(device);

// 1. Intersection test against a mask texture
const mask: MaskOptions = { mask: userMaskTexture };
processor.intersect(mask, splat).then((selection: Uint8Array) => {
    // selection[i] === 1 indicates splat i is inside the mask
    console.log('Selected splats:', selection.filter(v => v === 1).length);
});

// 2. Compute bounds for selected and visible splats
const selBound = new BoundingBox();
const visBound = new BoundingBox();
processor.calcBound(splat, selBound, visBound).then(() => {
    console.log('Selected AABB:', selBound.getMin(), selBound.getMax());
});

// 3. Retrieve world-space positions as Float32Array
processor.calcPositions(splat).then((positions: Float32Array) => {
    // Layout: [x0, y0, z0, w0, x1, y1, z1, w1, ...]
    console.log('First splat world position:', positions.slice(0, 3));
});

Summary

Frequently Asked Questions

What is the role of the DataProcessor class?

DataProcessor acts as a high-level façade that abstracts the complexity of GPU compute operations. It owns instances of Intersect, CalcBound, and CalcPositions, and exposes simple async methods that guarantee sequential execution through an internal promise chain, preventing texture access conflicts when multiple operations are queued.

How does SuperSplat prevent race conditions during GPU processing?

The implementation prevents race conditions by maintaining a processingPromise variable in src/data-processor/index.ts that chains all GPU operations. Each call to intersect(), calcBound(), or calcPositions() appends its work to this chain, ensuring that only one compute shader executes at a time and that texture resources are not accessed simultaneously by parallel operations.

What data formats do the GPU processors return?

Intersect returns a Uint8Array where each byte represents a boolean selection state (1 for selected, 0 for rejected). CalcPositions returns a Float32Array containing world-space coordinates in [x, y, z, w] tuples. CalcBound does not return raw arrays; instead, it mutates BoundingBox objects passed as arguments with the calculated min and max vectors for both selected and visible splat sets.

Why perform splat calculations on the GPU rather than the CPU?

SuperSplat processes millions of Gaussian splats interactively. CPU-based iteration would create frame-rate bottlenecks during selection, bound calculation, and coordinate transformation. By moving these operations to compute shaders that run on the GPU using drawQuadWithShader, the library leverages parallel processing across thousands of GPU cores while keeping the JavaScript main thread responsive for UI updates.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →