How JSAR Achieves High Performance Rendering: GPU-Centric Architecture and Batch Optimization

JSAR achieves high performance rendering by enforcing a strict draw call budget—often ≤10 per frame—through aggressive dynamic batching, instanced geometry rendering, and texture atlasing, while leveraging a thin WebGL wrapper and Skia GPU acceleration to minimize CPU overhead.

The JSAR runtime (m-creativelab/jsar-runtime) delivers smooth 60fps+ rendering for complex spatialized HTML content by treating the GPU as the primary execution target rather than an afterthought. By implementing a JSAR high performance rendering pipeline that prioritizes batch consolidation and state minimization, the engine eliminates traditional bottlenecks associated with DOM-to-3D translation.

The GPU-Centric Rendering Pipeline

At the core of JSAR's performance strategy is a thin WebGL/OpenGL ES backend that maps HTML-based UI elements directly to GPU primitives. This architecture ensures that the CPU acts merely as a coordinator, while the GPU handles the heavy lifting of rasterization and composition.

WebGL Context with Draw Call Limits

The runtime enforces a hard cap on GPU draw calls through the WEBGL_MAX_COUNT_PER_DRAWCALL constant defined in src/client/graphics/webgl_context.cpp. When a scene exceeds this budget, the system logs an error warning, forcing developers to optimize their content. This strict limit ensures that the renderer never submits more than the optimal number of commands to the GPU per frame.

// src/client/graphics/webgl_context.cpp
if (drawCallCount > WEBGL_MAX_COUNT_PER_DRAWCALL) {
  LOG_ERROR("Exceeded maximum draw calls per frame");
}

Batch and Instanced Rendering Strategies

To stay within the strict draw call budget, JSAR employs two complementary strategies: dynamic batching for heterogeneous UI elements and instanced rendering for repeated geometry.

Dynamic Batching of UI Elements

Each frame, the runtime collects HTML elements and merges them into a small set of vertex buffers. Located in src/client/builtin_scene/client_renderer.cpp, this system converts DOM nodes into textured quads that share identical GPU state, allowing hundreds of elements to render in a single draw call.

The batching process packs position, UV coordinates, and color data into shared Float32Array buffers on the TypeScript side before submission to the GPU.

Instanced Rendering for Repeated Geometry

For geometry that repeats—such as Gaussian splats, point clouds, or UI glyphs—JSAR uses glDrawArraysInstanced to render thousands of instances with a single draw call. The GaussianSplatting material in src/client/builtin_scene/materials/gaussian_splatting.cpp packs instance attributes (position, scale, color) into vertex buffers, achieving massive geometry throughput without CPU overhead.

// src/client/builtin_scene/materials/gaussian_splatting.cpp
glBindVertexArray(vao_);
glDrawArraysInstanced(GL_TRIANGLE_STRIP, 0, 4, instanceCount); // ≤ 1 draw call for many splats

Texture and State Optimization

Beyond geometry batching, JSAR minimizes GPU state changes through texture atlasing and render-state caching.

Texture Atlasing to Minimize Bind Changes

The runtime combines multiple UI images and font glyphs into a single GPU texture atlas. Implemented in src/client/builtin_scene/texture_altas.cpp, this technique eliminates texture binding switches during a frame, reducing driver overhead and keeping the render thread unblocked.

Render-State Caching

The RenderState class defined in src/xr/render_state.hpp tracks active WebGL layers, depth ranges, and field-of-view parameters. State changes are only emitted when a real difference is detected, preventing redundant GPU commands and ensuring that JSAR high performance rendering characteristics remain consistent across different XR scenarios.

Cross-Platform GPU Acceleration

JSAR abstracts graphics API differences while leveraging platform-specific optimizations for 2D and 3D content.

Skia Integration for 2-D Rasterization

For high-fidelity 2D canvas operations, JSAR integrates Skia (located in thirdparty/headers/skia), which provides GPU-accelerated text rendering, SVG rasterization, and path stroking. According to the implementation in thirdparty/headers/skia/src/core/SkCanvas.cpp, Skia draws directly into the same GPU surface used for 3D content, eliminating extra copy passes and maintaining the strict performance budget.

Cross-Platform Driver Abstraction

The build system selects the appropriate backend—OpenGL ES on Android, OpenGL on macOS, or other targets—at compile-time via CMakeLists.txt. This abstraction ensures that the same batching and instancing logic works across platforms without requiring per-platform code paths in the renderer.

Implementation Examples

The following snippets demonstrate how JSAR enforces performance constraints and prepares geometry for the GPU.

Enforcing Draw Call Limits in C++

The WebGL context actively monitors the draw call budget to prevent frame-time spikes.

// src/client/graphics/webgl_context.cpp
if (drawCallCount > WEBGL_MAX_COUNT_PER_DRAWCALL) {
  LOG_ERROR("Exceeded maximum draw calls per frame");
}

Preparing Batched Vertex Buffers in TypeScript

On the JavaScript side, the runtime packs element data into shared buffers before GPU submission.

// lib/runtime2/viewers/model3d.ts (simplified)
const vertices: Float32Array = new Float32Array(batchSize * VERTEX_SIZE);
for (let i = 0; i < elements.length; ++i) {
  const e = elements[i];
  // pack position, uv, colour into the shared buffer
  vertices.set(e.toVertexData(), i * VERTEX_SIZE);
}
gl.bufferData(gl.ARRAY_BUFFER, vertices, gl.DYNAMIC_DRAW);
gl.drawArraysInstanced(gl.TRIANGLES, 0, VERTEX_COUNT, elements.length);

Summary

Frequently Asked Questions

What is the maximum number of draw calls JSAR allows per frame?

JSAR enforces a hard limit defined by WEBGL_MAX_COUNT_PER_DRAWCALL in src/client/graphics/webgl_context.cpp. When this threshold is exceeded, the runtime logs an error warning, typically keeping the count at 10 or fewer draw calls per frame to ensure consistent 60fps performance.

How does JSAR handle thousands of HTML elements in 3D space?

The runtime uses dynamic batching implemented in src/client/builtin_scene/client_renderer.cpp to collect DOM nodes each frame and merge them into shared vertex buffers. This converts hundreds of individual HTML elements into textured quads that render in a single GPU draw call, minimizing CPU overhead.

What graphics APIs does JSAR support?

JSAR abstracts graphics APIs through a thin WebGL wrapper that maps to OpenGL ES on Android, OpenGL on macOS, and other platform-specific backends. The build system selects the appropriate driver at compile-time via CMakeLists.txt, ensuring the same batching and instancing logic works across all supported platforms.

How does JSAR optimize 2D rendering alongside 3D content?

JSAR integrates the Skia graphics library (located in thirdparty/headers/skia) for GPU-accelerated 2D rasterization. Skia renders text, SVG, and canvas content directly into the same GPU surface used for 3D geometry, eliminating expensive copy passes and maintaining the engine's strict draw call budget.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →