Performance Optimization Strategies in ArmorPaint: GPU-Accelerated Painting and Neural Inference
ArmorPaint achieves real-time 3D painting speeds by executing heavy workloads on the GPU through Metal and WebGPU backends, fusing neural network operations, and managing memory through custom pool allocators rather than standard heap allocation.
The open-source 3D painting software ArmorPaint (armory3d/armorpaint) delivers interactive brush strokes and AI-assisted tools by aggressively optimizing both GPU and CPU code paths. This article examines the specific performance optimization strategies in ArmorPaint as implemented in its Iris toolset and UI subsystems, covering everything from kernel fusion to memory layout.
GPU-First Rendering Architecture
Native Metal and WebGPU Backends
ArmorPaint offloads all heavy image-processing steps—convolution, blurring, upscaling, and neural-network inference—directly to the GPU. On macOS and iOS, the engine leverages Apple’s Metal framework, while WebAssembly builds target the WebGPU standard. This eliminates CPU-bound bottlenecks and unnecessary memory copies between host and device.
In base/tools/iris/iris_qwen3.c, the implementation keeps the entire transformer forward pass GPU-resident. The code avoids shuttling hidden states back to system RAM, ensuring that tensor operations remain on the accelerator until the final pixel is written.
Fused Transformer Kernels
Rather than launching separate shaders for each mathematical operation, ArmorPaint fuses layers to reduce kernel launch overhead. The file base/tools/iris/iris_transformer_flux.c contains a "GPU-optimized single block forward" implementation that merges the SiLU activation (sigmoid multiplied by input) with preceding linear transformations. This fusion cuts down on both compute latency and memory traffic.
/* SiLU(gate) * up – fused for better performance */
void fused_silu_gate(float *gate, float *up, int n) {
for (int i = 0; i < n; i++) {
gate[i] = gate[i] * (1.0f / (1.0f + expf(-gate[i]))) * up[i];
}
}
(source: base/tools/iris/iris_transformer_flux.c)
Low-Level Kernel Optimization
Tiled im2col and GEMM
For CPU-side matrix operations that support the GPU pipeline, ArmorPaint organizes image patches using an im2col transformation followed by General Matrix Multiply (GEMM). The file base/tools/iris/iris_kernels.c implements this with explicit tiling to maximize cache locality. By rearranging data so that inner loops traverse contiguous memory, the code dramatically improves CPU cache hit rates during convolution operations.
/* im2col + GEMM optimization with tiling */
void im2col_tiled(const float *img, float *col, int h, int w, int kh, int kw) {
int TILE = 32;
for (int i = 0; i < h; i += TILE) {
for (int j = 0; j < w; j += TILE) {
// Process tile to improve cache locality
// ...
}
}
}
(source: base/tools/iris/iris_kernels.c)
Flash-Attention Kernels
The same kernel file also provides Flash-Attention implementations that tile the attention matrix to avoid materializing full quadratic attention maps, reducing memory complexity from O(n²) to O(n) during AI-assisted painting features.
Memory Management Strategies
Pool-Based Allocation
To prevent heap fragmentation during bursty operations like rapid brush strokes or temporary JavaScript object creation, ArmorPaint uses a custom memory pool. The allocator in base/tools/amake/quickjs-amalgam.c reserves a large block upfront and serves temporary allocations from this arena, eliminating repetitive malloc/free syscalls.
/* Allocate a temporary buffer from the fast pool – avoids malloc overhead */
void *tmp = pool_alloc(&ctx->temp_pool, TMP_BUFFER_SIZE);
process_brush(tmp, brush_data);
pool_free(&ctx->temp_pool, tmp);
(source: base/tools/amake/quickjs-amalgam.c)
Compact Data Formats
Neural network weights and texture data are stored in bandwidth-efficient formats. As noted in iris_transformer_flux.c, the engine uses 16-bit floating-point (f16) weights for AI models and 8-bit packed masks for brush data. These compact representations reduce GPU memory pressure and improve cache line utilization.
CPU Concurrency and Profiling
Multithreaded Task Dispatch
While the GPU handles pixel shaders, CPU-side tasks such as asset decoding and UI updates run on a background thread pool. The dispatcher defined in paint/sources/ui/queue.c pushes independent painting operations—like brush stroke tessellation and mask updates—onto worker threads, keeping the main UI loop responsive at 60 FPS.
Debug Performance Panel
ArmorPaint includes a lightweight profiling overlay that developers can toggle without recompiling. In paint/sources/ui/tab_debug.c, the tab_debug_performance flag enables real-time display of GPU time, CPU time, and memory usage. When disabled, the build remains lean by compiling out instrumentation code.
/* Enable the performance debug panel */
static bool tab_debug_performance = false;
if (ui_panel(&tab_debug_performance, "Performance", false, false, false)) {
ui_label("GPU time: %.2f ms", performance.gpu_time);
ui_label("CPU time: %.2f ms", performance.cpu_time);
}
(source: paint/sources/ui/tab_debug.c)
Summary
- GPU-resident inference keeps AI models on the accelerator via
iris_qwen3.c, avoiding expensive host-device transfers. - Fused kernels in
iris_transformer_flux.ccombine activations and linear layers to reduce kernel launch overhead. - Tiled im2col implementations in
iris_kernels.coptimize CPU cache usage during convolution. - Memory pools in
quickjs-amalgam.celiminate heap fragmentation for temporary buffers. - Compact formats (f16 weights, 8-bit masks) reduce bandwidth and memory footprint.
- Multithreaded dispatch in
queue.cisolates heavy work from the UI thread. - Lazy evaluation and the debug panel in
tab_debug.callow on-demand profiling without release overhead.
Frequently Asked Questions
How does ArmorPaint maintain real-time performance during AI model inference?
ArmorPaint routes all transformer operations through GPU-resident pipelines defined in base/tools/iris/iris_qwen3.c. By keeping hidden states in GPU memory and using fused kernels from iris_transformer_flux.c, the engine avoids the latency penalties of PCIe transfers and CPU synchronization.
What mechanism prevents memory fragmentation during intensive painting sessions?
The engine uses a custom pool allocator implemented in base/tools/amake/quickjs-amalgam.c. Instead of calling malloc for every temporary buffer, the code requests memory from a pre-allocated arena and returns it to the pool after use, maintaining heap stability during bursty brush operations.
How can developers profile specific bottlenecks in the rendering pipeline?
Developers can enable the tab_debug_performance flag in paint/sources/ui/tab_debug.c to overlay GPU and CPU timing metrics directly in the viewport. This instrumentation is compiled out when the flag is disabled, ensuring zero overhead in production builds.
Why does ArmorPaint implement custom kernels instead of using standard BLAS libraries?
While ArmorPaint can fall back to platform BLAS for CPU operations, custom kernels in base/tools/iris/iris_kernels.c allow for aggressive tiling and fusion specific to painting workloads. These optimizations—such as Flash-Attention and fused SiLU gates—are not available in generic libraries and provide significant latency reductions for 3D texture resolution sizes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →