# Performance Optimization Strategies in ArmorPaint: GPU-Accelerated Painting and Neural Inference

> Discover ArmorPaint's performance optimization strategies. Learn how GPU acceleration, neural inference fusion, and custom memory allocators deliver real-time 3D painting.

- Repository: [Armory 3D/armorpaint](https://github.com/armory3d/armorpaint)
- Tags: performance
- Published: 2026-09-13

---

**ArmorPaint achieves real-time 3D painting speeds by executing heavy workloads on the GPU through Metal and WebGPU backends, fusing neural network operations, and managing memory through custom pool allocators rather than standard heap allocation.**

The open-source 3D painting software ArmorPaint (armory3d/armorpaint) delivers interactive brush strokes and AI-assisted tools by aggressively optimizing both GPU and CPU code paths. This article examines the specific **performance optimization strategies in ArmorPaint** as implemented in its Iris toolset and UI subsystems, covering everything from kernel fusion to memory layout.

## GPU-First Rendering Architecture

### Native Metal and WebGPU Backends

ArmorPaint offloads all heavy image-processing steps—convolution, blurring, upscaling, and neural-network inference—directly to the GPU. On macOS and iOS, the engine leverages Apple’s Metal framework, while WebAssembly builds target the WebGPU standard. This eliminates CPU-bound bottlenecks and unnecessary memory copies between host and device.

In [`base/tools/iris/iris_qwen3.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_qwen3.c), the implementation keeps the entire transformer forward pass GPU-resident. The code avoids shuttling hidden states back to system RAM, ensuring that tensor operations remain on the accelerator until the final pixel is written.

### Fused Transformer Kernels

Rather than launching separate shaders for each mathematical operation, ArmorPaint fuses layers to reduce kernel launch overhead. The file [`base/tools/iris/iris_transformer_flux.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_transformer_flux.c) contains a "GPU-optimized single block forward" implementation that merges the SiLU activation (sigmoid multiplied by input) with preceding linear transformations. This fusion cuts down on both compute latency and memory traffic.

```c
/* SiLU(gate) * up – fused for better performance */
void fused_silu_gate(float *gate, float *up, int n) {
    for (int i = 0; i < n; i++) {
        gate[i] = gate[i] * (1.0f / (1.0f + expf(-gate[i]))) * up[i];
    }
}

```

(source: [`base/tools/iris/iris_transformer_flux.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_transformer_flux.c))

## Low-Level Kernel Optimization

### Tiled im2col and GEMM

For CPU-side matrix operations that support the GPU pipeline, ArmorPaint organizes image patches using an `im2col` transformation followed by General Matrix Multiply (GEMM). The file [`base/tools/iris/iris_kernels.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_kernels.c) implements this with explicit tiling to maximize cache locality. By rearranging data so that inner loops traverse contiguous memory, the code dramatically improves CPU cache hit rates during convolution operations.

```c
/* im2col + GEMM optimization with tiling */
void im2col_tiled(const float *img, float *col, int h, int w, int kh, int kw) {
    int TILE = 32;
    for (int i = 0; i < h; i += TILE) {
        for (int j = 0; j < w; j += TILE) {
            // Process tile to improve cache locality
            // ...
        }
    }
}

```

(source: [`base/tools/iris/iris_kernels.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_kernels.c))

### Flash-Attention Kernels

The same kernel file also provides Flash-Attention implementations that tile the attention matrix to avoid materializing full quadratic attention maps, reducing memory complexity from O(n²) to O(n) during AI-assisted painting features.

## Memory Management Strategies

### Pool-Based Allocation

To prevent heap fragmentation during bursty operations like rapid brush strokes or temporary JavaScript object creation, ArmorPaint uses a custom memory pool. The allocator in [`base/tools/amake/quickjs-amalgam.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/amake/quickjs-amalgam.c) reserves a large block upfront and serves temporary allocations from this arena, eliminating repetitive `malloc`/`free` syscalls.

```c
/* Allocate a temporary buffer from the fast pool – avoids malloc overhead */
void *tmp = pool_alloc(&ctx->temp_pool, TMP_BUFFER_SIZE);
process_brush(tmp, brush_data);
pool_free(&ctx->temp_pool, tmp);

```

(source: [`base/tools/amake/quickjs-amalgam.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/amake/quickjs-amalgam.c))

### Compact Data Formats

Neural network weights and texture data are stored in bandwidth-efficient formats. As noted in [`iris_transformer_flux.c`](https://github.com/armory3d/armorpaint/blob/main/iris_transformer_flux.c), the engine uses 16-bit floating-point (f16) weights for AI models and 8-bit packed masks for brush data. These compact representations reduce GPU memory pressure and improve cache line utilization.

## CPU Concurrency and Profiling

### Multithreaded Task Dispatch

While the GPU handles pixel shaders, CPU-side tasks such as asset decoding and UI updates run on a background thread pool. The dispatcher defined in [`paint/sources/ui/queue.c`](https://github.com/armory3d/armorpaint/blob/main/paint/sources/ui/queue.c) pushes independent painting operations—like brush stroke tessellation and mask updates—onto worker threads, keeping the main UI loop responsive at 60 FPS.

### Debug Performance Panel

ArmorPaint includes a lightweight profiling overlay that developers can toggle without recompiling. In [`paint/sources/ui/tab_debug.c`](https://github.com/armory3d/armorpaint/blob/main/paint/sources/ui/tab_debug.c), the `tab_debug_performance` flag enables real-time display of GPU time, CPU time, and memory usage. When disabled, the build remains lean by compiling out instrumentation code.

```c
/* Enable the performance debug panel */
static bool tab_debug_performance = false;
if (ui_panel(&tab_debug_performance, "Performance", false, false, false)) {
    ui_label("GPU time: %.2f ms", performance.gpu_time);
    ui_label("CPU time: %.2f ms", performance.cpu_time);
}

```

(source: [`paint/sources/ui/tab_debug.c`](https://github.com/armory3d/armorpaint/blob/main/paint/sources/ui/tab_debug.c))

## Summary

- **GPU-resident inference** keeps AI models on the accelerator via [`iris_qwen3.c`](https://github.com/armory3d/armorpaint/blob/main/iris_qwen3.c), avoiding expensive host-device transfers.
- **Fused kernels** in [`iris_transformer_flux.c`](https://github.com/armory3d/armorpaint/blob/main/iris_transformer_flux.c) combine activations and linear layers to reduce kernel launch overhead.
- **Tiled im2col** implementations in [`iris_kernels.c`](https://github.com/armory3d/armorpaint/blob/main/iris_kernels.c) optimize CPU cache usage during convolution.
- **Memory pools** in [`quickjs-amalgam.c`](https://github.com/armory3d/armorpaint/blob/main/quickjs-amalgam.c) eliminate heap fragmentation for temporary buffers.
- **Compact formats** (f16 weights, 8-bit masks) reduce bandwidth and memory footprint.
- **Multithreaded dispatch** in [`queue.c`](https://github.com/armory3d/armorpaint/blob/main/queue.c) isolates heavy work from the UI thread.
- **Lazy evaluation** and the debug panel in [`tab_debug.c`](https://github.com/armory3d/armorpaint/blob/main/tab_debug.c) allow on-demand profiling without release overhead.

## Frequently Asked Questions

### How does ArmorPaint maintain real-time performance during AI model inference?

ArmorPaint routes all transformer operations through GPU-resident pipelines defined in [`base/tools/iris/iris_qwen3.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_qwen3.c). By keeping hidden states in GPU memory and using fused kernels from [`iris_transformer_flux.c`](https://github.com/armory3d/armorpaint/blob/main/iris_transformer_flux.c), the engine avoids the latency penalties of PCIe transfers and CPU synchronization.

### What mechanism prevents memory fragmentation during intensive painting sessions?

The engine uses a custom pool allocator implemented in [`base/tools/amake/quickjs-amalgam.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/amake/quickjs-amalgam.c). Instead of calling `malloc` for every temporary buffer, the code requests memory from a pre-allocated arena and returns it to the pool after use, maintaining heap stability during bursty brush operations.

### How can developers profile specific bottlenecks in the rendering pipeline?

Developers can enable the `tab_debug_performance` flag in [`paint/sources/ui/tab_debug.c`](https://github.com/armory3d/armorpaint/blob/main/paint/sources/ui/tab_debug.c) to overlay GPU and CPU timing metrics directly in the viewport. This instrumentation is compiled out when the flag is disabled, ensuring zero overhead in production builds.

### Why does ArmorPaint implement custom kernels instead of using standard BLAS libraries?

While ArmorPaint can fall back to platform BLAS for CPU operations, custom kernels in [`base/tools/iris/iris_kernels.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_kernels.c) allow for aggressive tiling and fusion specific to painting workloads. These optimizations—such as Flash-Attention and fused SiLU gates—are not available in generic libraries and provide significant latency reductions for 3D texture resolution sizes.