ArmorPaint's iris.c ML Compute Library: Interface with Vulkan and Metal Backends
iris.c serves as the core machine-learning inference engine in ArmorPaint, orchestrating FLUX-2 diffusion models through a platform-agnostic pipeline that delegates heavy GPU computations to specialized Vulkan or Metal backend modules via conditional compilation.
The iris.c library lies at the heart of the armory3d/armorpaint repository, enabling on-device text-to-image generation without external dependencies. It implements the complete ML pipeline—from tokenizing prompts with the Qwen-3 encoder to sampling latent representations—while abstracting GPU acceleration through a unified interface that supports both desktop Vulkan GPUs and Apple Silicon via Metal.
What is iris.c in ArmorPaint?
iris.c functions as a self-contained inference runtime located at base/tools/iris/iris.c. It manages the entire lifecycle of diffusion model execution, including architecture discovery, weight loading, text encoding, and latent decoding, all while maintaining strict memory budgets for consumer-grade hardware.
Core Inference Pipeline
The library implements a six-stage pipeline for generating images from text prompts. First, iris_load_dir (lines 88-105) reads the model directory structure, parsing model_index.json and configuration files to establish architecture parameters. It immediately loads the VAE weights but defers the heavy transformer and text encoder until needed. Next, iris_encode_text (lines 558-588) handles tokenization and runs the Qwen-3 text encoder to produce fixed-size text embeddings as float * arrays.
For image generation, the system uses iris_vae_encode and iris_vae_decode (lines 46-48) to convert between RGBA pixels and latent tensors. The iris_load_transformer_if_needed function (lines 220-240) then loads Flux transformer weights—either via standard safetensors or memory-mapped mmap—and executes the diffusion sampling process using Euler, CFG, or inpainting samplers.
Memory Budget Management
To accommodate limited GPU memory, iris.c implements aggressive resource accounting through attention_bytes and fit_refs_for_attention (lines 610-660). For Metal backends, it caps attention matrix sizes to the 4 GB MPS limit inherent to consumer Apple Silicon devices. When using Vulkan, it leverages flash-style attention algorithms that never materialize the full attention matrix in device memory. After each generation step, iris_release_text_encoder and iris_release_transformer (lines 158-176) explicitly free GPU caches and weight buffers to minimize peak memory consumption.
How iris.c Interfaces with the Vulkan Backend
When compiled with the USE_VULKAN macro, iris.c includes iris_vulkan.h and delegates all GPU operations to base/tools/iris/iris_vulkan.c.
Compute Shader Architecture
The Vulkan backend executes ML operations through GLSL-style compute shaders stored as .comp files (such as iris_vulkan_res_attn.comp and iris_vulkan_gemm.comp). At runtime, iris_vulkan.c compiles these shaders into SPIR-V modules (lines 150-190), creates Vulkan descriptor sets, and dispatches them through command buffers. This approach allows the transformer layers— including RMS normalization, Swish activations, and multi-head attention—to execute entirely on the GPU.
Weight Loading and Execution Flow
The function iris_vulkan_load_weights uploads model parameters into GPU buffers, while iris_vulkan_release_weight_cache clears these allocations when iris_release_transformer is called. During inference, the core library calls iris_vulkan_execute, passing the current iris_ctx pointer. The Vulkan module then sequences the shader pipeline to perform matrix multiplications, attention calculations, and residual connections required by the Flux architecture.
How iris.c Interfaces with the Metal Backend
For macOS and Apple Silicon devices, defining USE_METAL includes iris_metal.h and compiles base/tools/iris/iris_metal.c as the compute backend.
Metal Kernel Compilation
Unlike Vulkan's SPIR-V pipeline, the Metal backend utilizes .metal source files containing compute kernels generated from the same algorithmic descriptions as the Vulkan shaders. These kernels compile into a MTLLibrary at application startup. The function iris_metal_execute (lines 210-250) records commands using MTLComputeCommandEncoder, mapping each transformer layer to corresponding Metal compute functions.
Unified Memory Management
Because Metal on macOS uses unified memory architecture shared between CPU and GPU, iris.c must explicitly reset state to prevent memory growth across generation sessions. The iris_metal_reset function—invoked from iris_release_transformer and iris_release_text_encoder—clears command pools and weight caches. This reset mechanism ensures that the 4 GB memory budget remains available for subsequent inference tasks on memory-constrained devices like the M1 or M2 chips.
Cross-Platform Design Patterns
The abstraction layer in iris.c (lines 24-27) uses conditional compilation to select backends without modifying high-level inference logic:
#ifdef USE_VULKAN
#include "iris_vulkan.h"
#else
#include "iris_metal.h"
#endif
This design allows ArmorPaint to ship a single codebase that runs efficiently on both desktop GPUs (via Vulkan) and Apple Silicon (via Metal). The core library handles model topology and data flow, while backend-specific modules manage the low-level GPU queue submission, shader/kernel dispatch, and memory alignment requirements unique to each graphics API.
Summary
iris.catbase/tools/iris/iris.cimplements the complete FLUX-2 inference pipeline, including text encoding, VAE processing, and transformer sampling.- Backend abstraction occurs through conditional compilation macros (
USE_VULKANorUSE_METAL), allowing the same high-level code to target different GPU APIs. - Vulkan integration utilizes SPIR-V compute shaders compiled from
.compfiles, with weight management handled iniris_vulkan.c. - Metal integration employs
.metalcompute kernels and explicit state reset mechanisms to manage unified memory on Apple Silicon. - Memory safety is enforced through explicit resource cleanup functions (
iris_release_transformer,iris_release_text_encoder) and attention matrix size limits tailored to each backend's constraints.
Frequently Asked Questions
How does iris.c handle different GPU memory limits on Vulkan versus Metal?
iris.c calculates potential attention matrix sizes using attention_bytes and fit_refs_for_attention (lines 610-660). For Metal, it enforces a hard 4 GB limit to respect MPS constraints on consumer Apple Silicon devices. For Vulkan, it implements flash-style attention algorithms that avoid materializing complete attention matrices, allowing the system to run on GPUs with varying memory capacities without configuration changes.
Can iris.c switch between Vulkan and Metal at runtime?
No, backend selection occurs at compile time through the USE_VULKAN or USE_METAL preprocessor directives defined in base/tools/iris/iris.c (lines 24-27). The resulting binary contains only one backend implementation, ensuring minimal binary size and avoiding runtime overhead from dynamic dispatch mechanisms.
What file formats does iris.c use for model weights?
The library supports the safetensors format for standard file loading and memory-mapped (mmap) access for zero-copy weight initialization, as implemented in iris_load_transformer_if_needed (lines 220-240). Model architecture configurations are parsed from standard JSON files including model_index.json, transformer/config.json, and vae/config.json.
Where does the actual image decoding happen in the iris.c pipeline?
While transformer inference runs on the GPU via backend-specific shaders, VAE decoding (iris_vae_decode) executes on the CPU within base/tools/iris/iris_vae.c. This design choice keeps the GPU memory footprint low during the final RGBA reconstruction phase, as the VAE decoder requires significantly less computational overhead compared to the diffusion transformer.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →