# ArmorPaint's iris.c ML Compute Library: Interface with Vulkan and Metal Backends

> Explore ArmorPaint's iris.c ML compute library, a platform-agnostic engine interfacing with Vulkan and Metal backends for efficient GPU computation of FLUX-2 diffusion models.

- Repository: [Armory 3D/armorpaint](https://github.com/armory3d/armorpaint)
- Tags: internals
- Published: 2026-09-14

---

**[`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) serves as the core machine-learning inference engine in ArmorPaint, orchestrating FLUX-2 diffusion models through a platform-agnostic pipeline that delegates heavy GPU computations to specialized Vulkan or Metal backend modules via conditional compilation.**

The [`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) library lies at the heart of the **armory3d/armorpaint** repository, enabling on-device text-to-image generation without external dependencies. It implements the complete ML pipeline—from tokenizing prompts with the Qwen-3 encoder to sampling latent representations—while abstracting GPU acceleration through a unified interface that supports both desktop Vulkan GPUs and Apple Silicon via Metal.

## What is iris.c in ArmorPaint?

[`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) functions as a self-contained inference runtime located at [`base/tools/iris/iris.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris.c). It manages the entire lifecycle of diffusion model execution, including architecture discovery, weight loading, text encoding, and latent decoding, all while maintaining strict memory budgets for consumer-grade hardware.

### Core Inference Pipeline

The library implements a six-stage pipeline for generating images from text prompts. First, `iris_load_dir` (lines 88-105) reads the model directory structure, parsing [`model_index.json`](https://github.com/armory3d/armorpaint/blob/main/model_index.json) and configuration files to establish architecture parameters. It immediately loads the VAE weights but defers the heavy transformer and text encoder until needed. Next, `iris_encode_text` (lines 558-588) handles tokenization and runs the Qwen-3 text encoder to produce fixed-size text embeddings as `float *` arrays.

For image generation, the system uses `iris_vae_encode` and `iris_vae_decode` (lines 46-48) to convert between RGBA pixels and latent tensors. The `iris_load_transformer_if_needed` function (lines 220-240) then loads Flux transformer weights—either via standard `safetensors` or memory-mapped `mmap`—and executes the diffusion sampling process using Euler, CFG, or inpainting samplers.

### Memory Budget Management

To accommodate limited GPU memory, [`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) implements aggressive resource accounting through `attention_bytes` and `fit_refs_for_attention` (lines 610-660). For Metal backends, it caps attention matrix sizes to the 4 GB MPS limit inherent to consumer Apple Silicon devices. When using Vulkan, it leverages flash-style attention algorithms that never materialize the full attention matrix in device memory. After each generation step, `iris_release_text_encoder` and `iris_release_transformer` (lines 158-176) explicitly free GPU caches and weight buffers to minimize peak memory consumption.

## How iris.c Interfaces with the Vulkan Backend

When compiled with the `USE_VULKAN` macro, [`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) includes [`iris_vulkan.h`](https://github.com/armory3d/armorpaint/blob/main/iris_vulkan.h) and delegates all GPU operations to [`base/tools/iris/iris_vulkan.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_vulkan.c).

### Compute Shader Architecture

The Vulkan backend executes ML operations through GLSL-style compute shaders stored as `.comp` files (such as `iris_vulkan_res_attn.comp` and `iris_vulkan_gemm.comp`). At runtime, [`iris_vulkan.c`](https://github.com/armory3d/armorpaint/blob/main/iris_vulkan.c) compiles these shaders into SPIR-V modules (lines 150-190), creates Vulkan descriptor sets, and dispatches them through command buffers. This approach allows the transformer layers— including RMS normalization, Swish activations, and multi-head attention—to execute entirely on the GPU.

### Weight Loading and Execution Flow

The function `iris_vulkan_load_weights` uploads model parameters into GPU buffers, while `iris_vulkan_release_weight_cache` clears these allocations when `iris_release_transformer` is called. During inference, the core library calls `iris_vulkan_execute`, passing the current `iris_ctx` pointer. The Vulkan module then sequences the shader pipeline to perform matrix multiplications, attention calculations, and residual connections required by the Flux architecture.

## How iris.c Interfaces with the Metal Backend

For macOS and Apple Silicon devices, defining `USE_METAL` includes [`iris_metal.h`](https://github.com/armory3d/armorpaint/blob/main/iris_metal.h) and compiles [`base/tools/iris/iris_metal.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_metal.c) as the compute backend.

### Metal Kernel Compilation

Unlike Vulkan's SPIR-V pipeline, the Metal backend utilizes `.metal` source files containing compute kernels generated from the same algorithmic descriptions as the Vulkan shaders. These kernels compile into a `MTLLibrary` at application startup. The function `iris_metal_execute` (lines 210-250) records commands using `MTLComputeCommandEncoder`, mapping each transformer layer to corresponding Metal compute functions.

### Unified Memory Management

Because Metal on macOS uses unified memory architecture shared between CPU and GPU, [`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) must explicitly reset state to prevent memory growth across generation sessions. The `iris_metal_reset` function—invoked from `iris_release_transformer` and `iris_release_text_encoder`—clears command pools and weight caches. This reset mechanism ensures that the 4 GB memory budget remains available for subsequent inference tasks on memory-constrained devices like the M1 or M2 chips.

## Cross-Platform Design Patterns

The abstraction layer in [`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) (lines 24-27) uses conditional compilation to select backends without modifying high-level inference logic:

```c
#ifdef USE_VULKAN
    #include "iris_vulkan.h"
#else
    #include "iris_metal.h"
#endif

```

This design allows ArmorPaint to ship a single codebase that runs efficiently on both desktop GPUs (via Vulkan) and Apple Silicon (via Metal). The core library handles model topology and data flow, while backend-specific modules manage the low-level GPU queue submission, shader/kernel dispatch, and memory alignment requirements unique to each graphics API.

## Summary

- **[`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c)** at [`base/tools/iris/iris.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris.c) implements the complete FLUX-2 inference pipeline, including text encoding, VAE processing, and transformer sampling.
- **Backend abstraction** occurs through conditional compilation macros (`USE_VULKAN` or `USE_METAL`), allowing the same high-level code to target different GPU APIs.
- **Vulkan integration** utilizes SPIR-V compute shaders compiled from `.comp` files, with weight management handled in [`iris_vulkan.c`](https://github.com/armory3d/armorpaint/blob/main/iris_vulkan.c).
- **Metal integration** employs `.metal` compute kernels and explicit state reset mechanisms to manage unified memory on Apple Silicon.
- **Memory safety** is enforced through explicit resource cleanup functions (`iris_release_transformer`, `iris_release_text_encoder`) and attention matrix size limits tailored to each backend's constraints.

## Frequently Asked Questions

### How does iris.c handle different GPU memory limits on Vulkan versus Metal?

[`iris.c`](https://github.com/armory3d/armorpaint/blob/main/iris.c) calculates potential attention matrix sizes using `attention_bytes` and `fit_refs_for_attention` (lines 610-660). For Metal, it enforces a hard 4 GB limit to respect MPS constraints on consumer Apple Silicon devices. For Vulkan, it implements flash-style attention algorithms that avoid materializing complete attention matrices, allowing the system to run on GPUs with varying memory capacities without configuration changes.

### Can iris.c switch between Vulkan and Metal at runtime?

No, backend selection occurs at compile time through the `USE_VULKAN` or `USE_METAL` preprocessor directives defined in [`base/tools/iris/iris.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris.c) (lines 24-27). The resulting binary contains only one backend implementation, ensuring minimal binary size and avoiding runtime overhead from dynamic dispatch mechanisms.

### What file formats does iris.c use for model weights?

The library supports the **safetensors** format for standard file loading and **memory-mapped (mmap)** access for zero-copy weight initialization, as implemented in `iris_load_transformer_if_needed` (lines 220-240). Model architecture configurations are parsed from standard JSON files including [`model_index.json`](https://github.com/armory3d/armorpaint/blob/main/model_index.json), [`transformer/config.json`](https://github.com/armory3d/armorpaint/blob/main/transformer/config.json), and [`vae/config.json`](https://github.com/armory3d/armorpaint/blob/main/vae/config.json).

### Where does the actual image decoding happen in the iris.c pipeline?

While transformer inference runs on the GPU via backend-specific shaders, VAE decoding (`iris_vae_decode`) executes on the CPU within [`base/tools/iris/iris_vae.c`](https://github.com/armory3d/armorpaint/blob/main/base/tools/iris/iris_vae.c). This design choice keeps the GPU memory footprint low during the final RGBA reconstruction phase, as the VAE decoder requires significantly less computational overhead compared to the diffusion transformer.