How the ANE Dynamic Pipeline Avoids Recompilation When Weights Change

The ANE dynamic pipeline treats model weights as runtime data rather than compile-time constants by packing activations and weights into a shared IOSurface and using static MIL kernels that read specific slices via fixed offset parameters.

The Apple Neural Engine (ANE) typically requires kernel recompilation whenever model weights change, quickly exhausting the per-process limit of approximately 119 compilations. The dynamic pipeline implementation in the maderix/ANE repository eliminates this bottleneck by decoupling weight data from the compiled program structure, enabling unlimited training iterations without triggering the compiler.

The Recompilation Bottleneck in Static ANE Pipelines

Traditional ANE implementations embed weight values directly into the MIL (Machine Intermediate Language) program text. When gradients update the weights during training, the MIL source changes, forcing a call to compile_kern_mil_w and consuming one of the limited compilation slots. After roughly 119 updates, the process crashes or stalls, making continuous training impossible.

The IOSurface Strategy: Co-Locating Activations and Weights

The dynamic pipeline circumvents recompilation by treating weights as ordinary runtime data that occupies the same IOSurface buffer as activations.

Spatial Layout and Fixed Offsets

In training/training_dynamic/io.h, the helper io_write_dyn writes activation tensors followed by weight tensors into a single surface using a combined spatial dimension calculated as seq + oc[source]:

void io_write_dyn(IOSurfaceRef io, float *act, int ic, int seq, 
                  float *weights, int oc) {
    // Activations occupy the start of the spatial dimension
    io_write_fp16(io, act, ic * seq, ACT_SP_OFF);
    // Weights appended at fixed offset determined by sequence length
    io_write_fp16_at(io, weights, ic * oc, ACT_SP_OFF + ic * seq);
}

For transformer attention layers, specialized staging functions like stage_sdpa_fwd_weights and stage_wo_fwd_weights copy weight matrices into predetermined surface locations calculated from model constants (SEQ, DIM, Q_DIM) without modifying any MIL code[source].

Static MIL Kernels with Dynamic Slicing

Rather than baking weight values into the program, the pipeline generates static MIL text once that references data via fixed offsets.

Offset-Aware MIL Generation

In training/training_dynamic/mil_dynamic.h, the gen_dyn_matmul_mil function creates kernels that receive act_sp_off (activation offset) and w_sp_off (weight offset) parameters[source]. The helper gen_dyn_matmul constructs slice operations using these offsets:

NSString *gen_dyn_matmul(int ic, int oc, int seq, 
                         int act_sp_off, int w_sp_off) {
    // Slice operations index the IOSurface at runtime
    NSString *slice_w = [NSString stringWithFormat:
        @"slice_by_size:%d,%d,%d", w_sp_off, ic, oc];
    // No weight values are hard-coded in the returned MIL text
    return [NSString stringWithFormat:@"...", slice_w, ...];
}

The compiled kernel uses slice_by_size operations to extract weight slices at runtime. Because the MIL program contains only offset values and dimension constants (never floating-point weight data), the identical program text works for every training iteration.

The Training Workflow: Compile Once, Run Many

The decoupled architecture follows a strict separation between compilation and data staging:

  1. Compilation Phase (One-time): Call compile_kern_mil_w with the MIL text generated by gen_dyn_matmul_mil. This creates a reusable Kern object with fixed slice logic.

  2. Data Staging (Per-iteration): Before each forward pass, update the IOSurface with current weights using io_write_dyn for linear layers or stage_sdpa_fwd_weights for attention mechanisms.

  3. Evaluation (Per-iteration): Invoke ane_eval or ane_eval_req on the pre-compiled kernel. The ANE reads the updated weights directly from the IOSurface slices.

This workflow applies identically to both generic matrix multiplications and complex attention mechanisms, with staging functions adjusting for layer-specific dimensions while the underlying MIL kernel remains static.

Complete Implementation Example

The following pattern demonstrates the full workflow for a dynamic training step:

/* 1️⃣ Build MIL program once */
NSString *mil = gen_dyn_matmul_mil(/*ic=*/128, /*oc=*/256, /*seq=*/1024);
Kern *k = compile_kern_mil_w(mil, nil, ic_bytes, oc_bytes);

/* 2️⃣ Per-step: Pack current activations and weights */
float *activations = ...;  // [128, 1024]
float *weights = ...;      // [128, 256] - updated via backprop
io_write_dyn(k->ioIn, activations, 128, 1024, weights, 256);

/* 3️⃣ Run the pre-compiled kernel */
ane_eval(k);

/* 4️⃣ Retrieve output [oc, seq] */
float output[256*1024];
io_read_dyn(k->ioOut, output, 256, 1024);

Summary

  • Shared IOSurface buffer: Weights and activations occupy the same spatial buffer using a seq + oc layout, updated by host code before each evaluation.
  • Offset-based MIL kernels: The gen_dyn_matmul generator and attention equivalents create programs that reference weight data via fixed offsets (w_sp_off) rather than embedding values.
  • Single compilation: compile_kern_mil_w executes once per model architecture; subsequent training steps only rewrite the IOSurface payload.
  • Unlimited updates: By avoiding the ANE compiler limit (~119 compilations per process), the dynamic pipeline supports infinite weight updates during training.

Frequently Asked Questions

What is the ANE compiler limit and why does it matter?

The Apple Neural Engine enforces a hard limit of approximately 119 kernel compilations per process. In traditional training pipelines where each weight change triggers recompilation, this limit caps training at 119 steps. The dynamic pipeline circumvents this by compiling kernels only once, regardless of how many times weights are updated.

How does the IOSurface layout prevent weights from being compiled into the kernel?

The system reserves a combined spatial dimension where activations occupy the start of the buffer and weights follow at a predetermined offset. MIL kernels use slice_by_size operations with compile-time constants describing the layout (offsets and dimensions) but never contain the actual floating-point weight values. When weights change, only the IOSurface memory is rewritten; the MIL program remains bit-for-bit identical.

Can this approach handle complex architectures like transformers?

Yes. The repository includes specialized staging functions for attention mechanisms, including stage_sdpa_fwd_weights for scaled dot-product attention and stage_wo_fwd_weights for output projections. These functions calculate surface offsets based on model hyperparameters (SEQ, DIM, Q_DIM) and use the same IOSurface strategy to avoid recompilation across all transformer layers.

Which source files contain the core dynamic pipeline implementation?

The implementation spans three primary headers in training/training_dynamic/:

  • mil_dynamic.h generates MIL text with offset-aware slicing (gen_dyn_matmul, gen_sdpa_fwd_dynamic).
  • io.h provides IOSurface helpers for packing data (io_write_dyn, stage_sdpa_fwd_weights).
  • ane_runtime.h wraps the compilation and evaluation APIs (compile_kern_mil_w, ane_eval).

Model-specific constants reside in training/training_dynamic/models/*.h.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →