How the ANE Dynamic Pipeline Avoids Recompilation When Weights Change
The ANE dynamic pipeline treats model weights as runtime data rather than compile-time constants by packing activations and weights into a shared IOSurface and using static MIL kernels that read specific slices via fixed offset parameters.
The Apple Neural Engine (ANE) typically requires kernel recompilation whenever model weights change, quickly exhausting the per-process limit of approximately 119 compilations. The dynamic pipeline implementation in the maderix/ANE repository eliminates this bottleneck by decoupling weight data from the compiled program structure, enabling unlimited training iterations without triggering the compiler.
The Recompilation Bottleneck in Static ANE Pipelines
Traditional ANE implementations embed weight values directly into the MIL (Machine Intermediate Language) program text. When gradients update the weights during training, the MIL source changes, forcing a call to compile_kern_mil_w and consuming one of the limited compilation slots. After roughly 119 updates, the process crashes or stalls, making continuous training impossible.
The IOSurface Strategy: Co-Locating Activations and Weights
The dynamic pipeline circumvents recompilation by treating weights as ordinary runtime data that occupies the same IOSurface buffer as activations.
Spatial Layout and Fixed Offsets
In training/training_dynamic/io.h, the helper io_write_dyn writes activation tensors followed by weight tensors into a single surface using a combined spatial dimension calculated as seq + oc[source]:
void io_write_dyn(IOSurfaceRef io, float *act, int ic, int seq,
float *weights, int oc) {
// Activations occupy the start of the spatial dimension
io_write_fp16(io, act, ic * seq, ACT_SP_OFF);
// Weights appended at fixed offset determined by sequence length
io_write_fp16_at(io, weights, ic * oc, ACT_SP_OFF + ic * seq);
}
For transformer attention layers, specialized staging functions like stage_sdpa_fwd_weights and stage_wo_fwd_weights copy weight matrices into predetermined surface locations calculated from model constants (SEQ, DIM, Q_DIM) without modifying any MIL code[source].
Static MIL Kernels with Dynamic Slicing
Rather than baking weight values into the program, the pipeline generates static MIL text once that references data via fixed offsets.
Offset-Aware MIL Generation
In training/training_dynamic/mil_dynamic.h, the gen_dyn_matmul_mil function creates kernels that receive act_sp_off (activation offset) and w_sp_off (weight offset) parameters[source]. The helper gen_dyn_matmul constructs slice operations using these offsets:
NSString *gen_dyn_matmul(int ic, int oc, int seq,
int act_sp_off, int w_sp_off) {
// Slice operations index the IOSurface at runtime
NSString *slice_w = [NSString stringWithFormat:
@"slice_by_size:%d,%d,%d", w_sp_off, ic, oc];
// No weight values are hard-coded in the returned MIL text
return [NSString stringWithFormat:@"...", slice_w, ...];
}
The compiled kernel uses slice_by_size operations to extract weight slices at runtime. Because the MIL program contains only offset values and dimension constants (never floating-point weight data), the identical program text works for every training iteration.
The Training Workflow: Compile Once, Run Many
The decoupled architecture follows a strict separation between compilation and data staging:
-
Compilation Phase (One-time): Call
compile_kern_mil_wwith the MIL text generated bygen_dyn_matmul_mil. This creates a reusableKernobject with fixed slice logic. -
Data Staging (Per-iteration): Before each forward pass, update the IOSurface with current weights using
io_write_dynfor linear layers orstage_sdpa_fwd_weightsfor attention mechanisms. -
Evaluation (Per-iteration): Invoke
ane_evalorane_eval_reqon the pre-compiled kernel. The ANE reads the updated weights directly from the IOSurface slices.
This workflow applies identically to both generic matrix multiplications and complex attention mechanisms, with staging functions adjusting for layer-specific dimensions while the underlying MIL kernel remains static.
Complete Implementation Example
The following pattern demonstrates the full workflow for a dynamic training step:
/* 1️⃣ Build MIL program once */
NSString *mil = gen_dyn_matmul_mil(/*ic=*/128, /*oc=*/256, /*seq=*/1024);
Kern *k = compile_kern_mil_w(mil, nil, ic_bytes, oc_bytes);
/* 2️⃣ Per-step: Pack current activations and weights */
float *activations = ...; // [128, 1024]
float *weights = ...; // [128, 256] - updated via backprop
io_write_dyn(k->ioIn, activations, 128, 1024, weights, 256);
/* 3️⃣ Run the pre-compiled kernel */
ane_eval(k);
/* 4️⃣ Retrieve output [oc, seq] */
float output[256*1024];
io_read_dyn(k->ioOut, output, 256, 1024);
Summary
- Shared IOSurface buffer: Weights and activations occupy the same spatial buffer using a
seq + oclayout, updated by host code before each evaluation. - Offset-based MIL kernels: The
gen_dyn_matmulgenerator and attention equivalents create programs that reference weight data via fixed offsets (w_sp_off) rather than embedding values. - Single compilation:
compile_kern_mil_wexecutes once per model architecture; subsequent training steps only rewrite the IOSurface payload. - Unlimited updates: By avoiding the ANE compiler limit (~119 compilations per process), the dynamic pipeline supports infinite weight updates during training.
Frequently Asked Questions
What is the ANE compiler limit and why does it matter?
The Apple Neural Engine enforces a hard limit of approximately 119 kernel compilations per process. In traditional training pipelines where each weight change triggers recompilation, this limit caps training at 119 steps. The dynamic pipeline circumvents this by compiling kernels only once, regardless of how many times weights are updated.
How does the IOSurface layout prevent weights from being compiled into the kernel?
The system reserves a combined spatial dimension where activations occupy the start of the buffer and weights follow at a predetermined offset. MIL kernels use slice_by_size operations with compile-time constants describing the layout (offsets and dimensions) but never contain the actual floating-point weight values. When weights change, only the IOSurface memory is rewritten; the MIL program remains bit-for-bit identical.
Can this approach handle complex architectures like transformers?
Yes. The repository includes specialized staging functions for attention mechanisms, including stage_sdpa_fwd_weights for scaled dot-product attention and stage_wo_fwd_weights for output projections. These functions calculate surface offsets based on model hyperparameters (SEQ, DIM, Q_DIM) and use the same IOSurface strategy to avoid recompilation across all transformer layers.
Which source files contain the core dynamic pipeline implementation?
The implementation spans three primary headers in training/training_dynamic/:
mil_dynamic.hgenerates MIL text with offset-aware slicing (gen_dyn_matmul,gen_sdpa_fwd_dynamic).io.hprovides IOSurface helpers for packing data (io_write_dyn,stage_sdpa_fwd_weights).ane_runtime.hwraps the compilation and evaluation APIs (compile_kern_mil_w,ane_eval).
Model-specific constants reside in training/training_dynamic/models/*.h.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →