# How the ANE Dynamic Pipeline Avoids Recompilation When Weights Change

> Learn how the ANE dynamic pipeline bypasses recompilation when model weights change. Discover its innovative approach to runtime data handling for improved efficiency.

- Repository: [Manjeet Singh/ANE](https://github.com/maderix/ANE)
- Tags: internals
- Published: 2026-07-31

---

**The ANE dynamic pipeline treats model weights as runtime data rather than compile-time constants by packing activations and weights into a shared IOSurface and using static MIL kernels that read specific slices via fixed offset parameters.**

The Apple Neural Engine (ANE) typically requires kernel recompilation whenever model weights change, quickly exhausting the per-process limit of approximately 119 compilations. The dynamic pipeline implementation in the [maderix/ANE](https://github.com/maderix/ANE) repository eliminates this bottleneck by decoupling weight data from the compiled program structure, enabling unlimited training iterations without triggering the compiler.

## The Recompilation Bottleneck in Static ANE Pipelines

Traditional ANE implementations embed weight values directly into the MIL (Machine Intermediate Language) program text. When gradients update the weights during training, the MIL source changes, forcing a call to `compile_kern_mil_w` and consuming one of the limited compilation slots. After roughly 119 updates, the process crashes or stalls, making continuous training impossible.

## The IOSurface Strategy: Co-Locating Activations and Weights

The dynamic pipeline circumvents recompilation by treating **weights as ordinary runtime data** that occupies the same IOSurface buffer as activations.

### Spatial Layout and Fixed Offsets

In [`training/training_dynamic/io.h`](https://github.com/maderix/ANE/blob/main/training/training_dynamic/io.h), the helper `io_write_dyn` writes activation tensors followed by weight tensors into a single surface using a combined spatial dimension calculated as `seq + oc`[[source]](https://github.com/maderix/ANE/blob/main/training/training_dynamic/io.h#L78-L88):

```objc
void io_write_dyn(IOSurfaceRef io, float *act, int ic, int seq, 
                  float *weights, int oc) {
    // Activations occupy the start of the spatial dimension
    io_write_fp16(io, act, ic * seq, ACT_SP_OFF);
    // Weights appended at fixed offset determined by sequence length
    io_write_fp16_at(io, weights, ic * oc, ACT_SP_OFF + ic * seq);
}

```

For transformer attention layers, specialized staging functions like `stage_sdpa_fwd_weights` and `stage_wo_fwd_weights` copy weight matrices into predetermined surface locations calculated from model constants (`SEQ`, `DIM`, `Q_DIM`) without modifying any MIL code[[source]](https://github.com/maderix/ANE/blob/main/training/training_dynamic/io.h#L64-L71).

## Static MIL Kernels with Dynamic Slicing

Rather than baking weight values into the program, the pipeline generates **static MIL text once** that references data via fixed offsets.

### Offset-Aware MIL Generation

In [`training/training_dynamic/mil_dynamic.h`](https://github.com/maderix/ANE/blob/main/training/training_dynamic/mil_dynamic.h), the `gen_dyn_matmul_mil` function creates kernels that receive `act_sp_off` (activation offset) and `w_sp_off` (weight offset) parameters[[source]](https://github.com/maderix/ANE/blob/main/training/training_dynamic/mil_dynamic.h#L37-L45). The helper `gen_dyn_matmul` constructs slice operations using these offsets:

```objc
NSString *gen_dyn_matmul(int ic, int oc, int seq, 
                         int act_sp_off, int w_sp_off) {
    // Slice operations index the IOSurface at runtime
    NSString *slice_w = [NSString stringWithFormat:
        @"slice_by_size:%d,%d,%d", w_sp_off, ic, oc];
    // No weight values are hard-coded in the returned MIL text
    return [NSString stringWithFormat:@"...", slice_w, ...];
}

```

The compiled kernel uses `slice_by_size` operations to extract weight slices at runtime. Because the MIL program contains only **offset values** and **dimension constants** (never floating-point weight data), the identical program text works for every training iteration.

## The Training Workflow: Compile Once, Run Many

The decoupled architecture follows a strict separation between compilation and data staging:

1.  **Compilation Phase (One-time)**: Call `compile_kern_mil_w` with the MIL text generated by `gen_dyn_matmul_mil`. This creates a reusable `Kern` object with fixed slice logic.

2.  **Data Staging (Per-iteration)**: Before each forward pass, update the IOSurface with current weights using `io_write_dyn` for linear layers or `stage_sdpa_fwd_weights` for attention mechanisms.

3.  **Evaluation (Per-iteration)**: Invoke `ane_eval` or `ane_eval_req` on the pre-compiled kernel. The ANE reads the updated weights directly from the IOSurface slices.

This workflow applies identically to both generic matrix multiplications and complex attention mechanisms, with staging functions adjusting for layer-specific dimensions while the underlying MIL kernel remains static.

## Complete Implementation Example

The following pattern demonstrates the full workflow for a dynamic training step:

```objc
/* 1️⃣ Build MIL program once */
NSString *mil = gen_dyn_matmul_mil(/*ic=*/128, /*oc=*/256, /*seq=*/1024);
Kern *k = compile_kern_mil_w(mil, nil, ic_bytes, oc_bytes);

/* 2️⃣ Per-step: Pack current activations and weights */
float *activations = ...;  // [128, 1024]
float *weights = ...;      // [128, 256] - updated via backprop
io_write_dyn(k->ioIn, activations, 128, 1024, weights, 256);

/* 3️⃣ Run the pre-compiled kernel */
ane_eval(k);

/* 4️⃣ Retrieve output [oc, seq] */
float output[256*1024];
io_read_dyn(k->ioOut, output, 256, 1024);

```

## Summary

-   **Shared IOSurface buffer**: Weights and activations occupy the same spatial buffer using a `seq + oc` layout, updated by host code before each evaluation.
-   **Offset-based MIL kernels**: The `gen_dyn_matmul` generator and attention equivalents create programs that reference weight data via fixed offsets (`w_sp_off`) rather than embedding values.
-   **Single compilation**: `compile_kern_mil_w` executes once per model architecture; subsequent training steps only rewrite the IOSurface payload.
-   **Unlimited updates**: By avoiding the ANE compiler limit (~119 compilations per process), the dynamic pipeline supports infinite weight updates during training.

## Frequently Asked Questions

### What is the ANE compiler limit and why does it matter?

The Apple Neural Engine enforces a hard limit of approximately 119 kernel compilations per process. In traditional training pipelines where each weight change triggers recompilation, this limit caps training at 119 steps. The dynamic pipeline circumvents this by compiling kernels only once, regardless of how many times weights are updated.

### How does the IOSurface layout prevent weights from being compiled into the kernel?

The system reserves a combined spatial dimension where activations occupy the start of the buffer and weights follow at a predetermined offset. MIL kernels use `slice_by_size` operations with compile-time constants describing the layout (offsets and dimensions) but never contain the actual floating-point weight values. When weights change, only the IOSurface memory is rewritten; the MIL program remains bit-for-bit identical.

### Can this approach handle complex architectures like transformers?

Yes. The repository includes specialized staging functions for attention mechanisms, including `stage_sdpa_fwd_weights` for scaled dot-product attention and `stage_wo_fwd_weights` for output projections. These functions calculate surface offsets based on model hyperparameters (`SEQ`, `DIM`, `Q_DIM`) and use the same IOSurface strategy to avoid recompilation across all transformer layers.

### Which source files contain the core dynamic pipeline implementation?

The implementation spans three primary headers in `training/training_dynamic/`:
-   [`mil_dynamic.h`](https://github.com/maderix/ANE/blob/main/mil_dynamic.h) generates MIL text with offset-aware slicing (`gen_dyn_matmul`, `gen_sdpa_fwd_dynamic`).
-   [`io.h`](https://github.com/maderix/ANE/blob/main/io.h) provides IOSurface helpers for packing data (`io_write_dyn`, `stage_sdpa_fwd_weights`).
-   [`ane_runtime.h`](https://github.com/maderix/ANE/blob/main/ane_runtime.h) wraps the compilation and evaluation APIs (`compile_kern_mil_w`, `ane_eval`).

Model-specific constants reside in `training/training_dynamic/models/*.h`.