# How the MIL Compiler Works with `_ANEInMemoryModelDescriptor`: In-Memory ANE Compilation Explained

> Discover how the MIL compiler uses _ANEInMemoryModelDescriptor for in-memory ANE compilation. Convert MIL programs and weights to ANE kernels in RAM, ditching disk-based .mlmodelc bundles.

- Repository: [Manjeet Singh/ANE](https://github.com/maderix/ANE)
- Tags: internals
- Published: 2026-07-31

---

**The MIL compiler leverages the private `_ANEInMemoryModelDescriptor` class to convert Machine-Learning Intermediate Language (MIL) programs and weight blobs into executable Apple Neural Engine (ANE) kernels entirely in RAM, eliminating the need for disk-based `.mlmodelc` bundles.**

The `maderix/ANE` repository demonstrates this in-memory compilation pipeline through a low-level Objective-C wrapper that interfaces with private `AppleNeuralEngine.framework` APIs. By dynamically loading the `_ANEInMemoryModelDescriptor` and `_ANEInMemoryModel` classes, the system transforms textual MIL representations into hardware-accelerated kernels ready for inference.

## Loading Private ANE Symbols at Runtime

The compilation process begins by dynamically linking against the private `AppleNeuralEngine.framework` and obtaining class references for the descriptor and model objects.

In [`training/ane_runtime.h`](https://github.com/maderix/ANE/blob/main/training/ane_runtime.h), the initialization routine uses `dlopen` to load the framework and `NSClassFromString` to resolve the private classes:

```objc
dlopen("/System/Library/PrivateFrameworks/AppleNeuralEngine.framework/AppleNeuralEngine", RTLD_NOW);

Class g_ANEDesc  = NSClassFromString(@"_ANEInMemoryModelDescriptor");  // [ane_runtime.h L27]
Class g_ANEInMem = NSClassFromString(@"_ANEInMemoryModel");            // [ane_runtime.h L28]
Class g_ANEReq   = NSClassFromString(@"_ANERequest");                  // [ane_runtime.h L29]
Class g_ANEIO    = NSClassFromString(@"_ANEIOSurfaceObject");          // [ane_runtime.h L30]

```

These global class pointers are cached during `ane_init` and reused throughout the application lifecycle to minimize the overhead of repeated symbol lookups.

## Creating the Descriptor from MIL Text

The `_ANEInMemoryModelDescriptor` acts as a factory for compiling MIL programs. It exposes the selector `modelWithMILText:weights:optionsPlist:`, which accepts the raw MIL source as `NSData`, an optional weight dictionary, and a configuration plist.

```objc
NSDictionary *wdict = weightData ?
    @{@"@model_path/weights/weight.bin": @{@"offset": @0, @"data": weightData}} : nil;

id desc = ((id(*)(Class,SEL,id,id,id))objc_msgSend)(
    g_ANEDesc,
    @selector(modelWithMILText:weights:optionsPlist:),
    milText,    // NSData containing MIL string
    wdict,      // Weight map or nil
    nil);       // Options plist (nil for defaults)  // [ane_runtime.h L55-L62]

```

If the MIL syntax is invalid or the weights are incompatible, the descriptor returns `NULL` and compilation aborts immediately. The weight dictionary uses a specific key format (`@model_path/weights/weight.bin`) to map binary blobs to tensor constants declared in the MIL program.

## Instantiating the In-Memory Model

Once the descriptor is validated, the system creates an `_ANEInMemoryModel` instance using the `inMemoryModelWithDescriptor:` constructor. This object holds the compiled ANE program and manages the IOSurface handles for input and output tensors.

```objc
id mdl = ((id(*)(Class,SEL,id))objc_msgSend)(
    g_ANEInMem,
    @selector(inMemoryModelWithDescriptor:), desc);  // [ane_runtime.h L64-L66]

```

The model object remains in memory until explicitly unloaded, making it suitable for high-frequency inference scenarios where disk I/O would introduce unacceptable latency.

## Compiling and Loading the Kernel

The model must undergo two distinct phases before execution: compilation and loading. Both phases accept a **Quality of Service (QoS)** parameter and an options dictionary.

```objc
// Compilation phase
if (!((BOOL(*)(id,SEL,unsigned int,id,NSError**))objc_msgSend)(
        mdl, @selector(compileWithQoS:options:error:), 21, @{}, &err)) {
    // Handle compilation failure  // [ane_runtime.h L77-L80]
}

// Loading phase
if (!((BOOL(*)(id,SEL,unsigned int,id,NSError**))objc_msgSend)(
        mdl, @selector(loadWithQoS:options:error:), 21, @{}, &err)) {
    // Handle loading failure  // [ane_runtime.h L82-L86]
}

```

A QoS value of `21` corresponds to high-priority execution. The empty options dictionary allows the ANE runtime to select default optimization parameters based on the hardware generation.

## Wiring IOSurfaces for Zero-Copy I/O

The ANE hardware operates on **IOSurface** objects for zero-copy memory sharing between the CPU and Neural Engine. For each input and output tensor, the wrapper allocates an IOSurface of the exact byte size required by the tensor shape.

```objc
// Surface allocation helpers from ane_runtime.h L34-L42
IOSurfaceRef inputSurf  = ane_create_surface(inputBytes);
IOSurfaceRef outputSurf = ane_create_surface(outputBytes);

```

These surfaces are wrapped in `_ANEIOSurfaceObject` instances and attached to the request:

```objc
id ioObject = ((id(*)(Class,SEL,IOSurfaceRef))objc_msgSend)(
    g_ANEIO, @selector(objectWithIOSurface:), inputSurf);

```

The IOSurface backing stores remain mapped into process memory, allowing direct `memcpy` operations for feeding input data and retrieving results without additional buffer copies.

## Building and Executing ANE Requests

An `_ANERequest` object encapsulates the complete execution context, mapping specific IOSurface objects to input and output indices. The request is constructed using `requestWithInputs:inputIndices:outputs:outputIndices:weightsBuffer:perfStats:procedureIndex:`.

```objc
k->request = ((id(*)(Class,SEL,id,id,id,id,id,id,id))objc_msgSend)(
    g_ANEReq,
    @selector(requestWithInputs:inputIndices:outputs:outputIndices:weightsBuffer:perfStats:procedureIndex:),
    inputArray,   // NSArray of _ANEIOSurfaceObject
    inputIndices, // NSIndexSet
    outputArray,  // NSArray of _ANEIOSurfaceObject
    outputIndices,// NSIndexSet
    nil,          // weightsBuffer (optional)
    nil,          // perfStats
    @0);          // procedureIndex  // [ane_runtime.h L23-L26]

```

Execution occurs through the `evaluateWithQoS:options:request:error:` method on the loaded model:

```objc
BOOL success = ((BOOL(*)(id,SEL,unsigned int,id,id,NSError**))objc_msgSend)(
    k->model, @selector(evaluateWithQoS:options:request:error:),
    21, @{}, k->request, &err);  // [ane_runtime.h L42-L47]

```

## Data Flow and Resource Cleanup

Input data is written directly to the IOSurface base address using `IOSurfaceLock` and `memcpy`:

```objc
IOSurfaceLock(k->ioInputs[idx], 0, NULL);
memcpy(IOSurfaceGetBaseAddress(k->ioInputs[idx]), data, bytes);
IOSurfaceUnlock(k->ioInputs[idx], 0, NULL);

```

Similarly, outputs are read using `kIOSurfaceLockReadOnly` to prevent cache invalidation during CPU access. After inference completes, resources are released through `unloadWithQoS:error:` followed by `CFRelease` calls on the IOSurface objects and removal of temporary debug directories.

## Complete Compilation Example

The following example compiles a simple convolution MIL program and executes it on the ANE:

```objc
// 1. Define MIL program (Conv + Cast to FP32)
NSString *mil = @"program(1.3)\n"
                "{\n"
                "    input = tensor<fp16, [1,3,224,224]>(\"in\");\n"
                "    w = const<fp16>(\"weight.bin\");\n"
                "    conv = conv(input, w, pad=1, stride=1);\n"
                "    out = cast(conv, fp32);\n"
                "    output(out);\n"
                "}";

NSData *milData = [mil dataUsingEncoding:NSUTF8StringEncoding];
NSData *weights = [NSMutableData dataWithLength:9*3*3*4]; // 3x3 conv, 3 channels

// 2. Compile kernel (1 input, 1 output)
ANEKernel *k = ane_compile(milData, weights,
                           1, (size_t[]){1*3*224*224*2},   // FP16 input
                           1, (size_t[]){1*3*224*224*4});  // FP32 output

// 3. Run inference
float input[1*3*224*224] = {0};
ane_write_input(k, 0, input, sizeof(input));

if (!ane_eval(k)) abort();

float output[1*3*224*224];
ane_read_output(k, 0, output, sizeof(output));

// 4. Cleanup
ane_free(k);

```

## Summary

- **`_ANEInMemoryModelDescriptor`** serves as the factory class that compiles MIL text and weight dictionaries into ANE-compatible program descriptors without disk access.
- The compilation pipeline requires explicit **compile** and **load** phases using `compileWithQoS:options:error:` and `loadWithQoS:options:error:` before execution.
- **IOSurface** objects provide zero-copy memory mapping between CPU and ANE hardware, wrapped in `_ANEIOSurfaceObject` instances for the request builder.
- The `requestWithInputs:inputIndices:outputs:outputIndices:weightsBuffer:perfStats:procedureIndex:` method binds tensors to the execution context, evaluated via `evaluateWithQoS:options:request:error:`.
- All private framework interactions are encapsulated in [`training/ane_runtime.h`](https://github.com/maderix/ANE/blob/main/training/ane_runtime.h), with MIL generators available in [`training/ane_mil_gen.h`](https://github.com/maderix/ANE/blob/main/training/ane_mil_gen.h) and [`training/stories_mil.h`](https://github.com/maderix/ANE/blob/main/training/stories_mil.h).

## Frequently Asked Questions

### What is the role of `_ANEInMemoryModelDescriptor` in the MIL compiler?

`_ANEInMemoryModelDescriptor` is a private Objective-C class that functions as the entry point for in-memory compilation. It exposes the `modelWithMILText:weights:optionsPlist:` method, which parses MIL source code and weight blobs to produce a descriptor object. This descriptor is subsequently passed to `_ANEInMemoryModel` to generate the executable kernel, eliminating the need to write intermediate `.mlmodelc` files to disk.

### How does the ANE compiler handle weight data without file system access?

The compiler accepts weights through a dictionary parameter where keys follow the format `@model_path/weights/weight.bin` and values contain `NSData` objects with byte offsets. According to the implementation in [`ane_runtime.h`](https://github.com/maderix/ANE/blob/main/ane_runtime.h), these dictionaries map directly to `const` tensor declarations in the MIL program, allowing the runtime to resolve weight references from memory buffers rather than the file system.

### Why does the compilation process require both `compile` and `load` phases?

The two-phase process separates hardware code generation from resource allocation. The `compileWithQoS:options:error:` method translates the MIL program into ANE machine code and validates memory requirements, while `loadWithQoS:options:error:` allocates the actual hardware contexts and reserves IOSurface-backed memory. This separation allows for pre-compilation of models that can be loaded on-demand during latency-critical inference loops.

### What are the risks of using `_ANEInMemoryModelDescriptor` in production applications?

`_ANEInMemoryModelDescriptor` is part of the private `AppleNeuralEngine.framework`, which is not documented or sanctioned for App Store distribution. Applications using these APIs risk rejection from the App Store and potential breakage across iOS or macOS updates, as the class signatures and behavior may change without notice in future system releases.