How the MIL Compiler Works with `_ANEInMemoryModelDescriptor`: In-Memory ANE Compilation Explained

The MIL compiler leverages the private _ANEInMemoryModelDescriptor class to convert Machine-Learning Intermediate Language (MIL) programs and weight blobs into executable Apple Neural Engine (ANE) kernels entirely in RAM, eliminating the need for disk-based .mlmodelc bundles.

The maderix/ANE repository demonstrates this in-memory compilation pipeline through a low-level Objective-C wrapper that interfaces with private AppleNeuralEngine.framework APIs. By dynamically loading the _ANEInMemoryModelDescriptor and _ANEInMemoryModel classes, the system transforms textual MIL representations into hardware-accelerated kernels ready for inference.

Loading Private ANE Symbols at Runtime

The compilation process begins by dynamically linking against the private AppleNeuralEngine.framework and obtaining class references for the descriptor and model objects.

In training/ane_runtime.h, the initialization routine uses dlopen to load the framework and NSClassFromString to resolve the private classes:

dlopen("/System/Library/PrivateFrameworks/AppleNeuralEngine.framework/AppleNeuralEngine", RTLD_NOW);

Class g_ANEDesc  = NSClassFromString(@"_ANEInMemoryModelDescriptor");  // [ane_runtime.h L27]
Class g_ANEInMem = NSClassFromString(@"_ANEInMemoryModel");            // [ane_runtime.h L28]
Class g_ANEReq   = NSClassFromString(@"_ANERequest");                  // [ane_runtime.h L29]
Class g_ANEIO    = NSClassFromString(@"_ANEIOSurfaceObject");          // [ane_runtime.h L30]

These global class pointers are cached during ane_init and reused throughout the application lifecycle to minimize the overhead of repeated symbol lookups.

Creating the Descriptor from MIL Text

The _ANEInMemoryModelDescriptor acts as a factory for compiling MIL programs. It exposes the selector modelWithMILText:weights:optionsPlist:, which accepts the raw MIL source as NSData, an optional weight dictionary, and a configuration plist.

NSDictionary *wdict = weightData ?
    @{@"@model_path/weights/weight.bin": @{@"offset": @0, @"data": weightData}} : nil;

id desc = ((id(*)(Class,SEL,id,id,id))objc_msgSend)(
    g_ANEDesc,
    @selector(modelWithMILText:weights:optionsPlist:),
    milText,    // NSData containing MIL string
    wdict,      // Weight map or nil
    nil);       // Options plist (nil for defaults)  // [ane_runtime.h L55-L62]

If the MIL syntax is invalid or the weights are incompatible, the descriptor returns NULL and compilation aborts immediately. The weight dictionary uses a specific key format (@model_path/weights/weight.bin) to map binary blobs to tensor constants declared in the MIL program.

Instantiating the In-Memory Model

Once the descriptor is validated, the system creates an _ANEInMemoryModel instance using the inMemoryModelWithDescriptor: constructor. This object holds the compiled ANE program and manages the IOSurface handles for input and output tensors.

id mdl = ((id(*)(Class,SEL,id))objc_msgSend)(
    g_ANEInMem,
    @selector(inMemoryModelWithDescriptor:), desc);  // [ane_runtime.h L64-L66]

The model object remains in memory until explicitly unloaded, making it suitable for high-frequency inference scenarios where disk I/O would introduce unacceptable latency.

Compiling and Loading the Kernel

The model must undergo two distinct phases before execution: compilation and loading. Both phases accept a Quality of Service (QoS) parameter and an options dictionary.

// Compilation phase
if (!((BOOL(*)(id,SEL,unsigned int,id,NSError**))objc_msgSend)(
        mdl, @selector(compileWithQoS:options:error:), 21, @{}, &err)) {
    // Handle compilation failure  // [ane_runtime.h L77-L80]
}

// Loading phase
if (!((BOOL(*)(id,SEL,unsigned int,id,NSError**))objc_msgSend)(
        mdl, @selector(loadWithQoS:options:error:), 21, @{}, &err)) {
    // Handle loading failure  // [ane_runtime.h L82-L86]
}

A QoS value of 21 corresponds to high-priority execution. The empty options dictionary allows the ANE runtime to select default optimization parameters based on the hardware generation.

Wiring IOSurfaces for Zero-Copy I/O

The ANE hardware operates on IOSurface objects for zero-copy memory sharing between the CPU and Neural Engine. For each input and output tensor, the wrapper allocates an IOSurface of the exact byte size required by the tensor shape.

// Surface allocation helpers from ane_runtime.h L34-L42
IOSurfaceRef inputSurf  = ane_create_surface(inputBytes);
IOSurfaceRef outputSurf = ane_create_surface(outputBytes);

These surfaces are wrapped in _ANEIOSurfaceObject instances and attached to the request:

id ioObject = ((id(*)(Class,SEL,IOSurfaceRef))objc_msgSend)(
    g_ANEIO, @selector(objectWithIOSurface:), inputSurf);

The IOSurface backing stores remain mapped into process memory, allowing direct memcpy operations for feeding input data and retrieving results without additional buffer copies.

Building and Executing ANE Requests

An _ANERequest object encapsulates the complete execution context, mapping specific IOSurface objects to input and output indices. The request is constructed using requestWithInputs:inputIndices:outputs:outputIndices:weightsBuffer:perfStats:procedureIndex:.

k->request = ((id(*)(Class,SEL,id,id,id,id,id,id,id))objc_msgSend)(
    g_ANEReq,
    @selector(requestWithInputs:inputIndices:outputs:outputIndices:weightsBuffer:perfStats:procedureIndex:),
    inputArray,   // NSArray of _ANEIOSurfaceObject
    inputIndices, // NSIndexSet
    outputArray,  // NSArray of _ANEIOSurfaceObject
    outputIndices,// NSIndexSet
    nil,          // weightsBuffer (optional)
    nil,          // perfStats
    @0);          // procedureIndex  // [ane_runtime.h L23-L26]

Execution occurs through the evaluateWithQoS:options:request:error: method on the loaded model:

BOOL success = ((BOOL(*)(id,SEL,unsigned int,id,id,NSError**))objc_msgSend)(
    k->model, @selector(evaluateWithQoS:options:request:error:),
    21, @{}, k->request, &err);  // [ane_runtime.h L42-L47]

Data Flow and Resource Cleanup

Input data is written directly to the IOSurface base address using IOSurfaceLock and memcpy:

IOSurfaceLock(k->ioInputs[idx], 0, NULL);
memcpy(IOSurfaceGetBaseAddress(k->ioInputs[idx]), data, bytes);
IOSurfaceUnlock(k->ioInputs[idx], 0, NULL);

Similarly, outputs are read using kIOSurfaceLockReadOnly to prevent cache invalidation during CPU access. After inference completes, resources are released through unloadWithQoS:error: followed by CFRelease calls on the IOSurface objects and removal of temporary debug directories.

Complete Compilation Example

The following example compiles a simple convolution MIL program and executes it on the ANE:

// 1. Define MIL program (Conv + Cast to FP32)
NSString *mil = @"program(1.3)\n"
                "{\n"
                "    input = tensor<fp16, [1,3,224,224]>(\"in\");\n"
                "    w = const<fp16>(\"weight.bin\");\n"
                "    conv = conv(input, w, pad=1, stride=1);\n"
                "    out = cast(conv, fp32);\n"
                "    output(out);\n"
                "}";

NSData *milData = [mil dataUsingEncoding:NSUTF8StringEncoding];
NSData *weights = [NSMutableData dataWithLength:9*3*3*4]; // 3x3 conv, 3 channels

// 2. Compile kernel (1 input, 1 output)
ANEKernel *k = ane_compile(milData, weights,
                           1, (size_t[]){1*3*224*224*2},   // FP16 input
                           1, (size_t[]){1*3*224*224*4});  // FP32 output

// 3. Run inference
float input[1*3*224*224] = {0};
ane_write_input(k, 0, input, sizeof(input));

if (!ane_eval(k)) abort();

float output[1*3*224*224];
ane_read_output(k, 0, output, sizeof(output));

// 4. Cleanup
ane_free(k);

Summary

  • _ANEInMemoryModelDescriptor serves as the factory class that compiles MIL text and weight dictionaries into ANE-compatible program descriptors without disk access.
  • The compilation pipeline requires explicit compile and load phases using compileWithQoS:options:error: and loadWithQoS:options:error: before execution.
  • IOSurface objects provide zero-copy memory mapping between CPU and ANE hardware, wrapped in _ANEIOSurfaceObject instances for the request builder.
  • The requestWithInputs:inputIndices:outputs:outputIndices:weightsBuffer:perfStats:procedureIndex: method binds tensors to the execution context, evaluated via evaluateWithQoS:options:request:error:.
  • All private framework interactions are encapsulated in training/ane_runtime.h, with MIL generators available in training/ane_mil_gen.h and training/stories_mil.h.

Frequently Asked Questions

What is the role of _ANEInMemoryModelDescriptor in the MIL compiler?

_ANEInMemoryModelDescriptor is a private Objective-C class that functions as the entry point for in-memory compilation. It exposes the modelWithMILText:weights:optionsPlist: method, which parses MIL source code and weight blobs to produce a descriptor object. This descriptor is subsequently passed to _ANEInMemoryModel to generate the executable kernel, eliminating the need to write intermediate .mlmodelc files to disk.

How does the ANE compiler handle weight data without file system access?

The compiler accepts weights through a dictionary parameter where keys follow the format @model_path/weights/weight.bin and values contain NSData objects with byte offsets. According to the implementation in ane_runtime.h, these dictionaries map directly to const tensor declarations in the MIL program, allowing the runtime to resolve weight references from memory buffers rather than the file system.

Why does the compilation process require both compile and load phases?

The two-phase process separates hardware code generation from resource allocation. The compileWithQoS:options:error: method translates the MIL program into ANE machine code and validates memory requirements, while loadWithQoS:options:error: allocates the actual hardware contexts and reserves IOSurface-backed memory. This separation allows for pre-compilation of models that can be loaded on-demand during latency-critical inference loops.

What are the risks of using _ANEInMemoryModelDescriptor in production applications?

_ANEInMemoryModelDescriptor is part of the private AppleNeuralEngine.framework, which is not documented or sanctioned for App Store distribution. Applications using these APIs risk rejection from the App Store and potential breakage across iOS or macOS updates, as the class signatures and behavior may change without notice in future system releases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →