# How Magnitude Loads and Evicts AI Models Based on Memory Constraints

> Learn how Magnitude manages AI model loading and eviction to optimize memory. Discover its memory-aware system for stable performance under resource constraints.

- Repository: [Magnitude/magnitude](https://github.com/magnitudedev/magnitude)
- Tags: internals
- Published: 2026-09-05

---

**Magnitude employs a memory-aware allocation system that tracks the memory footprint of each loaded AI model and automatically evicts low-priority instances when system pressure exceeds safe thresholds, ensuring stable performance under constrained resources.**

Magnitude, an open-source inference orchestration system hosted at `magnitudedev/magnitude`, implements a sophisticated memory management layer that governs how AI models are loaded into system memory and selectively unloaded when resources run low. The architecture separates concerns between model allocation, hardware monitoring, and eviction policy, allowing developers to configure memory guards while the system handles automatic cleanup. This article examines the source code implementation of memory-constrained model loading, from the initial SDK request through to the ranking algorithms that decide which models to retain.

## Model Request and Allocation Process

When a client initiates a model load request, the system creates a structured allocation record that tracks memory domains and consumption limits before the model weights are ever moved into RAM.

### Entry Point in model-commands.ts

The loading sequence begins in [`packages/acn/src/model-commands.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/model-commands.ts), where the `loadModel` command receives requests from the SDK. This command delegates to `client.models.ensureModelInstance`, which validates the request against current resource availability. According to the source code, this entry point acts as the gatekeeper for all model instantiation requests, ensuring that memory accounting happens before physical allocation occurs.

### Creating the ModelInstanceAllocation

Inside [`packages/sdk/src/inference-projection.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/inference-projection.ts), the system constructs a **ModelInstanceAllocation** object that records the model's **memory domains** (`memoryDomainId`) and estimated memory footprint. The allocation schema references [`packages/icn-protocol/src/generated/schemas.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/icn-protocol/src/generated/schemas.ts), which defines the `memory_limit` and `memory_domains` fields that constrain how much RAM the instance may consume. This allocation record persists throughout the model's lifecycle, enabling the system to track exactly how much memory is committed to each loaded model.

## Detecting Memory Pressure

Before eviction occurs, the system must detect when available memory drops below operational thresholds.

### Hardware Monitoring in local-inference-hardware.ts

The ACN daemon continuously probes system metrics via [`packages/acn/src/local-inference-hardware.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/local-inference-hardware.ts), which interfaces with low-level hardware counters to measure available RAM. When metrics indicate that free memory has fallen below the configured safety margin, the daemon updates the instance lifecycle state. The test suite in [`packages/icn/src/instances/index.test.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/icn/src/instances/index.test.ts) confirms that this condition triggers a **`low_memory`** error code, signaling that the system cannot satisfy the current allocation request without freeing resources.

### Lifecycle State Transitions

Upon detecting pressure, the daemon marks affected instances with the lifecycle reason **`"memory_pressure"`**. This state transition serves as the internal signal that triggers the eviction workflow. As documented in [`design/inference/system-memory-management.md`](https://github.com/magnitudedev/magnitude/blob/main/design/inference/system-memory-management.md), this monitoring loop operates continuously, checking thresholds against the **memory domains** defined in [`info/icn/memory-abstractions.md`](https://github.com/magnitudedev/magnitude/blob/main/info/icn/memory-abstractions.md) to determine whether the current working set exceeds capacity.

## Eviction Policy and Ranking

Once memory pressure is confirmed, Magnitude applies a ranking algorithm to select which models to unload.

### The local-model-ranking-policy.ts Implementation

The eviction logic resides in [`packages/acn/src/local-model-ranking-policy.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/local-model-ranking-policy.ts), which implements the **ModelRankingPolicy** interface. The policy evaluates each loaded model across three dimensions: recency of usage, explicit user-provided priority rankings, and whether the model is "pinned" by an active session. Models with the lowest composite priority score are selected for eviction first, ensuring that frequently accessed or critical models remain resident in memory.

### Memory Domain Cleanup

When a model is chosen for eviction, the system unloads its weights from the in-memory cache (`MemoryStorage`) and releases the associated **memory domains** tracked in the `ModelInstanceAllocation` record. This cleanup process updates the allocation state to reflect freed resources, allowing the original request that triggered the pressure to proceed with its own allocation. If the evicted model is requested again later, the system treats it as a fresh load, reconstructing the allocation from scratch.

## Configuration and Customization

Developers can influence the memory management behavior through SDK configuration and custom policy implementations.

To load a model with automatic memory monitoring:

```typescript
import { createClient } from "@magnitudedev/sdk";

const client = createClient({ /* …config… */ });

// Triggers ModelCommands.loadModel and memory allocation tracking
await client.models.loadModel("gpt-4");

```

To handle low-memory scenarios gracefully:

```typescript
try {
  await client.models.loadModel("large-vision-model");
} catch (err) {
  if (err.code === "low_memory") {
    console.warn("Memory pressure detected. Try a smaller model or free resources.");
    // Implement fallback logic here
  }
}

```

To configure a hard memory limit for a session:

```typescript
export const sessionOptions = {
  memoryGuard: { 
    kind: "on", 
    maxBytes: 8_000_000_000  // 8 GB cap triggers eviction above this threshold
  }
};

```

To implement a custom eviction policy:

```typescript
import { ModelRankingPolicy } from "@magnitudedev/acn";

class CustomPolicy implements ModelRankingPolicy {
  rank(models) {
    // Prioritize recent usage and smaller memory footprints
    return models.sort((a, b) => {
      const recency = b.lastUsed - a.lastUsed;
      const efficiency = a.memoryBytes - b.memoryBytes;
      return recency + efficiency;
    });
  }
}

```

## Summary

- **Memory-aware allocation** begins at [`packages/acn/src/model-commands.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/model-commands.ts) with the `loadModel` command, which creates a `ModelInstanceAllocation` in [`packages/sdk/src/inference-projection.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/inference-projection.ts) to track memory domains.
- **Pressure detection** relies on [`packages/acn/src/local-inference-hardware.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/local-inference-hardware.ts) to monitor system RAM, triggering `"memory_pressure"` lifecycle states and **`low_memory`** error codes when thresholds are breached.
- **Eviction decisions** are handled by [`packages/acn/src/local-model-ranking-policy.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/local-model-ranking-policy.ts), which ranks models by usage patterns and priority, unloading the lowest-ranked candidates to free memory domains.
- **Configuration options** include session-level `memoryGuard` limits and custom `ModelRankingPolicy` implementations for domain-specific eviction logic.

## Frequently Asked Questions

### How does Magnitude detect when system memory is under pressure?

The ACN daemon continuously monitors hardware metrics through [`packages/acn/src/local-inference-hardware.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/local-inference-hardware.ts), comparing available RAM against configured safety thresholds. When available memory drops below these limits, the daemon marks instances with a `"memory_pressure"` lifecycle reason and propagates **`low_memory`** error codes to pending requests, as verified in [`packages/icn/src/instances/index.test.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/icn/src/instances/index.test.ts).

### What determines which AI model gets evicted when memory runs low?

The [`packages/acn/src/local-model-ranking-policy.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/acn/src/local-model-ranking-policy.ts) file implements a ranking algorithm that scores each loaded model based on recent usage history, explicit priority assignments, and whether the model is pinned by an active session. The model with the lowest composite score is evicted first, with its memory domains released back to the system pool.

### Can I prevent specific models from being evicted?

Yes. The ranking policy respects "pinned" status assigned to models associated with active sessions. By configuring your session to pin critical models or by implementing a custom `ModelRankingPolicy` that assigns infinite priority to specific model IDs, you can ensure those models remain resident even under severe memory pressure.

### How should applications handle low_memory errors from the Magnitude SDK?

Applications should catch errors where `err.code === "low_memory"` and implement fallback strategies such as loading smaller model variants, queuing requests until resources free up, or prompting users to close other applications. The SDK throws these errors specifically from the allocation phase in [`packages/sdk/src/inference-projection.ts`](https://github.com/magnitudedev/magnitude/blob/main/packages/sdk/src/inference-projection.ts) before memory exhaustion causes system instability.