# How MLX's Memory Allocator and Buffer Management Work: A Deep Dive into the Core Subsystem

> Explore MLX's memory allocator and buffer management. Understand how it handles CPU, Metal, and CUDA memory with specialized optimization for high performance.

- Repository: [ml-explore/mlx](https://github.com/ml-explore/mlx)
- Tags: deep-dive
- Published: 2026-06-18

---

**MLX abstracts all memory handling behind a backend-agnostic allocator API that uses a singleton pattern to return `Buffer` objects, with specialized implementations for CPU (simple heap allocation), Metal (caching with heap optimization), and CUDA (pooled unified memory with small-block optimization).**

The ml-explore/mlx repository implements a unified memory management layer that hides backend complexity behind a small, consistent C++ API. Understanding **MLX's memory allocator** and its buffer management strategies is essential for optimizing performance on Apple Silicon GPUs, NVIDIA hardware, and CPU-only environments.

## Core Allocator Architecture

### The Buffer Abstraction

At the heart of the system is `allocator::Buffer`, defined in [`mlx/allocator.h`](https://github.com/ml-explore/mlx/blob/main/mlx/allocator.h) (lines 11-31). This thin wrapper encapsulates a raw pointer while remembering the underlying allocation type. The key method is `raw_ptr()`, which returns a host-accessible pointer; on GPU backends, calling this may trigger an automatic move to unified memory to ensure CPU readability.

### The Allocator Interface

The pure-virtual base class `allocator::Allocator` (lines 33-50 in [`mlx/allocator.h`](https://github.com/ml-explore/mlx/blob/main/mlx/allocator.h)) defines the contract every backend must implement:

- **`malloc(size_t)`** – Allocates a new buffer of the requested size
- **`free(Buffer)`** – Returns a buffer to the allocator (may cache it)
- **`size(Buffer)`** – Queries the size of an existing buffer
- **`make_buffer(void*, size_t)`** – Creates a non-copying view of existing memory
- **`release(Buffer)`** – Releases a view created with `make_buffer`

### Global Accessor and Singleton Pattern

The global function `allocator()` returns a reference to the singleton allocator for the current active backend (CPU, Metal, or CUDA). Inline helper functions in [`mlx/allocator.h`](https://github.com/ml-explore/mlx/blob/main/mlx/allocator.h) (lines 52-73) forward to this singleton:

- `mlx::core::malloc`
- `mlx::core::free`
- `mlx::core::size`
- `mlx::core::make_buffer`
- `mlx::core::release`

This design ensures that switching from CPU to GPU execution requires only changing the device selection (e.g., `mlx::core::set_default_device(mlx::core::Device::gpu)`), not the allocation code.

## Backend-Specific Implementations

### CPU/No-GPU Allocator

The simplest implementation resides in [`mlx/backend/no_gpu/allocator.cpp`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/no_gpu/allocator.cpp). It performs standard `malloc` and `free` operations on the host heap without caching or special residency management.

### Metal Allocator with Caching and Heap Optimization

For Apple Silicon GPUs, [`mlx/backend/metal/allocator.h`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/metal/allocator.h) and [`mlx/backend/metal/allocator.cpp`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/metal/allocator.cpp) implement a sophisticated caching allocator featuring:

- **BufferCache**: Before allocating new memory, the allocator attempts `buffer_cache_.reuse_from_cache(size)`. If the cache would exceed `max_pool_size_`, it evicts the oldest entries.
- **Heap Optimization**: Allocations ≤ 256 B are placed in a Metal heap (`heap_`) to reduce driver overhead.
- **Residency Set**: Non-heap buffers are tracked in `residency_set_`, allowing the OS to maintain them resident when a *wired limit* is configured.
- **Memory Limits**: The allocator tracks active and peak memory while respecting configurable cache limits and memory limits.

### CUDA Allocator with Unified Memory and Pools

The CUDA backend in [`mlx/backend/cuda/allocator.cpp`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/cuda/allocator.cpp) implements a hybrid approach:

- **Small-Block Pool**: A pre-allocated pool of 4 × page-size blocks (`small_pool_size`) serves allocations ≤ 8 B without invoking the driver.
- **Page-Size-Aligned Pool**: Larger allocations use aligned memory pools.
- **Unified Memory Fallback**: `unified_malloc` selects `cudaMallocManaged` when all devices support managed memory; otherwise it falls back to `cudaMallocHost`.
- **CUDA Memory Pools**: When supported, `malloc_async` uses `cudaMallocAsync` while monitoring `cudaMemPoolGetAttribute` to prevent over-reservation.
- **Caching**: Like Metal, it maintains a `buffer_cache_` and tracks active/peak memory statistics.

## Buffer Lifecycle and Zero-Copy Operations

The standard lifecycle follows four distinct phases:

1. **Allocation**: `auto buf = mlx::core::malloc(bytes);` delegates to the backend's `malloc` implementation.
2. **Usage**: Access the opaque device pointer via `buf.ptr()`, or obtain a host-accessible pointer through `buf.raw_ptr()` (which may trigger synchronization on GPU backends).
3. **Deallocation**: `mlx::core::free(buf);` returns the buffer to the cache or releases it to the driver.
4. **Zero-Copy Views**: Wrap existing memory without copying using `auto view = mlx::core::make_buffer(ptr, size);`. These views must be explicitly released with `mlx::core::release(view);`.

## Practical Code Examples

```cpp
#include "mlx/allocator.h"

int main() {
    // Allocate 1 MiB on the default device (CPU, Metal, or CUDA)
    auto buf = mlx::core::malloc(1 << 20);

    // Get raw host pointer (moves to unified memory on GPU if necessary)
    void* host_ptr = buf.raw_ptr();
    std::memset(host_ptr, 0xAB, 1 << 20);

    // Query allocation size
    size_t sz = mlx::core::size(buf);  // Returns 1048576

    // Release back to allocator (cached if possible)
    mlx::core::free(buf);

    // Zero-copy example: wrap existing memory
    float existing[256];
    auto view = mlx::core::make_buffer(existing, sizeof(existing));
    // Use view in MLX operations...
    mlx::core::release(view);  // Required cleanup
}

```

This code runs unchanged across backends because the allocator singleton automatically dispatches to the correct implementation based on the current device context.

## Summary

- **MLX's memory allocator** uses a singleton-based design with a pure-virtual `Allocator` interface and lightweight `Buffer` wrappers defined in [`mlx/allocator.h`](https://github.com/ml-explore/mlx/blob/main/mlx/allocator.h).
- **Metal backend** ([`mlx/backend/metal/allocator.cpp`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/metal/allocator.cpp)) implements caching with `BufferCache`, heap optimization for small allocations (≤ 256 B), and residency set management for wired memory limits.
- **CUDA backend** ([`mlx/backend/cuda/allocator.cpp`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/cuda/allocator.cpp)) provides small-block pools (≤ 8 B), unified memory fallback (`cudaMallocManaged`), and integration with CUDA memory pools when available.
- **Zero-copy operations** allow wrapping existing host pointers via `make_buffer()` and `release()` without memory duplication.
- All backends share identical semantics through global functions like `mlx::core::malloc` and `mlx::core::free`, enabling portable code across CPU, Apple Silicon, and NVIDIA GPUs.

## Frequently Asked Questions

### What is the difference between `mlx::core::malloc` and standard `malloc`?

**`mlx::core::malloc`** delegates to the active backend's allocator singleton, returning an `allocator::Buffer` object rather than a raw pointer. On GPU backends, this may allocate device memory or unified memory, whereas standard `malloc` only allocates host heap memory. The MLX version also enables backend-specific optimizations like caching and residency management that standard allocation cannot provide.

### How does MLX handle memory allocation on Apple Silicon GPUs?

On Metal, MLX uses a **caching allocator** ([`mlx/backend/metal/allocator.cpp`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/metal/allocator.cpp)) that maintains a `BufferCache` to reuse recently freed buffers. It places small allocations (≤ 256 B) into a dedicated Metal heap to reduce driver overhead and tracks non-heap buffers in a residency set to support wired memory limits. Before allocating new memory, it always attempts to reclaim from the cache.

### What is zero-copy buffer management in MLX?

Zero-copy management allows wrapping existing host memory without duplication through **`mlx::core::make_buffer(void* ptr, size_t size)`**. This creates a non-owning `Buffer` view that can be used in MLX operations. Unlike `malloc`, which allocates new memory, `make_buffer` requires the caller to manage the underlying lifetime and must be paired with **`mlx::core::release(view)`** to clean up the wrapper object.

### How does the CUDA allocator optimize for small allocations?

The CUDA implementation ([`mlx/backend/cuda/allocator.cpp`](https://github.com/ml-explore/mlx/blob/main/mlx/backend/cuda/allocator.cpp)) maintains a **small-block pool** consisting of 4 × page-size blocks that serves allocations ≤ 8 B without invoking `cudaMalloc`. For larger requests, it uses page-size-aligned pools and can fall back to unified memory (`cudaMallocManaged`) when all devices support it. When CUDA memory pools are available, it uses `cudaMallocAsync` with attribute monitoring to prevent over-reservation.