How MLX's Memory Allocator and Buffer Management Work: A Deep Dive into the Core Subsystem

MLX abstracts all memory handling behind a backend-agnostic allocator API that uses a singleton pattern to return Buffer objects, with specialized implementations for CPU (simple heap allocation), Metal (caching with heap optimization), and CUDA (pooled unified memory with small-block optimization).

The ml-explore/mlx repository implements a unified memory management layer that hides backend complexity behind a small, consistent C++ API. Understanding MLX's memory allocator and its buffer management strategies is essential for optimizing performance on Apple Silicon GPUs, NVIDIA hardware, and CPU-only environments.

Core Allocator Architecture

The Buffer Abstraction

At the heart of the system is allocator::Buffer, defined in mlx/allocator.h (lines 11-31). This thin wrapper encapsulates a raw pointer while remembering the underlying allocation type. The key method is raw_ptr(), which returns a host-accessible pointer; on GPU backends, calling this may trigger an automatic move to unified memory to ensure CPU readability.

The Allocator Interface

The pure-virtual base class allocator::Allocator (lines 33-50 in mlx/allocator.h) defines the contract every backend must implement:

  • malloc(size_t) – Allocates a new buffer of the requested size
  • free(Buffer) – Returns a buffer to the allocator (may cache it)
  • size(Buffer) – Queries the size of an existing buffer
  • make_buffer(void*, size_t) – Creates a non-copying view of existing memory
  • release(Buffer) – Releases a view created with make_buffer

Global Accessor and Singleton Pattern

The global function allocator() returns a reference to the singleton allocator for the current active backend (CPU, Metal, or CUDA). Inline helper functions in mlx/allocator.h (lines 52-73) forward to this singleton:

  • mlx::core::malloc
  • mlx::core::free
  • mlx::core::size
  • mlx::core::make_buffer
  • mlx::core::release

This design ensures that switching from CPU to GPU execution requires only changing the device selection (e.g., mlx::core::set_default_device(mlx::core::Device::gpu)), not the allocation code.

Backend-Specific Implementations

CPU/No-GPU Allocator

The simplest implementation resides in mlx/backend/no_gpu/allocator.cpp. It performs standard malloc and free operations on the host heap without caching or special residency management.

Metal Allocator with Caching and Heap Optimization

For Apple Silicon GPUs, mlx/backend/metal/allocator.h and mlx/backend/metal/allocator.cpp implement a sophisticated caching allocator featuring:

  • BufferCache: Before allocating new memory, the allocator attempts buffer_cache_.reuse_from_cache(size). If the cache would exceed max_pool_size_, it evicts the oldest entries.
  • Heap Optimization: Allocations ≤ 256 B are placed in a Metal heap (heap_) to reduce driver overhead.
  • Residency Set: Non-heap buffers are tracked in residency_set_, allowing the OS to maintain them resident when a wired limit is configured.
  • Memory Limits: The allocator tracks active and peak memory while respecting configurable cache limits and memory limits.

CUDA Allocator with Unified Memory and Pools

The CUDA backend in mlx/backend/cuda/allocator.cpp implements a hybrid approach:

  • Small-Block Pool: A pre-allocated pool of 4 × page-size blocks (small_pool_size) serves allocations ≤ 8 B without invoking the driver.
  • Page-Size-Aligned Pool: Larger allocations use aligned memory pools.
  • Unified Memory Fallback: unified_malloc selects cudaMallocManaged when all devices support managed memory; otherwise it falls back to cudaMallocHost.
  • CUDA Memory Pools: When supported, malloc_async uses cudaMallocAsync while monitoring cudaMemPoolGetAttribute to prevent over-reservation.
  • Caching: Like Metal, it maintains a buffer_cache_ and tracks active/peak memory statistics.

Buffer Lifecycle and Zero-Copy Operations

The standard lifecycle follows four distinct phases:

  1. Allocation: auto buf = mlx::core::malloc(bytes); delegates to the backend's malloc implementation.
  2. Usage: Access the opaque device pointer via buf.ptr(), or obtain a host-accessible pointer through buf.raw_ptr() (which may trigger synchronization on GPU backends).
  3. Deallocation: mlx::core::free(buf); returns the buffer to the cache or releases it to the driver.
  4. Zero-Copy Views: Wrap existing memory without copying using auto view = mlx::core::make_buffer(ptr, size);. These views must be explicitly released with mlx::core::release(view);.

Practical Code Examples

#include "mlx/allocator.h"

int main() {
    // Allocate 1 MiB on the default device (CPU, Metal, or CUDA)
    auto buf = mlx::core::malloc(1 << 20);

    // Get raw host pointer (moves to unified memory on GPU if necessary)
    void* host_ptr = buf.raw_ptr();
    std::memset(host_ptr, 0xAB, 1 << 20);

    // Query allocation size
    size_t sz = mlx::core::size(buf);  // Returns 1048576

    // Release back to allocator (cached if possible)
    mlx::core::free(buf);

    // Zero-copy example: wrap existing memory
    float existing[256];
    auto view = mlx::core::make_buffer(existing, sizeof(existing));
    // Use view in MLX operations...
    mlx::core::release(view);  // Required cleanup
}

This code runs unchanged across backends because the allocator singleton automatically dispatches to the correct implementation based on the current device context.

Summary

  • MLX's memory allocator uses a singleton-based design with a pure-virtual Allocator interface and lightweight Buffer wrappers defined in mlx/allocator.h.
  • Metal backend (mlx/backend/metal/allocator.cpp) implements caching with BufferCache, heap optimization for small allocations (≤ 256 B), and residency set management for wired memory limits.
  • CUDA backend (mlx/backend/cuda/allocator.cpp) provides small-block pools (≤ 8 B), unified memory fallback (cudaMallocManaged), and integration with CUDA memory pools when available.
  • Zero-copy operations allow wrapping existing host pointers via make_buffer() and release() without memory duplication.
  • All backends share identical semantics through global functions like mlx::core::malloc and mlx::core::free, enabling portable code across CPU, Apple Silicon, and NVIDIA GPUs.

Frequently Asked Questions

What is the difference between mlx::core::malloc and standard malloc?

mlx::core::malloc delegates to the active backend's allocator singleton, returning an allocator::Buffer object rather than a raw pointer. On GPU backends, this may allocate device memory or unified memory, whereas standard malloc only allocates host heap memory. The MLX version also enables backend-specific optimizations like caching and residency management that standard allocation cannot provide.

How does MLX handle memory allocation on Apple Silicon GPUs?

On Metal, MLX uses a caching allocator (mlx/backend/metal/allocator.cpp) that maintains a BufferCache to reuse recently freed buffers. It places small allocations (≤ 256 B) into a dedicated Metal heap to reduce driver overhead and tracks non-heap buffers in a residency set to support wired memory limits. Before allocating new memory, it always attempts to reclaim from the cache.

What is zero-copy buffer management in MLX?

Zero-copy management allows wrapping existing host memory without duplication through mlx::core::make_buffer(void* ptr, size_t size). This creates a non-owning Buffer view that can be used in MLX operations. Unlike malloc, which allocates new memory, make_buffer requires the caller to manage the underlying lifetime and must be paired with mlx::core::release(view) to clean up the wrapper object.

How does the CUDA allocator optimize for small allocations?

The CUDA implementation (mlx/backend/cuda/allocator.cpp) maintains a small-block pool consisting of 4 × page-size blocks that serves allocations ≤ 8 B without invoking cudaMalloc. For larger requests, it uses page-size-aligned pools and can fall back to unified memory (cudaMallocManaged) when all devices support it. When CUDA memory pools are available, it uses cudaMallocAsync with attribute monitoring to prevent over-reservation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →