How MLX's Memory Allocator and Buffer Management Work: A Deep Dive into the Core Subsystem
MLX abstracts all memory handling behind a backend-agnostic allocator API that uses a singleton pattern to return Buffer objects, with specialized implementations for CPU (simple heap allocation), Metal (caching with heap optimization), and CUDA (pooled unified memory with small-block optimization).
The ml-explore/mlx repository implements a unified memory management layer that hides backend complexity behind a small, consistent C++ API. Understanding MLX's memory allocator and its buffer management strategies is essential for optimizing performance on Apple Silicon GPUs, NVIDIA hardware, and CPU-only environments.
Core Allocator Architecture
The Buffer Abstraction
At the heart of the system is allocator::Buffer, defined in mlx/allocator.h (lines 11-31). This thin wrapper encapsulates a raw pointer while remembering the underlying allocation type. The key method is raw_ptr(), which returns a host-accessible pointer; on GPU backends, calling this may trigger an automatic move to unified memory to ensure CPU readability.
The Allocator Interface
The pure-virtual base class allocator::Allocator (lines 33-50 in mlx/allocator.h) defines the contract every backend must implement:
malloc(size_t)– Allocates a new buffer of the requested sizefree(Buffer)– Returns a buffer to the allocator (may cache it)size(Buffer)– Queries the size of an existing buffermake_buffer(void*, size_t)– Creates a non-copying view of existing memoryrelease(Buffer)– Releases a view created withmake_buffer
Global Accessor and Singleton Pattern
The global function allocator() returns a reference to the singleton allocator for the current active backend (CPU, Metal, or CUDA). Inline helper functions in mlx/allocator.h (lines 52-73) forward to this singleton:
mlx::core::mallocmlx::core::freemlx::core::sizemlx::core::make_buffermlx::core::release
This design ensures that switching from CPU to GPU execution requires only changing the device selection (e.g., mlx::core::set_default_device(mlx::core::Device::gpu)), not the allocation code.
Backend-Specific Implementations
CPU/No-GPU Allocator
The simplest implementation resides in mlx/backend/no_gpu/allocator.cpp. It performs standard malloc and free operations on the host heap without caching or special residency management.
Metal Allocator with Caching and Heap Optimization
For Apple Silicon GPUs, mlx/backend/metal/allocator.h and mlx/backend/metal/allocator.cpp implement a sophisticated caching allocator featuring:
- BufferCache: Before allocating new memory, the allocator attempts
buffer_cache_.reuse_from_cache(size). If the cache would exceedmax_pool_size_, it evicts the oldest entries. - Heap Optimization: Allocations ≤ 256 B are placed in a Metal heap (
heap_) to reduce driver overhead. - Residency Set: Non-heap buffers are tracked in
residency_set_, allowing the OS to maintain them resident when a wired limit is configured. - Memory Limits: The allocator tracks active and peak memory while respecting configurable cache limits and memory limits.
CUDA Allocator with Unified Memory and Pools
The CUDA backend in mlx/backend/cuda/allocator.cpp implements a hybrid approach:
- Small-Block Pool: A pre-allocated pool of 4 × page-size blocks (
small_pool_size) serves allocations ≤ 8 B without invoking the driver. - Page-Size-Aligned Pool: Larger allocations use aligned memory pools.
- Unified Memory Fallback:
unified_mallocselectscudaMallocManagedwhen all devices support managed memory; otherwise it falls back tocudaMallocHost. - CUDA Memory Pools: When supported,
malloc_asyncusescudaMallocAsyncwhile monitoringcudaMemPoolGetAttributeto prevent over-reservation. - Caching: Like Metal, it maintains a
buffer_cache_and tracks active/peak memory statistics.
Buffer Lifecycle and Zero-Copy Operations
The standard lifecycle follows four distinct phases:
- Allocation:
auto buf = mlx::core::malloc(bytes);delegates to the backend'smallocimplementation. - Usage: Access the opaque device pointer via
buf.ptr(), or obtain a host-accessible pointer throughbuf.raw_ptr()(which may trigger synchronization on GPU backends). - Deallocation:
mlx::core::free(buf);returns the buffer to the cache or releases it to the driver. - Zero-Copy Views: Wrap existing memory without copying using
auto view = mlx::core::make_buffer(ptr, size);. These views must be explicitly released withmlx::core::release(view);.
Practical Code Examples
#include "mlx/allocator.h"
int main() {
// Allocate 1 MiB on the default device (CPU, Metal, or CUDA)
auto buf = mlx::core::malloc(1 << 20);
// Get raw host pointer (moves to unified memory on GPU if necessary)
void* host_ptr = buf.raw_ptr();
std::memset(host_ptr, 0xAB, 1 << 20);
// Query allocation size
size_t sz = mlx::core::size(buf); // Returns 1048576
// Release back to allocator (cached if possible)
mlx::core::free(buf);
// Zero-copy example: wrap existing memory
float existing[256];
auto view = mlx::core::make_buffer(existing, sizeof(existing));
// Use view in MLX operations...
mlx::core::release(view); // Required cleanup
}
This code runs unchanged across backends because the allocator singleton automatically dispatches to the correct implementation based on the current device context.
Summary
- MLX's memory allocator uses a singleton-based design with a pure-virtual
Allocatorinterface and lightweightBufferwrappers defined inmlx/allocator.h. - Metal backend (
mlx/backend/metal/allocator.cpp) implements caching withBufferCache, heap optimization for small allocations (≤ 256 B), and residency set management for wired memory limits. - CUDA backend (
mlx/backend/cuda/allocator.cpp) provides small-block pools (≤ 8 B), unified memory fallback (cudaMallocManaged), and integration with CUDA memory pools when available. - Zero-copy operations allow wrapping existing host pointers via
make_buffer()andrelease()without memory duplication. - All backends share identical semantics through global functions like
mlx::core::mallocandmlx::core::free, enabling portable code across CPU, Apple Silicon, and NVIDIA GPUs.
Frequently Asked Questions
What is the difference between mlx::core::malloc and standard malloc?
mlx::core::malloc delegates to the active backend's allocator singleton, returning an allocator::Buffer object rather than a raw pointer. On GPU backends, this may allocate device memory or unified memory, whereas standard malloc only allocates host heap memory. The MLX version also enables backend-specific optimizations like caching and residency management that standard allocation cannot provide.
How does MLX handle memory allocation on Apple Silicon GPUs?
On Metal, MLX uses a caching allocator (mlx/backend/metal/allocator.cpp) that maintains a BufferCache to reuse recently freed buffers. It places small allocations (≤ 256 B) into a dedicated Metal heap to reduce driver overhead and tracks non-heap buffers in a residency set to support wired memory limits. Before allocating new memory, it always attempts to reclaim from the cache.
What is zero-copy buffer management in MLX?
Zero-copy management allows wrapping existing host memory without duplication through mlx::core::make_buffer(void* ptr, size_t size). This creates a non-owning Buffer view that can be used in MLX operations. Unlike malloc, which allocates new memory, make_buffer requires the caller to manage the underlying lifetime and must be paired with mlx::core::release(view) to clean up the wrapper object.
How does the CUDA allocator optimize for small allocations?
The CUDA implementation (mlx/backend/cuda/allocator.cpp) maintains a small-block pool consisting of 4 × page-size blocks that serves allocations ≤ 8 B without invoking cudaMalloc. For larger requests, it uses page-size-aligned pools and can fall back to unified memory (cudaMallocManaged) when all devices support it. When CUDA memory pools are available, it uses cudaMallocAsync with attribute monitoring to prevent over-reservation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →