How to Benchmark MLX Operations and Compare Performance Against PyTorch

MLX provides a built-in Python benchmarking suite under benchmarks/python/ that measures operator latency with warm-up runs, forced evaluation via mx.eval, and direct comparison against PyTorch implementations.

The ml-explore/mlx repository includes a lightweight benchmarking framework designed specifically for its lazy evaluation model. Since MLX constructs computation graphs lazily, accurate performance measurement requires explicit graph materialization using mx.eval. The suite supports both CPU and Metal GPU backends, allowing you to quantify speed-ups against PyTorch across different tensor operations and data types.

Understanding the MLX Benchmarking Architecture

The benchmarking suite consists of three modular components that handle timing utilities, individual operator tests, and cross-framework comparisons.

Core Timing Utilities in time_utils.py

The foundation of the suite is benchmarks/python/time_utils.py, which implements the time_fn helper. This utility performs five warm-up iterations to prime the Metal runtime, then executes 100 timed iterations (or 10 for heavy operations like matmul). It uses Python's time.perf_counter for high-resolution wall-clock timing and forces evaluation with mx.eval to ensure kernels are actually executed rather than just queued in the computational graph.

Single-Operator Benchmarks in single_ops.py

For quick profiling of individual primitives, benchmarks/python/single_ops.py provides ready-to-run benchmarks covering arithmetic, reductions, reshapes, and neural network primitives. Each test generates random input tensors, optionally moves them to the target device using mx.set_default_device(mx.gpu) or mx.cpu, and reports average latency in milliseconds.

Comparative Driver in compare.py

The benchmarks/python/comparative/compare.py script orchestrates head-to-head performance tests. It spawns subprocesses running bench_mlx.py and bench_torch.py with identical parameters, parses their stdout as floating-point timings, and calculates relative speed-up using the formula (t_torch - t_mlx) / t_torch. Positive values indicate MLX is faster; negative values favor PyTorch.

Handling Lazy Evaluation in Timing Loops

MLX's lazy execution model requires specific handling to obtain accurate measurements. Before timing begins, the benchmark scripts run warm-up iterations to load kernels and mitigate first-run latency spikes. During the timed section, each iteration calls mx.eval on both the inputs and the operation result to force immediate materialization of the computation graph. Without this explicit synchronization, benchmarks would only measure graph construction overhead rather than actual kernel execution time.

Running MLX Benchmarks

You can execute benchmarks from the command line to measure specific operations or compare full workloads.

Benchmark Single Operators on CPU

To measure baseline CPU performance for basic operations like add or matmul:

python benchmarks/python/single_ops.py

This defaults to CPU execution and outputs timing lines such as:


Timing add ... 1.23845 msec
Timing matmul ... 4.51234 msec

Benchmark on Metal GPU

Switch to GPU acceleration by passing the --gpu flag, which internally calls mx.set_default_device(mx.gpu):

python benchmarks/python/single_ops.py --gpu

Compare MLX vs. PyTorch Performance

Use the comparative driver to calculate relative speed-up. For example, to benchmark the add operator with specific tensor dimensions:

python benchmarks/python/comparative/compare.py \
    --filter add --size 32x1024x1024

Sample output:


0.1523    add --size 32x1024x1024

This indicates MLX is approximately 15.2% faster than PyTorch for this operation.

Testing Different Data Types and Fused Operations

You can test multiple dtypes and fused implementations using additional flags:

python benchmarks/python/comparative/compare.py \
    -d float32 float16 \
    --filter linear --size 1024x1024 --fused

The -d flag iterates over specified data types, while --fused selects optimized implementations like mx.addmm for linear layers.

Extending the Benchmark Suite

To add a custom operator, define a function in bench_mlx.py following the existing pattern:

def custom_op(x):
    y = x
    for _ in range(100):
        y = mx.custom_op(y)
    mx.eval(y)
    return y

# In the __main__ block:

elif args.benchmark == "custom":
    print(bench(custom_op, x))

Create a matching implementation in bench_torch.py, then invoke compare.py with the new benchmark name to obtain comparative metrics.

Summary

  • Lazy evaluation requires explicit synchronization: Always call mx.eval on inputs and results to force kernel execution before measuring time.
  • Warm-up iterations matter: The suite runs 5 warm-up iterations to prime the Metal runtime and stabilize timings.
  • Modular architecture: Use time_utils.py for timing logic, single_ops.py for quick profiling, and compare.py for cross-framework analysis.
  • Device flexibility: Toggle between CPU and GPU benchmarks using mx.set_default_device() or command-line flags like --gpu.
  • Relative speed-up calculation: The comparative driver reports (t_torch - t_mlx) / t_torch, where positive values indicate MLX performance advantages.

Frequently Asked Questions

Why do MLX benchmarks require mx.eval calls?

MLX uses lazy evaluation to build computation graphs before execution. Without calling mx.eval, operations are only queued in the graph rather than executed. The benchmarking suite explicitly evaluates tensors to ensure measured time reflects actual kernel execution on the CPU or Metal GPU, not just graph construction overhead.

How many iterations does the MLX benchmark suite run?

The suite runs 5 warm-up iterations to initialize the runtime and load kernels, followed by 100 timed iterations for most operations. Heavy operations like matrix multiplication use 10 iterations to keep total runtime reasonable while maintaining statistical significance.

Can I benchmark custom MLX operators against PyTorch?

Yes. Add your operator to benchmarks/python/comparative/bench_mlx.py following the existing function signature pattern, create an equivalent PyTorch implementation in bench_torch.py, and run compare.py with your benchmark name. The comparative driver automatically calculates relative performance differences.

What does the output value mean in compare.py?

The output represents the relative speed-up of MLX compared to PyTorch, calculated as (t_torch - t_mlx) / t_torch. A positive value (e.g., 0.15) means MLX is 15% faster; a negative value indicates PyTorch is faster for that specific operation and configuration.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →