# How to Benchmark MLX Operations and Compare Performance Against PyTorch

> Benchmark MLX operations and compare performance directly against PyTorch using the built-in Python benchmarking suite. Measure operator latency, warm-up runs, and forced evaluation for accurate results.

- Repository: [ml-explore/mlx](https://github.com/ml-explore/mlx)
- Tags: performance
- Published: 2026-06-18

---

**MLX provides a built-in Python benchmarking suite under `benchmarks/python/` that measures operator latency with warm-up runs, forced evaluation via `mx.eval`, and direct comparison against PyTorch implementations.**

The ml-explore/mlx repository includes a lightweight benchmarking framework designed specifically for its lazy evaluation model. Since MLX constructs computation graphs lazily, accurate performance measurement requires explicit graph materialization using `mx.eval`. The suite supports both CPU and Metal GPU backends, allowing you to quantify speed-ups against PyTorch across different tensor operations and data types.

## Understanding the MLX Benchmarking Architecture

The benchmarking suite consists of three modular components that handle timing utilities, individual operator tests, and cross-framework comparisons.

### Core Timing Utilities in [`time_utils.py`](https://github.com/ml-explore/mlx/blob/main/time_utils.py)

The foundation of the suite is [`benchmarks/python/time_utils.py`](https://github.com/ml-explore/mlx/blob/main/benchmarks/python/time_utils.py), which implements the `time_fn` helper. This utility performs five warm-up iterations to prime the Metal runtime, then executes 100 timed iterations (or 10 for heavy operations like `matmul`). It uses Python's `time.perf_counter` for high-resolution wall-clock timing and forces evaluation with `mx.eval` to ensure kernels are actually executed rather than just queued in the computational graph.

### Single-Operator Benchmarks in [`single_ops.py`](https://github.com/ml-explore/mlx/blob/main/single_ops.py)

For quick profiling of individual primitives, [`benchmarks/python/single_ops.py`](https://github.com/ml-explore/mlx/blob/main/benchmarks/python/single_ops.py) provides ready-to-run benchmarks covering arithmetic, reductions, reshapes, and neural network primitives. Each test generates random input tensors, optionally moves them to the target device using `mx.set_default_device(mx.gpu)` or `mx.cpu`, and reports average latency in milliseconds.

### Comparative Driver in [`compare.py`](https://github.com/ml-explore/mlx/blob/main/compare.py)

The [`benchmarks/python/comparative/compare.py`](https://github.com/ml-explore/mlx/blob/main/benchmarks/python/comparative/compare.py) script orchestrates head-to-head performance tests. It spawns subprocesses running [`bench_mlx.py`](https://github.com/ml-explore/mlx/blob/main/bench_mlx.py) and [`bench_torch.py`](https://github.com/ml-explore/mlx/blob/main/bench_torch.py) with identical parameters, parses their stdout as floating-point timings, and calculates relative speed-up using the formula `(t_torch - t_mlx) / t_torch`. Positive values indicate MLX is faster; negative values favor PyTorch.

## Handling Lazy Evaluation in Timing Loops

MLX's lazy execution model requires specific handling to obtain accurate measurements. Before timing begins, the benchmark scripts run warm-up iterations to load kernels and mitigate first-run latency spikes. During the timed section, each iteration calls `mx.eval` on both the inputs and the operation result to force immediate materialization of the computation graph. Without this explicit synchronization, benchmarks would only measure graph construction overhead rather than actual kernel execution time.

## Running MLX Benchmarks

You can execute benchmarks from the command line to measure specific operations or compare full workloads.

### Benchmark Single Operators on CPU

To measure baseline CPU performance for basic operations like `add` or `matmul`:

```bash
python benchmarks/python/single_ops.py

```

This defaults to CPU execution and outputs timing lines such as:

```

Timing add ... 1.23845 msec
Timing matmul ... 4.51234 msec

```

### Benchmark on Metal GPU

Switch to GPU acceleration by passing the `--gpu` flag, which internally calls `mx.set_default_device(mx.gpu)`:

```bash
python benchmarks/python/single_ops.py --gpu

```

### Compare MLX vs. PyTorch Performance

Use the comparative driver to calculate relative speed-up. For example, to benchmark the `add` operator with specific tensor dimensions:

```bash
python benchmarks/python/comparative/compare.py \
    --filter add --size 32x1024x1024

```

Sample output:

```

0.1523    add --size 32x1024x1024

```

This indicates MLX is approximately 15.2% faster than PyTorch for this operation.

### Testing Different Data Types and Fused Operations

You can test multiple dtypes and fused implementations using additional flags:

```bash
python benchmarks/python/comparative/compare.py \
    -d float32 float16 \
    --filter linear --size 1024x1024 --fused

```

The `-d` flag iterates over specified data types, while `--fused` selects optimized implementations like `mx.addmm` for linear layers.

## Extending the Benchmark Suite

To add a custom operator, define a function in [`bench_mlx.py`](https://github.com/ml-explore/mlx/blob/main/bench_mlx.py) following the existing pattern:

```python
def custom_op(x):
    y = x
    for _ in range(100):
        y = mx.custom_op(y)
    mx.eval(y)
    return y

# In the __main__ block:

elif args.benchmark == "custom":
    print(bench(custom_op, x))

```

Create a matching implementation in [`bench_torch.py`](https://github.com/ml-explore/mlx/blob/main/bench_torch.py), then invoke [`compare.py`](https://github.com/ml-explore/mlx/blob/main/compare.py) with the new benchmark name to obtain comparative metrics.

## Summary

- **Lazy evaluation requires explicit synchronization**: Always call `mx.eval` on inputs and results to force kernel execution before measuring time.
- **Warm-up iterations matter**: The suite runs 5 warm-up iterations to prime the Metal runtime and stabilize timings.
- **Modular architecture**: Use [`time_utils.py`](https://github.com/ml-explore/mlx/blob/main/time_utils.py) for timing logic, [`single_ops.py`](https://github.com/ml-explore/mlx/blob/main/single_ops.py) for quick profiling, and [`compare.py`](https://github.com/ml-explore/mlx/blob/main/compare.py) for cross-framework analysis.
- **Device flexibility**: Toggle between CPU and GPU benchmarks using `mx.set_default_device()` or command-line flags like `--gpu`.
- **Relative speed-up calculation**: The comparative driver reports `(t_torch - t_mlx) / t_torch`, where positive values indicate MLX performance advantages.

## Frequently Asked Questions

### Why do MLX benchmarks require `mx.eval` calls?

MLX uses lazy evaluation to build computation graphs before execution. Without calling `mx.eval`, operations are only queued in the graph rather than executed. The benchmarking suite explicitly evaluates tensors to ensure measured time reflects actual kernel execution on the CPU or Metal GPU, not just graph construction overhead.

### How many iterations does the MLX benchmark suite run?

The suite runs 5 warm-up iterations to initialize the runtime and load kernels, followed by 100 timed iterations for most operations. Heavy operations like matrix multiplication use 10 iterations to keep total runtime reasonable while maintaining statistical significance.

### Can I benchmark custom MLX operators against PyTorch?

Yes. Add your operator to [`benchmarks/python/comparative/bench_mlx.py`](https://github.com/ml-explore/mlx/blob/main/benchmarks/python/comparative/bench_mlx.py) following the existing function signature pattern, create an equivalent PyTorch implementation in [`bench_torch.py`](https://github.com/ml-explore/mlx/blob/main/bench_torch.py), and run [`compare.py`](https://github.com/ml-explore/mlx/blob/main/compare.py) with your benchmark name. The comparative driver automatically calculates relative performance differences.

### What does the output value mean in [`compare.py`](https://github.com/ml-explore/mlx/blob/main/compare.py)?

The output represents the relative speed-up of MLX compared to PyTorch, calculated as `(t_torch - t_mlx) / t_torch`. A positive value (e.g., 0.15) means MLX is 15% faster; a negative value indicates PyTorch is faster for that specific operation and configuration.