# Monty Performance Comparison to CPython: Startup Latency and Execution Speed Analysis

> Discover the Monty performance comparison to CPython. Analyze microsecond cold-start latency and execution speeds 0.5x to 5x faster than CPython for high-frequency Python execution.

- Repository: [Pydantic/monty](https://github.com/pydantic/monty)
- Tags: performance
- Published: 2026-02-16

---

**Monty delivers microsecond-scale cold-start latency (≈60 µs) and steady-state execution speeds within 0.5× to 5× of native CPython, making it ideal for sandboxed, high-frequency Python execution.**

The `pydantic/monty` repository provides a Rust-based Python interpreter designed for secure, embedded execution. Unlike traditional CPython integrations that rely on subprocess calls or WebAssembly containers, Monty runs as an in-process library, fundamentally altering the performance characteristics for AI-generated code and microservice workloads.

## Cold-Start Latency: Microseconds vs. Milliseconds

Monty eliminates the process-spawning overhead that plagues typical CPython deployments. According to the benchmark script [[`scripts/startup_performance.py`](https://github.com/pydantic/monty/blob/main/scripts/startup_performance.py)](https://github.com/pydantic/monty/blob/main/scripts/startup_performance.py), a simple `Monty('1 + 1').run()` call completes in approximately **0.06 ms (60 microseconds)**.

By comparison, common CPython deployment strategies exhibit significantly higher latency:

- **Docker container**: ≈195 ms
- **Sub-process Python**: ≈30 ms  
- **Pyodide (WebAssembly)**: ≈2.8 seconds

This three-order-of-magnitude improvement in startup time makes Monty suitable for serverless functions and real-time code evaluation where every millisecond counts.

## Steady-State Execution Performance

Once initialized, Monty's bytecode interpreter delivers competitive performance against CPython through PyO3 bindings. The benchmark suite in [[`crates/monty/benches/main.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs)](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs) demonstrates a performance range of **5× faster to 5× slower** than CPython, depending on workload characteristics.

### Micro-Benchmark Results

The [`main.rs`](https://github.com/pydantic/monty/blob/main/main.rs) benchmark file defines several canonical micro-benchmarks executed via `run_monty` and `run_cpython` functions:

- **`add_two`** (1 + 2): ≈0.3 µs (Monty) vs. ≈1–2 µs (CPython)
- **`list_append`** (append 42): ≈0.4 µs (Monty) vs. ≈1–2 µs (CPython)  
- **`loop_mod_13`** (10,000-iteration loop): ≈6 µs (Monty) vs. comparable CPython times
- **`kitchen_sink`** (full-feature script): ≈50 µs (Monty) vs. ≈120 µs (CPython)

The **kitchen-sink** benchmark, located in [[`crates/monty/test_cases/bench__kitchen_sink.py`](https://github.com/pydantic/monty/blob/main/crates/monty/test_cases/bench__kitchen_sink.py)](https://github.com/pydantic/monty/blob/main/crates/monty/test_cases/bench__kitchen_sink.py), exercises most supported language features and shows Monty completing in roughly **40–50% of the time** required by CPython for equivalent work.

### Performance Variance Factors

Monty excels in **tight loops and arithmetic operations** where its zero-copy heap and lack of C-FFI overhead dominate. For **larger workloads** involving complex comprehensions or heavy object manipulation, the gap narrows as both interpreters spend proportionally more time in similar bytecode operations.

## Why Monty Achieves These Performance Characteristics

The performance advantages stem from architectural decisions in the Rust implementation:

**In-process, native Rust execution** eliminates the cross-process marshaling required by Docker or subprocess-based Python execution. The interpreter lives in the same memory space as the host application.

**Zero-copy data structures** via the custom heap implementation in [[`crates/monty/src/heap.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/src/heap.rs)](https://github.com/pydantic/monty/blob/main/crates/monty/src/heap.rs) use manual reference counting tuned for fast allocation and deallocation, reducing garbage collection pauses.

**Bytecode VM mirroring CPython** as implemented in [[`crates/monty/src/bytecode.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/src/bytecode.rs)](https://github.com/pydantic/monty/blob/main/crates/monty/src/bytecode.rs) parses Python source once into a compact representation. Subsequent iterations execute only the bytecode, similar to CPython’s interpreter loop but without the GIL or dynamic linking overhead.

**No dynamic linking to the CPython runtime** means all built-ins and standard-library shims are pure Rust functions (see `crate::builtins`), eliminating C-FFI call overhead entirely.

**Selective feature set** supporting only a subset of Python (no full stdlib, no third-party modules) keeps the dispatch path short and avoids complex import machinery that would otherwise slow down execution.

## Benchmark Methodology and Reproducibility

The performance data derives from rigorous benchmarking in [[`crates/monty/benches/main.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs)](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs), using the Criterion.rs framework for statistical rigor.

**Monty execution path** creates a `MontyRun` instance once via `MontyRun::new`, then calls `run_no_limits` repeatedly. This isolates execution cost from parsing overhead:

```rust
fn run_monty(bench: &mut Bencher, code: &str, expected: i64) {
    let ex = MontyRun::new(code.to_owned(), "test.py", vec![], vec![]).unwrap();
    let r = ex.run_no_limits(vec![]).unwrap();
    let int_value: i64 = r.as_ref().try_into().unwrap();
    assert_eq!(int_value, expected);
    bench.iter(|| {
        let r = ex.run_no_limits(vec![]).unwrap();
        let int_value: i64 = r.as_ref().try_into().unwrap();
        black_box(int_value);
    });
}

```

**CPython execution path** uses PyO3 bindings. The Python snippet is wrapped in a `def main():` function via `wrap_for_cpython` so the last expression becomes a return value. The function object is cached across iterations and invoked via `fun.call0(py)`:

```rust
fn run_cpython(bench: &mut Bencher, code: &str) {
    Python::with_gil(|py| {
        let wrapped = wrap_for_cpython(code);
        let fun: &PyFunction = py.eval(&wrapped, None, None).unwrap().extract().unwrap();
        bench.iter(|| {
            let result = fun.call0(py).unwrap();
            black_box(result);
        });
    });
}

```

Both benchmarks assert the correct result to prevent dead-code elimination and use `black_box` to inhibit compiler optimizations that would remove the timing loop.

## Summary

- **Monty achieves approximately 60 µs cold-start latency**, three orders of magnitude faster than Dockerized CPython (195 ms) and Pyodide (2.8 s).
- **Steady-state execution ranges from 5× faster to 5× slower than CPython**, with micro-benchmarks like arithmetic and list operations typically showing 3–5× speedups.
- **Architectural advantages** include in-process Rust execution, zero-copy heaps ([`crates/monty/src/heap.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/src/heap.rs)), and a bytecode VM ([`crates/monty/src/bytecode.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/src/bytecode.rs)) without CPython's GIL or FFI overhead.
- **Benchmarking methodology** in [`crates/monty/benches/main.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs) uses Criterion.rs with isolated parsing and execution phases to ensure reproducible, statistically valid results.

## Frequently Asked Questions

### How does Monty achieve sub-millisecond startup times compared to CPython?

Monty runs as an in-process Rust library rather than a separate process or container. According to [`scripts/startup_performance.py`](https://github.com/pydantic/monty/blob/main/scripts/startup_performance.py), a `Monty('1 + 1').run()` call completes in roughly 60 microseconds because it avoids the operating system overhead of spawning subprocesses, initializing Docker containers, or loading WebAssembly modules that CPython deployments typically require.

### Is Monty faster than CPython for all Python workloads?

No. The benchmarks in [`crates/monty/benches/main.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs) show a performance range of 5× faster to 5× slower depending on the workload. Monty excels at tight loops and arithmetic operations (often 3–5× faster) but may fall behind CPython on larger workloads involving complex object manipulation or operations that benefit from CPython's highly optimized C implementations for built-in types.

### What architectural decisions enable Monty's execution speed?

Three key implementation details in the `pydantic/monty` repository drive performance: the **zero-copy heap** in [`crates/monty/src/heap.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/src/heap.rs) using manual reference counting, the **bytecode VM** in [`crates/monty/src/bytecode.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/src/bytecode.rs) that mirrors CPython's approach without the GIL, and the **absence of dynamic linking** to the CPython runtime, eliminating C-FFI overhead by implementing built-ins as pure Rust functions.

### How are the Monty vs CPython benchmarks conducted?

The benchmark suite in [`crates/monty/benches/main.rs`](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs) uses the Criterion.rs framework to measure both interpreters. For Monty, it creates a `MontyRun` instance once and repeatedly calls `run_no_limits`. For CPython, it uses PyO3 bindings to cache a compiled function object and invoke it via `fun.call0(py)`. Both paths use `black_box` to prevent compiler optimizations and assert correct results to avoid dead-code elimination, ensuring statistically valid comparisons.