Monty Performance Comparison to CPython: Startup Latency and Execution Speed Analysis
Monty delivers microsecond-scale cold-start latency (≈60 µs) and steady-state execution speeds within 0.5× to 5× of native CPython, making it ideal for sandboxed, high-frequency Python execution.
The pydantic/monty repository provides a Rust-based Python interpreter designed for secure, embedded execution. Unlike traditional CPython integrations that rely on subprocess calls or WebAssembly containers, Monty runs as an in-process library, fundamentally altering the performance characteristics for AI-generated code and microservice workloads.
Cold-Start Latency: Microseconds vs. Milliseconds
Monty eliminates the process-spawning overhead that plagues typical CPython deployments. According to the benchmark script [scripts/startup_performance.py](https://github.com/pydantic/monty/blob/main/scripts/startup_performance.py), a simple Monty('1 + 1').run() call completes in approximately 0.06 ms (60 microseconds).
By comparison, common CPython deployment strategies exhibit significantly higher latency:
- Docker container: ≈195 ms
- Sub-process Python: ≈30 ms
- Pyodide (WebAssembly): ≈2.8 seconds
This three-order-of-magnitude improvement in startup time makes Monty suitable for serverless functions and real-time code evaluation where every millisecond counts.
Steady-State Execution Performance
Once initialized, Monty's bytecode interpreter delivers competitive performance against CPython through PyO3 bindings. The benchmark suite in [crates/monty/benches/main.rs](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs) demonstrates a performance range of 5× faster to 5× slower than CPython, depending on workload characteristics.
Micro-Benchmark Results
The main.rs benchmark file defines several canonical micro-benchmarks executed via run_monty and run_cpython functions:
add_two(1 + 2): ≈0.3 µs (Monty) vs. ≈1–2 µs (CPython)list_append(append 42): ≈0.4 µs (Monty) vs. ≈1–2 µs (CPython)loop_mod_13(10,000-iteration loop): ≈6 µs (Monty) vs. comparable CPython timeskitchen_sink(full-feature script): ≈50 µs (Monty) vs. ≈120 µs (CPython)
The kitchen-sink benchmark, located in [crates/monty/test_cases/bench__kitchen_sink.py](https://github.com/pydantic/monty/blob/main/crates/monty/test_cases/bench__kitchen_sink.py), exercises most supported language features and shows Monty completing in roughly 40–50% of the time required by CPython for equivalent work.
Performance Variance Factors
Monty excels in tight loops and arithmetic operations where its zero-copy heap and lack of C-FFI overhead dominate. For larger workloads involving complex comprehensions or heavy object manipulation, the gap narrows as both interpreters spend proportionally more time in similar bytecode operations.
Why Monty Achieves These Performance Characteristics
The performance advantages stem from architectural decisions in the Rust implementation:
In-process, native Rust execution eliminates the cross-process marshaling required by Docker or subprocess-based Python execution. The interpreter lives in the same memory space as the host application.
Zero-copy data structures via the custom heap implementation in [crates/monty/src/heap.rs](https://github.com/pydantic/monty/blob/main/crates/monty/src/heap.rs) use manual reference counting tuned for fast allocation and deallocation, reducing garbage collection pauses.
Bytecode VM mirroring CPython as implemented in [crates/monty/src/bytecode.rs](https://github.com/pydantic/monty/blob/main/crates/monty/src/bytecode.rs) parses Python source once into a compact representation. Subsequent iterations execute only the bytecode, similar to CPython’s interpreter loop but without the GIL or dynamic linking overhead.
No dynamic linking to the CPython runtime means all built-ins and standard-library shims are pure Rust functions (see crate::builtins), eliminating C-FFI call overhead entirely.
Selective feature set supporting only a subset of Python (no full stdlib, no third-party modules) keeps the dispatch path short and avoids complex import machinery that would otherwise slow down execution.
Benchmark Methodology and Reproducibility
The performance data derives from rigorous benchmarking in [crates/monty/benches/main.rs](https://github.com/pydantic/monty/blob/main/crates/monty/benches/main.rs), using the Criterion.rs framework for statistical rigor.
Monty execution path creates a MontyRun instance once via MontyRun::new, then calls run_no_limits repeatedly. This isolates execution cost from parsing overhead:
fn run_monty(bench: &mut Bencher, code: &str, expected: i64) {
let ex = MontyRun::new(code.to_owned(), "test.py", vec![], vec![]).unwrap();
let r = ex.run_no_limits(vec![]).unwrap();
let int_value: i64 = r.as_ref().try_into().unwrap();
assert_eq!(int_value, expected);
bench.iter(|| {
let r = ex.run_no_limits(vec![]).unwrap();
let int_value: i64 = r.as_ref().try_into().unwrap();
black_box(int_value);
});
}
CPython execution path uses PyO3 bindings. The Python snippet is wrapped in a def main(): function via wrap_for_cpython so the last expression becomes a return value. The function object is cached across iterations and invoked via fun.call0(py):
fn run_cpython(bench: &mut Bencher, code: &str) {
Python::with_gil(|py| {
let wrapped = wrap_for_cpython(code);
let fun: &PyFunction = py.eval(&wrapped, None, None).unwrap().extract().unwrap();
bench.iter(|| {
let result = fun.call0(py).unwrap();
black_box(result);
});
});
}
Both benchmarks assert the correct result to prevent dead-code elimination and use black_box to inhibit compiler optimizations that would remove the timing loop.
Summary
- Monty achieves approximately 60 µs cold-start latency, three orders of magnitude faster than Dockerized CPython (195 ms) and Pyodide (2.8 s).
- Steady-state execution ranges from 5× faster to 5× slower than CPython, with micro-benchmarks like arithmetic and list operations typically showing 3–5× speedups.
- Architectural advantages include in-process Rust execution, zero-copy heaps (
crates/monty/src/heap.rs), and a bytecode VM (crates/monty/src/bytecode.rs) without CPython's GIL or FFI overhead. - Benchmarking methodology in
crates/monty/benches/main.rsuses Criterion.rs with isolated parsing and execution phases to ensure reproducible, statistically valid results.
Frequently Asked Questions
How does Monty achieve sub-millisecond startup times compared to CPython?
Monty runs as an in-process Rust library rather than a separate process or container. According to scripts/startup_performance.py, a Monty('1 + 1').run() call completes in roughly 60 microseconds because it avoids the operating system overhead of spawning subprocesses, initializing Docker containers, or loading WebAssembly modules that CPython deployments typically require.
Is Monty faster than CPython for all Python workloads?
No. The benchmarks in crates/monty/benches/main.rs show a performance range of 5× faster to 5× slower depending on the workload. Monty excels at tight loops and arithmetic operations (often 3–5× faster) but may fall behind CPython on larger workloads involving complex object manipulation or operations that benefit from CPython's highly optimized C implementations for built-in types.
What architectural decisions enable Monty's execution speed?
Three key implementation details in the pydantic/monty repository drive performance: the zero-copy heap in crates/monty/src/heap.rs using manual reference counting, the bytecode VM in crates/monty/src/bytecode.rs that mirrors CPython's approach without the GIL, and the absence of dynamic linking to the CPython runtime, eliminating C-FFI overhead by implementing built-ins as pure Rust functions.
How are the Monty vs CPython benchmarks conducted?
The benchmark suite in crates/monty/benches/main.rs uses the Criterion.rs framework to measure both interpreters. For Monty, it creates a MontyRun instance once and repeatedly calls run_no_limits. For CPython, it uses PyO3 bindings to cache a compiled function object and invoke it via fun.call0(py). Both paths use black_box to prevent compiler optimizations and assert correct results to avoid dead-code elimination, ensuring statistically valid comparisons.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →