# How the pgrust Query Optimizer Works: A Deep Dive into the Rust Implementation

> Explore how the pgrust query optimizer leverages Rust's safety and modularity to rebuild PostgreSQL's planner, offering a unique perspective on query optimization.

- Repository: [Michael Malis/pgrust](https://github.com/malisper/pgrust)
- Tags: deep-dive
- Published: 2026-07-13

---

**The pgrust query optimizer re-implements PostgreSQL's planner in safe Rust using a modular crate-based architecture that preserves the original C code's pipeline while exposing distinct optimizer stages through an "owner seam" boundary system.**

The pgrust project (`malisper/pgrust`) provides a complete Rust rewrite of PostgreSQL's query execution engine, with the pgrust query optimizer serving as the core component that transforms parsed SQL into executable query plans. Unlike traditional database optimizers written in unsafe C, pgrust leverages Rust's memory safety guarantees while maintaining architectural parity with PostgreSQL's proven planner design. The optimizer operates through a collection of specialized `backend-optimizer-util` crates that each handle a specific phase of the planning process.

## Entry Point and the Planner Hook

When a query arrives, the top-level entry point is the planner hook installed by the `simple_query` module. In [`crates/backend/tcop/postgres/src/simple_query.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/tcop/postgres/src/simple_query.rs), the `pgss_planner` function forwards the parsed query to the Rust planner implementation:

```rust
pub fn pgss_planner<'mcx>(… ) -> … {
    planner_seams::standard_planner::call(mcx, parse, query_string,…)
}

```

This hook mirrors PostgreSQL's `standard_planner` interface. After parsing, the query enters the **optimizer arena**, a memory context that manages the entire planner lifecycle with deterministic allocation and deallocation patterns.

## The Optimizer Pipeline Stages

The pgrust query optimizer follows PostgreSQL's classic pipeline, implementing each stage as a separate crate under the `backend/optimizer/util/` directory. Each crate declares an **owner seam** (e.g., `backend-optimizer-util-var`) that provides a safe-Rust boundary mimicking PostgreSQL's modular architecture.

### Variable Handling and Target Lists

The first stage normalizes variable references and builds target structures. In [`crates/backend/optimizer/util/vars/src/var.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/optimizer/util/vars/src/var.rs), the optimizer flattens and normalizes `Var` nodes through functions like `flatten_group_exprs`. Subsequently, [`crates/backend/optimizer/util/vars/src/tlist.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/optimizer/util/vars/src/tlist.rs) constructs `PathTarget` objects that represent the query's target-list manipulation, equivalent to PostgreSQL's [`optimizer/util/tlist.c`](https://github.com/malisper/pgrust/blob/main/optimizer/util/tlist.c).

### Catalog Lookups and Statistics

Before costing operations, the optimizer retrieves table metadata. The [`crates/backend/optimizer/util/plancat/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/optimizer/util/plancat/src/lib.rs) crate handles catalog look-ups, fetching size, width, and statistics via functions such as `get_rel_data_width` and `get_typavgwidth`. These statistics feed directly into the cost model's selectivity calculations.

### Cost Estimation and Clause Simplification

The cost model resides in [`crates/backend/optimizer/util/clauses/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/optimizer/util/clauses/src/lib.rs). This crate performs constant folding, selectivity estimation (`eq_sel`, `range_sel`), and cost calculations using GUC-defined variables. The optimizer reads cost constants from [`crates/backend/utils/misc/guc_tables/src/vars.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/utils/misc/guc_tables/src/vars.rs):

```rust
pub static cpu_operator_cost: GucRealVar = GucSlot::new("cpu_operator_cost");

```

These GUC variables—including `seq_page_cost` and `random_page_cost`—determine the final cost assigned to each potential execution path.

### Join Ordering and Path Selection

Join ordering logic lives in [`crates/backend/optimizer/util/joininfo/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/optimizer/util/joininfo/src/lib.rs). The `generate_join_paths` function builds a graph of `RestrictInfo` objects, evaluates possible join orders, and selects the cheapest tree using the cost model:

```rust
pub fn generate_join_paths(root: &mut PlannerInfo, …) { … }

```

Concurrently, [`crates/backend/optimizer/util/pathnode/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/optimizer/util/pathnode/src/lib.rs) enumerates viable access paths—sequential scans, index scans, and bitmap indexes—assigning costs to each alternative before the cheapest path is selected for the final plan.

## Configuring the Optimizer with GUC Variables

You can tune the pgrust query optimizer at runtime by manipulating GUC variables exposed through the `guc_tables` module. This example demonstrates enabling sequential scans and adjusting CPU operator costs:

```rust
use pgrust::client::Client;
use pgrust::utils::guc_tables::vars::{enable_seqscan, cpu_operator_cost};

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut client = Client::connect("host=localhost dbname=test", None)?;

    enable_seqscan::set(true);
    cpu_operator_cost::set(0.001);

    let rows = client.query(
        "SELECT * FROM orders o JOIN customers c ON o.cust_id = c.id WHERE o.amount > 100",
        &[],
    )?;
    
    println!("Returned {} rows", rows.len());
    Ok(())
}

```

*Key files*: [`crates/backend/tcop/postgres/src/simple_query.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/tcop/postgres/src/simple_query.rs) (entry point), [`crates/backend/utils/misc/guc_tables/src/vars.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/utils/misc/guc_tables/src/vars.rs) (GUC handling).

## Debugging Query Plans

To inspect the optimizer's decisions, enable planner statistics logging before executing `EXPLAIN`:

```rust
use pgrust::client::Client;
use pgrust::utils::guc_tables::vars::log_planner_stats;

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut client = Client::connect("host=localhost dbname=test", None)?;

    log_planner_stats::set(true);

    client.query("EXPLAIN (VERBOSE) SELECT * FROM orders", &[])?;
    // Detailed planner statistics are logged server-side
    Ok(())
}

```

This invokes the same planner hook in [`crates/backend/tcop/postgres/src/simple_query.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/tcop/postgres/src/simple_query.rs), but with additional instrumentation enabled via the GUC system.

## Summary

- **The pgrust query optimizer** is a Rust rewrite of PostgreSQL's planner that maintains the original architecture while enforcing memory safety.
- **Entry point**: The `pgss_planner` hook in [`simple_query.rs`](https://github.com/malisper/pgrust/blob/main/simple_query.rs) forwards queries to the Rust implementation.
- **Modular pipeline**: Eight distinct stages (var handling, target-list creation, relid computation, catalog lookups, clause simplification, join info, path generation, and cost selection) each reside in separate crates under `backend/optimizer/util/`.
- **Cost model**: Uses PostgreSQL-compatible GUC variables (`cpu_operator_cost`, `seq_page_cost`) defined in `guc_tables` to evaluate path costs.
- **Join ordering**: The `joininfo` crate constructs `RestrictInfo` graphs and selects optimal join trees using cost-based heuristics.
- **Extensibility**: The "owner seam" architecture allows adding new optimizer modules without modifying core logic.

## Frequently Asked Questions

### How does the pgrust query optimizer differ from PostgreSQL's original C implementation?

The pgrust query optimizer preserves PostgreSQL's algorithmic logic and GUC-based configuration system but re-implements everything in safe Rust using a crate-based modular structure. While PostgreSQL uses a monolithic C codebase with manual memory management, pgrust employs the **optimizer arena** pattern with Rust's ownership model to eliminate buffer overflows and use-after-free vulnerabilities.

### What is the "owner seam" architecture in pgrust?

The owner seam architecture is pgrust's method for creating safe boundaries between optimizer components. Each crate (e.g., `backend-optimizer-util-var`) declares itself as an owner of specific functionality, exposing interfaces that other crates can call without accessing internal implementation details. This mirrors PostgreSQL's hook system but enforces Rust's compile-time safety guarantees while allowing unit testing of individual planner stages.

### Can I extend the pgrust query optimizer with custom logic?

Yes. According to the pgrust source code, you can extend the optimizer by adding new **seam crates** (e.g., `backend-optimizer-util-pathnode-seams`) that override or augment existing behavior without modifying core files like [`crates/backend/optimizer/util/pathnode/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/optimizer/util/pathnode/src/lib.rs). This design supports custom cost models and access methods while maintaining the stability of the base pipeline.

### How does pgrust handle GUC variables for plan tuning?

The optimizer reads GUC variables such as `enable_seqscan`, `cpu_operator_cost`, and `log_planner_stats` from the `guc_tables` module in [`crates/backend/utils/misc/guc_tables/src/vars.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/utils/misc/guc_tables/src/vars.rs). These variables control cost constants and feature flags at runtime, allowing you to force specific join algorithms or adjust cost weights to influence the planner's decisions without recompiling the database.