How the pgrust Query Optimizer Works: A Deep Dive into the Rust Implementation
The pgrust query optimizer re-implements PostgreSQL's planner in safe Rust using a modular crate-based architecture that preserves the original C code's pipeline while exposing distinct optimizer stages through an "owner seam" boundary system.
The pgrust project (malisper/pgrust) provides a complete Rust rewrite of PostgreSQL's query execution engine, with the pgrust query optimizer serving as the core component that transforms parsed SQL into executable query plans. Unlike traditional database optimizers written in unsafe C, pgrust leverages Rust's memory safety guarantees while maintaining architectural parity with PostgreSQL's proven planner design. The optimizer operates through a collection of specialized backend-optimizer-util crates that each handle a specific phase of the planning process.
Entry Point and the Planner Hook
When a query arrives, the top-level entry point is the planner hook installed by the simple_query module. In crates/backend/tcop/postgres/src/simple_query.rs, the pgss_planner function forwards the parsed query to the Rust planner implementation:
pub fn pgss_planner<'mcx>(… ) -> … {
planner_seams::standard_planner::call(mcx, parse, query_string,…)
}
This hook mirrors PostgreSQL's standard_planner interface. After parsing, the query enters the optimizer arena, a memory context that manages the entire planner lifecycle with deterministic allocation and deallocation patterns.
The Optimizer Pipeline Stages
The pgrust query optimizer follows PostgreSQL's classic pipeline, implementing each stage as a separate crate under the backend/optimizer/util/ directory. Each crate declares an owner seam (e.g., backend-optimizer-util-var) that provides a safe-Rust boundary mimicking PostgreSQL's modular architecture.
Variable Handling and Target Lists
The first stage normalizes variable references and builds target structures. In crates/backend/optimizer/util/vars/src/var.rs, the optimizer flattens and normalizes Var nodes through functions like flatten_group_exprs. Subsequently, crates/backend/optimizer/util/vars/src/tlist.rs constructs PathTarget objects that represent the query's target-list manipulation, equivalent to PostgreSQL's optimizer/util/tlist.c.
Catalog Lookups and Statistics
Before costing operations, the optimizer retrieves table metadata. The crates/backend/optimizer/util/plancat/src/lib.rs crate handles catalog look-ups, fetching size, width, and statistics via functions such as get_rel_data_width and get_typavgwidth. These statistics feed directly into the cost model's selectivity calculations.
Cost Estimation and Clause Simplification
The cost model resides in crates/backend/optimizer/util/clauses/src/lib.rs. This crate performs constant folding, selectivity estimation (eq_sel, range_sel), and cost calculations using GUC-defined variables. The optimizer reads cost constants from crates/backend/utils/misc/guc_tables/src/vars.rs:
pub static cpu_operator_cost: GucRealVar = GucSlot::new("cpu_operator_cost");
These GUC variables—including seq_page_cost and random_page_cost—determine the final cost assigned to each potential execution path.
Join Ordering and Path Selection
Join ordering logic lives in crates/backend/optimizer/util/joininfo/src/lib.rs. The generate_join_paths function builds a graph of RestrictInfo objects, evaluates possible join orders, and selects the cheapest tree using the cost model:
pub fn generate_join_paths(root: &mut PlannerInfo, …) { … }
Concurrently, crates/backend/optimizer/util/pathnode/src/lib.rs enumerates viable access paths—sequential scans, index scans, and bitmap indexes—assigning costs to each alternative before the cheapest path is selected for the final plan.
Configuring the Optimizer with GUC Variables
You can tune the pgrust query optimizer at runtime by manipulating GUC variables exposed through the guc_tables module. This example demonstrates enabling sequential scans and adjusting CPU operator costs:
use pgrust::client::Client;
use pgrust::utils::guc_tables::vars::{enable_seqscan, cpu_operator_cost};
fn main() -> Result<(), Box<dyn std::error::Error>> {
let mut client = Client::connect("host=localhost dbname=test", None)?;
enable_seqscan::set(true);
cpu_operator_cost::set(0.001);
let rows = client.query(
"SELECT * FROM orders o JOIN customers c ON o.cust_id = c.id WHERE o.amount > 100",
&[],
)?;
println!("Returned {} rows", rows.len());
Ok(())
}
Key files: crates/backend/tcop/postgres/src/simple_query.rs (entry point), crates/backend/utils/misc/guc_tables/src/vars.rs (GUC handling).
Debugging Query Plans
To inspect the optimizer's decisions, enable planner statistics logging before executing EXPLAIN:
use pgrust::client::Client;
use pgrust::utils::guc_tables::vars::log_planner_stats;
fn main() -> Result<(), Box<dyn std::error::Error>> {
let mut client = Client::connect("host=localhost dbname=test", None)?;
log_planner_stats::set(true);
client.query("EXPLAIN (VERBOSE) SELECT * FROM orders", &[])?;
// Detailed planner statistics are logged server-side
Ok(())
}
This invokes the same planner hook in crates/backend/tcop/postgres/src/simple_query.rs, but with additional instrumentation enabled via the GUC system.
Summary
- The pgrust query optimizer is a Rust rewrite of PostgreSQL's planner that maintains the original architecture while enforcing memory safety.
- Entry point: The
pgss_plannerhook insimple_query.rsforwards queries to the Rust implementation. - Modular pipeline: Eight distinct stages (var handling, target-list creation, relid computation, catalog lookups, clause simplification, join info, path generation, and cost selection) each reside in separate crates under
backend/optimizer/util/. - Cost model: Uses PostgreSQL-compatible GUC variables (
cpu_operator_cost,seq_page_cost) defined inguc_tablesto evaluate path costs. - Join ordering: The
joininfocrate constructsRestrictInfographs and selects optimal join trees using cost-based heuristics. - Extensibility: The "owner seam" architecture allows adding new optimizer modules without modifying core logic.
Frequently Asked Questions
How does the pgrust query optimizer differ from PostgreSQL's original C implementation?
The pgrust query optimizer preserves PostgreSQL's algorithmic logic and GUC-based configuration system but re-implements everything in safe Rust using a crate-based modular structure. While PostgreSQL uses a monolithic C codebase with manual memory management, pgrust employs the optimizer arena pattern with Rust's ownership model to eliminate buffer overflows and use-after-free vulnerabilities.
What is the "owner seam" architecture in pgrust?
The owner seam architecture is pgrust's method for creating safe boundaries between optimizer components. Each crate (e.g., backend-optimizer-util-var) declares itself as an owner of specific functionality, exposing interfaces that other crates can call without accessing internal implementation details. This mirrors PostgreSQL's hook system but enforces Rust's compile-time safety guarantees while allowing unit testing of individual planner stages.
Can I extend the pgrust query optimizer with custom logic?
Yes. According to the pgrust source code, you can extend the optimizer by adding new seam crates (e.g., backend-optimizer-util-pathnode-seams) that override or augment existing behavior without modifying core files like crates/backend/optimizer/util/pathnode/src/lib.rs. This design supports custom cost models and access methods while maintaining the stability of the base pipeline.
How does pgrust handle GUC variables for plan tuning?
The optimizer reads GUC variables such as enable_seqscan, cpu_operator_cost, and log_planner_stats from the guc_tables module in crates/backend/utils/misc/guc_tables/src/vars.rs. These variables control cost constants and feature flags at runtime, allowing you to force specific join algorithms or adjust cost weights to influence the planner's decisions without recompiling the database.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →