How the pgrust SQL Parser Handles Query Parsing: A Deep Dive into the PostgreSQL 18.3 Port
The pgrust SQL parser processes queries through a memory-safe Rust pipeline that wraps PostgreSQL 18.3's Bison grammar, using a lexer with one-token look-ahead to feed tokens into base_yyparse, then converting the resulting C-style parse trees into owned Rust structures via arena-based lifetimes.
The pgrust SQL parser is a Rust port of the PostgreSQL 18.3 parser maintained in the malisper/pgrust repository. It provides a safe, idiomatic interface to the same Bison grammar that powers PostgreSQL, converting raw SQL strings into typed abstract syntax trees while eliminating memory safety risks through explicit arena management. This article examines how the parser handles lexical analysis, grammar processing, and memory management based on the actual implementation in the source code.
The Parser Pipeline: From Raw SQL to Abstract Syntax Tree
The pgrust SQL parser follows a four-stage pipeline that mirrors PostgreSQL's architecture while wrapping each component in Rust's type system and memory safety guarantees.
Entry Point: raw_parser()
The public API surface begins at raw_parser() in crates/backend/parser/driver/src/lib.rs. This function serves as the primary interface for all SQL parsing operations:
pub fn raw_parser<'mcx>(
mcx: Mcx<'mcx>,
str_: &'mcx str,
mode: RawParseMode,
) -> PgResult<PgVec<'mcx, RawStmt<'mcx>>> {
gram::base_yyparse::call(mcx, str_, mode)
}
The function accepts a memory context arena (Mcx), the SQL string, and a parse mode. It returns a PgResult containing a PgVec of RawStmt nodes, each representing a parsed statement in the query string. This design replaces PostgreSQL's C-style error handling (long jumps) with Rust's explicit Result type.
Lexical Analysis with BaseLexer
Before grammar processing begins, the input string passes through BaseLexer, which wraps the core scanner (core_yylex) exposed via crates/backend/parser/scan_seams/src/lib.rs. The lexer provides:
- One-token look-ahead that enables merging of multi-word tokens (e.g., converting
NOT LIKEintoNOT_LA) - Unicode de-escaping through
str_udescapeandcheck_uescapecharfor handlingUIDENTandUSCONSTtokens - Mode token seeding when non-default parse modes require initial context
The lexer initialization pattern appears in driver/src/lib.rs:
let seed = mode_seed(mcx, mode); // optional CoreToken
let lexer = BaseLexer::new(mcx, scanbuf, seed);
Grammar Processing via base_yyparse
The Bison-generated grammar lives in crates/backend/parser/gram_core/src/lib.rs, which exposes the base_yyparse seam function. This component:
- Consumes tokens from the
BaseLexerinstance - Builds the raw parse tree using PostgreSQL's original grammar rules
- Returns a C-style
List *ofRawStmtstructures
The grammar seam bridges the C parser implementation (auto-translated from gram.y in crates/_support/pgrust/gram_c2rust_fgram/src/lib.rs) with Rust's ownership system, allowing incremental porting while preserving exact PostgreSQL semantics.
Memory Safety and Result Conversion
Unlike the C implementation that relies on longjmp for error handling and raw pointers for memory management, the pgrust SQL parser converts the grammar output into safe Rust structures:
- The
List *from the C parser gets wrapped intoPgVec<'mcx, RawStmt<'mcx>> - All nodes use arena-based lifetimes (
Mcx) tied to the parsing context - Syntax errors propagate as
Err(PgError)rather than triggering non-local jumps
This conversion happens within the base_yyparse wrapper, ensuring that users interact only with fully owned, properly typed Rust structures.
Parse Modes and Look-Ahead Tokens
The pgrust SQL parser supports multiple entry points via RawParseMode, defined in driver/src/lib.rs. Each mode determines the initial parser state:
RAW_PARSE_DEFAULT: Standard SQL query parsing (no seed token)RAW_PARSE_TYPE_NAME: Parses a single type name expression (seeds withMODE_TYPE_NAME, token 774)RAW_PARSE_PLPGSQL_EXPR: Parses PL/pgSQL expressions (seeds withMODE_PLPGSQL_EXPR, token 775)
When a non-default mode is requested, the mode_seed function provides an initial look-ahead token that places the parser in the correct semantic context. This mechanism enables the same grammar to handle diverse inputs ranging from complete SELECT statements to isolated type expressions.
Unicode Handling and Token Merging
The lexer handles PostgreSQL's Unicode escape syntax through dedicated functions in driver/src/lib.rs. When encountering UIDENT (Unicode identifier) or USCONST (Unicode string constant) tokens, the parser:
- Validates escape sequences via
check_uescapechar - Converts escaped values to plain identifiers using
str_udescape - Preserves the original PostgreSQL behavior for Unicode validation and identifier truncation
Additionally, the one-token look-ahead capability handles multi-word keywords common in SQL:2011 syntax. For example, the sequence WITH TIME becomes WITH_LA, allowing the grammar to disambiguate between WITH as a clause starter and WITH TIME ZONE as a modifier.
Practical Usage Examples
To parse a standard SQL query using the pgrust SQL parser:
use pgrust::backend::parser::driver::{raw_parser, RawParseMode};
use pgrust::mcx::Mcx;
// Create a memory arena (required by all parser functions)
let ctx = ::mcx::MemoryContext::new("example");
let mcx = ctx.mcx();
// Parse a simple SELECT
let sql = "SELECT id, name FROM users WHERE active = true;";
let parse_tree = raw_parser(mcx, sql, RawParseMode::RAW_PARSE_DEFAULT)
.expect("SQL parsing failed");
// `parse_tree` is a `PgVec<RawStmt>`; iterate over statements:
for stmt in parse_tree.iter() {
println!("Statement kind: {:?}", stmt.stmt.node_tag());
}
For parsing type names (useful when implementing type resolution):
use pgrust::backend::parser::driver::raw_parse_type_name;
let ty = raw_parse_type_name("numeric(10,2)".to_string())
.expect("type parsing failed");
println!("Parsed type: {:?}", ty);
Key Source Files and Architecture
The parser implementation spans several specialized crates that separate concerns between lexing, grammar, and FFI:
crates/backend/parser/driver/src/lib.rs: Public API (raw_parser), lexer setup, Unicode handling, and mode token managementcrates/backend/parser/gram_core/src/lib.rs: Safe wrapper around the Bison grammar (base_yyparse) and C-to-Rust conversion logiccrates/_support/pgrust/gram_c2rust_fgram/src/lib.rs: Auto-generated C-to-Rust translation of the originalgram.ygrammar filecrates/backend/parser/scan_seams/src/lib.rs: FFI seam exposing the core lexical scanner (core_yylex)crates/backend/parser/gram_seams/src/lib.rs: FFI seam exposingbase_yyparseto the driver cratecrates/backend/parser/gram_fgram/src/tests.rs: End-to-end test suite validating parsing behavior against PostgreSQL semantics
Summary
- The pgrust SQL parser provides a Rust-native interface to PostgreSQL 18.3's grammar through the
raw_parser()function inbackend/parser/driver/src/lib.rs. - Lexical analysis uses
BaseLexerto wrap the core scanner with one-token look-ahead for multi-word token merging and mode-specific seeding. - Grammar processing drives the Bison-generated
base_yyparsevia FFI seams, converting raw C structures into safe Rust types using arena lifetimes (Mcx). - Multiple parse modes (
RAW_PARSE_DEFAULT,RAW_PARSE_TYPE_NAME, etc.) allow the same grammar to handle diverse input contexts through initial look-ahead tokens. - Unicode and error handling follow PostgreSQL semantics but use Rust's
PgResulttype for safe error propagation instead of C-style long jumps.
Frequently Asked Questions
How does pgrust maintain compatibility with PostgreSQL's parser?
The pgrust SQL parser preserves compatibility by directly porting the PostgreSQL 18.3 Bison grammar and scanner logic into Rust. The gram_c2rust_fgram crate contains an auto-generated translation of the original gram.y file, ensuring that tokenization and grammar rules match the C implementation exactly. FFI seams bridge the generated C code with safe Rust wrappers, allowing incremental porting while maintaining semantic parity.
What is the role of the memory arena (Mcx) in parsing?
The Mcx (memory context) parameter in raw_parser() provides an arena-based allocation strategy that replaces PostgreSQL's memory contexts. All parsed nodes allocate storage within this arena, ensuring that the entire parse tree shares the same lifetime and can be freed efficiently when the context drops. This approach eliminates use-after-free risks while maintaining the performance characteristics of bulk allocation.
How does pgrust handle syntax errors safely?
Instead of using the C parser's longjmp mechanism for error recovery, the pgrust SQL parser catches error conditions at the FFI boundary and converts them into Err(PgError) results. The raw_parser() function returns PgResult<PgVec<'mcx, RawStmt<'mcx>>>, forcing callers to explicitly handle parse failures. This transformation occurs in the grammar seam layer (gram_core/src/lib.rs), insulating Rust code from C-style error handling.
What are parse modes and when should I use them?
Parse modes (RawParseMode) specify the syntactic context for parsing. Use RAW_PARSE_DEFAULT for standard SQL queries, RAW_PARSE_TYPE_NAME when parsing data type names (e.g., numeric(10,2)), and RAW_PARSE_PLPGSQL_EXPR for PL/pgSQL expressions. Each mode seeds the lexer with a specific initial token that puts the Bison parser in the correct state to recognize the expected grammar subset, enabling precise parsing of partial SQL fragments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →