pgrust Parser Architecture: How SQL Parsing Works in the PL/pgSQL Frontend
pgrust implements a recursive-descent parser that mirrors PostgreSQL's pl_gram.y grammar, using a custom scanner to tokenize PL/pgSQL-specific constructs and delegating SQL fragments to the core parser via raw-parse modes.
The pgrust project provides a Rust implementation of PostgreSQL's PL/pgSQL language frontend. Understanding the pgrust parser architecture reveals how it achieves full compatibility with PostgreSQL's grammar while leveraging Rust's type safety and performance characteristics. The parser transforms PL/pgSQL source code into an abstract syntax tree (AST) through a pipeline of specialized components that bridge the core SQL lexer with PL/pgSQL-specific tokenization.
Recursive-Descent Parser Architecture
The pgrust parser is organized as a recursive-descent implementation that translates each production of the original Bison grammar (pl_gram.y) into corresponding Rust methods. The architecture consists of five primary components that work together to drive the parsing process.
The parser driver (located in crates/pl/plpgsql/src/plpgsql_gram/src/parser.rs) manages the overall parsing state, holding the scanner instance and storing the current token's semantic value (yylval) and location (yylloc). The recursive-descent grammar implements each Bison production as a Rust method (e.g., pl_block, stmt_if, stmt_assign), handling the top-down parsing logic.
The scanner (crates/pl/plpgsql/src/plpgsql_scanner/src/lib.rs) wraps the core SQL lexer and adds PL/pgSQL-specific tokenization for variables, labels, and special operators. Seams (or comp_seams) provide the bridge to PostgreSQL's compiler helpers (pl_comp.c) for variable handling, statement IDs, and error reporting. Finally, the AST nodes defined in crates/pl/plpgsql/src/plpgsql.rs represent the resulting parse tree structure.
Parsing begins through the plpgsql_yyparse or plpgsql_yyparse_with_lineno entry points, which instantiate the parser, initialize the PlpgsqlScanner, and begin consuming tokens via repeated calls to yylex().
The PL/pgSQL Scanner and Token Flow
Tokenization in pgrust follows a multi-layered approach where the scanner acts as an intermediary between the core SQL lexer and the recursive-descent parser.
The flow proceeds through these stages:
- Core SQL lexer produces raw tokens (
IDENT, operators, literals) via internal seam functions. PlpgsqlScanner::plpgsql_yylex(inlib.rs) reclassifies tokens:- Identifiers matching reserved PL/pgSQL keywords become
K_*tokens (defined inRESERVED_PL_KEYWORDS). - Names resolving to PL/pgSQL variables return
T_DATUM. - Dotted identifiers that fail resolution become
T_CWORD. - Unresolved simple identifiers become
T_WORD. - Special operators like
<<,>>, and#are handled explicitly.
- Identifiers matching reserved PL/pgSQL keywords become
- The parser consumes these tokens through
yylex(), storing values inself.yylvaland locations inself.yylloc.
The scanner maintains a push-back stack to support lookahead operations, enabling the parser to emulate Bison's look-ahead semantics when determining which grammar production to enter.
Grammar Implementation as Rust Methods
Each grammar production from the original Bison file is implemented as a Rust method following a consistent structural pattern. These methods validate token sequences and build corresponding AST nodes.
The pl_block method demonstrates this pattern:
fn pl_block(&mut self) -> PgResult<PLpgSQL_stmt> {
// decl_sect → opt_block_label [decl_start [decl_stmts]]
let declhdr = self.decl_sect(pre_label)?;
// Expect K_BEGIN …
self.expect_token(K_BEGIN)?;
// Process procedure statements
let body = self.proc_sect()?;
// Exception handling …
let exceptions = self.exception_sect()?;
// Expect K_END …
self.expect_token(K_END)?;
// Build the AST node
Ok(PLpgSQL_stmt::Block(Box::new(PLpgSQL_stmt_block { … })))
}
The grammar methods are organized into functional groups:
- Declarations:
decl_sect,decl_statement,decl_varname, anddecl_datatypehandle theDECLAREblock and variable definitions. - Control structures:
stmt_if,stmt_loop,stmt_for, andstmt_whileimplement conditional logic and iteration. - SQL statements:
stmt_execsql,stmt_call, andstmt_performprocess embedded SQL commands. - Exception handling:
exception_sectandstmt_getdiagmanage error capture and diagnostic retrieval. - Expression parsing:
read_sql_expressionandread_sql_constructdelegate to the core SQL parser using specialized raw-parse modes.
Helper functions like opt_semi, push_back_token, peek, and syntax_at provide look-ahead capabilities and precise error reporting through ereport(ERROR, …) wrappers.
How SQL Parsing Is Achieved via Raw-Parse Modes
When the parser encounters SQL statements within PL/pgSQL blocks, it delegates to the core SQL parser through a mechanism called raw-parse modes. This ensures that pgrust handles standard SQL syntax with complete PostgreSQL compatibility while treating the fragments as black-box expressions within the PL/pgSQL AST.
The process works as follows:
- The parser calls
read_sql_constructorread_sql_expressionupon encountering SQL-statement tokens (e.g.,K_SELECT,K_INSERT, or leading identifiers). - These methods pass a **
RawParseMode** parameter (such asRAW_PARSE_DEFAULT,RAW_PARSE_PLPGSQL_ASSIGN1, orRAW_PARSE_PLPGSQL_ASSIGN3) to indicate how the core parser should interpret the fragment—whether as a full statement, an expression, or an assignment target. - The core SQL parser constructs a
PLpgSQL_exprnode containing the raw query string, the parsed tree (represented as boxed Rust structs mirroring PostgreSQL'sNode *), and the parse mode metadata. - The PL/pgSQL parser may then rewrite the query (e.g.,
perform_rewrite_queryforPERFORMstatements) or validate it viacheck_sql_expr.
The stmt_assign method illustrates this integration:
fn stmt_assign(&mut self) -> PgResult<PLpgSQL_stmt> {
let datum = self.yylval.wdatum.clone().unwrap();
let datum_dno = datum.datum.ok_or_else(|| internal_error(...))? as i32;
comp_seam::check_assignable::call(datum_dno, loc)?;
// Read the RHS expression using the appropriate RAW_PARSE mode
let (mut expr, _, _) = self.read_sql_construct(
';' as i32, // stop at semicolon
0, 0, // no special start/end tokens
";",
match nnames { … }, // choose RAW_PARSE_PLPGSQL_ASSIGN*
false, true,
)?;
comp_seam::mark_expr_as_assignment_source::call(&mut expr, datum_dno);
Ok(PLpgSQL_stmt::Assign(Box::new(PLpgSQL_stmt_assign { …, expr: Some(expr) })))
}
Here, the scanner supplies the variable token (T_DATUM), the parser validates assignability through the comp_seam bridge, and the core SQL parser handles the right-hand side expression up to the terminating semicolon.
Key Source Files and Responsibilities
The pgrust parser architecture spans several specific source files:
crates/pl/plpgsql/src/plpgsql_gram/src/parser.rs: Contains the full recursive-descent implementation of the PL/pgSQL grammar, including the parser driver and all grammar production methods.crates/pl/plpgsql/src/plpgsql_scanner/src/lib.rs: Implements thePlpgsqlScannerthat bridges the core SQL lexer to PL/pgSQL tokens, handling variable lookup and keyword re-classification.crates/pl/plpgsql/src/plpgsql.rs: Defines the AST structures (PLpgSQL_stmt_*,PLpgSQL_expr, etc.) used throughout the parsing pipeline.crates/pl/plpgsql/src/plpgsql_gram/src/lib.rs: Exposes the public API including theplpgsql_yyparseentry point.
Summary
- pgrust uses a recursive-descent parser architecture that manually implements PostgreSQL's
pl_gram.ygrammar as Rust methods. - The
PlpgsqlScannermediates between the core SQL lexer and PL/pgSQL grammar, handling variable resolution (T_DATUM) and keyword classification. - SQL fragments are parsed using raw-parse modes (
RAW_PARSE_PLPGSQL_ASSIGN*, etc.) that delegate to the core SQL parser while maintaining PL/pgSQL context. - The seams mechanism (
comp_seams) provides interoperability with PostgreSQL's compiler infrastructure for variable handling and error reporting. - All parsing state is managed through explicit structures (
yylval,yylloc) rather than global variables, following Rust's ownership model.
Frequently Asked Questions
How does pgrust maintain compatibility with PostgreSQL's PL/pgSQL grammar?
pgrust achieves compatibility by directly mirroring the grammar productions from PostgreSQL's pl_gram.y Bison file as Rust methods in parser.rs. Each Bison rule is translated into a corresponding recursive-descent function (e.g., pl_block, stmt_if), ensuring that the accepted syntax matches the original C implementation exactly. The scanner also replicates PostgreSQL's token classification logic, including the handling of reserved keywords and variable identifiers.
What is the role of the scanner in the pgrust parser architecture?
The scanner (crates/pl/plpgsql/src/plpgsql_scanner/src/lib.rs) serves as the tokenization layer that extends the core SQL lexer with PL/pgSQL-specific rules. It reclassifies identifiers into K_* keyword tokens, resolves names to T_DATUM (variables), T_WORD (unresolved identifiers), or T_CWORD (dotted identifiers), and handles special syntax like label brackets (<<label>>). It also maintains a push-back stack for lookahead, enabling the recursive-descent parser to make parsing decisions based on future tokens.
How does pgrust parse SQL statements embedded in PL/pgSQL code?
When encountering SQL statements, the parser calls read_sql_construct with a specific RawParseMode parameter. This mode (such as RAW_PARSE_DEFAULT or RAW_PARSE_PLPGSQL_ASSIGN1) tells the core SQL parser how to interpret the fragment—whether as a complete statement or a partial expression. The core parser returns a PLpgSQL_expr node containing the parsed SQL tree, which the PL/pgSQL parser then embeds into its AST. This delegation ensures that pgrust supports all PostgreSQL SQL syntax without reimplementing the SQL grammar.
Why did pgrust choose a recursive-descent parser over a parser generator?
The recursive-descent implementation allows pgrust to replicate PostgreSQL's pl_gram.y behavior while leveraging Rust's type safety and error handling. By writing the grammar as explicit Rust methods rather than using a parser generator like Bison, the project avoids external build dependencies and gains fine-grained control over error reporting (via yyerror and syntax_at) and memory management. This approach also makes the parsing logic more transparent to Rust developers and easier to debug with standard Rust tooling.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →