pgrust Parser Architecture: How SQL Parsing Works in the PL/pgSQL Frontend

pgrust implements a recursive-descent parser that mirrors PostgreSQL's pl_gram.y grammar, using a custom scanner to tokenize PL/pgSQL-specific constructs and delegating SQL fragments to the core parser via raw-parse modes.

The pgrust project provides a Rust implementation of PostgreSQL's PL/pgSQL language frontend. Understanding the pgrust parser architecture reveals how it achieves full compatibility with PostgreSQL's grammar while leveraging Rust's type safety and performance characteristics. The parser transforms PL/pgSQL source code into an abstract syntax tree (AST) through a pipeline of specialized components that bridge the core SQL lexer with PL/pgSQL-specific tokenization.

Recursive-Descent Parser Architecture

The pgrust parser is organized as a recursive-descent implementation that translates each production of the original Bison grammar (pl_gram.y) into corresponding Rust methods. The architecture consists of five primary components that work together to drive the parsing process.

The parser driver (located in crates/pl/plpgsql/src/plpgsql_gram/src/parser.rs) manages the overall parsing state, holding the scanner instance and storing the current token's semantic value (yylval) and location (yylloc). The recursive-descent grammar implements each Bison production as a Rust method (e.g., pl_block, stmt_if, stmt_assign), handling the top-down parsing logic.

The scanner (crates/pl/plpgsql/src/plpgsql_scanner/src/lib.rs) wraps the core SQL lexer and adds PL/pgSQL-specific tokenization for variables, labels, and special operators. Seams (or comp_seams) provide the bridge to PostgreSQL's compiler helpers (pl_comp.c) for variable handling, statement IDs, and error reporting. Finally, the AST nodes defined in crates/pl/plpgsql/src/plpgsql.rs represent the resulting parse tree structure.

Parsing begins through the plpgsql_yyparse or plpgsql_yyparse_with_lineno entry points, which instantiate the parser, initialize the PlpgsqlScanner, and begin consuming tokens via repeated calls to yylex().

The PL/pgSQL Scanner and Token Flow

Tokenization in pgrust follows a multi-layered approach where the scanner acts as an intermediary between the core SQL lexer and the recursive-descent parser.

The flow proceeds through these stages:

  1. Core SQL lexer produces raw tokens (IDENT, operators, literals) via internal seam functions.
  2. PlpgsqlScanner::plpgsql_yylex (in lib.rs) reclassifies tokens:
    • Identifiers matching reserved PL/pgSQL keywords become K_* tokens (defined in RESERVED_PL_KEYWORDS).
    • Names resolving to PL/pgSQL variables return T_DATUM.
    • Dotted identifiers that fail resolution become T_CWORD.
    • Unresolved simple identifiers become T_WORD.
    • Special operators like <<, >>, and # are handled explicitly.
  3. The parser consumes these tokens through yylex(), storing values in self.yylval and locations in self.yylloc.

The scanner maintains a push-back stack to support lookahead operations, enabling the parser to emulate Bison's look-ahead semantics when determining which grammar production to enter.

Grammar Implementation as Rust Methods

Each grammar production from the original Bison file is implemented as a Rust method following a consistent structural pattern. These methods validate token sequences and build corresponding AST nodes.

The pl_block method demonstrates this pattern:

fn pl_block(&mut self) -> PgResult<PLpgSQL_stmt> {
    // decl_sect → opt_block_label [decl_start [decl_stmts]]
    let declhdr = self.decl_sect(pre_label)?;
    // Expect K_BEGIN …
    self.expect_token(K_BEGIN)?;
    // Process procedure statements
    let body = self.proc_sect()?;
    // Exception handling …
    let exceptions = self.exception_sect()?;
    // Expect K_END …
    self.expect_token(K_END)?;
    // Build the AST node
    Ok(PLpgSQL_stmt::Block(Box::new(PLpgSQL_stmt_block { … })))
}

The grammar methods are organized into functional groups:

  • Declarations: decl_sect, decl_statement, decl_varname, and decl_datatype handle the DECLARE block and variable definitions.
  • Control structures: stmt_if, stmt_loop, stmt_for, and stmt_while implement conditional logic and iteration.
  • SQL statements: stmt_execsql, stmt_call, and stmt_perform process embedded SQL commands.
  • Exception handling: exception_sect and stmt_getdiag manage error capture and diagnostic retrieval.
  • Expression parsing: read_sql_expression and read_sql_construct delegate to the core SQL parser using specialized raw-parse modes.

Helper functions like opt_semi, push_back_token, peek, and syntax_at provide look-ahead capabilities and precise error reporting through ereport(ERROR, …) wrappers.

How SQL Parsing Is Achieved via Raw-Parse Modes

When the parser encounters SQL statements within PL/pgSQL blocks, it delegates to the core SQL parser through a mechanism called raw-parse modes. This ensures that pgrust handles standard SQL syntax with complete PostgreSQL compatibility while treating the fragments as black-box expressions within the PL/pgSQL AST.

The process works as follows:

  1. The parser calls read_sql_construct or read_sql_expression upon encountering SQL-statement tokens (e.g., K_SELECT, K_INSERT, or leading identifiers).
  2. These methods pass a ** RawParseMode** parameter (such as RAW_PARSE_DEFAULT, RAW_PARSE_PLPGSQL_ASSIGN1, or RAW_PARSE_PLPGSQL_ASSIGN3) to indicate how the core parser should interpret the fragment—whether as a full statement, an expression, or an assignment target.
  3. The core SQL parser constructs a PLpgSQL_expr node containing the raw query string, the parsed tree (represented as boxed Rust structs mirroring PostgreSQL's Node *), and the parse mode metadata.
  4. The PL/pgSQL parser may then rewrite the query (e.g., perform_rewrite_query for PERFORM statements) or validate it via check_sql_expr.

The stmt_assign method illustrates this integration:

fn stmt_assign(&mut self) -> PgResult<PLpgSQL_stmt> {
    let datum = self.yylval.wdatum.clone().unwrap();
    let datum_dno = datum.datum.ok_or_else(|| internal_error(...))? as i32;
    comp_seam::check_assignable::call(datum_dno, loc)?;

    // Read the RHS expression using the appropriate RAW_PARSE mode
    let (mut expr, _, _) = self.read_sql_construct(
        ';' as i32,               // stop at semicolon
        0, 0,                     // no special start/end tokens
        ";",
        match nnames { … },       // choose RAW_PARSE_PLPGSQL_ASSIGN*
        false, true,
    )?;
    comp_seam::mark_expr_as_assignment_source::call(&mut expr, datum_dno);

    Ok(PLpgSQL_stmt::Assign(Box::new(PLpgSQL_stmt_assign { …, expr: Some(expr) })))
}

Here, the scanner supplies the variable token (T_DATUM), the parser validates assignability through the comp_seam bridge, and the core SQL parser handles the right-hand side expression up to the terminating semicolon.

Key Source Files and Responsibilities

The pgrust parser architecture spans several specific source files:

Summary

  • pgrust uses a recursive-descent parser architecture that manually implements PostgreSQL's pl_gram.y grammar as Rust methods.
  • The PlpgsqlScanner mediates between the core SQL lexer and PL/pgSQL grammar, handling variable resolution (T_DATUM) and keyword classification.
  • SQL fragments are parsed using raw-parse modes (RAW_PARSE_PLPGSQL_ASSIGN*, etc.) that delegate to the core SQL parser while maintaining PL/pgSQL context.
  • The seams mechanism (comp_seams) provides interoperability with PostgreSQL's compiler infrastructure for variable handling and error reporting.
  • All parsing state is managed through explicit structures (yylval, yylloc) rather than global variables, following Rust's ownership model.

Frequently Asked Questions

How does pgrust maintain compatibility with PostgreSQL's PL/pgSQL grammar?

pgrust achieves compatibility by directly mirroring the grammar productions from PostgreSQL's pl_gram.y Bison file as Rust methods in parser.rs. Each Bison rule is translated into a corresponding recursive-descent function (e.g., pl_block, stmt_if), ensuring that the accepted syntax matches the original C implementation exactly. The scanner also replicates PostgreSQL's token classification logic, including the handling of reserved keywords and variable identifiers.

What is the role of the scanner in the pgrust parser architecture?

The scanner (crates/pl/plpgsql/src/plpgsql_scanner/src/lib.rs) serves as the tokenization layer that extends the core SQL lexer with PL/pgSQL-specific rules. It reclassifies identifiers into K_* keyword tokens, resolves names to T_DATUM (variables), T_WORD (unresolved identifiers), or T_CWORD (dotted identifiers), and handles special syntax like label brackets (<<label>>). It also maintains a push-back stack for lookahead, enabling the recursive-descent parser to make parsing decisions based on future tokens.

How does pgrust parse SQL statements embedded in PL/pgSQL code?

When encountering SQL statements, the parser calls read_sql_construct with a specific RawParseMode parameter. This mode (such as RAW_PARSE_DEFAULT or RAW_PARSE_PLPGSQL_ASSIGN1) tells the core SQL parser how to interpret the fragment—whether as a complete statement or a partial expression. The core parser returns a PLpgSQL_expr node containing the parsed SQL tree, which the PL/pgSQL parser then embeds into its AST. This delegation ensures that pgrust supports all PostgreSQL SQL syntax without reimplementing the SQL grammar.

Why did pgrust choose a recursive-descent parser over a parser generator?

The recursive-descent implementation allows pgrust to replicate PostgreSQL's pl_gram.y behavior while leveraging Rust's type safety and error handling. By writing the grammar as explicit Rust methods rather than using a parser generator like Bison, the project avoids external build dependencies and gains fine-grained control over error reporting (via yyerror and syntax_at) and memory management. This approach also makes the parsing logic more transparent to Rust developers and easier to debug with standard Rust tooling.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →