How the pgrust SQL Parser Handles Query Parsing: A Deep Dive into the PostgreSQL 18.3 Port

The pgrust SQL parser processes queries through a memory-safe Rust pipeline that wraps PostgreSQL 18.3's Bison grammar, using a lexer with one-token look-ahead to feed tokens into base_yyparse, then converting the resulting C-style parse trees into owned Rust structures via arena-based lifetimes.

The pgrust SQL parser is a Rust port of the PostgreSQL 18.3 parser maintained in the malisper/pgrust repository. It provides a safe, idiomatic interface to the same Bison grammar that powers PostgreSQL, converting raw SQL strings into typed abstract syntax trees while eliminating memory safety risks through explicit arena management. This article examines how the parser handles lexical analysis, grammar processing, and memory management based on the actual implementation in the source code.

The Parser Pipeline: From Raw SQL to Abstract Syntax Tree

The pgrust SQL parser follows a four-stage pipeline that mirrors PostgreSQL's architecture while wrapping each component in Rust's type system and memory safety guarantees.

Entry Point: raw_parser()

The public API surface begins at raw_parser() in crates/backend/parser/driver/src/lib.rs. This function serves as the primary interface for all SQL parsing operations:

pub fn raw_parser<'mcx>(
    mcx: Mcx<'mcx>,
    str_: &'mcx str,
    mode: RawParseMode,
) -> PgResult<PgVec<'mcx, RawStmt<'mcx>>> {
    gram::base_yyparse::call(mcx, str_, mode)
}

The function accepts a memory context arena (Mcx), the SQL string, and a parse mode. It returns a PgResult containing a PgVec of RawStmt nodes, each representing a parsed statement in the query string. This design replaces PostgreSQL's C-style error handling (long jumps) with Rust's explicit Result type.

Lexical Analysis with BaseLexer

Before grammar processing begins, the input string passes through BaseLexer, which wraps the core scanner (core_yylex) exposed via crates/backend/parser/scan_seams/src/lib.rs. The lexer provides:

  1. One-token look-ahead that enables merging of multi-word tokens (e.g., converting NOT LIKE into NOT_LA)
  2. Unicode de-escaping through str_udescape and check_uescapechar for handling UIDENT and USCONST tokens
  3. Mode token seeding when non-default parse modes require initial context

The lexer initialization pattern appears in driver/src/lib.rs:

let seed = mode_seed(mcx, mode);           // optional CoreToken
let lexer = BaseLexer::new(mcx, scanbuf, seed);

Grammar Processing via base_yyparse

The Bison-generated grammar lives in crates/backend/parser/gram_core/src/lib.rs, which exposes the base_yyparse seam function. This component:

  • Consumes tokens from the BaseLexer instance
  • Builds the raw parse tree using PostgreSQL's original grammar rules
  • Returns a C-style List * of RawStmt structures

The grammar seam bridges the C parser implementation (auto-translated from gram.y in crates/_support/pgrust/gram_c2rust_fgram/src/lib.rs) with Rust's ownership system, allowing incremental porting while preserving exact PostgreSQL semantics.

Memory Safety and Result Conversion

Unlike the C implementation that relies on longjmp for error handling and raw pointers for memory management, the pgrust SQL parser converts the grammar output into safe Rust structures:

  • The List * from the C parser gets wrapped into PgVec<'mcx, RawStmt<'mcx>>
  • All nodes use arena-based lifetimes (Mcx) tied to the parsing context
  • Syntax errors propagate as Err(PgError) rather than triggering non-local jumps

This conversion happens within the base_yyparse wrapper, ensuring that users interact only with fully owned, properly typed Rust structures.

Parse Modes and Look-Ahead Tokens

The pgrust SQL parser supports multiple entry points via RawParseMode, defined in driver/src/lib.rs. Each mode determines the initial parser state:

  • RAW_PARSE_DEFAULT: Standard SQL query parsing (no seed token)
  • RAW_PARSE_TYPE_NAME: Parses a single type name expression (seeds with MODE_TYPE_NAME, token 774)
  • RAW_PARSE_PLPGSQL_EXPR: Parses PL/pgSQL expressions (seeds with MODE_PLPGSQL_EXPR, token 775)

When a non-default mode is requested, the mode_seed function provides an initial look-ahead token that places the parser in the correct semantic context. This mechanism enables the same grammar to handle diverse inputs ranging from complete SELECT statements to isolated type expressions.

Unicode Handling and Token Merging

The lexer handles PostgreSQL's Unicode escape syntax through dedicated functions in driver/src/lib.rs. When encountering UIDENT (Unicode identifier) or USCONST (Unicode string constant) tokens, the parser:

  • Validates escape sequences via check_uescapechar
  • Converts escaped values to plain identifiers using str_udescape
  • Preserves the original PostgreSQL behavior for Unicode validation and identifier truncation

Additionally, the one-token look-ahead capability handles multi-word keywords common in SQL:2011 syntax. For example, the sequence WITH TIME becomes WITH_LA, allowing the grammar to disambiguate between WITH as a clause starter and WITH TIME ZONE as a modifier.

Practical Usage Examples

To parse a standard SQL query using the pgrust SQL parser:

use pgrust::backend::parser::driver::{raw_parser, RawParseMode};
use pgrust::mcx::Mcx;

// Create a memory arena (required by all parser functions)
let ctx = ::mcx::MemoryContext::new("example");
let mcx = ctx.mcx();

// Parse a simple SELECT
let sql = "SELECT id, name FROM users WHERE active = true;";
let parse_tree = raw_parser(mcx, sql, RawParseMode::RAW_PARSE_DEFAULT)
    .expect("SQL parsing failed");

// `parse_tree` is a `PgVec<RawStmt>`; iterate over statements:
for stmt in parse_tree.iter() {
    println!("Statement kind: {:?}", stmt.stmt.node_tag());
}

For parsing type names (useful when implementing type resolution):

use pgrust::backend::parser::driver::raw_parse_type_name;

let ty = raw_parse_type_name("numeric(10,2)".to_string())
    .expect("type parsing failed");
println!("Parsed type: {:?}", ty);

Key Source Files and Architecture

The parser implementation spans several specialized crates that separate concerns between lexing, grammar, and FFI:

Summary

  • The pgrust SQL parser provides a Rust-native interface to PostgreSQL 18.3's grammar through the raw_parser() function in backend/parser/driver/src/lib.rs.
  • Lexical analysis uses BaseLexer to wrap the core scanner with one-token look-ahead for multi-word token merging and mode-specific seeding.
  • Grammar processing drives the Bison-generated base_yyparse via FFI seams, converting raw C structures into safe Rust types using arena lifetimes (Mcx).
  • Multiple parse modes (RAW_PARSE_DEFAULT, RAW_PARSE_TYPE_NAME, etc.) allow the same grammar to handle diverse input contexts through initial look-ahead tokens.
  • Unicode and error handling follow PostgreSQL semantics but use Rust's PgResult type for safe error propagation instead of C-style long jumps.

Frequently Asked Questions

How does pgrust maintain compatibility with PostgreSQL's parser?

The pgrust SQL parser preserves compatibility by directly porting the PostgreSQL 18.3 Bison grammar and scanner logic into Rust. The gram_c2rust_fgram crate contains an auto-generated translation of the original gram.y file, ensuring that tokenization and grammar rules match the C implementation exactly. FFI seams bridge the generated C code with safe Rust wrappers, allowing incremental porting while maintaining semantic parity.

What is the role of the memory arena (Mcx) in parsing?

The Mcx (memory context) parameter in raw_parser() provides an arena-based allocation strategy that replaces PostgreSQL's memory contexts. All parsed nodes allocate storage within this arena, ensuring that the entire parse tree shares the same lifetime and can be freed efficiently when the context drops. This approach eliminates use-after-free risks while maintaining the performance characteristics of bulk allocation.

How does pgrust handle syntax errors safely?

Instead of using the C parser's longjmp mechanism for error recovery, the pgrust SQL parser catches error conditions at the FFI boundary and converts them into Err(PgError) results. The raw_parser() function returns PgResult<PgVec<'mcx, RawStmt<'mcx>>>, forcing callers to explicitly handle parse failures. This transformation occurs in the grammar seam layer (gram_core/src/lib.rs), insulating Rust code from C-style error handling.

What are parse modes and when should I use them?

Parse modes (RawParseMode) specify the syntactic context for parsing. Use RAW_PARSE_DEFAULT for standard SQL queries, RAW_PARSE_TYPE_NAME when parsing data type names (e.g., numeric(10,2)), and RAW_PARSE_PLPGSQL_EXPR for PL/pgSQL expressions. Each mode seeds the lexer with a specific initial token that puts the Bison parser in the correct state to recognize the expected grammar subset, enabling precise parsing of partial SQL fragments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →