# How the pgrust SQL Parser Handles Query Parsing: A Deep Dive into the PostgreSQL 18.3 Port

> Explore the malisper/pgrust SQL parser and its memory-safe Rust pipeline for PostgreSQL 18.3 queries. Learn how it uses Bison grammar, a lexer, and arena-based lifetimes for efficient parsing.

- Repository: [Michael Malis/pgrust](https://github.com/malisper/pgrust)
- Tags: deep-dive
- Published: 2026-07-13

---

**The pgrust SQL parser processes queries through a memory-safe Rust pipeline that wraps PostgreSQL 18.3's Bison grammar, using a lexer with one-token look-ahead to feed tokens into `base_yyparse`, then converting the resulting C-style parse trees into owned Rust structures via arena-based lifetimes.**

The pgrust SQL parser is a Rust port of the PostgreSQL 18.3 parser maintained in the malisper/pgrust repository. It provides a safe, idiomatic interface to the same Bison grammar that powers PostgreSQL, converting raw SQL strings into typed abstract syntax trees while eliminating memory safety risks through explicit arena management. This article examines how the parser handles lexical analysis, grammar processing, and memory management based on the actual implementation in the source code.

## The Parser Pipeline: From Raw SQL to Abstract Syntax Tree

The pgrust SQL parser follows a four-stage pipeline that mirrors PostgreSQL's architecture while wrapping each component in Rust's type system and memory safety guarantees.

### Entry Point: `raw_parser()`

The public API surface begins at `raw_parser()` in [`crates/backend/parser/driver/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/driver/src/lib.rs). This function serves as the primary interface for all SQL parsing operations:

```rust
pub fn raw_parser<'mcx>(
    mcx: Mcx<'mcx>,
    str_: &'mcx str,
    mode: RawParseMode,
) -> PgResult<PgVec<'mcx, RawStmt<'mcx>>> {
    gram::base_yyparse::call(mcx, str_, mode)
}

```

The function accepts a memory context arena (`Mcx`), the SQL string, and a parse mode. It returns a `PgResult` containing a `PgVec` of `RawStmt` nodes, each representing a parsed statement in the query string. This design replaces PostgreSQL's C-style error handling (long jumps) with Rust's explicit `Result` type.

### Lexical Analysis with `BaseLexer`

Before grammar processing begins, the input string passes through `BaseLexer`, which wraps the core scanner (`core_yylex`) exposed via [`crates/backend/parser/scan_seams/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/scan_seams/src/lib.rs). The lexer provides:

1. **One-token look-ahead** that enables merging of multi-word tokens (e.g., converting `NOT LIKE` into `NOT_LA`)
2. **Unicode de-escaping** through `str_udescape` and `check_uescapechar` for handling `UIDENT` and `USCONST` tokens
3. **Mode token seeding** when non-default parse modes require initial context

The lexer initialization pattern appears in [`driver/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/driver/src/lib.rs):

```rust
let seed = mode_seed(mcx, mode);           // optional CoreToken
let lexer = BaseLexer::new(mcx, scanbuf, seed);

```

### Grammar Processing via `base_yyparse`

The Bison-generated grammar lives in [`crates/backend/parser/gram_core/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/gram_core/src/lib.rs), which exposes the `base_yyparse` seam function. This component:

- Consumes tokens from the `BaseLexer` instance
- Builds the raw parse tree using PostgreSQL's original grammar rules
- Returns a C-style `List *` of `RawStmt` structures

The grammar seam bridges the C parser implementation (auto-translated from `gram.y` in [`crates/_support/pgrust/gram_c2rust_fgram/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/_support/pgrust/gram_c2rust_fgram/src/lib.rs)) with Rust's ownership system, allowing incremental porting while preserving exact PostgreSQL semantics.

### Memory Safety and Result Conversion

Unlike the C implementation that relies on `longjmp` for error handling and raw pointers for memory management, the pgrust SQL parser converts the grammar output into safe Rust structures:

- The `List *` from the C parser gets wrapped into `PgVec<'mcx, RawStmt<'mcx>>`
- All nodes use arena-based lifetimes (`Mcx`) tied to the parsing context
- Syntax errors propagate as `Err(PgError)` rather than triggering non-local jumps

This conversion happens within the `base_yyparse` wrapper, ensuring that users interact only with fully owned, properly typed Rust structures.

## Parse Modes and Look-Ahead Tokens

The pgrust SQL parser supports multiple entry points via `RawParseMode`, defined in [`driver/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/driver/src/lib.rs). Each mode determines the initial parser state:

- **`RAW_PARSE_DEFAULT`**: Standard SQL query parsing (no seed token)
- **`RAW_PARSE_TYPE_NAME`**: Parses a single type name expression (seeds with `MODE_TYPE_NAME`, token 774)
- **`RAW_PARSE_PLPGSQL_EXPR`**: Parses PL/pgSQL expressions (seeds with `MODE_PLPGSQL_EXPR`, token 775)

When a non-default mode is requested, the `mode_seed` function provides an initial look-ahead token that places the parser in the correct semantic context. This mechanism enables the same grammar to handle diverse inputs ranging from complete `SELECT` statements to isolated type expressions.

## Unicode Handling and Token Merging

The lexer handles PostgreSQL's Unicode escape syntax through dedicated functions in [`driver/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/driver/src/lib.rs). When encountering `UIDENT` (Unicode identifier) or `USCONST` (Unicode string constant) tokens, the parser:

- Validates escape sequences via `check_uescapechar`
- Converts escaped values to plain identifiers using `str_udescape`
- Preserves the original PostgreSQL behavior for Unicode validation and identifier truncation

Additionally, the one-token look-ahead capability handles multi-word keywords common in SQL:2011 syntax. For example, the sequence `WITH TIME` becomes `WITH_LA`, allowing the grammar to disambiguate between `WITH` as a clause starter and `WITH TIME ZONE` as a modifier.

## Practical Usage Examples

To parse a standard SQL query using the pgrust SQL parser:

```rust
use pgrust::backend::parser::driver::{raw_parser, RawParseMode};
use pgrust::mcx::Mcx;

// Create a memory arena (required by all parser functions)
let ctx = ::mcx::MemoryContext::new("example");
let mcx = ctx.mcx();

// Parse a simple SELECT
let sql = "SELECT id, name FROM users WHERE active = true;";
let parse_tree = raw_parser(mcx, sql, RawParseMode::RAW_PARSE_DEFAULT)
    .expect("SQL parsing failed");

// `parse_tree` is a `PgVec<RawStmt>`; iterate over statements:
for stmt in parse_tree.iter() {
    println!("Statement kind: {:?}", stmt.stmt.node_tag());
}

```

For parsing type names (useful when implementing type resolution):

```rust
use pgrust::backend::parser::driver::raw_parse_type_name;

let ty = raw_parse_type_name("numeric(10,2)".to_string())
    .expect("type parsing failed");
println!("Parsed type: {:?}", ty);

```

## Key Source Files and Architecture

The parser implementation spans several specialized crates that separate concerns between lexing, grammar, and FFI:

- **[`crates/backend/parser/driver/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/driver/src/lib.rs)**: Public API (`raw_parser`), lexer setup, Unicode handling, and mode token management
- **[`crates/backend/parser/gram_core/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/gram_core/src/lib.rs)**: Safe wrapper around the Bison grammar (`base_yyparse`) and C-to-Rust conversion logic
- **[`crates/_support/pgrust/gram_c2rust_fgram/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/_support/pgrust/gram_c2rust_fgram/src/lib.rs)**: Auto-generated C-to-Rust translation of the original `gram.y` grammar file
- **[`crates/backend/parser/scan_seams/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/scan_seams/src/lib.rs)**: FFI seam exposing the core lexical scanner (`core_yylex`)
- **[`crates/backend/parser/gram_seams/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/gram_seams/src/lib.rs)**: FFI seam exposing `base_yyparse` to the driver crate
- **[`crates/backend/parser/gram_fgram/src/tests.rs`](https://github.com/malisper/pgrust/blob/main/crates/backend/parser/gram_fgram/src/tests.rs)**: End-to-end test suite validating parsing behavior against PostgreSQL semantics

## Summary

- **The pgrust SQL parser** provides a Rust-native interface to PostgreSQL 18.3's grammar through the `raw_parser()` function in [`backend/parser/driver/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/backend/parser/driver/src/lib.rs).
- **Lexical analysis** uses `BaseLexer` to wrap the core scanner with one-token look-ahead for multi-word token merging and mode-specific seeding.
- **Grammar processing** drives the Bison-generated `base_yyparse` via FFI seams, converting raw C structures into safe Rust types using arena lifetimes (`Mcx`).
- **Multiple parse modes** (`RAW_PARSE_DEFAULT`, `RAW_PARSE_TYPE_NAME`, etc.) allow the same grammar to handle diverse input contexts through initial look-ahead tokens.
- **Unicode and error handling** follow PostgreSQL semantics but use Rust's `PgResult` type for safe error propagation instead of C-style long jumps.

## Frequently Asked Questions

### How does pgrust maintain compatibility with PostgreSQL's parser?

The pgrust SQL parser preserves compatibility by directly porting the PostgreSQL 18.3 Bison grammar and scanner logic into Rust. The `gram_c2rust_fgram` crate contains an auto-generated translation of the original `gram.y` file, ensuring that tokenization and grammar rules match the C implementation exactly. FFI seams bridge the generated C code with safe Rust wrappers, allowing incremental porting while maintaining semantic parity.

### What is the role of the memory arena (Mcx) in parsing?

The `Mcx` (memory context) parameter in `raw_parser()` provides an arena-based allocation strategy that replaces PostgreSQL's memory contexts. All parsed nodes allocate storage within this arena, ensuring that the entire parse tree shares the same lifetime and can be freed efficiently when the context drops. This approach eliminates use-after-free risks while maintaining the performance characteristics of bulk allocation.

### How does pgrust handle syntax errors safely?

Instead of using the C parser's `longjmp` mechanism for error recovery, the pgrust SQL parser catches error conditions at the FFI boundary and converts them into `Err(PgError)` results. The `raw_parser()` function returns `PgResult<PgVec<'mcx, RawStmt<'mcx>>>`, forcing callers to explicitly handle parse failures. This transformation occurs in the grammar seam layer ([`gram_core/src/lib.rs`](https://github.com/malisper/pgrust/blob/main/gram_core/src/lib.rs)), insulating Rust code from C-style error handling.

### What are parse modes and when should I use them?

Parse modes (`RawParseMode`) specify the syntactic context for parsing. Use `RAW_PARSE_DEFAULT` for standard SQL queries, `RAW_PARSE_TYPE_NAME` when parsing data type names (e.g., `numeric(10,2)`), and `RAW_PARSE_PLPGSQL_EXPR` for PL/pgSQL expressions. Each mode seeds the lexer with a specific initial token that puts the Bison parser in the correct state to recognize the expected grammar subset, enabling precise parsing of partial SQL fragments.