Supported Cypher Clauses in Codebase-Memory-MCP: Complete Reference Guide

Codebase-Memory-MCP implements a read-only subset of the Neo4j Cypher language supporting 25+ clauses including MATCH, WHERE, RETURN, UNION, and aggregate functions, with all token definitions centralized in src/cypher/cypher.h.

Codebase-Memory-MCP is an open-source graph memory system designed for static code analysis. It ships with a lightweight, in-process Cypher query engine that parses and executes read-only queries against an embedded graph store. The supported subset focuses on safe, analytical operations while deliberately excluding data-modifying commands.

Core Query Structure Clauses

The engine recognizes fundamental Cypher building blocks defined as token types in src/cypher/cypher.h.

Pattern Matching and Filtering

  • MATCH – TOK_MATCH (line 23) initiates pattern-matching expressions against the graph.
  • WHERE – TOK_WHERE (line 24) filters matched elements using boolean predicates.
  • WITH – TOK_WITH (line 36) enables sub-queries and variable passing between query parts, supporting read-only pipelines.

Projection and Results

  • RETURN – TOK_RETURN (line 25) projects selected expressions or columns from the match.
  • AS – TOK_AS (line 31) renames expressions (e.g., COUNT(g) AS cnt).
  • DISTINCT – TOK_DISTINCT (line 32) removes duplicate rows from result projections.

Sorting, Pagination, and Set Operations

Ordering and Limits

  • ORDER BY – Implemented via TOK_ORDER (line 26) and TOK_BY (line 27) to sort result sets.
  • ASC / DESC – TOK_ASC (line 38) and TOK_DESC (line 39) specify sorting directions.
  • LIMIT – TOK_LIMIT (line 28) caps the number of returned rows.
  • SKIP – TOK_SKIP (line 46) offsets the result set for pagination.

Advanced Result Handling

  • UNION – TOK_UNION (line 47) combines results from two MATCH blocks (read-only only).
  • UNWIND – TOK_UNWIND (line 48) expands a list into individual rows.

Boolean Logic and Conditional Expressions

Logical Operators

  • AND – TOK_AND (line 29) for boolean conjunction inside WHERE clauses.
  • OR – TOK_OR (line 30) for boolean disjunction.
  • NOT – TOK_NOT (line 37) for predicate negation.
  • XOR – TOK_XOR (line 45) for exclusive-or logical operations.

Conditional Logic

  • CASE … WHEN … THEN … ELSE … END – Implemented via TOK_CASE (line 63) through TOK_END (line 68), enabling conditional expressions within RETURN or WHERE clauses.

Predicates and Functions

String and Comparison Predicates

  • CONTAINS – TOK_CONTAINS (line 34) performs substring searches (e.g., x CONTAINS "foo").
  • STARTS WITH – Built from TOK_STARTS (line 35) combined with WITH for string prefix matching.
  • IN – TOK_IN (line 42) tests membership in lists (e.g., x IN [1,2,3]).
  • IS NULL – Uses TOK_IS (line 43) paired with TOK_NULL_KW (line 44) for null checks.

Aggregate Functions

The engine treats aggregate functions as keywords tokenized in src/cypher/cypher.h:

  • COUNT – TOK_COUNT (line 33)
  • SUM – TOK_SUM
  • AVG – TOK_AVG
  • MIN – TOK_MIN_KW
  • MAX – TOK_MAX_KW
  • collect – TOK_COLLECT

String Transformation Functions

  • toLower – TOK_TOLOWER
  • toUpper – TOK_TOUPPER
  • toString – TOK_TOSTRING

Query Engine Architecture

The Cypher implementation operates in four distinct stages defined in src/cypher/cypher.c:

  1. Lexing – cbm_lex() tokenizes the raw query string into the token types defined in src/cypher/cypher.h.
  2. Parsing – cbm_cypher_parse() constructs an abstract syntax tree (AST) from the token stream, enforcing grammar rules such as mandatory MATCH and RETURN sequences.
  3. Planning – The AST transforms into an execution plan that traverses the in-memory graph store (cbm_store_t).
  4. Execution – cbm_cypher_execute() runs the plan and produces a cbm_cypher_result_t iterable by the caller.

The test suite in tests/test_cypher.c validates each clause individually through tests like TEST(cypher_parse_where_regex) and TEST(cypher_exec_return_order_limit).

Practical Query Examples

Basic Pattern Matching

cbm_cypher_result_t r = {0};
int rc = cbm_cypher_execute(store,
    "MATCH (f:Function) RETURN f.name, f.qualified_name",
    "example", 0, &r);

Filtering with Regular Expressions and String Predicates

rc = cbm_cypher_execute(store,
    "MATCH (f:Function) "
    "WHERE f.name =~ \".*Order.*\" "
    "  AND f.path CONTAINS \"src/\" "
    "RETURN f.name",
    "example", 0, &r);

Aggregation with Sorting and Pagination

rc = cbm_cypher_execute(store,
    "MATCH (f)-[:CALLS]->(g) "
    "RETURN f.name, COUNT(g) AS cnt "
    "ORDER BY cnt DESC "
    "SKIP 10 LIMIT 5",
    "example", 0, &r);

Combining Results with UNION

rc = cbm_cypher_execute(store,
    "MATCH (n) RETURN n.name "
    "UNION "
    "MATCH (m) RETURN m.name",
    "example", 0, &r);

Expanding Lists with UNWIND

rc = cbm_cypher_execute(store,
    "UNWIND [1,2,3] AS x RETURN x",
    "example", 0, &r);

Limitations and Unsupported Clauses

While the lexer recognizes keywords such as CREATE, DELETE, and MERGE for syntax completeness, these write-operation clauses are deliberately unsupported. Attempting to use them triggers a parse error, ensuring the engine remains safe for read-only analytical workloads against the immutable code graph.

Summary

  • Codebase-Memory-MCP supports a read-only subset of Cypher defined in src/cypher/cypher.h with 25+ token types including MATCH, WHERE, RETURN, and aggregate functions.
  • The query pipeline implements four stages—lexing (cbm_lex), parsing (cbm_cypher_parse), planning, and execution (cbm_cypher_execute).
  • Sorting and pagination are supported via ORDER BY, LIMIT, and SKIP tokens.
  • Set operations UNION and UNWIND enable complex result combinations and list expansions.
  • Write clauses (CREATE, DELETE, MERGE) are recognized but explicitly rejected to maintain read-only safety.

Frequently Asked Questions

Does Codebase-Memory-MCP support write operations like CREATE or DELETE?

No. While the lexer tokenizes keywords such as CREATE, DELETE, and MERGE, the parser explicitly rejects them to maintain a read-only safety model. This design prevents accidental modification of the in-process graph store during static analysis operations.

How does the engine handle pagination in Cypher queries?

Pagination combines the SKIP and LIMIT clauses. TOK_SKIP (defined at line 46 in src/cypher/cypher.h) offsets the result set, while TOK_LIMIT (line 28) caps the number of rows returned. For example, SKIP 10 LIMIT 5 returns rows 11 through 15.

What aggregate functions are available in Codebase-Memory-MCP?

The engine supports COUNT, SUM, AVG, MIN, MAX, and collect, each defined as distinct token types (TOK_COUNT, TOK_SUM, TOK_AVG, TOK_MIN_KW, TOK_MAX_KW, TOK_COLLECT). These functions operate within RETURN clauses to compute grouped aggregations across matched patterns.

Is the Cypher implementation compatible with Neo4j?

Codebase-Memory-MCP implements a subset of Neo4j Cypher focused on read-only graph traversal. Queries using supported clauses like MATCH, WHERE, RETURN, and UNION will execute identically, but data-modifying statements and advanced Neo4j-specific extensions are not supported.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →