How Code-Graph-RAG Generates Cypher Queries for Its RAG System

Code-Graph-RAG generates Cypher queries through a hybrid architecture that combines static query templates for retrieval operations with dynamic query builders that construct parameterized MERGE and CREATE statements during graph ingestion.

Code-Graph-RAG is an open-source retrieval-augmented generation (RAG) system that transforms codebases into graph structures stored in Memgraph. To power its RAG capabilities, the system must efficiently generate Cypher queries for both data ingestion and context retrieval. The implementation in vitali87/code-graph-rag achieves this through a two-stage approach defined primarily in codebase_rag/cypher_queries.py and orchestrated by the MemgraphIngestor class in codebase_rag/graph_service.py.

The Two-Stage Cypher Generation Architecture

The system separates query generation into distinct static and dynamic components, optimizing for both runtime performance and structural flexibility.

Static Query Catalog for Retrieval Operations

Pre-defined Cypher strings handle common RAG retrieval tasks without runtime string construction overhead. These constants reside in codebase_rag/cypher_queries.py and target specific access patterns. For example, CYPHER_FIND_BY_QUALIFIED_NAME (lines 69-74) retrieves entities by their fully qualified identifier, while CYPHER_TRACE_CALLABLES (lines 84-89) traces call relationships through the codebase.

These templates utilize parameterized placeholders such as $qn and $prefix, allowing the same query structure to handle different inputs safely through Memgraph's parameter binding. When the RAG component processes user queries, it selects the appropriate static template and injects runtime values via the fetch_all method.

Dynamic Query Builders for Graph Ingestion

During initial codebase parsing, the system constructs Cypher programmatically to accommodate arbitrary node labels and property dictionaries. The cypher_queries.py file provides several builder functions that return parameterized strings:

  • build_merge_node_query (lines 86-88): Generates MERGE statements with dynamic label insertion to prevent duplicate nodes.
  • build_create_relationship_query (lines 19-33): Constructs relationship creation Cypher with parameterized type and property mappings.
  • build_nodes_by_ids_query (lines 63-71): Creates batched node lookup queries for efficient retrieval.
  • wrap_with_unwind (lines 58-60): Transforms single-row operations into batch-capable statements by prepending UNWIND $rows AS row.

These builders accept Python data structures and return Cypher strings that separate code from data, preventing injection vulnerabilities while handling heterogeneous graph schemas.

From Codebase to Graph: The Ingestion Pipeline

The transformation of source code into graph nodes relies on aggressive batching strategies that minimize database round-trips.

Structure Extraction and Batching

The StructureProcessor class in codebase_rag/parsers/structure_processor.py traverses repositories to identify packages, folders, files, and language entities. Rather than executing individual Cypher statements per discovery, it accumulates operations through MemgraphIngestor methods:

self.ingestor.ensure_node_batch(cs.NodeLabel.PACKAGE, {...})
self.ingestor.ensure_relationship_batch(
    parent_identifier,
    cs.RelationshipType.CONTAINS_PACKAGE,
    (cs.NodeLabel.PACKAGE, cs.KEY_QUALIFIED_NAME, package_qn)
)

These calls stage data in memory until batch thresholds trigger flush operations.

Converting Batches to Cypher Statements

When ingestor.flush_all() is invoked, MemgraphIngestor in codebase_rag/graph_service.py processes accumulated data through _flush_node_label_group. This method selects the appropriate builder—build_merge_node_query for node labels or build_create_relationship_query for relationships—and aggregates rows using wrap_with_unwind:

query = build_merge_node_query(label, id_key)
self._execute_batch_on(target_conn, query, batch_rows)

The UNWIND clause enables single-query batch insertion, significantly improving ingestion performance compared to individual transactions. The _execute_batch_on method handles the actual parameter binding and execution against the Memgraph connection.

Retrieval Queries for RAG Operations

During the retrieval phase, Code-Graph-RAG leverages its static query catalog to fetch context for LLM prompts. The fetch_all method in graph_service.py executes parameterized versions of the static templates:

rows = ingestor.fetch_all(
    CYPHER_FIND_BY_QUALIFIED_NAME,
    {"qn": "myproj.services.UserService.get_user"}
)

Results are converted to Python dictionaries via _cursor_to_results before being passed to the language model for answer generation, completing the RAG loop.

Practical Implementation Examples

The following patterns demonstrate actual usage of the Cypher generation system within the codebase:

Ingesting a new function node:

ingestor.ensure_node_batch(
    cs.NodeLabel.FUNCTION,
    {
        cs.KEY_QUALIFIED_NAME: "myproj.services.UserService.get_user",
        cs.KEY_NAME: "get_user",
        cs.KEY_PATH: "services/user_service.py",
        cs.KEY_START_LINE: 12,
        cs.KEY_END_LINE: 30,
    },
)

Flushing batched data to trigger dynamic query construction:

ingestor.flush_all()  # Internally calls build_merge_node_query and wrap_with_unwind

Retrieving context for RAG responses:

result = ingestor.fetch_all(
    CYPHER_TRACE_CALLABLES,
    {"prefix": "myproj.controllers"}
)

Summary

  • Static templates in cypher_queries.py provide optimized, parameterized Cypher for common retrieval patterns like CYPHER_FIND_BY_QUALIFIED_NAME and CYPHER_TRACE_CALLABLES.
  • Dynamic builders such as build_merge_node_query and build_create_relationship_query construct ingestion queries from runtime Python data structures with proper parameterization.
  • Batch processing via wrap_with_unwind minimizes database round-trips during the ingestion of large codebases by combining multiple operations into single UNWIND statements.
  • The MemgraphIngestor class orchestrates both ingestion and retrieval, executing generated Cypher against Memgraph and returning structured results to the RAG pipeline.

Frequently Asked Questions

How does Code-Graph-RAG prevent Cypher injection attacks?

The system uses parameterized queries exclusively. Dynamic builder functions return Cypher strings with placeholder variables (e.g., $qn, $id), while actual values are passed separately through fetch_all or _execute_batch_on in graph_service.py. This separation of code and data ensures malicious inputs cannot alter query structure.

What is the difference between MERGE and CREATE queries in the ingestion pipeline?

build_merge_node_query generates MERGE statements that match existing nodes or create them if absent, preventing duplicates during incremental updates. Conversely, build_create_relationship_query produces CREATE statements for relationships, assuming the connecting nodes already exist from prior node ingestion steps.

Can the static Cypher queries be customized for specific codebases?

Yes. The CYPHER_* constants in cypher_queries.py are standard Python strings that can be modified or extended. Developers can add new retrieval patterns following the existing naming convention and parameter structure, then reference them in custom RAG components without changing the core ingestion logic in graph_service.py.

How does batching improve Cypher query performance?

The wrap_with_unwind function transforms single-row operations into batch-capable statements by prepending UNWIND $rows AS row. This allows MemgraphIngestor to transmit hundreds of nodes or relationships in a single database transaction, reducing network latency and transaction overhead compared to executing individual Cypher statements per graph element.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →