# How Code-Graph-RAG Generates Cypher Queries for Its RAG System

> Learn how Code-Graph-RAG generates Cypher queries using static templates and dynamic builders for efficient graph data ingestion and retrieval in its RAG system.

- Repository: [Vitali Avagyan/code-graph-rag](https://github.com/vitali87/code-graph-rag)
- Tags: how-to-guide
- Published: 2026-08-18

---

**Code-Graph-RAG generates Cypher queries through a hybrid architecture that combines static query templates for retrieval operations with dynamic query builders that construct parameterized MERGE and CREATE statements during graph ingestion.**

Code-Graph-RAG is an open-source retrieval-augmented generation (RAG) system that transforms codebases into graph structures stored in Memgraph. To power its RAG capabilities, the system must efficiently generate Cypher queries for both data ingestion and context retrieval. The implementation in `vitali87/code-graph-rag` achieves this through a two-stage approach defined primarily in [`codebase_rag/cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py) and orchestrated by the `MemgraphIngestor` class in [`codebase_rag/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_service.py).

## The Two-Stage Cypher Generation Architecture

The system separates query generation into distinct static and dynamic components, optimizing for both runtime performance and structural flexibility.

### Static Query Catalog for Retrieval Operations

Pre-defined Cypher strings handle common RAG retrieval tasks without runtime string construction overhead. These constants reside in [`codebase_rag/cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/cypher_queries.py) and target specific access patterns. For example, `CYPHER_FIND_BY_QUALIFIED_NAME` (lines 69-74) retrieves entities by their fully qualified identifier, while `CYPHER_TRACE_CALLABLES` (lines 84-89) traces call relationships through the codebase.

These templates utilize parameterized placeholders such as `$qn` and `$prefix`, allowing the same query structure to handle different inputs safely through Memgraph's parameter binding. When the RAG component processes user queries, it selects the appropriate static template and injects runtime values via the `fetch_all` method.

### Dynamic Query Builders for Graph Ingestion

During initial codebase parsing, the system constructs Cypher programmatically to accommodate arbitrary node labels and property dictionaries. The [`cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/cypher_queries.py) file provides several builder functions that return parameterized strings:

- **`build_merge_node_query`** (lines 86-88): Generates `MERGE` statements with dynamic label insertion to prevent duplicate nodes.
- **`build_create_relationship_query`** (lines 19-33): Constructs relationship creation Cypher with parameterized type and property mappings.
- **`build_nodes_by_ids_query`** (lines 63-71): Creates batched node lookup queries for efficient retrieval.
- **`wrap_with_unwind`** (lines 58-60): Transforms single-row operations into batch-capable statements by prepending `UNWIND $rows AS row`.

These builders accept Python data structures and return Cypher strings that separate code from data, preventing injection vulnerabilities while handling heterogeneous graph schemas.

## From Codebase to Graph: The Ingestion Pipeline

The transformation of source code into graph nodes relies on aggressive batching strategies that minimize database round-trips.

### Structure Extraction and Batching

The `StructureProcessor` class in [`codebase_rag/parsers/structure_processor.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/parsers/structure_processor.py) traverses repositories to identify packages, folders, files, and language entities. Rather than executing individual Cypher statements per discovery, it accumulates operations through `MemgraphIngestor` methods:

```python
self.ingestor.ensure_node_batch(cs.NodeLabel.PACKAGE, {...})
self.ingestor.ensure_relationship_batch(
    parent_identifier,
    cs.RelationshipType.CONTAINS_PACKAGE,
    (cs.NodeLabel.PACKAGE, cs.KEY_QUALIFIED_NAME, package_qn)
)

```

These calls stage data in memory until batch thresholds trigger flush operations.

### Converting Batches to Cypher Statements

When `ingestor.flush_all()` is invoked, `MemgraphIngestor` in [`codebase_rag/graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/codebase_rag/graph_service.py) processes accumulated data through `_flush_node_label_group`. This method selects the appropriate builder—`build_merge_node_query` for node labels or `build_create_relationship_query` for relationships—and aggregates rows using `wrap_with_unwind`:

```python
query = build_merge_node_query(label, id_key)
self._execute_batch_on(target_conn, query, batch_rows)

```

The `UNWIND` clause enables single-query batch insertion, significantly improving ingestion performance compared to individual transactions. The `_execute_batch_on` method handles the actual parameter binding and execution against the Memgraph connection.

## Retrieval Queries for RAG Operations

During the retrieval phase, Code-Graph-RAG leverages its static query catalog to fetch context for LLM prompts. The `fetch_all` method in [`graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_service.py) executes parameterized versions of the static templates:

```python
rows = ingestor.fetch_all(
    CYPHER_FIND_BY_QUALIFIED_NAME,
    {"qn": "myproj.services.UserService.get_user"}
)

```

Results are converted to Python dictionaries via `_cursor_to_results` before being passed to the language model for answer generation, completing the RAG loop.

## Practical Implementation Examples

The following patterns demonstrate actual usage of the Cypher generation system within the codebase:

**Ingesting a new function node:**

```python
ingestor.ensure_node_batch(
    cs.NodeLabel.FUNCTION,
    {
        cs.KEY_QUALIFIED_NAME: "myproj.services.UserService.get_user",
        cs.KEY_NAME: "get_user",
        cs.KEY_PATH: "services/user_service.py",
        cs.KEY_START_LINE: 12,
        cs.KEY_END_LINE: 30,
    },
)

```

**Flushing batched data to trigger dynamic query construction:**

```python
ingestor.flush_all()  # Internally calls build_merge_node_query and wrap_with_unwind

```

**Retrieving context for RAG responses:**

```python
result = ingestor.fetch_all(
    CYPHER_TRACE_CALLABLES,
    {"prefix": "myproj.controllers"}
)

```

## Summary

- **Static templates** in [`cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/cypher_queries.py) provide optimized, parameterized Cypher for common retrieval patterns like `CYPHER_FIND_BY_QUALIFIED_NAME` and `CYPHER_TRACE_CALLABLES`.
- **Dynamic builders** such as `build_merge_node_query` and `build_create_relationship_query` construct ingestion queries from runtime Python data structures with proper parameterization.
- **Batch processing** via `wrap_with_unwind` minimizes database round-trips during the ingestion of large codebases by combining multiple operations into single `UNWIND` statements.
- The `MemgraphIngestor` class orchestrates both ingestion and retrieval, executing generated Cypher against Memgraph and returning structured results to the RAG pipeline.

## Frequently Asked Questions

### How does Code-Graph-RAG prevent Cypher injection attacks?

The system uses parameterized queries exclusively. Dynamic builder functions return Cypher strings with placeholder variables (e.g., `$qn`, `$id`), while actual values are passed separately through `fetch_all` or `_execute_batch_on` in [`graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_service.py). This separation of code and data ensures malicious inputs cannot alter query structure.

### What is the difference between MERGE and CREATE queries in the ingestion pipeline?

`build_merge_node_query` generates `MERGE` statements that match existing nodes or create them if absent, preventing duplicates during incremental updates. Conversely, `build_create_relationship_query` produces `CREATE` statements for relationships, assuming the connecting nodes already exist from prior node ingestion steps.

### Can the static Cypher queries be customized for specific codebases?

Yes. The `CYPHER_*` constants in [`cypher_queries.py`](https://github.com/vitali87/code-graph-rag/blob/main/cypher_queries.py) are standard Python strings that can be modified or extended. Developers can add new retrieval patterns following the existing naming convention and parameter structure, then reference them in custom RAG components without changing the core ingestion logic in [`graph_service.py`](https://github.com/vitali87/code-graph-rag/blob/main/graph_service.py).

### How does batching improve Cypher query performance?

The `wrap_with_unwind` function transforms single-row operations into batch-capable statements by prepending `UNWIND $rows AS row`. This allows `MemgraphIngestor` to transmit hundreds of nodes or relationships in a single database transaction, reducing network latency and transaction overhead compared to executing individual Cypher statements per graph element.